Google DeepMind Blog · June 24, 2026

Introducing computer use in Gemini 3.5 Flash

Why it matters

Google integrated computer use into Gemini 3.5 Flash so agents can act across browser, mobile, and desktop environments. Optional enterprise safeguards can require confirmation for sensitive actions or stop a task when indirect prompt injection is detected.

My takeaway: Built-in computer use expands the model’s action surface. Pair model-level defenses with sandboxing, least-privilege access, explicit approval for irreversible actions, and independent monitoring that can halt compromised sessions.
Keep exploring

More curated notes connected through Agent Security and Prompt Injection.

OWASP GenAI Security Project · guide

OWASP Top 10 for Agentic Applications for 2026

OWASP's community guide organizes agentic-system risk into ten categories, including goal hijacking, tool misuse, identity and privilege abuse, memory poisoning, insecure inter-agent communication, cascading failures, and rogue-agent behavior. It provides a shared taxonomy and mitigation starting point rather than a certification checklist or evidence that a deployed system is secure.

OECD.AI Wonk · guide

A five-step roadmap to closing the AI evaluation gap

The roadmap addresses evaluation results that overstate real-world performance or fail to transfer across deployment contexts. Its five steps balance standardized and local tests, evaluate throughout the lifecycle, build qualified assurance and communication capacity, tailor tests to each value-chain actor and technology, and use a coordinated, trusted process for updating methods.

OpenAI News · guide

The Defender’s Window

OpenAI describes a staged program for AI-assisted defense: use agents to review code and infrastructure, triage alerts, enumerate attack paths, and validate security invariants while retaining strong isolation and least privilege. Its recommended rollout starts with internet-facing services and vulnerability backlogs, moves security review into CI, requires focused fixes and regression tests, and expands from read-only triage to narrowly bounded automation only after teams build evidence and confidence.