AI Engineer · October 3, 2026

Managed agent sandboxes: separate conversation state, files, and credentials

Managed agent sandboxes: separate conversation state, files, and credentials video thumbnail
Why it matters

Ivan Leo explains the distinction between conversation continuity and a persistent execution environment in Gemini’s managed-agent APIs. An interaction identifier carries conversational state; an environment identifier selects the sandbox and its files. Primary documentation adds explicit environment lifecycle controls and a network proxy that can attach credentials to configured destinations without placing those secrets inside the sandbox. These are separate controls: reusing files does not recreate conversation context, and hiding a credential does not prevent an agent from misusing the authority it grants. Network destinations, token scope, retained state, and environment cleanup still need deliberate configuration.

My takeaway: In a test project, resume conversation and environment IDs independently. Use a narrowly scoped test credential, probe an unapproved destination, inspect agent-visible logs for secrets, and delete the environment when the workflow finishes.
Keep exploring

More curated notes connected through Agent Security and AI Engineering.

OWASP GenAI Security Project · guide

OWASP Top 10 for Agentic Applications for 2026

OWASP's community guide organizes agentic-system risk into ten categories, including goal hijacking, tool misuse, identity and privilege abuse, memory poisoning, insecure inter-agent communication, cascading failures, and rogue-agent behavior. It provides a shared taxonomy and mitigation starting point rather than a certification checklist or evidence that a deployed system is secure.

Microsoft Security Blog · guide

AI vulnerability research: measure reproducible findings and completed fixes

Microsoft’s FORGE account describes the work between a model’s vulnerability claim and a useful repair: reusable builds, duplicate removal, reachability checks, project-specific verification, reproducible triggers and regression tests. Structured rejection reasons help improve later searches. The useful operational measure is the flow of findings that survive verification and reach a fix, rather than the number of candidates generated. Reported successful-case costs exclude parts of screening, failed attempts and human work, so they are not the total cost of operating this pipeline.

OpenAI News · framework

Frontier training safety cases: connect evidence to enforced pause and rollback controls

OpenAI proposes training-run safety cases combining alignment evaluations, containment and monitoring with explicit operational ownership. Concrete measures include immutable transcripts, held-out incident tests, checks for evaluation gaming, response deadlines and fail-closed monitoring. Independent internal challenge, leadership vetoes and tracking downstream uses support stopping a run and reversing affected work. The article describes recommendations still being implemented, rather than audited proof that every safeguard already operates or that residual risk has been eliminated.