METR ยท July 28, 2026

How independent researchers could investigate AI propensities after misalignment incidents

Why it matters

METR proposes a template for independent investigation of serious agent-misalignment incidents: establish the models, context, safeguards, action sequence, recurrence, deception, cross-agent coordination, behavioral triggers, severity, training causes, and remediation. Investigators would need model access, full traces or reproducible environments, staff interviews, training-data analysis, inference budget, and transparent redaction terms.

My takeaway: Pre-negotiate incident-investigation access before deployment. Preserve prompts, memory, tool traces, checkpoints, safeguards, and training lineage; enable controlled reproduction and ablation; give investigators relevant staff and compute; and publish scope, access, and redaction limits alongside findings so boards and the public can judge the level of assurance.