Why it matters
Your agent passes offline evals at 90%. You ship. Production immediately finds failure modes your eval never saw. Sound familiar? The culprit is almost always the same: the "customer" in your offline eval is an off-the-shelf LLM that sounds nothing like your real users, and your synthetic test set doesn't capture how m
My takeaway: Build Evals That Actually Matter - Nick Ung, Lyft is an agent-security signal. The practical read is that autonomy, memory, tool permissions, and third-party integrations are the control surface that needs threat modeling and monitoring.