Why it matters
METR organizes agent-capability measures around performance as a function of expenditure, comparing fixed-budget scores, cost to reach a score, returns to test-time scaling, human-equivalent time and expenditure horizons, and human-relative cost. It explains when familiar benchmark scores break down—particularly when performance keeps improving with more inference or human benchmarks saturate—and notes that full cost, reliability, coverage, and elicitation choices affect the result.
My takeaway: Do not compare agents on pass rate alone when test-time compute varies. Plot score against total expenditure, state the elicitation budget, pair capability with cost and reliability, and use a human-grounded metric only where the human comparison remains meaningful.