Why it matters
METR introduces the “expenditure horizon”: the budget at which a human and an AI agent produce equal gains on an optimization problem. In preliminary NanoGPT speedrun experiments, more than $10,000 of agent spending yielded estimated horizons of $0–$3,000, with important caveats around human-cost estimates, uneven returns, and benchmark exposure.
My takeaway: This makes agent capability an economic return curve rather than a binary pass/fail score. Evaluations should report token costs, experiment compute, human labor, uncertainty, and hybrid human-agent performance so apparent gains can be compared with the resources required to produce them.