@penzlik
Anthropic shipped Opus 5 on July 24. Everyone reposted the benchmark chart. Fewer people noticed what the chart actually is.
It is not a score. It is a cost curve.
Same price as Opus 4.8: $5 per million in, $25 out. On ARC-AGI-3, built from problems the model has not seen, it went from 1.5% to 30.2%. Opus 4.8 shipped two months ago.
On OSWorld 2.0 it beats Fable 5's best result at about a third of the cost. Fable 5 is Anthropic's own top tier. They undercut themselves and published the chart that proves it.
The frontier stopped competing on who is smartest. It now competes on who closes the task cheaper. Opus 5 ships with an effort dial: low, medium, high, max. You are not buying a model anymore. You are buying reasoning by the unit.
One line missing from the graphic: context is still 200k. For long-running agents that is the constraint that actually bites, and nothing on that chart measures it.