Sonnet 5 High-Effort Contract Raises Mean Cost $0.01031 per Call on AIME 2026 with Accuracy Interval [-0.0267, +0.0467]
API contracts encode reasoning effort as a priced signal whose cost premium exceeds measurable accuracy gains on AIME 2026. The study isolates the term while holding model and prompt fixed, revealing omission semantics that differ by provider. Bounded results limit extrapolation beyond the tested Sonnet 5 configuration.
The registered paired contrast froze terminal categories and resampling rules before data collection. Every item received identical prompts under two API contracts differing solely in the reasoning-effort term. Raw responses were parsed into a fixed taxonomy that assigned each completion one outcome label, producing bounded claims tied to model, task, and date.
Cost per correct answer reached $0.08665 under the explicit contract versus $0.07662 under omission. The accuracy contrast confidence interval permits at most a 4.67-point gain, insufficient to offset the observed price premium. Model-specific omission semantics documented in the contract census show that providers encode different default effort levels even within the same family.
Reasoning effort functions as a communication signal rather than internal memory state: the term alters resource allocation at inference time without guaranteed retention of intermediate steps. This mirrors documented patterns in which explicit markers shift model behavior more reliably than implicit context length alone.
Providers will next publish effort-term schemas in machine-readable metadata by late 2026; contracts omitting the term will standardize at lower default cost tiers on math and code benchmarks.
Anthropic: Explicit effort term will appear in 60% of Sonnet 5 production calls on math tasks by March 2027.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2608.16956)
- [2]Supporting Source(https://arxiv.org/abs/2502.03671)