Source-linked AI summary
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
Yeabin Moon
TL;DR
The paper asks what buyers receive when they explicitly set reasoning effort versus omit it while holding the model fixed. In a registered Sonnet 5 contrast, explicit high effort cost $0.01031 more per call without a detected accuracy difference, though the interval permits gains up to 4.67 percentage points.
Problem
Existing evidence does not isolate effort-setting effects within a fixed model or establish whether nominally identical omission controls share semantics.
Method
A registered paired contrast held Sonnet 5, prompts, 30 AIME 2026 items, output rail, service product, and price schedule fixed across explicit-high and omitted requests.
Results
$0.01031 per call higher cost accompanied explicit high effort, while the accuracy contrast was +0.0133 and no accuracy difference was detected.
Takeaways & Limitations
Sonnet high and Sonnet omitted are distinct request contracts within one model, extending cost-aware routing with a controlled contract choice.
Takeaways & Limitations
The cohort contains only 30 curated items from one competition-mathematics task family, so interpretation beyond comparable contest problems requires exchangeability.
Abstract
from arXiv · showhide
API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls. Mean delivered cost was \$0.01031 per call higher under the explicit-high contract than under the omitted contract [+\$0.00204, +\$0.01974]. The corresponding accuracy contrast was +0.0133 [-0.0267, +0.0467]; we did not detect an accuracy difference, and the interval permits a gain of up to 4.67 percentage points that this design cannot rule out. Cost per correct answer was \$0.08665 under the high-effort contract and \$0.07662 under the omitted contract, as registered point estimates. A dated contract census, Models-API metadata, and preregistered raw-response probes further documented model-specific omission semantics, including within a provider; claims remained at documentation grade when raw structure was indeterminate. The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined; the resulting claims are bounded to the model, task, and collection date studied.
Introduction
The study treats API pricing as a model-specific contract and tests how omitting an effort term changes delivered cost and accuracy while holding Sonnet 5 fixed. In a registered paired contrast, explicit high effort cost more, whereas the accuracy difference was not detected and could still include a modest gain.
- Motivation: API inference cost varies with completion length and hidden reasoning tokens, so the printed model price does not equal the price of a completed task.Prior evidence found realized costs reversed listed-token-price rankings in roughly one in three of 336 model-task comparisons.
- Open question: Prior work changed configured models and analyzed cost separately from quality, leaving effort-cost effects within a fixed model unresolved.Its paid attempts were not classified into mutually exclusive delivered outcomes such as correct, wrong, rail-exhausted, other no-answer, and provider failures.
- Study design: The registered primary design contrasts Sonnet 5 with explicit high effort against Sonnet 5 with effort omitted, with effort and thinking treated as separate request fields.In the omitted cell, both fields were unspecified; documented defaults were high effort and adaptive thinking, while disabling thinking required an explicit thinking: {"type": "disabled"} request.
- Findings: $0.01031 per call was the mean delivered-cost increase under explicit high effort, while the accuracy contrast was +0.0133 and no accuracy difference was detected.The registered interval permits a gain of up to 4.67 percentage points, so the result does not establish that high effort buys nothing.
Related work
Prior work frames deployment as cost-aware model selection and evaluates capability alongside inference expense. This paper applies those established ideas within one commercial model, comparing reasoning-effort contracts rather than proposing a new metric.
- Cost-aware deployment: Cost-aware deployment treats model selection as an economic decision, using cost-constrained utility maximization and cascades that escalate queries when cheaper responses are inadequate.FrugalGPT treats available services and per-query prices as routing inputs; this paper changes the chosen object by varying a reasoning-effort term within one model.
- Cost-adjusted evaluation: Cost-adjusted evaluation combines task success with inference cost, while this paper divides total delivered cost by correct outcomes within one commercial model.The cost-per-correct quantity is presented as an existing family member, not a new metric, and is applied across two reasoning-effort conditions.
- Prior evidence on realized costs: The Price Reversal Phenomenon reports that listed token prices misorder realized workload costs in roughly one-third of 336 pinned-v2 comparisons.The reported reversals are attributed to thinking-token volume in single-turn tasks and additional turn and context effects in agentic settings.
- Prior evidence on realized costs: That work also reports substantial cost variation across repeated calls to fixed queries, which this paper treats as an established result rather than a contribution.The cited design runs each model at a single reasoning setting, as far as the supplied passage states.
Methods
The study defines comparisons as dated contract cells and evaluates them on a frozen 30-item AIME 2026 cohort with registered repeated-call and cross-model designs. Prespecified parsing, terminal categories, cost reconstruction, evidence grading, and item-clustered uncertainty procedures constrain interpretation to observed contract properties and item-level data.
- Contract design: A contract cell combines requested and served models, effort term or omission, output rail, service product, request mode, prompt, price schedule, and collection date.This definition treats served-model or service-tier mismatches as observable contract deviations and prevents model-specific findings from becoming provider-wide claims.
- Contract design: The main grid used four registered cells across 30 AIME 2026 items, with five calls per item for both Sonnet contracts and one call per item for Terra and Fable.The paired Sonnet contrast is the repeated-call estimand; Terra and Fable provide single-pass cross-model reference points.
- Contract verification: Raw echoed effort, thinking structure, and disabled-thinking request outcomes upgraded only the contract property directly identified, while absent structure with missing or zero token fields remained indeterminate.The fixed evidence hierarchy prevented ambiguous raw responses from upgrading documented omission behavior.
- Outcome measurement: Each attempted call received exactly one terminal category: correct, wrong, no_answer_rail, no_answer_other, or provider_failure.Valid final claims are classified as correct or wrong regardless of stop state, while no-answer and provider-failure categories follow fixed precedence rules.
- Statistical analysis: 10,000 item-clustered percentile-bootstrap draws resampled 30 items with replacement while carrying every call and compared cell belonging to each selected item.Cell means weight item-level values equally, and paired contrasts are formed within shared items before averaging; the analysis seed was 20260713.
- Statistical analysis: The 30-item cohort was eligible for median, p90, p95, and maximum reporting, but not p99 or 90%-CVaR under the preregistered item thresholds.The analysis refused ineligible tail quantities even when software could calculate sample analogues.
Results
Under the registered Sonnet 5 contracts, explicit high effort increased mean delivered cost without a detected accuracy difference, while omission semantics and other outcomes remained model-, task-, and date-specific. The results also document reasoning-token delivery, rail utilization, and descriptive latency within their registered scopes.
- Accuracy and cost: $0.01031 per call higher mean delivered cost under explicit-high than omitted, with accuracy contrast +0.0133 and cost per correct answer $0.08665 versus $0.07662.The cost interval was [+$0.00204, +$0.01974], and the accuracy interval was [-0.0267, +0.0467]; cost-per-correct values were registered point estimates.
- Contract-control evidence: Documentation assigned omitted effort to high for Claude Sonnet 5, while its omitted probe showed no thinking block and zero thinking tokens, leaving omission documentation-grade.The raw evidence could not distinguish adaptive-zero from disabled thinking; a disabled-thinking-plus-high request was accepted.
- Reasoning-token delivery: 140 of 150 Sonnet omitted calls delivered positive reasoning-token counts versus 141 of 150 explicit-high calls, with positive-call medians of 3,133.5 versus 3,359 tokens.These post hoc descriptive quantities show thinking occurred in most omitted calls but do not establish equivalence or a causal effect.
- Rail utilization: No main-grid call ended in no_answer_rail, and none of 360 calls reached 99% rail utilization; Sonnet high reached 50% utilization in 7/150 calls versus 4/150 for Sonnet omitted.At least 90% utilization occurred in 1/150 Sonnet high calls and 0/150 Sonnet omitted calls.
- Latency: Median latency was 17.1 s under Sonnet omitted and 16.7 s under Sonnet high in the registered first-20-item sequential subsample.Latency results were descriptive, excluded the later bounded-concurrency phase, and were not interpreted as causal effects of effort.
Discussion
The registered within-model comparison found higher delivered cost under Sonnet 5’s explicit-high contract without detecting an accuracy difference, while the accuracy interval still permits a gain buyers might value. The discussion limits interpretation to dated, model-specific contracts and emphasizes item-specific failure modes, omission semantics, and descriptive cost-effectiveness rankings.
- Primary result: Up to 4.67 percentage points of accuracy gain remains compatible with the interval and could be valued at the observed premium.The registered cost-per-correct point estimate was also higher under Sonnet high.
- Failure modes: Repeated calls yielded item-specific contract behaviors, including cheap wrong answers, censored no-answers, and one costly correct computation.These observations do not establish how prevalent the behaviors are.
- Failure modes: Visible wrong answers characterized Terra’s unsuccessful calls, while Sonnet also produced no-text no-answer outcomes in the observed cells.This descriptive distinction does not establish provider-wide prevalence or a causal provider difference.
- Comparative interpretation: Terra omitted, Sonnet omitted, Sonnet high, and Fable omitted formed the descriptive cost-per-correct ordering, without implying dominance.No separate cost-per-correct intervals were registered, and the underlying cost and accuracy intervals overlap.
- Contract scope: Omission and thinking-control semantics varied by exact model-contract, including within one provider, so omission cannot be treated as a provider-wide default.Positive thinking structure verifies realized thinking, whereas absence on an easy response remains asymmetric evidence.
Limitations
The study’s limitations concern benchmark freshness and scope, asymmetric repetition and small-item tail resolution, restricted noncausal latency summaries, and the narrow GPT-5.4-mini bridge. Billed hidden computation is measurable in volume and price but remains mechanistically opaque.
- Benchmark scope and freshness: The 30-item AIME 2026 cohort comes from one competition-mathematics task family, while benchmark contamination, cutoff metadata, and later model updates remain uncertain.The 2025-to-2026 comparison is described as a freshness diagnostic rather than proof that AIME 2026 was uncontaminated.
- Sampling and repetition: Five repeats per item were collected for Sonnet high and omitted, but Terra omitted and Fable omitted were observed once per item; 30 items limit tail resolution.The registered plan permits median, p90, p95, and maximum estimates but refuses p99 with fewer than 100 items or 90%-CVaR with few items.
- Latency scope: Latency summaries cover only the first 20 item-distinct sequential calls, are descriptive rather than causal, and cannot separate request-form effects from systematic dispatch order.Sonnet high was dispatched before omitted rather than in randomized or counterbalanced order.
- Cross-model bridge and observability: The GPT-5.4-mini bridge contains five non-contemporaneous items and supports only a dated descriptive anchor, while hidden-computation content remains unavailable despite measurable volume and price.The bridge supports no equivalence, population-accuracy, or cost- and latency-tail inference.
Data, code, and spend disclosure
The study release documents the pinned AIME 2026 dataset, licensing, redistribution limits, and released files. Total API expenditure through the July 18 main collection was $33.043144, reconciled against provider dashboards at cent granularity.
- Dataset and release: AIME 2026 was loaded from MathArena/aime_2026 at pinned Hub revision d2de22f3c656b4f56cf8981212186377d1e23bc3 under CC BY-NC-SA 4.0.The release contains item identifiers, contract configurations, row-level measurements, and derived terminal outcomes.
- Dataset and release: The study release excludes AIME problem text and reference answers, while HEADLINES is not redistributed under CC BY-NC-ND 4.0.The numerical release consists of release_rows.jsonl and cell_statistics.json, among other files.
- Spend reconciliation: $33.043144 was the total API expenditure through the July 18 main collection, including $31.755307 on Anthropic and $1.287837 on OpenAI.The total comprised precollection runs, July 16 contract probes, the main session, and a $0.007600 Anthropic charge from an interrupted non-streaming request.
- Spend reconciliation: Main-session row totals matched both provider dashboards at cent granularity.The supplemental charge was dashboard-confirmed in results/spend_reconciliation_20260722.json.