Source-linked AI summary
It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories
Yigit Utku Bulut
TL;DR
The paper asks whether reasoning traces reveal genuine breakthrough states and early outcome information, noting that existing measurements omit controls matched to those claims. It introduces restart-controlled truncation and difficulty-controlled prediction analyses, finding that breakthroughs are usually budget artifacts and pooled early-signal performance does not establish within-attempt information. The authors therefore argue for matched restart, question-only, or within-problem controls when interpreting trajectory-level claims.
Problem
Reasoning-trace claims about breakthroughs and early-legible outcomes lack counterfactual controls that distinguish prefix value from budget fit and within-attempt information from problem difficulty.
Method
The paper compares continuation solve rates with matched-budget from-scratch restart curves and evaluates early signals using difficulty-controlled, within-problem analyses of model and public-corpus data.
Results
Exactly 1 of 178 cells is prefix-limited, while matched-budget continuation beats restarting in 9 of 9 comparisons; pooled early-signal AUROCs are explained by difficulty and are chance-level within problems at t=4.
Takeaways & Limitations
High pooled probe AUROCs should be treated as difficulty measurements until question-only baselines or within-problem evaluations demonstrate within-attempt outcome information.
Takeaways & Limitations
The instrumented evidence is limited to two small models, one benchmark family, 4-bit quantization, and generated-token matching rather than FLOPs or latency.
Abstract
from arXiv · showhide
Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the continuation budget from when a prefix carries value that fresh computation cannot buy, comparing per-anchor continuation solve rates against from-scratch restart curves at matched total generated-token budget. Applied to 178 problem-model cells (89 MATH problems x two small open models, an outcome-blind but difficulty-targeted cohort), exactly 1 of 178 cells survives as prefix-limited; restart dose-response separates a compute-starved model from a capability-limited one; and wherever the matched budget lies inside the restart grid, continuing the model's own prefix beats restarting (9 of 9) -- predominantly compute compression rather than expanded reachability. Second, a pre-registered, difficulty-controlled test finds no detectable outcome information in early-window internal signals beyond a problem-difficulty baseline, and two generation-free analyses of public corpora show why this control is needed: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations -- inside the published probe range -- and a closely matched reconstruction of the closest published early-window positive recovers a comparable pooled result (0.849) while within problem it is statistically indistinguishable from chance at all ten anchors (0.496 at t=4); a post-hoc within-targeted probe finds only a small average residual, concentrated in three low-failure problems. High pooled probe AUROCs cannot by themselves establish within-attempt information; a question-only baseline or within-problem evaluation is required.
1 Introduction
The paper argues that claims about breakthrough moments and early-legible outcomes require counterfactual controls matched to the claim. Applying restart and difficulty controls, it finds breakthroughs are mostly budget artifacts and pooled early-signal performance can reflect problem difficulty rather than within-attempt information.
- Motivation: Two prominent interpretations of reasoning traces—breakthrough moments and early-legible outcomes—lack counterfactual controls at the level of the claim.Breakthrough claims need matched fresh computation, while prediction claims need question-only and within-problem controls.
- Controls and approach: The study compares per-anchor continuation solve rates with from-scratch restart curves at matched total generated-token budget across 178 problem–model cells.The cohort contains 89 MATH problems and two small open models selected by outcome-blind intermediate-difficulty screens.
- Restart findings: Exactly 1 of 178 cells survives as prefix-limited, indicating that apparent breakthroughs are usually explained by continuation-budget fit rather than prefix-locked value.Three cells crossed the advantage margin, but an in-grid restart already solved two; 98 post-freeze cells added no new cases.
- Restart findings: Restart dose–response distinguishes a compute-starved model, whose success rises from 0% to 79% with budget, from a capability-limited model that does not improve within the measured range.This separates budget failure from capability failure rather than treating solvability as binary.
- Restart findings: At matched total token budget, continuing the model’s own prefix beats restarting in 9 of 9 comparisons, supporting compute compression as the predominant role of accumulated reasoning.A larger restart budget reaches the same threshold in 11/13 cases and the prefix’s own rate in 9/13.
- Prediction findings: Early internal signals show no detectable outcome information beyond difficulty, while public analyses reach pooled AUROCs of 0.873 and 0.849 but are chance-level within problems at t=4.The within-targeted residual is small on average and concentrated in three low-failure problems.
2 A Restart-Controlled Probe of Reasoning Value
The restart-controlled probe truncates model trajectories, measures continuation solvability, and compares it with matched-budget fresh restarts. Its design distinguishes budget-fit time from prefix-value time while using replication, censoring rules, and preregistered decision procedures to stabilize classifications.
- Setup: The study instruments two small open-weight models on MATH with token-level logits and hidden states, using outcome-blind intermediate-difficulty cohort selection.The models are Gemma-4 E4B and Ministral-3 3B, run locally in thinking mode with 4-bit quantization.
- Truncation probe: The truncation probe samples four continuations from each trajectory prefix under a 1,024-token reasoning budget plus a 512-token answer reserve.Anchors are log-spaced from 16 to 8,192 tokens, and the first stable solve-rate crossing defines budget-fit time TF.
- Restart control: Restart curves estimate from-scratch solve rates at budgets of 1,024, 2,048, 4,096, and 8,192 generated tokens.Rates are interpolated log-linearly between grid points, with four attempts per grid budget.
- Restart control: The matched-budget prefix advantage is adv(t) = p̂(t; B) − R̂(t + B), and prefix-value time TV is the earliest stable anchor meeting both solve-rate and advantage margins.The primary advantage margin is δ = 0.5, with sensitivity analyses at 0.25 and 0.75.
- Robustness and preregistration: Replication enlarges ambiguous cells and strengthens terminal events, reducing optimistic single-shot labels before classifications are derived.All ambiguous 1–3-of-4 cells were expanded to eight attempts; pooling removed six events and added one.
- Robustness and preregistration: Decision rules and the confirmatory prediction protocol were committed before governed outcomes were observed, with gate failures recorded rather than relabeled.Post-hoc analyses are explicitly labeled in the amendment reporting them.
3 Results I: Breakthroughs Under Control
Across 178 problem–model cells, restart controls show that nearly all apparent breakthroughs are explained by continuation-budget sufficiency rather than prefix-locked value. Restart dose–response also separates compute-starved from capability-limited failures, while many problems remain probabilistically intermediate.
- Most measured breakthroughs are budget artifacts: Exactly 1 of 178 cells is prefix-limited, making genuine prefix-locked value rare under the study’s frozen taxonomy.The count is descriptive for an outcome-blind but difficulty-targeted cohort; repeated problems across models are not independent.
- Restart dose–response separates failure modes: Ministral-3’s restart solve rate rises from 0% to 79% with budget, whereas Gemma-4 remains flat at 10–12%.The measured restart budgets span 1,024 to 8,192 tokens, separating compute-starved from capability-limited failures within this range.
- Accumulated reasoning is predominantly compute compression: At exactly matched budgets, continuing the model’s prefix beats restarting in all 9 comparisons, but larger restarts reach the threshold in 11/13 cells.The matched-prefix advantage is therefore primarily compute compression rather than expanded reachability within the measured range.
- A persistent band of intermediate solvability: A third of ambiguous cells remain intermediate after eight attempts, with 55 of 148 landing at 3–5 successes.The independent second half of attempts reproduces this intermediate band, despite wide eight-attempt intervals.
4 Results II: A Difficulty-Controlled Null for Early Signals
The paper tests whether early internal signals predict outcomes beyond problem difficulty and finds no detectable incremental information in the confirmatory or post-hoc analyses. Public-corpus controls show that pooled probe performance can arise from difficulty, while within-problem evaluation is near chance.
- 4.1 A pre-registered, difficulty-controlled null: Neither confirmatory endpoint showed a detectable gain from early-window dynamics beyond the question-only difficulty baseline.Primary AUROC changed from 0.83 to 0.85, ∆=+0.026 [−0.054,+0.167]; secondary AUROC changed from 0.88 to 0.78, ∆=−0.090 [−0.213,+0.033].
- 4.1 A pre-registered, difficulty-controlled null: No tested forecast window from 128 to 2,048 tokens yielded a detectable gain in the post-hoc sweep.
- 4.1 A pre-registered, difficulty-controlled null: Difficulty alone nearly solved the Gemma-4 prediction task but left substantial unexplained variation for Ministral-3.The difficulty baseline reached test AUROC 1.0 for Gemma-4 and 0.69–0.73 for Ministral-3 across the two endpoints.
- 4.3 The ceiling exists in the literature’s own regime: A trace-blind leave-one-out difficulty proxy reached AUROC 0.873 on 192,315 DeepSeek-R1 generations, within the published internal-probe range.The proxy reads neither the reasoning trace nor, for 92% of problems, more than a single binary observation.
- 4.4 Dissecting a published probe positive in the reasoning-model regime: The reconstructed early-window positive reached pooled AUROC 0.849 at t=4 but was statistically indistinguishable from chance within problem at every anchor.Within-problem AUROC was 0.496 [0.466,0.527] under the pre-registered failure-count weighting and 0.515 [0.481,0.562] under exact discordant-pair weighting.
- 4.4 Dissecting a published probe positive in the reasoning-model regime: Post-hoc within-attempt information was small on average and concentrated in three low-failure problems.A problem-centered probe was at chance at t=4, while a per-problem oracle found AUROC 0.78–0.92 in three problems.
5 Related Work
Prior work uses truncation and resampling to study reasoning value, but these probes often lack matched-budget restart controls and difficulty-controlled within-problem evaluation. This paper reframes the questions around compute compression versus reachability and outcome information versus problem difficulty.
- Counterfactual controls: Matched-budget restart controls distinguish prefix value from solutions becoming feasible only because the continuation budget increased.The paper identifies this as the missing control for interpreting truncation crossings.
- Process supervision: Intermediate rollout values remain common after eight attempts, challenging single-crossing and monotone-value assumptions used in some process-supervision schemes.Roughly a third of ambiguous states remain intermediate, with success between 0.3 and 0.6.
- Sequential versus parallel compute: 9/9 exactly matched comparisons favor continuing the model’s prefix, while an 8,192-token restart reaches threshold in 11/13 cases.These results characterize accumulated reasoning primarily as compression rather than expanded reachability.
- Sequential versus parallel compute: Restart dose–response curves separate compute-starved failures that dissolve with budget from capability-limited failures that remain unsolved.This provides a diagnostic interpretation for sequential-versus-parallel test-time compute comparisons.
- Outcome prediction: Published outcome probes often omit difficulty controls, allowing pooled performance to reflect which problem is being attempted rather than how the attempt is progressing.The paper calls for question-only baselines and within-problem evaluation.
6 Discussion and Limitations
The discussion argues that restart curves and within-problem controls change how trajectory breakthroughs and pooled outcome probes should be interpreted. It also bounds the conclusions to the studied models, benchmark, measurements, and underpowered interior-timing analyses.
- Discussion: A four-point restart curve can identify whether a deployment is compute-starved or capability-limited.For compute-starved models, abandoning a long prefix typically costs a budget multiple; for capability-limited models, neither continuation nor restart helps.
- Discussion: Pooled probe AUROCs should be treated as difficulty measurements until question-only baselines or within-problem evaluations demonstrate within-attempt information.The within-problem dissection provides the sharpest evidence, while the preregistered cohort null is supporting evidence.
- Discussion: Within-targeted information is small on average and concentrated in a few low-failure problems, unlike the difficulty signal reflected in pooled positives.Rare failures on easier problems are more legible from early tokens, whereas common failures on hard problems are not.
- Limitations: The study’s main limitation is scope: two small 4-bit-quantized models, one benchmark family, limited attempts, and token-based rather than FLOP- or latency-matched budgets.The public dissection also uses full-precision generations rescored through a 4-bit model and inherits unknown sampling temperature.
- Future work: Future work should apply the frozen instrument to a large RL-trained reasoner, where interior events would directly revisit the timing question.The paper also proposes studying the shape of the value curve rather than only its threshold crossings.
A Protocol Amendments and Pre-Registration Timeline
The study froze its protocols and amendments before observing governed outcomes, then documented staged responses to budget, censoring, cohort, and early-signal issues. It also fixed the token-level measurements and summary features used for probing.
- Pre-registration and amendments: Every decision rule was frozen in a dated amendment before the outcomes it governed were observed.The amendments and their gate resolutions were archived with the accompanying code repository.
- Breakthrough forecasting protocol: The breakthrough protocol fixed anchors, four continuations, a 1,024 + 512-token budget, threshold τ = 0.75, censoring rules, and preassigned research splits.The target was within-trajectory forecasting of P(TF ≤t + k | features through t).
- Protocol amendments: A1–A3 addressed censoring, budget sensitivity, and cohort expansion; A2 falsified Ministral-3 labels while validating Gemma-4’s early crossings.At the larger budget, 16-token prefixes solved for Ministral-3, whereas Gemma-4’s early crossings remained valid.
- Restart control: The restart-controlled A5 protocol froze restart curves, matched-budget advantage, δ = 0.5, regime precedence, and interpolation sensitivity.Its pilot gate scored 7/8 against the required 8/8 and was recorded as a failure before the permitted enlargement resolved the borderline cell.
- Cohort construction and label hardening: Expansion proceeded through staged, outcome-blind waves with numeric gates, while ambiguity enlargement trimmed labels by removing six events and adding one.The wave-1 gate failed because screening selected 16K terminal solvability while probes tested 1,024-token continuations; the wave-3 screen was repaired before outcomes existed.
- Confirmatory test: The confirmatory early-signal test froze endpoints, features, model class, folds, power gates, and success criteria; both endpoints failed, and the horizon endpoint was not fit.The forecast-point sweep was explicitly labeled post-hoc.
- Signal construction: Token-level instrumentation used next-token distributions and final-layer hidden states, with entropy, top-1/top-2 margin, and trajectory statistics summarized over t ≤512.The fifteen frozen summaries included means, standard deviations, robust slopes, surprisal max-rise, divergence statistics, hidden-state changes, and spectral summaries of Jensen–Shannon divergence.
B.2 The budget falsification that motivated the restart control
Budget sensitivity showed that apparent breakthroughs can disappear when continuation budgets increase, motivating a restart-controlled comparison. The effect differed across models: Ministral-3’s crossings dissolved, whereas Gemma-4’s remained.
- Budget sensitivity: 16-token Ministral-3 prefixes solved at the 4,096-token budget, dissolving their apparent mid-trajectory crossings.The same development-cohort cells were re-probed with paired seeds.
- Budget sensitivity: Gemma-4’s early crossings remained pinned at the larger budget, supporting their interpretation as genuine rather than budget artifacts.The figure contrasts the models under the same budget manipulation.
- Noise hardening: Terminal replication confirmed 13 of 16 Ministral-3 threshold-censored candidates and rejected Gemma-4’s single candidate.Pooling ambiguous cells to eight attempts changed labels on 7 + 6 of 49 + 49 trajectories, mostly through losses.
- Noise hardening: Noise hardening removed six events and added one across the project, trimming rather than inflating the expansion-cohort labels.Gemma-4’s changes were mostly interval shifts.
B.4 Sensitivity of TV labels
Sensitivity analyses show that TV labels are robust to branch-based validation and conservative interpolation, while changing the advantage threshold alters only a small number of prefix-value events.
- Independent-half check: 58% of 148 independently enlarged cells fell in the intermediate 1–3-success range across branches 4–7.The branch counts were never used for selection, countering concentration at the extremes expected from pure selection artifacts.
- Threshold sensitivity: At δ = 0.5, 3 of 178 cells were prefix-value events; δ = 0.25 yielded 7, while δ = 0.75 yielded 2.All additions at δ = 0.25 occurred inside budget-limited cells of the budget-elastic model.
- Interpolation and cross-model robustness: The conservative upper-envelope restart estimate agreed with primary labels on 178 of 178 cells.Collapsed regime classes agreed across models for 42 of 89 problems.
B.5 Confirmatory test: level-free variants and per-model detail
Level-free and per-model analyses provide descriptive detail around the confirmatory test, whose table reports that no variant met the success criterion. The horizon endpoint lacked sufficient power for model fitting, while numerical and teacher-forcing checks passed.
- Level-free variants: No level-free variant in Table 1 met the success criterion after difficulty level was dropped from both feature sets.The table reports held-out test AUROCs for both endpoints.
- Per-model detail: The primary endpoint baseline scores were 1.00 for Gemma-4 and 0.69 for Ministral-3, with Ministral-3 early signals at 0.76.These are per-model test-split descriptives, not per-model claims.
- Per-model detail: The secondary endpoint scores were 1.00 for Gemma-4 and 0.73 for Ministral-3.The reported sample sizes were n=11 and n=14, respectively.
- Power gate: The horizon endpoint had 21 project-wide interior events and 4 in the test split, below the frozen power gate, so no model was fit.The gate required five problem groups per class per fold.
- Probe evaluation: The hidden-state probe table reports pooled and within-problem metrics across anchors, with problem-disjoint out-of-fold prediction and failure-count weighting.The pooled subset contains 1,024 trajectories at t=4, while the within-problem analysis covers 22 mid-band problems.
- Robustness checks: Float32 reproduced float64 with zero NaN out-of-fold predictions, and teacher-forcing fidelity reached mean top-1 agreement 0.906.The 10th-percentile agreement was 0.859; scikit-learn overflow warnings were cosmetic.
B.7 Post-hoc within-problem robustness checks
Post-hoc within-problem checks find heterogeneous, limited early signal: strong separability occurs in a few low-failure problems, while high-failure problems show none.
- Heterogeneous signal: 84 of 2,178 failures come from three strongly separable problems: 20, 22, and 13.At t=4, oracle AUROCs are 0.92, 0.84, and 0.78, respectively.
- Heterogeneous signal: Problems 20 and 22 have failures four to six times shorter than successes, indicating an early answer-without-reasoning mode.Median lengths are 2,364 versus 15,321 and 1,299 versus 5,168 characters.
- Aggregate interpretation: The high-failure problems dominating the failure-weighted mean show no separability, so within-attempt information is small on average and concentrated in rare-failure cases.A transferable component appears only from t ≈32, whereas the rare-failure cases are visible from the first tokens.
- Reported endpoints: Table 3 reports failure-weighted means, per-problem oracles, the largest oracle, and counts of problems above 0.6, while Table 4 lists both endpoints at every forecast point.Table 4 also marks the pre-registered evaluation and notes that larger prefixes shrink and select the sample.
B.9 Reproducibility details
The paper fixes generation, verification, cohort construction, classifier, validation, confidence-interval, and exclusion procedures to support reproducible analyses.
- Generation and verification: Both models use temperature 0.6, top-p 0.95, top-k 20, thinking mode, prompt v1, and a 16,384-token base-trajectory cap.Probe and restart branches use deterministic hash-derived seeds, with checkpoints pinned to released readiness manifests.
- Generation and verification: Final answers are extracted and scored by numeric or symbolic equivalence; GPQA uses the extracted option letter, and unparseable samples score incorrect.The reported unparseable rate is 3.0%.
- Cohort construction: Development and supplement problems come from a 100-problem level-balanced pool selected by frozen outcome-blind rules, while expansion screening uses six fresh 3,072-token attempts.The screen selects problems with 1–5 verified successes and spans two models and three seeds.
- Analysis configuration: Classifier analyses use standardized balanced logistic regression, problem-grouped stratified folds, and PCA capped at 128 components for the public-dump probe battery.Instrumented analyses use C = 0.1; public-dump analyses use scikit-learn defaults with C = 1.
- Uncertainty and exclusions: Confidence intervals use percentile bootstraps over 2,000 problem-clustered draws, with 1,000 draws for Appendix B.7 post-hoc checks.
- Uncertainty and exclusions: Fifteen of 178 cells ending before the 512-token forecast window are excluded from confirmatory endpoints under the frozen inclusion rule.The exclusions comprise 10 training, 4 validation, and 1 test cell.