Source-linked AI summary

Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence

Fabricio F Costa

arXiv:2608.14903v1cs.AI

TL;DR

Frontier-AI forecasts depend on measurements being operational, comparable, jointly observed, and adjusted for changing instruments. This paper audits those preconditions using a frozen event-centric record of systems, benchmarks, events, and sources. It finds sparse measurement joins, scale- and protocol-dependent benchmark links, and concentrated provenance, concluding that dated forecasts must make their measurement system explicit.

  • Problem

    Frontier-AI forecasts lack consistently operational, comparable, jointly observed measurements that account for changing evaluation instruments.

  • Method

    The paper audits these preconditions using a frozen, curated event-centric record of selected systems, versioned benchmarks, graded events, and source records.

  • Results

    Public measurements are sparsely joined, benchmark links depend on scale and design, and 52 of 71 substantive events (73.2%) come from one programme.

  • Takeaways & Limitations

    A defensible dated forecast should explicitly specify its target, joins, benchmark links, protocols, and evidence independence.

  • Takeaways & Limitations

    The benchmark bridges are small and observational, and the MMLU/MMLU-Pro comparison changes construct-relevant content and protocol.

Abstract

from arXiv · show

Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief. This paper audits whether the public measurement record supports those connections before another trend is fitted. I construct a frozen, event-centric record through 12 August 2026 with 62 selected systems, 12 versioned benchmarks, seven capability or impact criteria, 144 graded events, 27 source records, and 408 typed relations. The record is an audit sample, not a census. Only seven systems jointly observe estimated training compute and a METR 50 percent task horizon. Training compute is absent for 19 of 27 closed systems, including every selected closed release from 2026, while none of the 35 open-weight systems has a METR horizon observation. Benchmark succession creates a second break: a seven-system link from METR Time Horizon 1.0 to 1.1 has a log-scale slope of 1.206 (95 percent CI 1.021 to 1.390), whereas a six-system MMLU to MMLU-Pro comparison appears shift-like under logit and probit links but not under linear or logarithmic links. The observed bridges have about 80 percent power only for slope departures near 25 percent. Provenance is concentrated: 52 of 71 substantive quantitative events, or 73.2 percent, come from one measurement programme, and 76.1 percent are laboratory releases. A review of 56 methodological and empirical sources identifies 16 complementary measurement directions spanning resources, inference budgets, reliability, agentic work, safety, human preference, field outcomes, and forecast backtesting. No direction supplies a replacement scalar. The result is not that frontier AI forecasting is impossible, but that a defensible dated forecast is a claim about a versioned measurement system with explicit joins, protocols, links, and source dependence, not merely a fitted curve or calendar date.

1 Introduction

Quantitative frontier-AI forecasts are meaningful only when targets, observations, progression variables, uncertainty models, and measurement instruments are comparable and explicitly defined. This paper audits those preconditions through a reproducible event-centric record, identifying measurement constraints rather than a software failure.

  • Forecasting preconditions: Forecasts select indicators, estimate rates of change, and solve for threshold-crossing dates, but those dates require operational targets and comparable observations.The uncertainty model must also account for changes in the measuring instrument.
  • Measurement problem: The attempted progression analysis could not be identified without strong additional assumptions because capability and resource records were sparsely joined, benchmark versions changed scale, and evaluation protocols changed.The paper characterizes this as a measurement finding, not a software failure.
  • Contributions: The paper’s three contributions are a reproducible event-centric audit, quantification of sparse joint observability, benchmark-linking uncertainty, and provenance concentration, and a measurement portfolio with study designs.Releases, results, revisions, corrections, contamination findings, field experiments, and source assertions are separately timestamped and versioned.
  • Evidence infrastructure: The semantic graph is infrastructure that makes compute-capability links, benchmark-revision changes, and dependence on one programme answerable from a shared evidence ledger.Its role is to support questions about the evidence record rather than serve as the paper’s main scientific claim.

2 From a Score to a Forecastable Quantity

A benchmark score becomes forecastable only when treated as a dated observation produced by a model, instrument, version, protocol, resource budget, and scoring procedure. A defensible forecast must specify its estimand, progression axis, threshold rule, version links, protocol record, and uncertainty from provenance and source dependence.

  • Measurement representation: A benchmark result is an observation generated by a model, instrument, protocol, and time-varying measurement setup—not an intrinsic model property.The record includes benchmark family and version, prompting or agent protocol, inference-resource budget, judge or scoring procedure, and observation date.
  • Forecast requirements: A dated forecast adds a target estimand, progression axis, and threshold rule, with predictor and outcome jointly observed or missingness modelled.Changing benchmark versions, inference budgets, or scaffolding can move scores without changing the underlying system.
  • Forecast requirements: Benchmark succession requires linking function observations that place versions on a defensible common scale, rather than concatenating scores.The protocol and resource choices must remain part of the measurement record because they can alter observed performance.
  • Measurement design: The recommended method is to state the estimand, design linking observations, and preserve the protocol that makes the scale interpretable.Psychometric aggregation can help when response designs and anchors support a latent scale, but small, clustered, or non-normal model samples can destabilize item and ranking inferences.
  • Audit questions: The audit applies stricter standards than merely drawing a regression line, including explicit constructs, sufficient joint observations, version links, protocol records, and uncertainty for provenance and source dependence.These requirements also include replication and source dependence in the uncertainty representation.

3 Method and Data Audit

The study freezes a deliberately non-census public evidence record, normalizes it into versioned events and relations, and audits whether the joins, scale continuity, protocol dependence, and provenance support forecastability. It stress-tests candidate links and power while mapping external measurement alternatives rather than fitting a single replacement metric.

  • Sampling and cutoff: 12 August 2026 is the evidence cutoff for 62 selected systems—27 closed and 35 open-weight—chosen for compute–capability analysis and release importance, not as a complete catalogue.A separate audit records four known pre-cutoff systems absent from the analytic set.
  • Corpus construction: 144 events, 12 benchmark-version records, seven criterion records, 27 source records, 269 nodes, and 408 typed relations form the normalized audit corpus.Every substantive quantitative event has a source identifier and an evidence grade.
  • Audit design: The audit tests joint observability, benchmark-version bridges, scale continuity, protocol dependence, and provenance rather than treating the event graph as the scientific output itself.Training compute, METR horizons, capability results, and peer-reviewed measurements are recorded per model to reveal whether candidate variables can enter the same analysis.
  • Benchmark linking: β = 1 denotes a shift-like relationship on the chosen scale, whereas β̸ = 1 denotes shape change; ordinary least squares with leave-one-out sensitivity is used as a transparent diagnostic.The linking stress tests compare METR Time Horizon 1.0 with 1.1 on a base-2 logarithmic scale and MMLU with MMLU-Pro under four representations.
  • Power analysis: 80% power is calculated for a two-sided α = 0.05 test using the noncentral t distribution, with projected sample sizes holding observed residual noise and anchor-score dispersion fixed.These are design illustrations rather than universal rules, and anchor placement enters the power calculation.
  • Source and literature audit: 56 external sources—29 peer-reviewed and 27 other sources—are coded into 16 measurement directions and normalized into a machine-readable direction–source map.The source audit excludes administrative release rows from provenance-concentration statistics and defines substantive quantitative events separately.

4 Result 1: The Intended Measurement Join Is Sparse

The intended compute–horizon measurement join is sparse and structurally concentrated in closed systems: only seven systems jointly observe training compute and METR p50, while no open-weight system does. This missingness is consequential because the observed systems are historically clustered, and absent compute is tied to disclosure practices rather than plausibly random.

  • Sparse compute–horizon join: Only seven systems jointly observe training compute and METR p50, and no open-weight system supplies the intended compute–horizon join.Training compute is available for 43 of 62 systems, but 19 of 27 closed systems lack an estimate; all 35 open-weight systems have one, while METR measurements occur only in the closed block.
  • Sparse compute–horizon join: The seven jointly observed systems are historically clustered rather than designed across developers, architectures, or access regimes.A descriptive log–log slope can be fitted, but its smoothness does not identify whether compute, algorithms, data, inference effort, model family, or evaluation protocol caused the change.
  • Missingness mechanism: Public compute missingness is unlikely to be ignorable because disclosure relates to access regime, developer practice, and time.Treating absent compute as random would convert an institutional disclosure pattern into a statistical assumption.
  • Missingness mechanism: Controlled training suites such as Pythia provide dense checkpoints and fixed data order for identifying training dynamics, unlike frontier-release records.The contrast underscores that the selected public record lacks an equivalent controlled design.

5 Result 2: Changing Instruments, Concentrated Provenance

The audit finds that benchmark versions and protocols are changing measurement instruments, while the quantitative evidence is concentrated in one programme and laboratory releases. These features limit claims of longitudinal progress and independent replication from the record.

  • Changing Instruments: Benchmark results are versioned events because methodology changes, corrections, contamination findings, measurement ceilings, and replacements affect longitudinal inference.Refreshes and repairs require explicit linking, and item lifecycle is part of the measurement process.
  • Concentrated Provenance: 52 of 71 substantive events (73.2%) come from the METR Time Horizon 1.1 programme, while 54 (76.1%) are laboratory releases.The remaining events comprise 62 benchmark results, four field experiments, four substrate estimates, and one efficiency observation; venue classes also include ten (14.1%) preprints, six (8.5%) peer-reviewed publications, and one (1.4%) vendor source.
  • Changing Instruments: Protocol history can change scores even without benchmark revision, as measurements vary with item order, reasoning mode, persona, conversation history, or the LLM judge.A fixed candidate-response set can receive different scores when the LLM judge is replaced.
  • Concentrated Provenance: The 73.2% concentration does not imply 52 independent replications because shared tasks, protocols, and model-fitting choices can induce correlated error.The percentages describe this audit rather than the full evaluation literature, and the concentration is not presented as a criticism of METR.

6 Result 3: Benchmark Links Depend on Scale and Design

Benchmark links are sensitive to score scale and benchmark design: the METR revision departs from a pure additive log-scale shift, while MMLU-to-MMLU-Pro results vary across link functions and constructs. The observed designs have limited power, supporting designed anchor panels and formal linking methods rather than concatenated leaderboards.

  • METR revision link: 1.206 is the fitted log2 slope linking METR Time Horizon 1.0 to 1.1 across seven systems, with a 95% CI of 1.021–1.390.Leave-one-out slopes range from 1.126 to 1.275, and the revision is not well represented by a pure additive shift on the log scale.
  • Cross-benchmark stress test: 0.976 is the MMLU-to-MMLU-Pro logit slope across six common systems, compared with probit 1.043, linear 1.312, and logarithmic 1.984.MMLU-Pro changes item selection, answer options, and reasoning demands, making this a cross-benchmark stress test rather than a claim of construct invariance.
  • Statistical power: Roughly 80% power was available only for slope departures near 25%; the METR bridge had about 63% power, while the MMLU logit bridge had about 6%.Under the same-noise and same-spread assumption, detecting a 10% departure would require approximately 23–31 common systems.
  • Design implications: Benchmark revisions need designed anchor panels spanning prior score ranges and model families, with uncertainty on both axes.Psychometric linking, benchmark-agreement analysis, and fixed-parameter calibration are more appropriate foundations than concatenating versioned leaderboards.

7 Beyond One Ruler: A Measurement Portfolio

The section argues for a measurement portfolio rather than a replacement benchmark or single AGI score, because frontier-AI progress spans scientifically distinct layers with different evidence properties. It identifies complementary directions and recurring design requirements for making forecasts more valid, joinable, reliable, and externally grounded.

  • Measurement portfolio: A portfolio of measurements is preferable to any single benchmark because frontier-AI progress comprises scientifically distinct layers with different roles.The proposed approach is explicitly not another benchmark that replaces all others.
  • Measurement portfolio: Public availability, joinability, protocol control, and external validity are distinct properties, so the directions should not be averaged into one AGI score.Examples include longitudinally rich but protocol-sensitive human preference, externally strong but difficult-to-attribute field outcomes, selectively missing closed-frontier compute, and non-monotonic safety endpoints.
  • Measurement portfolio: Sixteen complementary directions span resources, behavior, construct validity, human references, safety, deployment, and field outcomes.Figure 5 presents literature-based design assessments and states the potential forecasting role of each direction.
  • Design requirements: Nine recurring design requirements include anchored calibration, open-weight sentinels, elicitation envelopes, resource frontiers, reliability curves, fresh item streams, causal field outcomes, transfer tests, and replicated provenance.The requirements also call for multi-lab replication with assertion-level provenance, supported by model cards, dataset sheets, and formal provenance standards.

8 Implications for Dated Forecasts

A dated frontier-AI forecast is the endpoint of a measurement chain linking resources, deployed-system behavior, and outcomes, with failures arising from missingness, protocol dependence, benchmark drift, and source dependence. Defensible forecasts must expose their measurement choices and where extrapolation begins rather than treating them as invisible.

  • Measurement chain: Forecasts link resources, deployed-system configuration, behavior, institutional outcomes, and a versioned endpoint to a date; four failure modes can break these links.The failure modes are structured missingness, protocol dependence, benchmark drift, and source dependence.
  • Measurement chain: Every dated claim inherits choices about its target, ruler, version link, data join, source dependence, and validation rule.A forecast is the end of a measurement chain, not its beginning.
  • Forecast packet: A defensible forecast packet should specify the operational estimand, benchmark lifecycle and linking function, model configuration, inference budget, joint observations, missingness model, provenance, replications, and frozen backtest.Backtesting should use proper scoring rules or interval coverage after measurement changes.
  • Transparent extrapolation: Provocative forecasting remains compatible with imperfect measurement, provided the forecast exposes where extrapolation begins instead of treating choices as invisible.Release-date models, agent benchmarks, and expert surveys may remain useful despite missing compute, changing benchmarks, or differing criteria.

9 Limitations

The audit’s limitations concern sample coverage, benchmark comparability, review design, public-only evidence, and the limits of event graphs for causal inference. These constraints bound what the record can establish about frontier AI forecasting.

  • 9 Limitations: The 62 systems are a curated analytic sample, not a population estimate, so coverage percentages describe this record rather than all frontier releases through 12 August 2026.Known pre-cutoff omissions are listed, but that list is not guaranteed exhaustive.
  • 9 Limitations: Benchmark bridges are small and observational, while the MMLU/MMLU-Pro comparison changes construct-relevant content and protocol rather than constituting psychometric equating.OLS assumes the old transformed score is measured without error; errors-in-variables methods would be preferable if comparable uncertainty existed for both versions.
  • 9 Limitations: The portfolio review is source-audited and structured, but it is neither a registered systematic review nor a meta-analysis.Its zero-to-three ratings are reasoned design assessments intended to expose tradeoffs, not rank research programmes.
  • 9 Limitations: Only public evidence is represented, excluding potentially denser private evaluations, internal training records, undisclosed inference budgets, and proprietary deployment outcomes.Such evidence would not repair the public record used for independent forecasting, but it limits claims about what developers themselves can infer.
  • 9 Limitations: Event semantics improve traceability, not causal identification: graphs can expose missing joins or source dominance but cannot provide counterfactuals, invariant constructs, or independent replications.The graph therefore diagnoses evidential gaps without resolving them.

10 Conclusion

The audit concludes that frontier AI forecasting lacks a common ruler because resource and capability measurements are sparsely joined, benchmark revisions can change scale, link choices affect diagnosis, and evidence is concentrated. Defensible forecasts therefore require a versioned measurement portfolio and explicit claims about targets, joins, revisions, protocols, and evidence independence.

  • The audited record lacks a common ruler: only seven systems join resource and capability tables, benchmark versions can change scale, link choice affects diagnosis, and one programme supplies most measurements.These limitations motivate moving the measurement system into the forecast rather than treating it as background.
  • A versioned measurement portfolio is needed instead of one purported scalar, spanning resources, efficiency, inference budgets, reliability, agentic work, psychometrics, validity, safety, preferences, field outcomes, and backtesting.Each layer addresses a different question and carries a different uncertainty structure.
  • A dated forecast is a composite scientific claim requiring an operational target, joined observations, linked revisions, controlled protocols, and sufficiently independent evidence.Without these claims, forecast dates may be more precise than what is actually being measured along the way.

Data and Code Availability

The reproducibility package provides the data, provenance records, analysis code, and a single entry point needed to rebuild the paper’s outputs. A clean run regenerates the normalized data, analyses, figures, bibliography, and manuscript PDF.

  • Reproducibility package: The package includes 26 CSV tables, an RDF graph, a 27-source verification log, a cutoff audit, and a 56-source literature table.It also provides the normalized direction–source map, 16-direction measurement portfolio, and nine cross-cutting design requirements.
  • Reproduction: A clean run rebuilds the normalized data, graph, linking and power analyses, all six figures, bibliography, and manuscript PDF.The package includes analysis and figure scripts plus one reproduction entry point.
Loading 2608.14903v1…