Source-linked AI summary
A Statistical Audit of Physical AI Benchmark Redundancy
Zaruhi Navasardyan, Hrant Davtyan
TL;DR
Physical AI evaluation lacks a common, sufficiently overlapping benchmark basis, and overlapping benchmarks can double-count shared abilities in reported averages. The paper audits a 51-model, 12-benchmark matrix, quantifies redundancy, and selects a smaller suite that preserves most of the full suite’s utility while supporting Bradley–Terry ranking. Its audit finds substantial, compressible redundancy and is applicable beyond physical AI.
Problem
Physical AI lacks a common evaluation basis because model reports use hand-picked benchmark sets with little overlap, while relationships among benchmarks and possible double-counting remain unmeasured.
Method
The paper assembles a 51-model, 12-benchmark score matrix, measures redundancy through benchmark relationships, and greedily selects benchmarks using dispersion and marginal information before fitting a Bradley–Terry ranking.
Results
Four selected benchmarks retain 78.5% of the utility of all 12, while collapsing two substitute pairs changes 22 of 51 model positions by at least three places.
Takeaways & Limitations
The audit treats benchmark suites as measurement instruments, showing that redundancy can be compressed and that the procedure can apply to any field with sufficiently overlapping benchmark-level scores.
Takeaways & Limitations
The evidence is observational, and none of the 12 benchmarks measures downstream physical-system task success, so transfer to manipulation or navigation cannot be established.
Abstract
from arXiv · showhide
Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct a matrix of 51 models on 12 physical AI benchmarks, selected from a registry of 51 benchmarks and 152 models by reporting density, combining scores from model cards and benchmark papers with our own evaluation runs under each benchmark's official protocol. We measure how much information the benchmarks share and show quantitative evidence of Redundancy. Redundancy affects reported rankings: collapsing the two substitute pairs into single columns moves 22 of 51 models by three or more places under an equally weighted average. We then select benchmarks greedily under a utility combining score dispersion with variance not explained by the already-selected set, and obtain a four-benchmark subset retaining 78.5\% of the utility of all 12, on which we fit a Bradley--Terry ranking. The procedure requires only benchmark-level scores with sufficient overlap and is not specific to physical AI.
1 INTRODUCTION
Physical AI evaluation lacks a common comparison axis because model reports use sparse, hand-picked benchmark sets, while overlapping benchmarks can double-count shared abilities. The paper addresses this by building a dense matrix and auditing redundancy and suite sufficiency.
- Motivation: Public model reports leave the model × benchmark matrix mostly empty, preventing direct comparison on a shared leaderboard.Reported headline averages can also silently double-count abilities when selected benchmarks overlap.
- Motivation: Pointing, 3D layout reasoning, and relational question answering are repeatedly evaluated across multiple benchmarks.Examples include Point-Bench, RefSpatial, RoboSpatial-Pointing, and Where2Place across recent model reports.
- Research gap: No prior work had quantified how much unique signal a physical-AI benchmark adds beyond benchmarks already in use.New benchmarks may address saturation, leakage, or shortcuts, but their distinctiveness premise was not checked.
- Approach: The authors assemble scores for 51 models on 12 physical-AI benchmarks and use the matrix to quantify redundancy and select a compact suite.Scores combine model cards, benchmark papers, and evaluations under official protocols.
- Research questions: The study asks how much benchmarks share information, how redundancy inflates pooled averages, and how small a suite can remain while separating models.These questions are framed as RQ1 Redundancy and RQ2 Sufficiency.
2 THE BENCHMARK–MODEL MATRIX
The benchmark–model matrix is constructed from a broader registry using density, recency, and diversity criteria, then completed with targeted evaluations under official protocols. Scores are standardized for multivariate analysis while rank correlations remain the default pairwise measure.
- Benchmark selection: Twelve benchmarks are selected from 51 using density, recency, and diversity criteria.Density requires reporting for at least 5 candidate models; recency reflects repeated use in recent reports.
- Score processing: Benchmark means range from 30 to 82 points, so raw score comparisons can confound difficulty with information.The paper therefore z-scores benchmark columns before multivariate steps and retains Spearman correlations for pairwise analysis.
- Benchmark selection: The selected suite spans pointing, single-image relational reasoning, and multi-view or video settings.The design aims to cover diverse tasks and settings used by physical-AI benchmarks.
- Model selection: The final model set contains 51 models selected from 152 registry models, each with scores on at least 8 of 12 benchmarks.Models span releases from 2024 to 2026 and 15 providers.
- Matrix completion: Published scores are supplemented with 159 targeted evaluations using official code, prompts, decoding, and answer-parsing rules where available.The runs fill missing matrix cells rather than constitute a systematic reproduction study.
3 REDUNDANCY (RQ1)
The 12-benchmark suite contains substantial positive redundancy, including two highly correlated substitute pairs and benchmarks predictable from peers. Collapsing substitutes changes many model rankings, showing that suite composition affects aggregate positions.
- Pairwise redundancy: An average pairwise Spearman ρ of 0.487 shows that all 12 benchmark correlations are positive.Spearman correlations compare model rankings independently of benchmark difficulty.
- Pairwise redundancy: EMBSPATIAL ↔ CV-BENCH correlate at ρ = 0.876, while WHERE2PLACE ↔ REFSPATIAL-BENCH correlate at ρ = 0.860.Both pairs are identified as substitutes with correlations above 0.8.
- Pairwise redundancy: BLINK forms an individual dendrogram branch, while ERQA + REALWORLDQA and REALWORLDQA + OMNISPATIAL show ρ = 0.758 and 0.748.These patterns distinguish a highly unique benchmark from other notable associations.
- Ranking consequences: Equal weighting gives pointing 2/12 of the aggregate weight and video-spatial reasoning 1/12 because abilities are represented by different numbers of benchmarks.This illustrates how correlated benchmark suites can disproportionately weight repeatedly measured capabilities.
- Ranking consequences: 22 of 51 models move by at least three ranking places after the two substitute pairs are collapsed into single columns.MIMO-EMBODIED-7B drops 9 places, while GPT-4O gains 9 places under the recomputed arithmetic average.
4 A MINIMAL BENCHMARK SUITE (RQ2)
The paper greedily selects a compact benchmark suite by combining model discrimination with information not already explained by selected benchmarks. Four benchmarks retain 78.5% of the full suite’s utility and support a Bradley–Terry ranking.
- Selection utility: The utility of candidate benchmark b is the product of discrimination and marginal information not explained by selected set S.Discrimination uses the Gini coefficient; marginal information is 1 − R2 from an OLS regression of b on S.
- Selection procedure: Forward selection starts with the benchmark that separates models most sharply, then recomputes every remaining candidate’s utility after each addition.The procedure runs through all 12 benchmarks and reads the stopping point from the cumulative-utility curve rather than fixing suite size in advance.
- Selected suite: 78.5% of the all-12 utility is reached by REFSPATIAL-BENCH, MINDCUBE, VSI-BENCH, and BLINK.WHERE2PLACE enters fifth at 85.0% cumulative, while the remaining seven benchmarks share the final 15.0%.
- Selected suite: WHERE2PLACE is passed over despite the suite’s second-highest Gini because REFSPATIAL-BENCH already reproduces most of its variance.This illustrates why marginal information depends on the benchmarks already selected rather than on benchmark quality in isolation.
- Selected suite: The selected four cover precise localization, spatial consistency across limited views, spatial reasoning over video, and multi-image perceptual primitives.The utility also balances MINDCUBE’s hub-like predictive coverage with BLINK’s isolated evidence.
- Ranking models: A Bradley–Terry model ranks the 51 models using 4,231 equal-weight binary benchmark votes from the four selected benchmarks.Each benchmark votes for the higher-scoring model whenever both models have been evaluated on it.
- Ranking models: Three of the top ten models are post-trained for embodied or spatial tasks, while scale improves physical ability within a training recipe but not between recipes.HY-EMBODIED-0.5 MOT-4B-A2B ranks sixth above substantially larger models.
5 LIMITATIONS
The audit is constrained by aggregate benchmark scores, possible score heterogeneity and noise, limited sample size, and design choices that are not guaranteed optimal or causally interpretable.
- Benchmark-level analysis: Item-level redundancy and measurement error cannot be localized directly because the analysis uses aggregate benchmark-level scores.The field does not provide item-level response data, although item-level audits could localize redundancy and estimate measurement error.
- Matrix score verification and possible heterogeneity and noise: The matrix combines published scores with own runs without a systematic reproduction study, and peer unpredictability may reflect distinct signal or noise.Separating these explanations requires repeated evaluations under resampled prompts, decoding seeds, and parsing rules, which current reporting does not provide.
- Sample size: Fifty-one models across 12 benchmarks may be thin for some statistical analyses.The authors therefore emphasize findings supported by statistical significance.
- Compact-suite selection: The four-benchmark suite is a defensible greedy choice, not an optimum, and its recovered core depends on the utility and discrimination measure.The core is robust to which benchmark starts selection but remains untested against a different definition of the discrimination measure.
- Observational evidence only: The evidence is observational, so the study cannot claim that the dominant axis causes benchmark performance or that residual physical signal transfers to manipulation or navigation.None of the 12 benchmarks measures downstream task success on a physical system.
6 CONCLUSION
The paper audits physical-AI benchmarks as a measurement instrument using a 51-model, 12-benchmark matrix. It finds substantial, partly general redundancy and compresses the suite to four benchmarks while preserving most discriminating power.
- Audit scope: The audit treats a 12-benchmark physical-AI suite as a measurement instrument rather than a scoreboard, using 51 models and published scores plus own evaluation runs.The procedure is based on a matrix assembled from model reports and new evaluations.
- Redundancy: Correlations are uniformly positive, averaging 0.487, and roughly half of shared information is general vision–language capability rather than physical-AI-specific signal.This quantifies both substantial overlap and the contribution of a broader capability shared with general benchmarks.
- Sufficiency: Four benchmarks selected for discrimination and marginal uniqueness retain 78.5% of the suite’s discriminating power.The compact subset demonstrates that much of the suite’s non-redundant signal can be preserved with fewer benchmarks.
- Generality: The audit requires only benchmark-level scores with sufficient model overlap and is not specific to physical AI.Its stated input structure supports application to other fields.
A COMPLETE MODEL LIST AND SCORE SOURCES
The complete model list records metadata for all 51 models and documents how their 12 benchmark scores were sourced, distinguishing published values from the authors’ own runs.
- Score sources: 405 of 564 filled cells are grouped as published, while the remaining 159 cells are the authors’ own runs.Published cells include direct transcriptions, medians of conflicting values, and one value borrowed from a twin model.
- Model metadata: Table 4 lists every model with its provider, parameter count, release date, base checkpoint, and how its twelve scores were obtained.The table covers the full matrix of 51 models and 12 benchmarks.
B STRUCTURE: HOW THE PHYSICAL SUITE RELATES TO GENERAL CAPABILITY
The physical suite’s dominant shared axis closely tracks general vision–language capability, but residual analysis reveals benchmark relationships that remain specific to physical abilities.
- Shared structure: 55.2% of variance is explained by the first principal component, with all 12 benchmarks loading positively from 0.48 to 0.88.The lowest loading is REALWORLDQA at 0.48 and the highest is CV-BENCH at 0.88.
- General capability alignment: Physical PC1 correlates with general-anchor PC1 at Spearman ρ = 0.952, including 0.950 with MMSTAR and 0.942 with VIDEO-MME.PC1 is estimated from the physical benchmarks alone before being compared with general anchors.
- Residual structure: Mean pairwise |ρ| falls from 0.487 to 0.250 after removing the external general axis, while correlations above 0.5 drop from 34 of 66 pairs to 6.This residualization shows that roughly half of benchmark-to-benchmark agreement reflects general capability.
- Specific relationships: Three relationships survive residualization at nearly full strength: the pointing pair, the 2D-perception pair, and ERQA ↔ MINDCUBE.The pointing and 2D-perception relationships are identified as substitute pairs measuring the same specific abilities.
- Specific relationships: VSI-BENCH and BLINK shift from raw-score ρ = 0.002 to residual ρ = −0.363, indicating a trade-off after conditioning on general capability.Among models of equal general capability, video-spatial reasoning strength trades off against multi-image perception strength.
C FORWARD SELECTION DETAILS
The paper evaluates a greedy benchmark-selection path by tracking marginal utility across all 12 benchmarks and testing whether the compact suite depends on its initial seed.
- Selection path: The last four benchmarks contribute 4.4% of the total utility between them under U(b | S) = g(b) (1 − R2(b ∼ S)).Here, g measures raw score spread and R2 measures how well the candidate is fit by the selected set.
- Seed robustness: Forcing a core benchmark into the opening slot yields the same four-benchmark set in a different order.The four-benchmark utility remains between 77.8% and 78.3%, versus 78.5% for the unconstrained path.
- Seed robustness: Seeding with a substitute lowers four-benchmark utility to 74.5%, while rejected-benchmark seeds yield 69.2–71.6%.The substitute case retains both pointing benchmarks, displacing BLINK to fifth place.
- Seed robustness: MINDCUBE and VSI-BENCH enter the first four in every forced-start run, and REFSPATIAL-BENCH is never excluded.Even forcing WHERE2PLACE, its substitute, does not keep REFSPATIAL-BENCH out of the first four.
- Interpretation: Because forward selection is greedy and has no optimality guarantee, the core is a compact suite under this utility rather than a unique optimum.The robustness experiment tests only the opening-pick degree of freedom.