Source-linked AI summary

Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit

Wassim Tenachi, Yashar Hezaveh, Laurence Perreault Levasseur, Pierre-Luc Bacon

arXiv:2609.02766v1cs.LGastro-ph.IM

TL;DR

The paper asks whether tabular foundation models learn physics in their priors, evaluates four TFMs against six baselines on regenerated data from 316 physical equations, and finds a split result. TFMs dominate interpolation even after tuning, but their priors cannot represent noiseless mechanisms or exploit physical dimensional structure, limiting their status as physical models.

  • Problem

    The paper asks whether TFMs’ Bayesian priors contain physics, given limited evidence beyond interpolation and no current pretraining on physical laws with units or continuous targets.

  • Method

    The authors evaluate four TFMs against six baselines on regenerated datasets sampled from 316 physical equations across in-domain, extrapolation, noise, and sampling regimes.

  • Results

    TFMs lead across strata under default, tuned, and ensembled evaluation, while failing to represent noiseless mechanisms and exploit dimensional structure.

  • Takeaways & Limitations

    TFMs are effective amortised interpolators of physical data, but acting as physical models would require priors that contain physics.

Abstract

from arXiv · show

Tabular foundation models (TFMs) learn to fill in tables the way language models fill in text, and tables are arguably the format in which most physical measurement arrives. Did they learn any physics in the process? They are Bayesian by construction, so the question is what their prior contains. We probe it directly, evaluating four of them (TabPFN-3, TabICLv2, TabDPT and Real-TabPFN-2.5) against six baselines on datasets sampled from 316 physical equations, in and out of domain. TFMs dominate, out of the box and after tuning. But we show that their prior can represent neither a noiseless mechanism nor physical units, which is why they interpolate physics without yet being able to act as physical models.

1 Introduction

The paper asks whether tabular foundation models learn physical structure from their table-based pretraining, treating their Bayesian prior as the object of study. Existing evidence emphasizes interpolation, while physical-law data and extrapolation remain insufficiently tested.

  • Motivation: Tables capture much physical measurement, yet tabular foundation models have received less attention than language and vision models.The paper motivates asking whether skimming large collections of tables can produce representations of the physical world.
  • What TFMs are: TFMs learn masked table entries from context and solve new tasks through one-pass in-context learning without gradient updates.Most models are pretrained on synthetic structural causal models, while others use real-table corpora.
  • Open question: TFMs remain weak at extrapolation, especially on physical signals, despite broad applications that have focused mainly on interpolation.This raises whether extrapolation of arbitrary tables requires learning something about the physics generating them.
  • Open question: The physics regime lacks prior coverage because benchmark suites omit deterministic-function data and current TFMs lack pretraining on physical laws with units or continuous targets.The proximity of equation-generated data to model-training ecosystems also motivates a contamination audit.
  • Research question: Because TFMs are Bayesian by construction, the paper tests what prior their posterior predictions encode using data sampled from known physical laws.Such data provides a clean setting for probing whether the learned prior contains physics.

2 Protocol

The protocol regenerates tables from physical equations under controlled sampling, noise, and domain regimes, then compares four TFMs with six tabular baselines using common rank-based evaluation. It reports both default and budget-matched tuned settings, including optional ensembling.

  • Data: The evaluation uses Feynman, LSR-Transform, and LSR-Synth data comprising 316 physical equations, with every table regenerated from its source equation.Regeneration leaves sample count, noise, and sampling domain under experimental control and avoids reusing published tables.
  • Regimes: Training points come from each equation’s sampling box, while test points come from shells at scale k=1 or k=2 outside it.The design sweeps 50, 200, or 2000 samples; noise levels 0, 0.01, or 0.1; and three seeds, with 2000 fixed test points.
  • Baselines and evaluation: Four open-weight TFMs are compared with six state-of-the-art tabular methods using NMSE and mean rank within each task-and-regime cell.Mean rank handles raw-error differences spanning orders of magnitude across equations.
  • Fairness: Default evaluation uses library settings, while tuned evaluation gives trained baselines 25 random-search configurations and optionally applies greedy post-hoc ensembling.TFM ensembling averages forward passes over input permutations and serves as the comparable computational choice.

3 Results & Analysis

The experiments find no detectable real-data contamination, while TFMs retain a performance lead under tuning and ensembling. However, their priors fail to represent noiseless mechanisms, exploit dimensional structure, and extrapolate physical signals reliably.

  • Contamination: Both real-data candidates fall inside the unexposed TFM band at every context size, producing no detectable contamination effect.The comparison measures seen-versus-unseen performance against models with no pretraining; effects below roughly one rank position would be invisible at this pool size.
  • Dimensional structure: TabPFN-3 degrades as dimensional reducibility increases, whereas trained baselines do not.Across models, error tracks raw column count rather than the number of dimensionless groups, despite physical laws often having lower effective dimensionality.
  • Noiseless mechanisms: At σ = 0, every TFM’s predictive width plateaus while RMSE continues falling, although at σ = 0.1 calibration width is as expected.The result indicates a failure specific to the noiseless limit rather than a general estimator deficiency.
  • Performance: TabPFN-3 leads the trained baselines across every stratum by mean rank, including tuned and ensembled settings except one largest-context comparison against its cheapest setting.The margin narrows out of domain at the largest context size, where both conditions favoring trained models apply.
  • Extrapolation: Out of domain, all four TFMs lose the damped oscillator’s oscillation within roughly half a period, though TabPFN-3 and TabICLv2 recover its decaying envelope.Their predictions relax toward a constant rather than continuing the phase, while CatBoost holds the last observed value.

4 Conclusion

TFMs interpolate physical data well without pretraining on physical laws, but their priors cannot represent noiseless mechanisms or exploit dimensional structure. They therefore remain amortised interpolators rather than physical models.

  • TFMs interpolate physical data well without having been pretrained on a physical law.
  • Their priors cannot represent a noiseless mechanism or exploit dimensional structure, properties present in every physical law.
  • TFMs are excellent amortised interpolators, but a physical model would require a prior containing physics.
  • Pretraining on continuous physical targets is proposed as a natural way to test whether TFMs can acquire such a prior.

A Extrapolating a damped harmonic oscillator

On a noiseless damped harmonic oscillator, all four TFMs fail to continue the oscillation reliably beyond the sampled range. Some preserve envelope or wave-like uncertainty traces, but these signals are not phase-locked.

  • Extrapolation: All four TFMs relax to a constant beyond the sampled range, roughly within half a period, rather than continuing the oscillation.TabPFN-3 and TabICLv2 decay toward zero, while CatBoost holds the last observed value.
  • Extrapolation: TabICLv2 shows wave-like predictive-density banding and TabDPT continues the signal for roughly a quarter period before flattening.Both effects are described as suggestive rather than periodic extrapolation.
  • Extrapolation: The apparent oscillatory banding is not phase-locked: TabICLv2 ranks 98th and TabPFN-3 157th among 201 shifted continuations.The unshifted true continuation would score best if the model had learned the oscillation’s phase.
  • Predictive uncertainty: Inside the sampled range, distributional TFMs report nonzero uncertainty despite noiseless data, whereas the GP reports exactly zero width.Reported widths are 0.010 for TabICLv2 and 0.019–0.020 for the two TabPFN variants, normalized.
  • Predictive uncertainty: TabDPT does not report zero uncertainty because it emits no predictive distribution.

B Trained baselines

The trained-baseline comparison covers six methods fitted separately for each task. Only the Gaussian process provides a closed-form predictive density; the other baselines return point estimates.

  • Trained baselines: Six trained baseline methods are fitted per task across deep and classical model families.
  • Predictive output: Only the Gaussian process emits a predictive distribution among the trained baselines.
  • Predictive output: The remaining baselines return point estimates and therefore cannot appear in uncertainty comparisons.

C Protocol details

The protocol fixes the extrapolation split, seed-based resampling, validation allocation, preprocessing disclosures, cost accounting, and reproducibility procedures. These choices distinguish interpolation from extrapolation and make failures and computational costs explicit.

  • The extrapolation shell: OOD test points are drawn from an outer shell around the training box, with interior points rejected to prevent interpolation leakage.All axes are extended simultaneously using per-axis box extension.
  • Seeds in place of cross-validation: Three independent seeds resample both training and test points, and the median is reported instead of cross-validation folds.This is intended to remain robust to a single catastrophic draw.
  • Validation carve-out: Tuned trained baselines fit on 0.8 nsamples because 20% is reserved for model selection, while TFMs condition on all samples.At nsamples=50, trained models see 40 rows versus 50 for TFMs.
  • Preprocessing: Column order is randomized per seed, model-specific preprocessing is retained and recorded, and metrics are computed in the original target space.
  • Cost accounting: Runtime is split into fit and predict phases, while pretraining cost is excluded and hardware is recorded for later filtering.
  • Reproducibility: Runs are keyed by resolved-configuration hashes and store configuration, software, hardware, and timing metadata for resumable reproduction.The main grid spans 316 tasks, three sample sizes, three noise levels, two split scales, three seeds, and ten models per protocol.

D Per-model benchmark results

Figure 6 unpacks benchmark ranks across all model–pipeline entities, showing stable family winners, limited ensemble gains, context-dependent tuning benefits, and a distinct TabDPT pattern.

  • Which model wins each family: TabPFN-3 is the best TFM in every stratum under both library-default and tuned settings.
  • Which model wins each family: GP beats Ridge everywhere, while CatBoost is the only gradient-boosted entrant.
  • Which model wins each family: TabM is the strongest tuned MLP, but RealMLP wins in several default, out-of-domain, and ensembled settings.
  • Ensembling contributes almost nothing: Post-hoc greedy selection measurably improves only RealMLP; median changes for TabM, CatBoost, GP, and Ridge are indistinguishable from zero.
  • Ensembling contributes almost nothing: Increasing ensemble-selection rounds from 25 to 200 produces results identical to four decimal places, with selection saturating at four to ten configurations.
  • Where tuning does help: Tuning improves every trained baseline more at the largest context size than at the smallest, except Ridge at σ = 0.
  • The validation carve-out: Tuned baselines fit on 0.8 nsamples because they hold out 20% for configuration selection, whereas TFMs condition on all samples.
  • A note on TabDPT: TabDPT performs worse than synthetic-prior counterparts on LSR-Synth and low-dimensional tasks, without an explanation in the paper.

E Pretraining corpora of the real-data TFMs

The audited real-data TFM corpora contain physics-related classification data but no analytic-function regression targets, limiting their direct exposure to physical laws.

  • TabDPT’s corpus contains 122 OpenML datasets: 93 classification and 28 regression tasks.
  • All six TabDPT datasets labelled Physics/astronomy are classification tasks, and none of nine Deterministic and simulated datasets has a continuous target.
  • No TabDPT regression target is generated by an analytic function.
  • Real-TabPFN-2.5’s 43-dataset corpus contains no real-world regression table; every entry is a classification task.
  • Its sole function-generated entry, fried, is a binarized Friedman #1 target thresholded at its mean.
  • Both corpora include Sloan Digital Sky Survey catalogues only as star/galaxy/quasar classification.
Loading 2609.02766v1…