Source-linked AI summary

When Does AI for PDEs Yield Scientific Evidence?

Wenshuo Wang

arXiv:2608.22504v1cs.AI

TL;DR

AI-for-PDE benchmarks mainly measure predictive or approximation accuracy, although scientific applications use outputs as evidence for claims. The paper formalizes claim-conditioned support and extends forward-simulation and inverse-problem benchmarks with paired accuracy- and evidence-based evaluations. It finds that the two objectives can select different methods and explains when and why those selections align or diverge.

  • Problem

    Existing AI-for-PDE benchmarks primarily evaluate agreement with reference targets or governing constraints, while scientific applications use outputs to support specified claims under assumptions and evidence standards.

  • Method

    The paper formalizes numerical accuracy and claim-conditioned evidential support, then adds prespecified claim contracts and evidence evaluations to PDEBench and PDEInvBench tasks.

  • Results

    Numerical accuracy and claim-conditioned support select different methods in all three headline comparisons, with differences persisting across most tested variants.

  • Takeaways & Limitations

    When AI outputs serve as scientific evidence, evidential support should be evaluated as a distinct endpoint rather than inferred from numerical accuracy alone.

  • Takeaways & Limitations

    The conclusions are bounded by operational, noncanonical evaluation targets, two PDE benchmarks with finitely many claims and methods, and marginal conformal guarantees under exchangeability.

Abstract

from arXiv · show

Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy. In physics research, however, AI outputs often serve as evidence for scientific claims. These two objectives are not equivalent: the former measures an output's agreement with a reference target or satisfaction of governing constraints; the latter asks whether, given a specified object of study, scientific claim, assumptions, and evidence standard, the output provides sufficient evidence for that claim. To bridge this gap, we extend a widely used PDE-simulation benchmark and a comprehensive benchmark for PDE inverse problems to enable, for the first time in AI for PDEs, evaluation of whether and to what extent model outputs support specified scientific claims. Our results show that numerical accuracy and evidential support can rank models differently, explain when and why they do so, and reveal that existing benchmarks can favor methods whose outputs provide weaker support for the scientific claims of interest. Together, we formalize, empirically demonstrate, and explain this evaluation--use mismatch in AI for PDEs.

1 Introduction

AI-for-PDE benchmarks mainly measure predictive or approximation accuracy, while scientific applications use model outputs as evidence for specified claims. This paper formalizes the distinction, extends forward and inverse benchmarks with claim-specific evaluations, and finds that accuracy and evidential support can select different methods.

  • AI-for-PDE benchmarks predominantly evaluate agreement with reference objects or governing constraints, whereas scientific use asks whether outputs support scientific claims.The paper illustrates this distinction with forward singularity scenarios and inverse rheological inference.
  • The paper formalizes numerical accuracy q(A; R, K) separately from claim-conditioned evidential support W(A; O, C, H, E).q concerns reference agreement or constraint satisfaction; W concerns sufficient evidence for a claim about an object under assumptions and an evidence standard.
  • Optimizing numerical accuracy can favor methods whose outputs provide weaker support for the scientific claims they are ultimately used to address.
  • The benchmark extensions retain original equations, datasets, task definitions, and accuracy metrics while adding prespecified scientific claims and claim-specific evidence evaluations.Evaluations cover quantity-of-interest and event claims for solution fields, plus range, regime, and regional-structure claims for inferred parameters or fields.
  • The experiments find that evaluation coverage is near nominal and informative, while numerical accuracy and claim-conditioned support select different methods across both benchmark extensions.The selection difference persists across most tested variants and is traced to claim-boundary proximity, error behavior, and evidence constraints.
  • The conclusions remain conditional on the studied benchmarks and prespecified claims, assumptions, and evidence standards.

2 Evaluation in AI-for-PDE Benchmarks and Use of Model Outputs as Scientific Evidence

AI-for-PDE evaluation research and AI-enabled scientific applications center on different outputs and goals. Existing forward and inverse benchmarks primarily assess predictive or approximation quality, whereas scientific applications interpret AI-generated PDE outputs as evidence for scoped claims under explicit assumptions and checks.

  • AI-for-PDE evaluation research centers on reusable benchmarks, protocols, and empirical evaluations, while scientific applications center on conclusions about specified physical systems or mathematical objects.In applications, AI-generated PDE-related outputs are integral to the supporting analysis.
  • Table 1 organizes forward and inverse tasks by scored output, reference target, tested constraint, and reported criterion.The synthesis distinguishes forward modes F1–F4 and inverse tasks I1–I6 according to evaluated outputs and input–output relations.
  • Forward benchmarks favor methods reproducing reference fields or trajectories, while inverse benchmarks favor agreement with withheld references, constraints, uncertainty criteria, or reference structure.Both families therefore predominantly assess predictive or approximation quality through reference agreement or constraint satisfaction.
  • Use of Model Outputs as Scientific Evidence: A forward application uses PINN-computed self-similar fields and a scaling exponent as evidence concerning consistency with an Euler singularity scenario.Residual, robustness, exponent-agreement, and known-solution checks support consistency under stated assumptions, not proof of finite-time blow-up.
  • Use of Model Outputs as Scientific Evidence: An inverse application uses inferred viscosity fields and derived strain-rate, stress, and stress-exponent estimates as evidence for rheological claims about Antarctic ice shelves.
  • Use of Model Outputs as Scientific Evidence: Across both applications, evidence evaluation interprets AI-generated PDE outputs under assumptions linking them to a specified object and applies claim-relevant checks.

3 The Evaluation–Use Mismatch

The paper distinguishes accuracy from evidential support by giving them different inputs, evaluated objects, and semantics. It then shows that aggregating these endpoints over the same benchmark distribution can induce different method selections and influence methodological development.

  • q measures agreement with references or tested constraints, whereas W measures whether an output and its evaluation record sufficiently support a scoped claim under assumptions and an evidence standard.W equals 1 only when prespecified acceptance criteria are met; it is not a claim-independent evidence score.
  • The two targets are not generally interchangeable: numerical accuracy alone does not determine claim-conditioned evidential support without a claim-specific bridge.A certified threshold test is presented as a restricted case where support reduces to thresholded certified accuracy.
  • The method-selection analysis fixes a benchmark distribution over task–claim pairs and contexts, then compares accuracy and support estimands for each method.
  • Accuracy aggregation Q_m evaluates artifact quality, whereas support aggregation S_m evaluates claim support under the same frozen distribution and claim contexts.Because the endpoints aggregate different properties, maximizing Q_m can induce a different selection than maximizing S_m.
  • The mismatch can propagate from benchmark evaluation into which methods the community selects, improves, and treats as progress.

4 Extending Evaluation from Accuracy to Evidential Support

The benchmark extensions preserve the original PDEBench and PDEInvBench tasks and accuracy evaluations while adding claim-conditioned evidence evaluations. This enables direct comparison of numerical accuracy with support for scientific claims under prespecified objects, assumptions, and evidence standards.

  • Benchmark extensions: The extensions add claim-conditioned evidence evaluations to PDEBench and PDEInvBench without changing their underlying tasks.The original equations, datasets, task definitions, eligible support, protocols, artifacts, and computational budgets are retained; only the evaluation target changes.
  • Evidence operationalization: A claim is supported when an independently calibrated evidence region satisfies the prespecified containment relation for its acceptance set.Under exchangeability and finite scores, split-conformal calibration supplies finite-sample marginal coverage; unsupported intermediate cases remain UNKNOWN.
  • Accuracy versus evidence: Numerical accuracy determines evidential support only in the restricted case where a sound claim-specific error bound certifies the claim threshold.Without that bridge, better numerical accuracy does not generally determine whether the scientific claim is supported.
  • Forward simulation: Forward-simulation contracts evaluate regional quantities, threshold exceedance, transport or flux, and event claims from predicted fields or trajectories.PDEBench retains reference-based diagnostics while adding two PDE-appropriate claim contracts for each audited configuration lineage.
  • Inverse inference: Inverse-inference contracts evaluate physically meaningful ranges, directional thresholds or regimes, and regional structure in recovered parameters or coefficient fields.Scalar targets use range and regime claims, while Darcy contracts use regional phase fractions and contrasts.

5 When and Why Accuracy and Evidential Support Rank Methods Differently

The section tests whether accuracy and evidential support select the same methods, then explains divergences through claim-boundary proximity, error direction, and evidence-region width. Across benchmark extensions, the two objectives often select different methods, with controlled analyses supporting the proposed geometric mechanism.

  • The evaluation asks whether evidence behavior is valid, whether Q and S select the same methods, and whether measurable claim, artifact, and region properties explain differences.
  • 5.1 Coverage, Informative Verdicts, and Artifact Sensitivity: 94.6–96.8% coverage on PDEBench and 93.8–97.1% on PDEInvBench indicate near-nominal empirical coverage with low directional error and invalidity.All-assigned wrong-direction rates are 0.26% and 0.54%, while invalid rates are 0.18% and 0.49%, respectively.
  • 5.1 Coverage, Informative Verdicts, and Artifact Sensitivity: All 70 PDEBench and 34 PDEInvBench contracts pass: clear oracle cases are supported or refuted, while boundary-near cases remain UNKNOWN at rates of 85.4–100% and 83.3–100%.This pattern shows that the endpoint resolves claims when evidence is separated from the boundary while retaining uncertainty near it.
  • 5.2 Changing the Evaluation Target Changes Method Selection: On identical supports, Q and S select different methods in all three headline comparisons: OmniArch-B versus VCNeF, SC-FNO versus FNOPE, and ResNet/FNOPE versus DeepONet-inclusive Q winners.Across 16 primary records, winner sets are disjoint in 9, match in 4, and remain unresolved or overlapping in 3.
  • 5.2 Changing the Evaluation Target Changes Method Selection: 80.6% of frozen variants retain disjoint Q- and S-winner sets, compared with 6.9% agreeing and 12.5% unresolved.The result comes from 58, 5, and 9 of 72 variants, respectively.
  • 5.3 Boundary Proximity, Error Direction, and Evidence Width Predict Divergence: Adverse error localization, direction, inverse signed bias, and wider evidence reduce focal support by 0.181, 0.147, 0.132, and 0.209, respectively, while ∆Q remains within frozen equivalence bands.Panel C's mechanism model reaches log loss 0.681 and balanced accuracy 0.734 across 52 holdouts, outperforming Q-based baselines.

6 Discussion

The discussion positions the benchmark extensions as operational tests of claim-conditioned evidence, while bounding their interpretation by prespecified contracts, studied benchmarks, and finite-sample assumptions.

  • Evaluation scope: AI-for-PDE evaluation benchmarks and AI-enabled scientific applications address different endpoints: output accuracy versus support for claims about specified physical systems or mathematical objects.The paper distinguishes benchmark-centered evaluation research from applications where AI-generated PDE outputs support scientific conclusions.
  • Interpretation: Accuracy-only evaluation can select methods whose outputs provide weaker support for the scientific claims those outputs are intended to support.The paper treats evidential support as a distinct evaluation endpoint rather than inferring it from numerical accuracy.
  • Operationalization: The extensions preserve original benchmark tasks and add auditable claim contracts, evidence evaluations, and frozen configurations for forward and inverse settings.The design records benchmark objects, variables, units, claims, assumptions, applicability, calibration, and checking procedures.
  • Interpretation: Claim provenance combines published scientific uses with benchmark-specific operational bindings, so registered boundaries and masks are not presented as universal physical thresholds.The paper separates literature provenance, benchmark definitions, fixed rules, and target-specific operational choices.
  • Prespecification: Claims were frozen before evaluated accuracy, support, and rankings, although fit and validation material informed a limited set of declared design-pilot choices.The stated protection is the absence of feedback from final-calibration or paired-evaluation outcomes to claim bindings.
  • Limitations: The conclusions are bounded by the studied registry and assignment rule, finitely many claims and methods, operational definitions, and marginal exchangeability-based guarantees.The evaluation does not establish exhaustive coverage, universal truth, broad generalization, or individual-case and distribution-shift guarantees.

A.5 Controls, Endpoints, Statistical Analysis, and Release

The evaluation uses frozen controls, endpoints, aggregation, and uncertainty procedures to test evidential support without changing benchmark-native accuracy measures. It also documents implementation and release boundaries while preserving invalid, unavailable, and diagnostic outcomes.

  • Controls: Three oracle regimes test separated focal-true, separated complement-true, and boundary-near cases across all claim contracts.Controls assess resolution of both claim directions while preserving uncertainty near claim boundaries.
  • Controls: Artifact corruptions and provenance audits test sensitivity to submitted content, leakage, reference exposure, evaluator recentering, reconstruction, and solver calls.Native-region dispersion perturbations remain confined to a secondary track.
  • Endpoints: The evidential endpoint Sm is the frozen macro-average of W on common support, with all assigned units retained and abstention unable to improve support.Unavailable finite q values instead make comparison records unavailable for Qm.
  • Statistical Analysis: Atomic observations are reduced within independent groups, while PDEBench averages applicable claims within lineages and macro-averages PDE families equally.This prevents reused trajectories, windows, claims, or overlapping regions from inflating effective sample size.
  • Statistical Analysis: Endpoint comparisons use paired hierarchical bootstraps with 10,000 draws, resampling benchmark-specific families, lineages or targets, regimes, groups, and reporting seeds.Deterministic procedures remain fixed and reporting seeds are not treated as independent scientific cases.
  • Release: The appendix supplies implementation specifications and frozen configurations, but benchmark data, runs, checkpoints, and aggregate paper outcomes are not distributed.The repository generates machine-readable tables and figures from completed production records.
  • Release: The appendix supports three conclusions: endpoint validity, changed method selection, and predictable boundary–direction– evidence-width divergence.These analyses provide the additional results needed to support the paper’s conclusions.

B.1 Endpoint Validity

Validity audits show that the claim-conditioned endpoint is calibrated, discriminates oracle artifacts from controls, and preserves provenance safeguards. It resolves separated claims while remaining uncertain near claim boundaries.

  • Calibration and direction: Every contract has 96/96 supported focal-true and 96/96 refuted complement-true separated cases; simultaneous worst-case lower bound 92.1%; boundary-near unknown range 41/48–48/48.The separated cases resolve both directions, whereas boundary-near cases remain frequently UNKNOWN.
  • Specificity: On 13,440 records per control, correct-resolution counts are 958, 722, 489, and 267; oracle–control separations are 0.929 [0.902, 0.951], 0.946 [0.919, 0.967], 0.964 [0.938, 0.979], and 0.980 [0.955, 0.990].The specificity audit compares oracle artifacts with null, random, shuffled, and task-swapped controls.
  • Calibration and direction: Coverage 94.6–96.8%; all-assigned wrong-direction 63/24,360 (0.259%); invalid 44/24,360 (0.181%).These are PDEBench validity-audit rates.
  • Provenance: Zero prohibited exposures or evaluator-side solver calls; 11/11 nominal tracks pass.The provenance audit covers the PDEInvBench endpoint.
  • Specificity: On 6,528 records per control, counts are 558, 397, 291, and 169; separations are 0.915 [0.879, 0.941], 0.939 [0.908, 0.961], 0.955 [0.927, 0.975], and 0.974 [0.949, 0.987].The inverse-benchmark specificity results compare oracle and control artifacts.
  • Overview: Across 104 contracts and controls, the completed validity audits report near-nominal coverage and low aggregate error rates.The audit includes 31,336 submitted W rows across PDEBench and PDEInvBench, with wrong-direction and invalid rows retained in denominators.

B.2 Paired Leaderboards and Primary Selections

Paired accuracy and evidential-support leaderboards select different methods in many benchmark records, including cases where global accuracy is nearly equivalent. The divergence reflects claim-relevant coordinates and evidential uncertainty rather than global error alone.

  • Leaderboard construction: The leaderboard tables report simultaneous 95% intervals for benchmark-native Q and evidential-support S, with rank columns separated from confidence-set selections.The paired endpoint comparisons use the frozen common support.
  • Primary selections: The three headline selections differ: PDEBench dynamics selects {OmniArch-B} under Q and {VCNeF} under S; inverse scalar selects {SC-FNO} and {FNOPE}; inverse Darcy selects {ResNet, DeepONet} and {FNOPE}.The 16 claim- and regime-specific records show where aggregate selections agree, differ, or remain unresolved.
  • Primary selections: The winner sets are disjoint in 9/16 records, match in 4/16, and overlap or remain unresolved in 3/16.Selections use simultaneous winner confidence sets rather than point estimates alone.
  • Primary selections: ∆Q = −0.00043 [−0.00137, 0.00054] while VCNeF has ∆S = 0.176 [0.121, 0.231] higher focal support in the shallow-water comparison.The accuracy difference satisfies the frozen equivalence interval, while the support contrast exceeds the registered minimum meaningful contrast of 0.05.
  • Mechanism: ResNet has global field error 0.0006 but regional errors 0.041/0.037 and half-width 0.052, whereas FNOPE has global error 0.0186, regional errors 0.018/0.021, and half-width 0.031.A small phase-boundary displacement can affect regional claim coordinates and uncertainty more than global relative error.

B.3 Robustness

Robustness analyses show that differing accuracy and support selections persist across most eligible one-factor variants. The effect remains substantial across the primary records with disjoint winner sets.

  • Robustness: Winner sets remain disjoint in 58/72 eligible variants (80.6%; simultaneous 95% interval 70.4–88.6%) across nine primary records with disjoint winner sets.The remaining variants match in 5 cases and remain unresolved in 9.

B.4 Natural Association and Interventions

Matched-accuracy interventions support a mechanism in which claim-directed or localized artifact error and wider evidence regions can increase divergence between numerical and evidential outcomes. Across 384 full-method-panel recomputations, the registered comparisons remained within frozen equivalence bands, while descriptive frequencies require dependence-aware interpretation.

  • Study design: 384 unmodified full-method-panel recomputations tested 16 primary-record specifications across componentwise budgets, evidence levels, and reporting views.Joint-stratum cutoffs were frozen using independent design-stage validation before paired evaluation.
  • Interpretation: The reported frequencies are descriptive because recomputations within a family or target are dependent, and no single adverse factor is sufficient across every benchmark component.This limits how the intervention frequencies should be generalized or interpreted as universal effects.
  • Intervention design: Matched-accuracy diagnostics changed one registered factor while preserving artifact schema, common support, non-target summaries, and the validation-frozen Q-equivalence band.The prespecified design assigned 120 localization, 120 forward-direction, 108 inverse-bias, and 132 evidence-width diagnostics.
  • Artifact-side mechanisms: Moving equal-energy error into the claim region or reversing its direction reduced support despite equivalent global accuracy, with reversed inverse signed bias producing the analogous result.These interventions isolate localization and direction as artifact-side mechanisms that can separate endpoints.
  • Equivalence control: Every ∆Q interval lay wholly inside its frozen equivalence band across the registered synthetic manipulation families.The result supports the intended matched-accuracy comparisons rather than differences caused by departures from the equivalence criterion.
  • Evidence-side mechanism: Expanding evidence width was tested as a distinct intervention, linking endpoint divergence to sensitivity to evidence-region calibration.The supplied result passage introduces this sensitivity but does not report its completed quantitative outcome.

B.5 Held-Out Prediction and Secondary Scope

Held-out evaluation tests whether pre-verdict mechanism summaries predict endpoint agreement beyond validation accuracy and task difficulty. The mechanism model performs better on unseen families and targets, while the conclusions remain probabilistic and limited to the studied benchmark components, claim contracts, evidence standards, and method panels.

  • Held-out evaluation: 52 whole-family or inverse-target holdouts tested generalization across dynamic, PINN, PDEBench-Darcy, scalar-inverse, and inverse-Darcy components.All methods, claims, seeds, and artifacts from each held-out family or target remained in the same outer fold.
  • Predictor: The mechanism predictor used cluster-equal-weighted ridge multinomial logistic regression with pre-verdict proxies for boundary clearance, directed error, localization, and calibrated width.Inputs excluded paired-evaluation margins, interval endpoints, W, and features that reconstruct winner sets.
  • Predictive results: 0.253 [0.143, 0.367] lower multiclass log loss and 0.176 [0.071, 0.277] higher balanced accuracy were achieved versus validation Q plus task difficulty.Thus the mechanism summaries add predictive information beyond validation accuracy and task difficulty.
  • Scope of prediction: The mechanism features predicted whether endpoint selections agree, differ, or remain unresolved on unseen families and targets beyond validation accuracy and task difficulty.This claim is probabilistic and restricted to the benchmark components, claim contracts, evidence standards, and method panels studied here.
  • Secondary scope: Native posterior or evidence outputs remained separate from common-wrapper leaderboards, with correct-resolution gains ranging from 0.044 [0.009, 0.080] to 0.093 [0.049, 0.137].The reported gains cover scalar FNOPE, Darcy FNOPE, Darcy iFNO, and Darcy DGenNO tracks.
Loading 2608.22504v1…