Source-linked AI summary
Shared Physics Responses Recover Hidden Rankings in Neural Operator Libraries
Hanbing Liang, Fujun Liu
TL;DR
Deployment must rank competing neural-operator predictions without computing an unavailable high-fidelity reference. The paper uses a finite-library decision quotient and one anchor-based linearized physics response to construct a shared task proxy, recovering over 99.6% of pairwise preferences and frequently improving direct predictions. It also provides sufficient certification conditions for strongly monotone discretizations, while identifying margins, anchor composition, and operator-data discrepancy as boundaries on reliability.
Problem
Deployment provides candidate fields and governing equations but not the high-fidelity solution needed to rank predictions for the current input and task.
Method
The method constructs a shared library anchor, propagates its complete physical defect once through a linearized solve, and compares the resulting task proxy with every candidate.
Results
The shared task proxy achieved 100% non-tie pair recovery in the reaction–diffusion census and recovered 99.59% of non-tied preferences across eight tested libraries.
Takeaways & Limitations
Accurate candidate ranking can persist without accurate full-task reconstruction, supporting both input-specific routing and panel-level reusable model selection.
Takeaways & Limitations
Certification for dataset or continuum targets requires an additional operator-data discrepancy bound, and empirical reliability remains dependent on margins and library composition.
Abstract
from arXiv · showhide
Selecting the optimal neural-operator prediction during deployment is challenging when high-fidelity reference solutions are unavailable. We demonstrate that under a squared Hilbert-space loss, ranking a finite model library depends strictly on the low-dimensional span of candidate differences, allowing us to score all models simultaneously using a single anchor-based linearized response of the governing equation. This shared physical diagnostic accurately recovered over 99.6\% of pairwise preferences and 99.0\% of optimal checkpoints across diverse Fourier and convolutional operator libraries for fluid, reaction-diffusion, and wave dynamics. Furthermore, the corrected physical proxy frequently outperformed the best individual candidates, and we establish computable sufficient conditions that rigorously certify exact decisions for strongly monotone discretizations. By exploiting the local dynamical response rather than raw defect magnitude, this framework enables the reliable and highly efficient deployment of scientific surrogates without requiring ground-truth data.
1 Introduction
The paper addresses reference-free ranking of surrogate predictions when deployment lacks a high-fidelity solution. It shows that candidate comparisons use a low-dimensional difference span and can be estimated with one shared physics response, achieving high recovery across diverse libraries while remaining sensitive to margins, composition, and admissibility.
- Motivation: Deployment lacks the high-fidelity solution needed to rank multiple surrogate predictions for the current input and scientific task.The paper frames this as reference-free finite-library ranking using only inputs, candidate predictions, and the governing PDE system.
- Finite-library decision quotient: Candidate preferences under squared Hilbert-space loss depend only on the span of candidate differences, whose dimension is at most the library size minus one.This decision quotient requires less information than reconstructing the hidden physical solution.
- Empirical recovery: 99.59% of non-tied pairwise preferences and 99.01% of tolerant top-1 choices were recovered across eight Fourier and convolutional operator libraries.The libraries crossed two equations, two architectures, and two seeded ensembles under controlled input shifts; the raw residual recovered 67.83% of pairs and 42.66% of tolerant top-1 choices.
- Cross-system validation: The proxy recovered 95.57% of exact top-1 decisions across eight Sine–Gordon libraries and preserved accurate comparisons under operator–data mismatch.In the compressible-flow probe, it recovered 10/10 comparisons despite improving direct proxy use over the best neural-operator prediction in only 4/10 cases.
- Reliability boundaries: Reliable decisions depend on candidate margins, library composition, and physical admissibility rather than raw defect magnitude alone.Extreme candidates reduced exact top-1 recovery with an arithmetic anchor, while backtracking increased feasible shock cases from 22 of 35 to 35 of 35.
3 Discussion
The discussion frames finite-library ranking as a lower-dimensional alternative to reconstructing unavailable solutions, implemented through one shared physics response. It reports strong empirical utility while identifying margins, anchor composition, admissibility, and evaluation-panel representativeness as operational boundaries.
- Discussion: Finite-library ranking requires only the candidate-difference comparison subspace, rather than full reconstruction of the hidden physical solution.The comparison subspace has dimension at most the library size minus one.
- Discussion: 100% non-tie recovery in the reaction-diffusion census exceeded candidate-wise variational estimators’ 82.7% recovery and 59.8% baseline.Candidate-specific correction-magnitude proxies produced almost identical choices to the shared calculation.
- Discussion: One shared physics response supports both query-time routing and panel-level checkpoint selection, changing post-inference physics cost from one solve per candidate to one library-wide solve.Overall cost still depends on the governing operator, discretization, and serving workflow.
- Discussion: Reliable deployment is constrained by small candidate margins, library-dependent anchor construction, shock-related admissibility, and whether an evaluation panel represents the deployment population.Neither arithmetic-mean nor medoid anchors was uniformly superior across tested datasets.
- Discussion: Strongly monotone reaction-diffusion discretizations admit rigorous residual-based certification, while future work targets mixed libraries, non-Hilbert objectives, multifidelity data, and severe stress tests.The paper’s broader aim is to identify the hidden-solution features necessary for distinguishing available predictions.
4 Methods
The method constructs one anchor-based, linearized physical response and reuses its task-space proxy to rank every candidate. Experiments compare shared propagation with residual, candidate-specific, validation, and centrality baselines across controlled PDE libraries and deployment settings.
- Experimental design: The experiments form eight libraries from two PDEs, two architectures, and two seeded ensembles, using 480 physical inputs and 1,920 input–library cases.Libraries contain epoch-20 and epoch-40 checkpoints from four training runs, with eligibility and matched-resampling rules specified.
- Shared physics ranking: The algorithm lifts candidates to a common state space, constructs an anchor, evaluates its complete defect, solves one linearized correction, and maps the corrected state to task space.The resulting q is shared across all N candidate comparisons.
- Anchor construction: The eight Burgers and reaction-diffusion libraries used arithmetic-mean anchors, while later comparisons used geometric medoids defined by candidate distances.These are distinct framework variants rather than universally interchangeable choices.
- Shared physics ranking: For linear task maps, qlin and qeval coincide exactly, so the shared task proxy can be evaluated as either the linearized or corrected-state task output.Nonlinear task maps retain distinct linearized and evaluated proxies.
- Baselines: Baselines include direct residual energy, candidate-specific linearized correction magnitude, validation loss, and squared distance to a candidate mean or medoid.The candidate-specific control uses one correction per candidate, unlike the shared response.
- PDE tasks: The Burgers task uses trajectories generated from ten observed frames through a fixed 91-step autoregressive rollout, with terminal-state task evaluation.The governing equation is evaluated with the prescribed initial state and complete defect formulation.
Code availability
The paper provides code and tests for reproducing its numerical analyses, including scripts for all six main figures.
- Code availability: Reproduction code, tests, and scripts for all six main figures are available in the cited source package.The package is identified by a repository URL and commit dafb2a0.
Ethics declarations
The supplied material develops finite-library decision geometry, observability, noise amplification, and restricted-domain limitations. It also states that exact recovery depends on the candidate-difference subspace, exposure conditions, and stability relative to decision margins.
- Finite-library decision quotient: Theorem 1 states that projecting the target onto D determines exact gaps, pair signs, weak ordering, and tie-broken top-1 decisions, with dim D ≤ N −1.Here D is the span of candidate differences.
- Finite-library decision quotient: Globally exact recovery requires observations whose kernel is orthogonal to the candidate-difference subspace, with finite-dimensional observation size at least dim D.The result applies to exact gaps, signs, weak ordering, and fixed tie-broken top-1 decisions.
- Restricted domains: On restricted physical domains, the kernel condition is necessary only when the domain exposes the relevant decision directions.A non-exposing domain can make sign and top-1 recoverable from less information than exact gaps.
- Continuous statistics: Continuous exact-gap statistics require output dimension at least the comparison dimension r, even without continuity, linearity, or measurability restrictions on the decoder.The proof uses injectivity of relative gaps on D.
- Stable recovery: Stable comparison recovery requires a finite observability constant and a bounded factorization recovering the comparison projection from observations.Identifiability alone does not guarantee stable recovery.
- Noise amplification: Under bounded observation noise, deterministic minimax error is controlled by the comparison observability constant, while restricted manifolds or non-spherical noise can yield smaller constants.The theorem concerns unrestricted state classes and full deterministic noise balls.
S2 Stable task factorization and nonlinear decision radii
Stable task factorization characterizes when a task comparison can be recovered from the linearized physics defect, while nonlinear localization converts uncertainty into certifiable ranking decisions.
- Stable task factorization: J = TL is equivalent to a bounded task factorization, or equivalently ∥Jh∥Z ≤ κ∥Lh∥Y for every h.The optimal factorization norm is κ∗(J | L).
- Stable task factorization: For scalar tasks, factorization reduces to the adjoint equation L∗z = J, so state identifiability may fail while a comparison direction remains observable.
- Nonlinear decision radii: When ηv = 0, the task correction is second-order accurate in a C1,1 setting once a localization engine supplies an anchor-to-root radius.The paper supplies such a computable radius for the strongly monotone reaction-diffusion operator.
- Nonlinear decision radii: A unique local zero and compact projected uncertainty set enable computable pair, top-1, regret, and abstention decisions when the stated constants and supports are available.The local Newton enclosure requires invertibility, nonlinear bounds, and a radius contained in the domain.
- Decision certification: A pair is certified when its decision interval lies strictly on one side of zero; a candidate is certified top-1 when all its pair intervals favor it.
- Panel selection: The bounded-i.i.d. concentration result is not directly applicable to the tested panel settings, which require clustered analysis because sampling is stratified and dependent.The paper distinguishes this theoretical result from the panel-selection studies.
S3 Periodic viscous Burgers: functional setting and proofs
The periodic viscous Burgers analysis establishes the functional setting and well-posedness needed for linearized and nonlinear correction identities, then characterizes their accuracy and sharpness.
- Functional setting and well-posedness: The periodic viscous Burgers equation has a unique strong solution for every u0 ∈ H1(T) and f ∈ L2(I; L2(T)).The solution bounds depend on ν, T, ∥u0∥H1, and ∥f∥L2.
- Functional setting and well-posedness: For every fixed v, the linearized operator Lv is a bounded isomorphism, supporting the correction solves used by the framework.
- Correction identities: The exact secant and plug-in identities yield candidate corrections and task estimates, with localization based on a condition such as 4αvηv < 1.
- Scope of the analysis: The Burgers localization theorem is not evaluated as a decision-certificate chain and does not support the paper’s hidden-reference accuracy claim.
- Self-midpoint correction: Under uniform inverse bounds, the self-midpoint error satisfies cn − e = O(∥e∥^(n+2)), while cubic order is specific to one self-midpoint step.Repeated Newton iterations can converge faster, so the result is not an oracle-budget optimum.
- Order-sharp rank reversals: A constructed perturbation makes the first true error smaller but its diagnostic score larger, proving the second-order norm-ranking margin is sharp for this decoder.
- Order-sharp rank reversals: A cubic-gap rank reversal also occurs for a fixed decoder, but it is not a PDE-wide impossibility independent of residual representation or computational budget.
S4 Sharp decision calculus and theory-only extensions
The sharp decision calculus expresses ranking uncertainty through projections onto candidate-difference directions and extends this framework to Bregman losses, anchor composition, and multi-anchor certification.
- Projected uncertainty: For each pair, the attainable gap set equals the proxy gap plus twice the uncertainty projected onto the candidate difference dij.
- Projected uncertainty: A candidate defeats another for every admissible target exactly when the corresponding projected gap condition holds, and unique top-1 status follows by satisfying it against every competitor.
- Projected uncertainty: For a ball uncertainty set, the sharp pairwise margin radius is 2ε∥dij∥M, with equality attained in every nonzero pair direction.
- Uncertainty composition: Independent uncertainty sources combine by Minkowski sum, allowing physics, algebraic-solve, discretization, branch, and operator–data discrepancies to be represented together.
- Bregman and smooth-loss coordinates: For Bregman losses, the exact decision coordinate is the projection of ∇Φ(t) onto the candidate-difference span, subject to open-domain exposure conditions.
- Multi-anchor extensions: Multi-anchor interval intersections can certify pairwise and top-1 decisions, but this construction is theory-only and was not evaluated in the reported studies.
S5 Operator extensions under finite training observations
Finite training observations cannot determine deployment behavior in hidden directions, while adjoint and nonlinear analyses clarify when shared physics diagnostics remain exact, sharp, or scope-limited.
- Finite-observation lifting: Finite observations permit continuous operator extensions that match all training features while changing the output at an unobserved deployment input.
- Pairwise comparison structure: Under squared loss with a bounded linear task, quadratic terms cancel and the unknown truth enters each pairwise comparison as an affine functional.
- Adjoint estimation: A direct adjoint estimator computes every linear-task pairwise point estimate simultaneously through one shared primal correction, whereas the conservative adjoint envelope was too loose to select models.
- Non-Hilbert extensions: For general smooth losses, the local gradient-difference span remains useful, but the exact N − 1 comparison-subspace reduction is special to squared Hilbert distance.
- Decoder dependence: Rank-reversal order depends on the decoder and residual coordinate, so high-order curvature is not an invariant of the PDE zero set.
- Scope boundaries: Residual zero does not identify a designated branch; the tested PDEs instead use fixed unique numerical branches in their regimes.
S6 PDE diagnostics and numerical contracts
The study specifies implementation-level physics defects, propagation schemes, task metrics, reference resolutions, and numerical tolerances across Burgers, reaction–diffusion, Sine–Gordon, and PDEBench systems.
- Burgers: Burgers uses periodic Fourier-grid trajectories, complete initial-and-time-step PDE defects, integrating-factor RK4 propagation, and terminal physical periodic L2 loss.The common task grid has 384 points in the two-PDE studies.
- Reaction–diffusion: Reaction–diffusion defects include boundary and interior terms, with exact boundary correction and preconditioned conjugate-gradient Jacobian solves.Interior residuals use conservative diffusion coefficients and the reaction factor λ + 3γw^2.
- Sine–Gordon: Sine–Gordon uses complete displacement-and-velocity velocity-Verlet defects, exact lower-triangular Jacobian action, forward substitution, and terminal phase-space physical L2.Candidate fields are Fourier-lifted to the common task grid.
- PDEBench: PDEBench rolls out three U-Net checkpoints autoregressively from ten observed frames to 101 frames on a 128×128 homogeneous-Neumann grid.Each model maps the same ten-frame, 20-channel history to one two-field frame.
- Numerical contracts: Raw-residual comparisons use matching candidates, inputs, reference states, and explicit block, time, and quadrature weighting rather than an unspecified continuous residual norm.Propagation differs computationally because it includes a linearized solve.
- Reference and solver checks: Two-PDE references use independent coarse/fine calculations, while PDEBench numerical trajectories and complete linearized-constraint closures were checked against tight residual tolerances.The PDEBench maximum relative step residual was 9.999 × 10^-11 and maximum closure was 2.310 × 10^-10.
S7 Candidate generation and evaluation sets
Candidate libraries span multiple PDEs, architectures, training runs, checkpoints, and deliberately shifted evaluation populations, with explicit admission, disjointness, and tolerance rules.
- Candidate generation: Eight cells crossed PDE, FNO/CNO architecture, and ensemble replicate, with 64 candidates from 32 runs and checkpoints at epochs 20 and 40.Admission required finite metrics, relative L2 error at most one, applicable hard conditions, and complete libraries.
- Evaluation sets: Each PDE evaluation used 240 balanced inputs across frequency, parameter, and combined shifts designed to test distributional generalization.The shifted inputs were disjoint from training, validation, and development populations.
- Selection panels: Small-panel selection reused eight two-PDE libraries across nested budgets of 3, 6, 12, and 24, with clustered resampling preserving shared-input dependencies.Each PDE used eight balanced 24-input panels plus a disjoint 240-input deployment population.
- Additional libraries: A separate study trained 16 new libraries and 128 candidates with disjoint runs, seeds, checkpoints, training samples, validation samples, and evaluation populations.It evaluated 480 instance-wise inputs, 2,304 selection inputs, and 1,920 evaluation inputs.
- Primary outcomes: The arithmetic-mean shared state achieved 97.9883% strict pairwise and 94.5833% exact top-1 accuracy, versus 71.7736% and 41.0677% for raw residual.One reaction–diffusion CNO library was markedly less accurate because two checkpoints were outliers under shift.
- Additional benchmarks: The Sine–Gordon study used eight FNO/CNO libraries and 240 new balanced inputs, while PDEBench compared three released U-Net training variants without fine-tuning.PDEBench inputs came from identifiers 0901–0999, absent from gradient training.
S8 Metrics, aggregation, and uncertainty
The evaluation distinguishes strict, tolerant, tie-aware, and top-1 metrics, aggregates dependent observations with clustered bootstrap procedures, and reports strong propagated-diagnostic performance with explicit scope limits.
- Metrics: Pairwise gaps use canonical ordering, positive values favor the first candidate, and primary accuracy conditions on hidden-reference non-ties under 10^-6 tolerance.Top-1 accepts a checkpoint within 10^-6 of the minimum, with exact top-1 reported separately.
- Metrics: The PDEBench study uses strict within-input pair comparisons and exact checkpoint-identity top-1, treating candidate pairs as dependent rather than independent samples.Its bootstrap interval is degenerate and is not interpreted as population uncertainty.
- Comparative results: Propagation-minus-raw differences were 29.2711 percentage points pairwise and 55.4732 top-1 in the 16-library comparison.The corresponding confidence intervals were [28.4673, 30.0791] and [51.0410, 59.9219].
- Pairwise results: Under matched 10^-6 tolerances, propagation labeled 99.5480% of all pairs correctly and recovered 4,413 of 4,437 reference ties, versus 62.2321% and one tie for raw residual.The primary all-pair three-way counts were 49,176/53,760 for propagation and 33,457/53,760 for raw residual under method-specific tolerances.
- Regret: Mean span-normalized regret was 2.4048 × 10^-4 for propagation and 0.274746 for raw residual, with medians of 0 and 0.128197.Performance varied by library: Burgers architecture groups showed no improvement in the reported mean normalized regret.
- Aggregation: Pooled clustered-bootstrap means were 99.7063% pairwise and 98.9549% top-1, compared with pooled count ratios of 107,205/107,520 and 3,800/3,840.The pooled estimators are not ratios of cell totals.
- Selection and correction: The corrected shared output matched or improved the best checkpoint’s task loss in all 1,920 primary cases, while shared and candidate-specific correction selected the same checkpoint in 191/192 cases.The remaining pair disagreements favored the shared method by 48 correct decisions to 40.
- Scope boundary: Latency measurements cover physics-component scaling only; end-to-end serving latency also depends on inference, data movement, scheduling, and the surrounding system.This is the stated scope boundary for computational-cost interpretation.
S10 Cross-resolution and locality analyses
Cross-resolution analysis derives exact drift-based diagnostic identities and certified interval screens, while empirical locality and predictive analyses reveal architecture-dependent stability and limits to universal thresholds.
- Resolution sensitivity: For the reaction–diffusion 256-to-512 transition, cells had 0, 1, 4, and 1 tolerance-label changes, with no top-1 changes.Every changed pair was within the bottom 0.70% of its cell’s margin distribution; continuum rates require more resolutions and asymptotic analysis.
- Cross-resolution identity: Coarse and fine outputs are lifted into one fixed Hilbert task space, where diagnostic gaps are expanded using shared-diagnostic and candidate grid-to-grid drifts.The expansion includes linear drift terms and a remaining bilinear term.
- Cross-resolution identity: The drift expansion is an exact algebraic identity, so “first order” and “second order” denote polynomial degree rather than a small-drift approximation.The identity closes on every row with maximum error 9.1593 × 10^-16 in the two populations.
- Certification: The interval screen certifies a pair sign or tie only when the interval lies wholly outside or inside the prescribed tolerance region; otherwise it abstains.A top-1 decision passes when one first-order winner is certified against every competitor.
- Certification: The exact cross-resolution stability screen had pooled coverages of 65.4464% and 65.0800%, with every covered decision algebraically guaranteed to agree.Uncovered decisions abstain, and coverage does not measure hidden-reference ranking.
- Locality and prediction: The first-order gap-change predictor improved pooled FNO AUROC by 0.0662 and AUPRC from 0.3639 to 0.5060, but reduced CNO AUROC by 0.0342 and AUPRC from 0.5724 to 0.5455.The contrasting FNO and CNO trends preclude a common empirical threshold in these data.
- Locality and prediction: Candidate-difference subspace angles did not explain pair flips, and a task-adjoint residual envelope had a median envelope-to-gap ratio of approximately 1.98 × 10^5.The envelope was therefore far too loose for useful decisions, while low-mode energy alone did not establish effective rank.
S11 Operator-aligned selective certification for cubic reaction– diffusion
This section develops exact selective certificates for ranking candidates under a strongly monotone reaction–diffusion discretization. The certificates bound operator-aligned pairwise gaps, account for projected dataset discrepancy, and abstain when sufficient conditions fail.
- Global residual enclosure: The theorem certifies the unique discrete solution of the strongly monotone reaction–diffusion operator and bounds the error of any approximate state.The result is formulated for the exact same-grid operator zero.
- Pairwise certification: Operator-aligned pair bounds guarantee the proxy and exact pairwise gaps have the same sign whenever the proxy gap exceeds its uncertainty bound.This yields a directly checkable condition for reliable pairwise decisions.
- Top-1 certification: Strict and tolerant top-1 certificates identify a unique minimizer or guarantee loss within τ of the best candidate; otherwise, the procedure abstains.The conditions apply to exact same-grid discrete-solution losses.
- Computational cost: One residual inverse action and r comparison-basis inverse actions evaluate all pairwise bounds, eliminating candidate-specific corrections while making certificate cost comparison-rank dependent.The implementation reuses the common matrix and scales with the comparison subspace dimension.
- Transfer limitation: Transfer from the discrete operator zero to dataset truth requires an additional projected discrepancy enclosure, which was unavailable for PDEBench and exploratory flow datasets.Discrepancy components orthogonal to the candidate-difference space do not affect squared-loss decisions.
- Numerical scope: The reported numerical comparisons use approximate operator zeros without propagating solver uncertainty, and floating-point roundoff is not rigorously enclosed.A strict floating-point certificate would require validated arithmetic such as outward rounding.
S12 Multi-objective transfer and proxy mechanism
The study transfers one corrected trajectory proxy across multiple PDEBench objectives and examines how correction changes full-state versus comparison-subspace accuracy. Metric-matched rescoring improves rankings broadly, while the mechanism differs between systems and nonlinear metrics remain outside the exact Hilbert-space identity.
- Objective-matched ranking: 88/99 high-band restricted-quadrant winners were recovered, versus 69/99 for the fixed terminal score, 23/99 for common-state rescoring, and 7/99 for raw defect energy.This objective produced five complete candidate orders across 99 inputs.
- Multi-objective transfer: 2,938/2,970 pair preferences, 979/990 top-1 decisions, and 959/990 complete orders were correct across ten correlated objectives with metric-matched rescoring.The corrected trajectory had no larger truth error than the best public candidate in all 990 metric–input cells.
- Metric scope: The normalized-RMSE plug-in is not one fixed truth-independent metric on both sides, so the Hilbert-space identity does not apply to it.Roots, maxima, normalization, and spectral reductions are treated as empirical transfer tests.
- Partial correction: At α = 1, aggregate pairwise accuracy rose from 51.9% to 98.1% for Sine–Gordon and from 54.9% to 98.9% for PDEBench.Overshooting to α = 1.25 reduced these accuracies to 96.9% and 93.5%, respectively.
- Partial correction: At α = 1, median full-state error contracted by factors of 1,835.7 in Sine–Gordon and 9.86 in PDEBench relative to the common state.The shared path reused the same correction at every tested α.
- Proxy mechanism: PDEBench comparison-subspace error contracted 16.1-fold versus 9.86-fold for full trajectory error, whereas Sine–Gordon success was primarily a full-proxy effect.Exact pair-gap drift identities closed to 2.11 × 10^-14 in the squared-Hilbert tasks.
S13 Candidate-wise variational residual baselines
This section compares stability-compatible variational residual estimators with raw residual ranking and probes corrected proxies in compressible-flow settings. Shared correction substantially improves reaction–diffusion selection, but flow results remain conditional on the operator, common state, and tested library.
- Reaction–diffusion results: The shared corrected proxy recovered all 42,761 non-tie pair decisions and all 1,920 tolerance-optimal choices in the reaction–diffusion controls.Its three exact top-1 misses had maximum regret 4.19 × 10^-9.
- Estimator comparison: The common-energy and candidate-adapted estimators selected the same top-1 candidate in all 1,920 cases and differed on only two of 53,760 pair preferences.The adapted construction nearly doubled median component time without decision improvement.
- Certification limitation: Variational scores are individual error upper bounds rather than pairwise certificates, so choosing the smallest majorant does not prove the smallest true error.The comparison is limited to the stationary identity-task reaction–diffusion population and post-inference diagnostic timings.
- Two-dimensional flow: In the two-dimensional flow probe, correction improved full-state and comparison-direction errors in every case, but perpendicular error improved in only four cases.The candidate-difference operator had rank one for the two-candidate library.
- One-dimensional flow: In the one-dimensional shock study, backtracking improved full-state error in 33/35 cases and comparison-subspace error in 34/35, with median factors 2.04 and 5.20.Comparison contraction exceeded full-state contraction in 31/35 cases.
- Flow scope: The best fixed selector reached 94.3% top-1 versus 91.4% for backtracking, and strong-shock conclusions require an operator-aligned generator study.The released-truth defect quantified rather than removed the operator mismatch.