Source-linked AI summary
Beyond Field Accuracy: Two-Axis Diagnosis of Inverse-PINN Parameter Error
Yifan Zhang, Qian Tao
TL;DR
Inverse PINNs can recover a field accurately while misestimating the scientific parameter, and field error alone does not explain this divergence. The paper introduces a two-axis post-training diagnosis separating finite-sample resolution from signed parameter preference and finds that its displacement tracks residual-profile minima and delivered parameter error across controlled synthetic tests.
Problem
Inverse PINNs can show divergent field and coefficient accuracy, leaving parameter recovery insufficiently characterized by field accuracy alone.
Method
The paper combines matched-forward refitting with a frozen-field, residual-view-aligned local score displacement and endpoint-consistency check.
Results
r = .945–.982 for frozen-profile minima and r = .994 for delivered signed log-error, with 237/240 directions correct across controlled synthetic tests.
Takeaways & Limitations
The two coordinates route follow-up toward observations, residual evidence, or endpoint delivery rather than treating parameter error as additive components.
Takeaways & Limitations
The diagnostic assumes a differentiable quadratic residual view with fixed field, metric, and weights, and larger coupled systems require separate validation.
Abstract
from arXiv · showhide
Inverse physics-informed neural networks (PINNs) can reconstruct a field accurately while returning an incorrect physical parameter. We introduce a two-axis post-training diagnosis that separates finite-sample resolution under a specified observation-and-estimation protocol from the signed parameter preference encoded by the final learned field and residual metric. The first axis repeatedly fits noisy observations with a matched forward estimator. At known synthetic truth, the second freezes the field and residual view and computes a local score displacement toward a nearby residual-profile minimum. Endpoint consistency then tests whether joint training delivers that preference under the same final view. Across three synthetic one-dimensional, scalar-parameter PDEs, matched-forward mean absolute relative error ranges from 2.34 percent to 17.46 percent. The displacement tracks frozen-profile minima across locked seeds, architectures, and fresh-noise retraining (r from .945 to .982), and it tracks delivered signed log-error in 240 fresh-noise RBA runs (r = .994; 237/240 correct directions). A coupled two-parameter Darcy check validates the full matrix calculation. The axes are complementary diagnostic coordinates, not additive error components or a deployable oracle-free estimator. Together, they route follow-up work toward observations, residual evidence, or endpoint delivery.
1 Introduction
Inverse-PINN field accuracy does not determine coefficient recovery, motivating a post-training diagnosis that separates matched-forward finite-sample resolution, frozen-field residual preference, and endpoint delivery. Across controlled synthetic PDE studies, the proposed displacement tracks residual minima and delivered signed parameter error, while matched-forward error varies across problems.
- Motivation: Lower field error can coincide with either improved or degraded coefficient estimates, separating field accuracy from coefficient recovery.The learned neural field is treated as an intermediate representation, while the inferred coefficient is the scientific target.
- Diagnostic framework: The diagnosis connects repeated matched-forward performance, signed final-residual preference, and endpoint consistency as distinct measurements.Matched-forward fitting provides a finite-sample reference for the specified estimator and observation protocol; endpoint consistency tests agreement between returned parameter and residual preference.
- Metric-aware score geometry: DM = −BM/HM combines residual magnitude with signed residual–sensitivity alignment to approximate a nearby profile displacement.BM is the signed residual–sensitivity score, and HM is local parameter strength; the matrix formulation is checked numerically for two coupled parameters.
- Controlled validation: r = .945–.982: DM tracks near-truth minima of the same residual objective across locked seeds, architectures, and fresh-noise blocks.The displacement tracks these minima more closely than field error or residual MSE.
- Controlled validation: r = .994, with 237/240 correct directions: DM tracks returned signed log-error in a view-aligned fresh-noise replay.Matched-forward error ranges from 2.34% to 17.46% across the three synthetic one-dimensional scalar-parameter PDEs, while paired neural–forward excess changes sign.
2 One Error, Two Scientific Objects
The paper separates inverse-PINN error into observation-axis precision from a specified matched-forward protocol and learned-field-axis signed parameter preference from a frozen residual profile. Endpoint consistency then measures whether joint training delivers that field-specific preference under the same residual view and weighting convention.
- Observation axis: The observation axis estimates repeated-noise parameter precision for a specified matched-forward estimator, data protocol, PDE, solver, observation layout, numerical budget, and noise model.Repeated generation and fitting targets the finite-sample estimand, with fixed observation layouts and noise assumptions.
- Learned-field axis: The learned-field axis measures signed near-truth preference from a frozen residual profile built from one realized dataset.The field, residual metric, and learned weights are frozen before evaluating the profile.
- Learned-field axis: The displacement is a local score-based diagnostic: it has a unique linearized minimizer when HM > 0, whereas HM = 0 leaves the linearized profile flat.Agreement with the nonlinear basin depends on curvature and higher-order terms, so frozen-profile scans test the local approximation.
- Endpoint consistency: Endpoint consistency separates the frozen profile preference from the endpoint-delivery gap under the same final residual view and weight convention used in training.The local displacement approximates the profile-preference term when the selected minimum lies in the same local basin.
- Learned-field axis: Residual-metric parameter strength, remaining residual, and signed alignment jointly determine local parameter displacement.A small aligned residual can produce more movement than a larger nearly orthogonal residual; signed modal contributions may also cancel.
3 Two-Axis Experimental Design
The study uses three linked measurements to separate observation resolution, learned-field preference and endpoint delivery, while comparing RBA with matched-forward errors. It evaluates this design across three scalar one-dimensional PDEs under a common observation and training protocol, with locked and fresh-noise confirmation blocks.
- Experimental design: Three linked measurements assess frozen-field preference and endpoint delivery, repeated-noise matched-forward observation resolution, and same-data RBA-versus-matched-forward parameter errors.These measurements form the study’s two-axis experimental design.
- Benchmarks and protocol: The benchmarks are Burgers viscosity p⋆= 0.01, Buckley–Leverett mobility ratio p⋆= 2, and Allen–Cahn reaction coefficient p⋆= 5.Each PDE uses four fixed layouts of 200 random observation sites with Gaussian noise of standard deviation 0.3 std(u).
- Residual views: Residual interventions vary rank, scale, and adaptive emphasis through pointwise, PatchEvidence, Gaussian, and RBA views to isolate mode visibility from signed parameter effect.PatchEvidence has rank at most 64 on 2,000 residual coordinates, while full-rank Gaussian retains the complete operator.
- Frozen-field diagnosis: Frozen-field diagnosis freezes the learned field and final residual view, then compares the signed displacement of an 81-point profile minimum with the trained parameter.The diagnosis computes BM, HM, DM, ∥r∥M, ρM, and eigenbasis contributions at known truth.
- Confirmation blocks: The evaluation uses 2,400 matched-forward estimates and 240 paired fresh-noise RBA comparisons, with confirmation across seed holdout, architectures, retraining, and coupled Darcy profiles.Seeds 8–9 define development, while seed-10–11 holdout, two architectures, 240 fresh RBA retrainings, and 20 Darcy profiles provide progressively different confirmation.
4 Results: Preference and Resolution
Field accuracy and parameter accuracy separate across controls, while score displacement closely predicts frozen-profile preference and delivered endpoint direction when residual views align. Matched-forward resolution and pipeline excess reveal distinct PDE-specific error profiles, and the coupled Darcy check validates the full matrix displacement.
- Preference and resolution: Field and parameter errors are nearly uncorrelated (r = .057; PDE-centered r = .220), while kernel-metric field gains produce parameter responses ranging from improvement to deterioration.On Allen–Cahn, lower field error accompanies larger parameter error.
- Preference and resolution: 72/72 development profiles have correct signed directions, with magnitude correlation 0.960; on 24 pointwise profiles, absolute-profile-error correlation is 0.966 versus 0.525 for field relative L2 and 0.139 for residual MSE at truth.Agreement transfers across locked seeds, architectures, and fresh-noise RBA fields.
- Preference and resolution: r = .994 signed-log correlation and 237/240 correct directions show that displacement tracks delivered endpoint error across 240 fresh-noise RBA retrainings under the corresponding frozen weighted residual views.The aligned score also correlates with delivered absolute errors at r = .983 [.977,.987] with 1.02-percentage-point MAE.
- Preference and resolution: Median cosine .999994 and relative error 9.50% for the full Darcy displacement beat 79.45% for the diagonal readout, with the full calculation winning 20/20 coupled cases.The check used 20 saved two-dimensional Darcy fields and a continuously refined two-dimensional frozen-profile minimum.
- Preference and resolution: 4.18% mean error on the original 12 benchmark realizations fell at the 37th percentile of repeated-noise distributions, versus an 8.83% balanced expectation.The 2,400-fit matched-forward reference agreed with the local Fisher prediction and achieved near-nominal profile coverage.
- Preference and resolution: 10.27% and 10.41% noiseless bias under coarse solvers reduced nominal coverage to 89.75% and 72.5%, while paired RBA-minus-forward excess changed sign for Burgers and differed by PDE.Allen–Cahn had the tightest resolution and largest positive excess; Burgers had weakest resolution and negative excess; Buckley–Leverett was intermediate with positive excess.
5 Related Work
Related work spans inverse-PINN formulation, parameter recoverability and uncertainty, frozen-field coefficient preference, and residual-view or attribution methods. These strands provide context for diagnosing accurate fields that nevertheless yield incorrect physical parameters.
- Inverse-PINN training and formulation: PINNs support joint state and parameter learning, while prior studies diagnose gradient imbalance and loss-geometry failures during training.Generic multi-task gradient surgery and PINN-specific balancing or adaptation alter update directions.
- Matched recovery and uncertainty: Recoverability research uses practical-identifiability and Cramér–Rao analyses, while Bayesian, conformalized, and Neyman-inversion methods quantify uncertainty.WALDO constructs finite-sample frequentist confidence regions by Neyman inversion.
- Frozen-field coefficient preference: Two-stage equation-error estimators and neural-field extensions show that small field error can coexist with substantial parameter error.Frozen-PINN residual sweeps recover encoded coefficients or expose fixed-field landscapes.
- Residual views and attribution: Residual formulations include weak variational residuals, conservative domain decomposition, goal-oriented weighting, and influence-based tracing of predictions or losses.Goal-oriented analysis has informed adaptive PINN sampling.
6 Using the Two-Axis Audit
The two-axis audit organizes evidence from repeated-noise matched-forward resolution, displaced residual preference, and endpoint-delivery gaps when parameter error is scientifically material. Known-truth audits route follow-up experiments while preserving the distinction between local residual preference and truth-centered parameter error.
- Known-truth audit: Known truth enables three stages: estimate matched-forward resolution, scan frozen residual profiles, then test endpoint consistency on identical fresh datasets.The protocol declares the forward, observation, solver, noise, and matched-estimator specifications before comparing both estimators.
- Decision routing: Matched resolution routes observation redesign, whereas ∆M routes residual-evidence decisions; tight resolution and small ∆M redirect attention to endpoint delivery.Follow-up may also target model specification or training budget.
- Authority and limits: The displacement is authoritative only for a differentiable quadratic residual view with the field, M, and learned weights fixed.Small or singular eigenvalues of HM limit the minimum-norm displacement to the identified subspace, while nonlinear-profile disagreement signals curvature, another basin, or scan-domain limits.
- Decision semantics: The audit coordinates are not additive delivered-error components and require application-declared tolerances for matched-forward resolution and tolerated frozen-profile shift.In coupled problems, the same reading applies along identified eigen-directions of HM.
- Unknown-truth use: Without known truth, candidate-centered scans measure local stability and curvature relative to the reported candidate, not distance from physical truth.Independent forward fits still quantify resolution under the declared observation and noise protocol.
7 Conclusion
Inverse-PINN parameter error comprises finite-sample resolution and signed parameter preference, diagnosed through matched repeated-noise estimation, local residual-profile displacement, and endpoint consistency. Within validated synthetic PDE settings, these local empirical axes form an attribution workflow rather than a global or oracle-free estimator.
- Two-axis diagnosis: Inverse-PINN parameter error separates finite-sample resolution under a specified protocol from signed preference encoded by the learned field and residual metric.Matched repeated-noise estimation measures resolution, while DM = −BM/HM characterizes local near-truth residual-profile preference.
- Action-oriented audit: Field error and residual MSE alone do not identify the source of parameter error because the audit distinguishes observation resolution, residual preference, and endpoint delivery.Endpoint consistency tests whether training delivers the preference encoded by the fixed learned field and residual metric.
- Action-oriented audit: Weak matched resolution motivates better observations or a revised parameter scope, whereas material frozen-field displacement motivates inspecting residual magnitude, signed alignment, and parameter strength.A small displacement with a displaced endpoint instead directs attention to optimization, model specification, or nonlocal behavior.
- Scope and limits: The diagnostic is a truth-centered local empirical oracle, not a global guarantee or deployable oracle-free estimator, validated on three synthetic one-dimensional scalar-parameter PDEs and one coupled two-parameter Darcy check.Within that scope, reading the three quantities together provides a practical attribution workflow.
Reproducibility Statement … Development and Block Authority
The paper diagnoses inverse-PINN parameter error through complementary finite-sample resolution and signed residual-based parameter preference, while distinguishing claim scope, evidence authority, and reproducibility materials. The appendix links protocols, checks, provenance, and archival code/data to audited evidence blocks.
- Reproducibility Statement: The appendix aligns equations, blockwise protocols, per-case tables, numerical checks, and claim boundaries with corresponding evidence sections.Its guide uses the same notation, evidence-block names, caption style, and claim boundaries as the main paper.
- Reproducibility Statement: An archival code-and-data release contains analysis scripts, staged evidence, selected checkpoints, and environment locks, but is excluded from the arXiv TeX source archive.Reported replication units remain specific to development, score-ratio holdout, architecture, and fresh-noise evidence blocks.
- Appendix Guide: The appendix guide organizes supporting material into claim scope, benchmark specification, implementation protocol, score geometry, finite-sample recoverability, and falsification sections.It provides direct links to supporting specifications and checks while summarizing the main interpretation in the two-axis audit.
- A Claim Scope and Evidence Map: Two complementary coordinates separate finite-sample resolution under the observation protocol from signed parameter preference supported by a fixed field’s residual evidence.The first is a repeated-sample reference; the second is conditional on a field trained from realized, potentially noisy observations.
- Development and Block Authority: Residual interventions preceded the diagnosis study, while the signed-score readout was developed on existing seed-8–9 fields and is therefore truth-informed rather than method-holdout development.Seeds 10–11 provide a readout-only holdout.
- Development and Block Authority: Architecture transfer locks the readout on existing architecture fields, fresh RBA noise retrains the unchanged pipeline, and the Darcy block applies the locked procedure.These blocks distinguish transfer, fresh-noise retraining, and coupled-parameter validation roles.
- Development and Block Authority: Score-ratio holdout governs only the new diagnostic readout, not field training.This authority boundary limits how holdout evidence supports the diagnostic claim.
- Development and Block Authority: Claims are restricted to audited benchmarks, and the vector claim is supported by one Darcy parameterization.The claim-to-evidence map makes this scope explicit.
B Benchmark and Observation Specification · C Implementation and Statistical Protocol
The study evaluates three declared scalar PDE benchmarks under matched discretization and specified observation layouts, then applies standardized inverse-PINN implementations, residual-profile diagnostics, and clustered statistical reporting. The protocol fixes training, profiling, and bootstrap procedures while separating matched-forward estimation from coarser-solver stress testing.
- B Benchmark and Observation Specification: Burgers uses ν⋆ = .01 on x ∈[−1, 1], t ∈[0, 1], with sinusoidal initial data and homogeneous Dirichlet boundaries.The declared equation is ut + uux = νuxx, with u(x, 0) = −sin(πx) and u(−1, t) = u(1, t) = 0.
- B Benchmark and Observation Specification: Buckley–Leverett uses M⋆ = 2 on x ∈[0, 1], t ∈[0, .5], with discontinuous initial data and fixed endpoint values.Its flux is fM(u) = u2/[u2 + M(1 −u)2], with u(x, 0) = 1{x < .25}, u(0, t) = 1, and u(1, t) = 0.
- B Benchmark and Observation Specification: Allen–Cahn uses λ⋆ = 5 on x ∈[−1, 1], t ∈[0, 1], with quadratic-cosine initial data and homogeneous Neumann boundaries.The equation is ut −.001uxx + λu(u2 −1) = 0, with u(x, 0) = x2 cos(πx).
- B Benchmark and Observation Specification: Each fine reference uses 201 spatial and 2,001 temporal nodes, while the matched forward estimator uses the same discretization and the coarser solver serves as a discrepancy stress test.Four observation layouts use layout seeds 8–11 and sample 200 fine-grid sites; each neural run also samples 2,000 residual sites.
- C Implementation and Statistical Protocol: The shared scalar implementation uses 200 observations, 2,000 fixed residual sites, a width-64, depth-3 tanh MLP, and 5,000 Adam steps at learning rate 10−3.Initial and true parameters are (.02, .01), (1, 2), and (2, 5) for Burgers, Buckley–Leverett, and Allen–Cahn, respectively.
- C Implementation and Statistical Protocol: PCGrad, PatchEvidence, and RBA define distinct residual-training procedures, with PCGrad jointly updating field and parameter from step 1 and PatchEvidence using warm-up and patch residuals.RBA uses cumulative normalized absolute-residual weights with decay .999 and increment .001; aligned Fresh RBA freezes the saved weights from the last evaluated training loss.
- C Implementation and Statistical Protocol: Frozen profiles use 81 log-spaced candidates and a centered log-parameter finite difference of 10−3, with robustness checks spanning 81 and 161 candidates and steps 5 × 10−4, 10−3, and 2 × 10−3.Gaussian and rank controls use independently min–max normalized coordinates, fixed bandwidth 1/(8 6) = .051031, and reported truncation ranks 64, 256, or full.
- C Implementation and Statistical Protocol: Associations are Pearson correlations, within-PDE correlations subtract each PDE mean, and percentile cluster bootstraps preserve declared observation-layout or residual-metric clusters.Reported analyses use 20,000 draws for RBA profile, endpoint, paired, and matched-forward audits, and 10,000 draws for Darcy; blocks are not pooled across replication units.
D Score Geometry and Implementation … Solver-Discrepancy Control
The paper develops a frozen-field score geometry, tests its numerical and architectural robustness, and separates residual-metric preference from finite-sample recoverability and endpoint delivery. Operator interventions, exact-condition controls, matched-forward references, and solver-discrepancy checks show that improved fields or residual norms do not by themselves identify or deliver the correct parameter.
- Frozen-Field Derivation: HM > 0 gives a unique Gauss–Newton displacement, whereas HM = 0 leaves the linearized profile flat and the parameter direction unidentified.The exact nonlinear minimum additionally depends on CM and higher-order terms; full rank of M is unnecessary.
- Metric and Spectral Conventions: M = I/Nf for pointwise evidence, while PatchEvidence uses M = A⊤A/K with K ≤64 and rank at most 64 over 2,000 residual coordinates.The Gaussian controls use independently min–max normalized coordinates with fixed σ = .051031; reported band shares are sensitivities of HM, not observation-map Fisher information.
- Numerical Evaluation and Local-Approximation Replication Ladder: 81-point frozen profiles over the three declared log-parameter ranges use centered finite differences with step 10−3, while robustness crosses 5 × 10−4, 10−3, and 2 × 10−3.Predicted error is 100| exp(−BM/HM) −1|, whereas actual frozen-profile error uses the discrete profile minimizer and known truth defines the diagnostic.
- Architecture Transfer and Fresh RBA Noise: .98174 is the aligned score-to-profile correlation across 240 fresh RBA runs, with within-PDE .98200 and noise-cluster interval [.97840,.98567].Architecture transfer across locked PCGrad/PatchEvidence fields gives a pooled PDE–seed interval of [.918,.984].
- Delivered-Endpoint Consistency: .99434 is the signed-log correlation between DM and log(ˆptrain/p⋆), with agreement in 237/240 directions; the frozen profile is retrospective because RBA changes weights during training.The pointwise secondary readout reaches signed-log r = .99449 and 234/240 direction agreement, while final-weight and pointwise minima coincide in 230/240 cases.
- Operator Interventions and Falsification: 100% is the FRKE assignment under its normalized quadratic form on adversarial zero-mean hard-patch residuals, versus 0.684% retained by rank-64; full rank removes cancellation modes but not BM.Changing Gaussian rank from 64 to 256 improves 3/8 Buckley and 4/8 Burgers profiles but 0/8 Allen–Cahn, so no universal cutoff is supported.
- Residual-Sensitivity Selector and RBA Replay; Held-Out Prediction Pilot: Uninformative Common Minimum: 27.89% is the residual-sensitivity selector’s mean parameter error, compared with 17.38% for RBA and 13.58% for FRKE, because coverage controls HM rather than signed BM.Final RBA weights correlate most with current |r| at .797, and replayed weighted profiles share the unweighted 81-point minimizer in 6/6 cases; the held-out pilot also selects normalized location .6875 for all three PDEs.
I 2D Two-Parameter Numerical-Fidelity Check
A locked 2D Darcy audit validates the coupled two-parameter displacement calculation without retraining, showing that the full local step closely matches the continuously refined frozen-profile displacement. The check’s empirical content is full-step fidelity, not the algebraic diagonal-versus-full identity.
- Audit setup: 20 existing 2D Darcy fields were audited with locked seeds 10–19, using PCGrad and PatchEvidence without retraining or checkpoint selection.The fields came from a protocol with 256 noisy observations, 2,048 interior samples, and 512 boundary samples.
- Profile optimization: All bounded nonlinear-least-squares starts converged to the same interior minimum, with maximum near-optimal multi-start spread below 10^-9.The profile used an 81-by-81 grid over the locked box [−1.5, 1.5] × [−1, 2.5].
- Numerical comparison: The full-matrix displacement matched the continuously refined frozen-profile displacement with seed-clustered median cosine intervals [.999992,.999997] and full relative-error intervals [.0921,.0982].The diagonal alternative replaces H by its diagonal, whereas the full calculation retains coupling |H01|/√H00H11.
- Full-step fidelity: The full step reduced truth-profile loss by a median factor .00739, with nearly exact direction but 9.5% remaining magnitude error.This tests fidelity to the continuously refined non-affine profile rather than the diagonal-versus-full algebraic identity.
- Coordinate-view validation: In the box-normalized coordinate view, the full step was favored in 20/20 fields, with median errors .0954 versus .8749.This provides an additional numerical check of the coupled displacement calculation on the audited fields.
J Recoverability-Calibrated Output Pilot
This section presents a truth-centered 90% coverage feasibility pilot using two-way cross-fitting over fresh noise replicas, not generic candidate-wise inversion or deployable confidence sets. It also reports RBA inflation factors and emphasizes that pooled pilot behavior provides neither equal conditional coverage nor guarantees for unseen candidates.
- Pilot design: 20 fresh noise replicas and two-way cross-fitting produced 240 truth-centered pilot cases across PDEs and layouts.Each fold pooled 40 calibration cases, or 10 per layout for stratified calibration.
- Pilot design: The finite-sample calibration rank for 90% target coverage is ⌈(ncal + 1) × .90⌉.The pilot is local and truth-centered rather than a validation of generic candidate-wise inversion.
- RBA calibration: 1.25, 4.22, and 7.75 forward standard errors are the RBA inflation factors for Burgers, Buckley–Leverett, and Allen–Cahn, respectively.RBA operationalizes the two axes but is not more width-efficient than plain calibration and is not proposed as a generic confidence-set method.
- Coverage limits: Pooled 90% coverage does not imply equal conditional coverage across PDEs or calibration layouts.The reported empirical behavior is from the truth-centered pilot only.
- Coverage limits: Replica-level splitting preserves correlated layouts within PDE–replica clusters, but the pilot is neither a simultaneous guarantee across PDEs nor a theorem for unseen candidates.The even/odd split is made by noise replica rather than by layout.