Source-linked AI summary
Computer model validation with functional output
M. J. Bayarri, J. O. Berger, J. Cafeo, G. Garcia-Donato, F. Liu, J. Palomo, R. J. Parthasarathy, R. Paulo, J. Sacks, D. Walsh
TL;DR
Computer-model validation must address irregular functional outputs and uncertainty in model inputs while assessing whether predictions are accurate enough for intended use. The paper uses wavelet representations and hierarchical validation of wavelet coefficients, then applies the approach to a vehicle-suspension stress-analysis problem. The analyses estimate bias and uncertainty and report successful bias-corrected predictions in several altered or new settings, while noting limitations of the modular approach and uncertain bias modeling.
Problem
Existing validation methods do not directly handle highly irregular functional data or uncertainty from unmeasured manufacturing variations, both of which arise in engineering applications.
Method
The paper represents irregular functional data with wavelets, applies a hierarchical scalar validation methodology to wavelet coefficients, and transforms back to compare model and field outputs.
Results
Bias-corrected predictions successfully extrapolated to a vehicle with additional mass, while the pure model prediction was not successful in the critical Region 1.
Takeaways & Limitations
The analyses show that the framework can support predictions in altered or new settings using computer-model and field-data information together with bias and uncertainty estimates.
Takeaways & Limitations
The modular approach is motivated by confounding among calibration parameters, the bias function, and error variance, while uncertain bias modeling complicates formal justification of a full Bayesian analysis.
Abstract
from arXiv · showhide
A key question in evaluation of computer models is Does the computer model adequately represent reality? A six-step process for computer model validation is set out in Bayarri et al. [Technometrics 49 (2007) 138--154] (and briefly summarized below), based on comparison of computer model runs with field data of the process being modeled. The methodology is particularly suited to treating the major issues associated with the validation process: quantifying multiple sources of error and uncertainty in computer models; combining multiple sources of information; and being able to adapt to different, but related scenarios. Two complications that frequently arise in practice are the need to deal with highly irregular functional data and the need to acknowledge and incorporate uncertainty in the inputs. We develop methodology to deal with both complications. A key part of the approach utilizes a wavelet representation of the functional data, applies a hierarchical version of the scalar validation methodology to the wavelet coefficients, and transforms back, to ultimately compare computer model output with field output. The generality of the methodology is only limited by the capability of a combination of computational tools and the appropriateness of decompositions of the sort (wavelets) employed here. The methods and analyses we present are illustrated with a test bed dynamic stress analysis for a particular engineering system.
1. Introduction.
The paper extends a six-step computer-model validation framework to irregular functional outputs, uncertain inputs, and prediction in altered settings. It focuses on assessing predictive accuracy using model and field data together, with tolerance bounds for possible prediction error.
- The six-step framework defines the problem, establishes evaluation criteria, designs experiments, approximates model output, combines field and computer-run data, and feeds results back into model revision.
- Irregular functional outputs cannot always be handled by treating time as an additional input because this may become computationally intractable.The paper identifies irregular functional data as a technically challenging extension beyond scalar or smooth functional outputs.
- Uncertainty from unmeasured manufacturing variations in tested components can be crucial to incorporate into engineering analyses.
- The paper considers prediction in altered or new settings where field data may be unavailable.
- Validation emphasizes whether predictions are accurate enough for intended use, reporting tolerance bounds rather than a binary model-validity judgment.
- The approach combines Gaussian-process response-surface approximations with Bayesian representations of model bias and uncertainty.
2. The test bed.
The test bed applies the validation framework to time-varying suspension loads measured in field tests and generated by an ADAMS computer model. It includes uncertain manufacturing inputs, calibration parameters, curve registration, and a constrained design of computer runs.
- The case study predicts loads over time on a vehicle suspension system subjected to stressful events such as hitting a pothole.
- The initial study has seven unmeasured system parameters with nominal values x_nom and unknown manufacturing variations δ*.
- Field and model curves are registered so that their peaks and valleys occur at corresponding locations.
- The ADAMS model is a finite-element-based code for complex dynamic behavior and contains two calibration parameters representing damping levels.
- Field tests recorded load histories at two suspension sites from seven drives of a single vehicle under Condition A inputs.
- Figure 1 compares model output with registered field output for Site 1 in Regions 1 and 2, marking reference peak locations used for registration.
- A 65-point Latin hypercube design plus a nominal center point was used for model runs; one failed run was deleted, leaving 65 usable design points.
3. Formulation, statistical model and assumptions.
The formulation represents real functional responses as calibrated computer-model output plus bias, while incorporating measurement and input uncertainty. Wavelet decomposition makes irregular functional data tractable by analyzing retained coefficients, approximating model coefficients with Gaussian processes, and reconstructing the original curves.
- 3.1. Formulation.: Measurement errors are modeled as independent mean-zero Gaussian processes, with additional uncertainty allowed for some inputs.The model distinguishes measurement error from uncertainty in unmeasured physical inputs.
- 3.1. Formulation.: Real response is modeled as calibrated computer-model output plus an additive bias, while calibration parameters affect only the model output.The calibration value may be physically true or treated as a fitted value when used as a tuning parameter.
- 3.1. Formulation.: The Bayesian unknowns include model output, calibration and physical input deviations, bias, and the measurement-error covariance, but irregularity and dimensionality require simplification.The paper uses a basis representation, specifically wavelets, to make the analysis feasible.
- 3.2. Wavelet decomposition.: Wavelet decomposition represents model and field curves with shared basis elements, after thresholding retains a manageable coefficient set while maintaining adequate reconstruction accuracy.The test-bed reconstruction was similarly accurate across field and model-run curves.
- 3.2. Wavelet decomposition.: Wavelet measurement-error coefficients are assumed independent across coefficients, despite potentially differing variances, to make later computations feasible.Residual field and simulated error processes exhibit similar correlation patterns in the test-bed illustration.
- 3.2. Wavelet decomposition.: The analysis treats retained wavelet coefficients separately and recombines them to obtain estimates and uncertainty for the original functional responses.The same basis is assumed to represent reality and the bias function accurately.
- 3.3. GASP approximation.: Each retained model wavelet coefficient is approximated as a Gaussian process over the computer-model inputs and calibration parameters using computer-run data.The Bayes predictor is the posterior mean of each coefficient conditional on the available model runs.
- 3.4. Other prior specifications.: Bias coefficients receive hierarchical priors by resolution level, with Gaussian priors and strong shrinkage toward zero.The shrinkage reflects the common computer-modeling assumption that biases are zero, making detected bias more credible under the prior.
4. Estimation and analysis.
The analysis combines wavelet-coefficient modeling with Bayesian posterior computation, while using approximations and modularization to make the high-dimensional validation problem tractable. Posterior inference relies on separate coefficient-level estimation, replicate-based variance posteriors, and MCMC sampling.
- Wavelet representation: 289 wavelet coefficients are used to represent the test-bed functional output.
- Coefficient-level estimation: A full Bayesian treatment would require 5780 parameters, so each coefficient is treated separately and its model parameters are estimated by maximum likelihood.The test bed has 65 computer-experiment design points.
- Predictive distribution: The resulting GASP predictive distribution is defined using the estimated correlation function and plug-in maximum-likelihood parameter estimates.
- Modular Bayesian analysis: The analysis modularizes variance estimation by using replicate observations to determine posteriors for the σ2_i parameters rather than integrating them into the global posterior.This choice is motivated by substantial confounding among calibration parameters, the bias function, and variance parameters.
- Modular Bayesian analysis: The bias model’s uncertainty makes a full Bayesian treatment difficult, supporting the simpler modular approach despite the absence of a formal argument.
- Posterior computation: MCMC sampling uses 200,000 iterations and retains every 200th draw, producing a final sample of 1000 posterior draws.Careful proposal selection was technically crucial for suitable mixing.
5. Results.
The analysis estimates bias and reality from posterior wavelet coefficients, quantifies uncertainty, and applies the resulting framework to altered and new vehicle settings.
- 5.1. Estimates of δ∗,u∗: Only x5 and x6 have posteriors substantially different from their priors, while x5 concentrates at its allowed-range endpoint.This raises the possibility that x5 is functioning as a tuning parameter, motivating additional modularization.
- 5.2. Estimation of bias and reality: The estimated bias function differs significantly from zero, especially near 8.7 and 9.1.The estimate is accompanied by pointwise tolerance bounds derived from posterior quantiles.
- 5.3. Predicting a new run; same system, same vehicle: Bias-corrected predictions of reality are reported with associated uncertainty bands, alongside pure model predictions and field runs.Uncertainty for prediction is formed from posterior draws of reconstructed reality curves.
- 5.3. Predicting a new run; same system, same vehicle: Predictions for new field runs have wider uncertainty bounds because field-run variability is added to the prediction process.The wider bounds reflect prediction of a field run rather than reality alone.
- 5.4.1. Same vehicle, different load: For added mass, translating the original bias-corrected prediction by D(t) successfully extrapolated in critical Region 1, whereas the pure model prediction was unsuccessful.The comparison uses predictions and tolerance bands for both regions, together with the added-mass field run.
- 5.4.2. New vehicle of the same type: For a new vehicle of the same type, uncertainty increases because new manufacturing variations are drawn from their prior and predictions target field runs.The framework also carries validation assumptions to new settings; for Condition B, multiplicative and additive predictions were similar in one comparison, while multiplicative predictions were better for Site 2.
6. Site 2 analyses.
Site 2 analyses use the same validation procedure as Site 1 while revealing site-specific posterior behavior and improved bias-corrected predictions. Multiplicative predictions are noticeably better than additive predictions under Condition B.
- Site 2 analyses: Site 2 analyses proceed in the same way as the Site 1 analyses, but posterior distributions differ because each site has limited data.The calibration parameters and δ are shown in Figure 11.
- Site 2 analyses: The estimated bias function and tolerance bounds accompany bias-corrected and pure model predictions for Site 2.These are presented in Figures 12 and 13.
- Site 2 analyses: Bias-corrected prediction performs noticeably better than pure model prediction for Site 2.Figure 13 displays both predictions with their corresponding tolerance bounds.
APPENDIX A: DATA REGISTRATION AND WAVELET DECOMPOSITION
The appendix registers irregular field curves before wavelet decomposition so that shared basis elements capture common peaks and valleys. A thresholded union of retained elements is then used for both sites.
- Data registration: Field curves are aligned because their important peaks and valleys occur at different locations across tests.The curves are registered so the major features occur at common locations, while model runs are not aligned because their variation is not attributed to random fluctuation.
- Data registration: The model-run average serves as the reference curve for piecewise linear alignment of field curves.The appendix describes registration around major events and their locations on the time axis.
- Data registration: The registration uses a dyadic grid over the interval [0, 65], with linear interpolation supplying outputs between grid points.The test bed uses q = 12 and K = 65 computer runs.
- Data registration: The plotted time coordinate is a convenient scaling of realigned time rather than literal time.This convention is used for the registered Site 1 data.
- Wavelet decomposition: Wavelet coefficients are retained at levels 0–3 and selected at higher levels by an upper-2.5% magnitude threshold.The retained basis functions are unified across model-run and field curves for subsequent analyses.
- Wavelet decomposition: The combined analysis uses 289 retained wavelet basis elements from the two sites.Site 1 contributes 231 elements and Site 2 contributes 213 before forming the combined union.
APPENDIX B: THE MCMC ALGORITHM
The MCMC algorithm samples uncertain calibration, input, variance, and wavelet-coefficient quantities through Metropolis–Hastings and conditional normal draws. Mixed proposal distributions address both flat and concentrated posteriors.
- MCMC algorithm: The algorithm draws δ*, u*, τ², and related quantities from posterior distributions before sampling wavelet coefficients.Given these draws, it samples field and model wavelet coefficients from their specified distributions.
- Conditional sampling: Conditional coefficient updates are implemented as draws from normal distributions with specified means and variances.This applies separately to the field and model wavelet coefficients.
- MCMC algorithm: Metropolis–Hastings updates propose τ² locally and update δ and u using an acceptance ratio.Each proposal is accepted with probability min(1,ρ).
- Proposal design: A 50–50 mixture combines uniform-over-support proposals with local proposals for parameters having concentrated posteriors.The local proposals are centered on the previous value and use a maximum step of 0.05.
- MCMC algorithm: The chain cycles through the Metropolis–Hastings steps 200 times before saving values for later analysis steps.This reduces the effect of the typically high correlation between successive iterations.