Source-linked AI summary
Robust Bayesian Inference for Unnormalized Models with Mixed-Domain Data
Jiongran Wang, Debdeep Pati, Anirban Bhattacharya
TL;DR
Doubly-intractable models hinder Bayesian inference, and misspecification can leave uncertainty poorly calibrated. The paper proposes SME-BETEL, combining score matching estimating equations with Bayesian ETEL and extending it to mixed-domain data. Its theory establishes calibrated credible sets under misspecification, while simulations and ozone data demonstrate the construction’s utility.
Problem
Parameter-dependent normalizing constants obstruct standard Bayesian inference, while existing uncertainty quantification may be poorly calibrated under model misspecification.
Method
SME-BETEL combines score matching estimating equations with Bayesian ETEL and develops a boundary-adapted score matching criterion for mixed-domain data.
Results
SME-BETEL credible sets are asymptotically calibrated to the sampling variability of the score matching estimator, with valid frequentist coverage under model misspecification.
Takeaways & Limitations
The framework provides a calibration-free route to robust Bayesian inference for unnormalized models, including mixed-domain spatial preferential models without evaluating spatial normalizing integrals.
Takeaways & Limitations
Future work must better characterize the robustness–efficiency tradeoff and develop score matching objectives for more complex mixed-domain supports.
Abstract
from arXiv · showhide
Many statistical models involve parameter-dependent normalizing constants that are computationally intractable, creating substantial obstacles to standard Bayesian inference. Although existing likelihood-based algorithms can often circumvent these constants, their uncertainty quantification may be poorly calibrated under model misspecification. To address these challenges, we propose SME-BETEL, a semiparametric Bayesian framework that combines score matching estimating equations with Bayesian exponentially tilted empirical likelihood. The resulting posterior avoids evaluation of normalizing constants and does not require learning-rate calibration. We establish consistency and asymptotic normality of the score matching estimator, and prove a Bernstein-von Mises theorem for the SME-BETEL posterior. These results show that SME-BETEL credible sets are asymptotically calibrated to the sampling variability of the score matching estimator, yielding valid frequentist coverage under model misspecification. We further develop a new score matching criterion for mixed-domain data, extending SME-BETEL to models whose observations combine components from different sample spaces. This construction enables robust Bayesian inference for mixed-domain doubly-intractable models, including preferential sampling models with an intractable spatial normalizing integral. Simulation studies show that SME-BETEL remains competitive under correct specification and substantially improves uncertainty quantification under misspecification. An ozone-monitoring application further demonstrates the utility of the mixed-domain construction for spatial preferential modeling.
1 Introduction
Doubly-intractable models make standard Bayesian inference difficult because parameter-dependent normalizing constants are intractable, while misspecification can undermine uncertainty calibration. SME-BETEL combines score matching with Bayesian ETEL to avoid these constants and provide robust inference, including for mixed-domain data.
- Doubly-intractable models require inference with parameter-dependent normalizing constants that make posterior odds computationally intractable.
- Existing methods can circumvent normalizing constants, but Bayesian uncertainty quantification based on a fully specified working model may be poorly calibrated under misspecification.
- Score matching bypasses normalizing constants, but generalized Bayesian procedures based on score-type losses typically require learning-rate tuning that affects uncertainty quantification.
- SME-BETEL uses score matching first-order conditions as Bayesian ETEL moment conditions, avoiding normalizing constants and learning-rate calibration while targeting calibrated posterior uncertainty.
- The mixed-domain extension combines Euclidean responses with bounded-domain locations and uses a boundary-adapted score matching objective for doubly-intractable spatial models.
- The framework is demonstrated through preferential sampling simulation and eastern United States ozone-monitoring data, with implementation in moderately high-dimensional spatial models.
2 SME-BETEL: Bayesian ETEL Based on Score Matching First-Order Conditions
SME-BETEL turns score matching criteria into Bayesian ETEL moment conditions, avoiding intractable normalizing constants. Its modular construction extends to mixed-domain observations by adapting score matching near bounded-domain boundaries.
- 2.1 Background: Score matching estimates parameters in unnormalized models using the data gradient of the log density, which does not depend on the normalizing constant.
- 2.1 Background: The sample score matching estimator minimizes the empirical score matching criterion formed from score components and their data-coordinate derivatives.
- 2.2 Methodology: SME-BETEL uses the first-order optimality conditions of the score matching objective as moment conditions in Bayesian ETEL for robust doubly-intractable inference.
- 2.2 Methodology: Bayesian ETEL constructs a surrogate likelihood from probability weights satisfying the sample moment condition and closest to uniform empirical weights in KL divergence.
- 2.2 Methodology: The standard construction requires the origin to lie in the interior of the moment vectors’ convex hull; otherwise proposed parameters are discarded during MH sampling.
- 2.3 Support-Adapted SME-BETEL: For mixed-domain observations, the method factors the model into conditional Euclidean and marginal bounded-domain components, then tapers score matching near the boundary.
- 2.3 Support-Adapted SME-BETEL: The computable mixed-domain criterion removes the unknown data score and avoids evaluating the spatial normalizing constant in preferential sampling models.
3 Simulation Study
The simulation evaluates SME-BETEL for misspecified preferential sampling with response-location pairs on mixed domains. SME-BETEL consistently outperforms the PRD method in recovering the latent spatial surface while avoiding the spatial normalizing integral.
- 3 Simulation Study: Preferential sampling jointly models responses y ∈ R and locations s ∈ D, where locations may depend on the underlying spatial process.The sampling density can involve an intractable spatial normalizing integral, making this a mixed-domain doubly-intractable problem.
- 3.1 Misspecified Preferential Sampling Model: The working model uses a response-location factorization with spatial surfaces represented through low-rank kernel convolution approximations.The simulation estimates regression, association, variance, spatial-range, and kernel-weight parameters from the approximation.
- 3.1 Misspecified Preferential Sampling Model: Misspecification arises because the data-generating response distribution has heavy-tailed noncentral Student t errors, whereas the working model assumes Gaussian errors.The PRD comparator also requires numerical approximation of the location-density normalizing integral.
- 3.1 Misspecified Preferential Sampling Model: The study uses 50 replications with sample sizes n = 250 and n = 400, evaluating predicted-surface bias, MSE, and MAD over spatial grids.Each method uses posterior sampling, and the reported metrics are averaged over grid locations.
- 3.1 Misspecified Preferential Sampling Model: Bias, MSE, and MAD are lower for SME-BETEL than for PRD at both sample sizes, with MSE and MAD decreasing as n increases.The authors describe these results as more accurate latent-surface recovery and practical robustness under misspecification.
4 Theoretical Properties of SME-BETEL
The theoretical analysis establishes asymptotic properties for SME-BETEL under regularity conditions, including a Bernstein–von Mises approximation centered at the score matching estimator. Its posterior variance matches the estimator’s sampling variance, supporting frequentist-valid credible sets under misspecification.
- Theoretical setup: The theory studies the original score matching objective for i.i.d. observations on R^m and analyzes the SME-BETEL posterior through a Bernstein–von Mises theorem.The framework builds on Bayesian ETEL asymptotics.
- Assumptions: The assumptions require a positive continuous prior density at θ∗, a unique interior moment-condition solution on compact Θ, smooth integrable estimating functions, and nonsingular identification matrices.These conditions are supplemented by a tail condition controlling the log-ETEL criterion away from θ∗.
- Theoretical results: The SME-BETEL posterior converges in total variation to a Gaussian distribution centered at the sample score matching estimator bθ_n with variance n^-1(Γ^⊤Δ^-1Γ)^-1.The estimator is defined as the minimizer of the sample score matching criterion.
- Theoretical results: The limiting posterior variance matches the asymptotic sampling variance of the score matching estimator under both correct specification and misspecification.Therefore, SME-BETEL credible sets are asymptotically valid frequentist confidence sets for θ∗.
- Misspecification: In a misspecified location model, the usual Bayesian 95% credible interval has about 56% asymptotic coverage, whereas SME-BETEL retains asymptotically valid 95% coverage.In 500 simulations with n = 50, empirical coverage was 0.55 for usual Bayes and 0.92 for SME-BETEL.
- Correct specification: Under correct specification, both usual Bayesian and SME-BETEL credible intervals are asymptotically valid, while SME-BETEL can have larger posterior variance and comparable estimation accuracy.The larger variance is attributed to score matching ignoring information in the normalizing constant.
5 Eastern United States Ozone Data
The ozone analysis applies SME-BETEL to a preferential-sampling geostatistical model for 631 monitoring locations and compares it with noninformative sampling. SME-BETEL indicates informative sampling and generally predicts higher ozone levels than the noninformative model, especially in northern New England.
- The analysis re-analyzes 631 monitoring locations with a preferential-sampling geostatistical model for eastern United States ozone data.The model uses rescaled spatial coordinates and a 106-dimensional parameter vector.
- SME-BETEL uses the mixed-domain score matching method without requiring explicit evaluation of the sampling-density normalizing constant.
- The 95% credible interval for the preferential-sampling parameter a excludes zero, indicating evidence of informative sampling.
- SME-BETEL generally predicts higher ozone levels than NIS, with the largest discrepancies in northern New England around Vermont, New Hampshire, and Maine.The right panel displays predicted values under NIS minus those under SME-BETEL; the difference is negative across most of the spatial domain.
6 Conclusion
The paper proposes SME-BETEL as a modular semiparametric Bayesian framework combining score matching estimating equations with Bayesian ETEL for doubly-intractable models. Its theory provides calibrated uncertainty under misspecification, while its mixed-domain extension broadens the supported data structures.
- SME-BETEL combines score matching estimating equations with Bayesian ETEL to avoid parameter-dependent normalizing constants while retaining calibrated posterior uncertainty.
- The framework is modular: deriving suitable score matching objectives and moment functions allows Bayesian ETEL to be applied across model classes and data supports.
- The mixed-domain score matching objective extends SME-BETEL to observations whose components lie on different supports.
- Under suitable regularity conditions, the score matching estimator is consistent and asymptotically normal, and the SME-BETEL posterior satisfies a Bernstein–von Mises theorem.
- SME-BETEL provides calibrated uncertainty quantification under model misspecification, while future work includes studying robustness–efficiency tradeoffs and constructing objectives for complex supports.
Appendix to “Robust Bayesian Inference for Unnormalized Models with Mixed-Domain Data”
The supplementary appendix contains theoretical proofs, supplementary simulations, model-specific formulas, and computational details.
- Section S1 contains proofs for mixed-domain score matching, estimator asymptotics, and the SME-BETEL Bernstein–von Mises theorem.
- Section S2 presents supplementary simulation studies and numerical details.
- Section S3 collects model-specific formulas and computational details, while Section S4 presents the Metropolis–Hastings algorithm used.
S1.1 Proof of Proposition 1
The mixed-domain score matching derivation handles Euclidean responses and bounded-domain locations within one objective. A boundary-vanishing weight removes boundary contributions, producing a tractable criterion that excludes the unknown data score and spatial normalizing integral.
- Mixed-domain observations x=(y,s) combine an unconstrained Euclidean response y with a location s on compact domain D.
- A smooth weight function h(s) that vanishes on the boundary of D removes the boundary contribution during integration by parts.
- The derivation factorizes the joint data and working-model densities into conditional response and marginal location components.
- The resulting objective depends only on derivatives of the working model and excludes the unknown data-generating score.
- When the location density is proportional to exp{ξ(s)}, derivatives with respect to s depend only on ξ(s;θ), so the intractable spatial normalizing integral does not appear.
S1.2 Consistency of Sample Score Matching Estimator
Under identification, regularity, positivity, boundary, and Lipschitz conditions, the sample score matching estimator is consistent for the unique population score matching target.
- The population criterion J(θ) has a unique minimizer θ∗, identifying the score matching target.
- Finite score-function moments, vanishing boundary terms, positive densities, and Lipschitz conditions provide the regularity framework for consistency.The Lipschitz condition controls parameter variation through square-integrable envelopes.
- The proof establishes uniform convergence of the sample score matching objective to its population counterpart over compact Θ.The argument uses symmetrization, Rademacher processes, covering numbers, and Dudley’s entropy integral.
- Under Assumptions S1–S5, i.i.d. observations, compact Θ, and a finite second moment, bθn converges in probability to θ∗.The target θ∗ is the unique minimizer of J(θ).
- Continuity and unique minimization convert convergence of objective values into convergence of the estimators.The proof separates the objective difference into empirical-process terms and uses compactness away from θ∗.
S1.3 Asymptotic normality
The score matching estimator is asymptotically normal under the consistency conditions and additional regularity assumptions, with a sandwich covariance determined by the estimating equations.
- Under Theorem S1’s conditions and main-text Assumptions 2–4, the score matching estimator has an asymptotically normal distribution.The proof uses the empirical first-order condition, Taylor expansion, a uniform law of large numbers, and the multivariate central limit theorem.
- The limiting covariance is (Γ⊤∆−1Γ)−1, the inverse Godambe information for the score matching estimating equations.∆ and Γ are defined through the estimating-equation variance and derivative conditions.
- Nonsingularity of ∆ and Γ ensures the covariance representation is well defined.
S1.4 Proof of Theorem 1
Theorem 1 is proved by locally expanding the SME-BETEL posterior around the score matching estimator, controlling the posterior mass in three regions, and obtaining the stated asymptotic limit.
- The proof follows a Bernstein–von Mises argument for Bayesian ETEL, adapted to a posterior defined directly on Θ.The moment function is generated from the score matching objective’s first-order condition.
- A localization result places θt = bθn + th/√n in a neighborhood where the local ETEL expansions are valid.The analysis uses h = √n(θ − bθn) and a logarithmic local radius.
- The local posterior expansion combines first- and second-order ETEL derivatives, with normalized Hessian leading term V = Γ⊤∆−1Γ.Multiplier and remainder terms are shown to be asymptotically negligible.
- The proof partitions the integration domain into A1n, A2n, and A3n to control local, intermediate, and tail contributions.The intermediate-region contribution is op(1), while the local region is controlled by a Gaussian envelope.
- Combining the regional bounds establishes convergence of the normalized posterior and completes Theorem 1.The limiting normalizing constant is strictly positive under the stated assumptions.
S2 Supplementary Simulation Studies
Supplementary simulations evaluate SME-BETEL in misspecified and correctly specified models, including autonormal, truncated normal, Bingham, and mixed-domain settings.
- Misspecified models: Under misspecified autonormal models, SME-BETEL maintains coverage close to nominal across field sizes, unlike Standard Bayes and DMH.Coverage declines substantially for the likelihood-based methods, especially for αh; αh is affected by the restriction βh = βv.
- Misspecified models: Under bounded-support misspecification, SME-BETEL achieves coverage close to nominal 95% across cases and sample sizes, while DMH consistently undercovers.The study uses a bounded-domain score matching objective.
- Correctly specified models: Under correct autonormal specification, all methods achieve coverage near nominal 95%, while Standard Bayes and DMH have slightly smaller MAD and MSE.SME-BETEL remains competitive but is slightly more dispersed because score matching omits normalizing-constant information.
- Correctly specified models: Correctly specified Bingham and exponential graphical-model simulations show performance patterns similar to the autonormal study.
- Support adaptations: For nonnegative data, log-transformed score matching and the alternative nonnegative-data approach produce similar bias, MSE, and MAD.
- Support adaptations: The mixed-domain construction uses a smooth cubic weight that equals one away from the boundary and tapers to zero within distance ϵ.The weight function is required to vanish on the compact-domain boundary.