Source-linked AI summary
A Kernel Test of Goodness of Fit
Kacper Chwialkowski, Heiko Strathmann, Arthur Gretton
TL;DR
The paper addresses goodness-of-fit testing when the target density is known only through its log gradient and target integrals are unavailable. It constructs an RKHS Stein discrepancy with a V-statistic and wild-bootstrap calibration for independent and dependent samples, then applies the test across sampling and density-estimation settings. The reported results establish bootstrap-based null calibration and near-certain rejection under alternatives, while experiments examine approximate MCMC and nonparametric density-estimation quality.
Problem
Goodness-of-fit testing is needed to determine whether samples match a target distribution when target-density integrals, including the normalisation constant, are unavailable.
Method
The paper defines an RKHS Stein discrepancy, estimates it with a closed-form V-statistic using the target log gradient and kernel, and calibrates tests with wild bootstrap for independent and dependent samples.
Results
Under the null, nB_n approximates nV_n for bootstrap quantiles, while under the alternative V_n dominates B_n, resulting in almost sure rejection; experiments also study random-feature model fit.
Takeaways & Limitations
The test provides a practical way to assess target fit and sample quality in approximate MCMC and nonparametric density estimation without target-distribution integrals.
Takeaways & Limitations
The approach assumes τ-mixing dependence and technical conditions; earlier competing Stein metrics also require complex constructions and nontrivial lower bounds outside some settings.
Abstract
from arXiv · showhide
We propose a nonparametric statistical test for goodness-of-fit: given a set of samples, the test determines how likely it is that these were generated from a target density function. The measure of goodness-of-fit is a divergence constructed via Stein's method using functions from a Reproducing Kernel Hilbert Space. Our test statistic is based on an empirical estimate of this divergence, taking the form of a V-statistic in terms of the log gradients of the target density and the kernel. We derive a statistical test, both for i.i.d. and non-i.i.d. samples, where we estimate the null distribution quantiles using a wild bootstrap procedure. We apply our test to quantifying convergence of approximate Markov Chain Monte Carlo methods, statistical model criticism, and evaluating quality of fit vs model complexity in nonparametric density estimation.
1. Introduction
The paper develops an RKHS-based Stein goodness-of-fit test that avoids target-distribution integrals and supports both independent and dependent samples. It addresses practical testing and sample-quality assessment, including approximate MCMC, while building on limitations of earlier approaches.
- The paper targets testing whether samples from q match a reference density p known only up to its normalisation constant.
- Earlier Stein discrepancies can be difficult to apply because they require complex function classes, linear programs, and nontrivial lower bounds.
- Existing Sobolev Stein discrepancies pose an unresolved testing problem because their asymptotic behaviour does not readily yield p-values or user-specified acceptance thresholds.
- The proposed goodness-of-fit test uses a Stein discrepancy over an RKHS function class and estimates it with a closed-form, quadratic-time V-statistic.
- Because the Stein operator requires only the gradient of the log target density, the method avoids target-density integrals, including the normalisation constant.
- The paper formulates tests for uncorrelated and correlated samples and applies them to approximate MCMC, statistical model criticism, and nonparametric density estimation.
2. Test Definition: Statistic and Threshold
The test measures goodness of fit through an RKHS Stein discrepancy and estimates its null threshold with a wild bootstrap, supporting both independent and correlated samples.
- Stein discrepancy: The method applies a Stein operator to an RKHS function class whose target expectation is zero, defining the goodness-of-fit discrepancy.The operator uses target-density derivatives and is constructed so the discrepancy compares empirical expectations with zero target expectations.
- Stein discrepancy: The RKHS formulation computes the discrepancy from observed samples without requiring samples from, or integrals under, the target distribution.The representation uses the norm of an empirical expectation involving ξp, avoiding direct access to target samples.
- Statistic: The squared discrepancy has a closed-form expression in terms of hp, provided Ehp(Z, Z) < ∞.This connects the RKHS norm to a kernel-based pairwise quantity and enables direct estimation from samples.
- Statistic: A quadratic-time V-statistic estimates the squared discrepancy from the observed samples.The estimator is formed from sample pairs and has the same basic V-statistic structure as earlier kernel tests.
- Threshold: Wild bootstrap samples estimate null-distribution quantiles and account for dependence through an auxiliary sign-changing Markov chain.For i.i.d. data, the sign-change probability is set to an = 0.5.
- Threshold: Under the null, the bootstrapped statistic nBn approximates nVn, while under the alternative Vn dominates Bn and yields almost sure rejection.The test rejects when Vn exceeds the empirical 1 − α bootstrap quantile.
3. Proofs of the Main Results
The proofs establish that the RKHS Stein discrepancy is well defined and discriminates probability measures, while the V-statistic and wild bootstrap have valid null and alternative behavior under dependence assumptions.
- The Stein operator has zero expectation under the target measure, providing the foundation for the discrepancy construction.
- The RKHS feature ξp is well defined under integrability conditions, and its inner product yields the kernelized Stein quantity hp.
- With a C0-universal kernel, the discrepancy is zero only when the observed and target distributions coincide.
- For τ-mixing observations, Lipschitz continuity and finite-moment conditions imply that nVn converges weakly under the null hypothesis.
- Under the null, bootstrap and statistic quantiles converge, whereas under the alternative Bn tends to zero and Vn tends to a positive constant.
- Excessive wild-bootstrap dependence can make Bn too low, causing Vn > Bn frequently and producing an overly conservative test.
4. Experiments
Experiments assess calibration and power under temporal dependence, compare the test with alternatives, and apply it to GP model criticism, approximate MCMC, and nonparametric density estimation.
- Student’s t vs. normal: Temporal correlation makes calibration and power sensitive to the wild-bootstrap parameter, while thinning plus adjustment preserves both more effectively.For correlated samples, the recommended procedure thins by 20, sets a_n = 0.1, and uses at least max(500k, d100) data points.
- Comparing to a parametric test in increasing dimensions: The Baringhaus–Henze and Stein-based tests are compared by test power across increasing sample sizes.The experiment uses standard-normal null samples and Student’s t alternatives across dimensions.
- Statistical model criticism on Gaussian processes: The GP model-criticism experiment finds the fitted Gaussian noise model unlikely to explain held-out data, even with n = 41 test points.The test statistic falls in an upper quantile of the bootstrapped null distribution.
- Bias quantification in approximate MCMC: For austerity MCMC, the test identifies ϵ = 0.4 as a good approximation of the true stationary distribution while remaining parsimonious in likelihood evaluations.Larger ϵ lowers expected computational cost because the algorithm tolerates a greater probability of an incorrect acceptance decision.
- Convergence in non-parametric density estimation: In nonparametric density estimation, the test tracks fit as data or random-feature complexity increases and detects over-smoothing in the finite-feature model.The p-value distribution is uniform at N = 5000 for the infinite-dimensional model, while over-smoothing remains detectable in the random-feature approximation.
5. Proofs
The proofs establish the Stein operator’s null expectation and the statistic’s positive-definiteness and asymptotic behavior under the null and alternative, under kernel, integrability, and mixing assumptions.
- Under the target density, the Stein operator has expectation zero for every function in the considered function class.
- Integration by parts yields the zero-expectation result when p(x)f_i(x) vanishes at infinity and Fubini–Tonelli integrability holds.
- The asymptotic and bootstrap results require boundedness, Lipschitz continuity, moment, stationarity, and mixing-related assumptions.
- The kernel-based statistic is positive definite and degenerate under the null because the Stein feature expectation is zero.
- Under the alternative hypothesis, V_n converges to a positive constant when the kernel’s zero comportment is positive.
6. MCMC convergence testing
This section connects τ-mixing to established dependence conditions and explains how such conditions support convergence testing for dependent Markov-chain samples.
- Strong mixing coefficients: τ-mixing is useful for dependent samples because many models, including Markov chains, are known to be strongly mixing.
- Mixing relationships: Inequalities relating τ-mixing to strong-mixing coefficients allow classical Markov-chain results to establish τ-mixing rates.
- Moment bounds: Moment conditions simplify tail-function bounds through Markov’s inequality, yielding an upper bound for the generalized inverse tail function.
- τ-mixing: Causal functions of stationary sequences, iterated random functions, Markov chains, and expanding maps can be τ-mixing under suitable assumptions.
- Markov chains: Harris recurrent and aperiodic Markov chains satisfy absolute regularity, while geometric ergodicity implies geometric decay of dependence coefficients.