Source-linked AI summary

Model Selection Principles in Misspecified Models

Jinchi Lv, Jun S. Liu

arXiv:1005.5483v2math.STstat.ME

TL;DR

Model selection in misspecified models lacks criteria that directly account for the discrepancy between the working and true models. The paper derives generalized Bayesian and KL-based criteria for misspecified GLMs, including a criterion that combines quasi-likelihood, dimensionality, and misspecification penalties. Numerical studies report advantages over classical criteria in correctly specified and misspecified settings, while accurate estimation of covariance-contrast quantities remains a limitation.

  • Problem

    Classical AIC and BIC do not explicitly account for model misspecification, although misspecification is common and evaluating selection then requires comparing models by discrepancy from the true distribution.

  • Method

    The paper derives asymptotic expansions of Bayesian and KL-divergence principles in misspecified GLMs and proposes GBIC, GAIC, and GBICp criteria.

  • Results

    Numerical studies demonstrate advantages of the new criteria for model selection in both correctly specified and misspecified models.

  • Takeaways & Limitations

    GBICp decomposes into negative maximum quasi-log-likelihood plus separate penalties for model dimensionality and model misspecification.

  • Takeaways & Limitations

    Estimating the trace and log-determinant of the covariance contrast matrix is nontrivial, especially in high dimensions, and more accurate estimation requires further theoretical study.

Abstract

from arXiv · show

Model selection is of fundamental importance to high dimensional modeling featured in many contemporary applications. Classical principles of model selection include the Kullback-Leibler divergence principle and the Bayesian principle, which lead to the Akaike information criterion and Bayesian information criterion when models are correctly specified. Yet model misspecification is unavoidable when we have no knowledge of the true model or when we have the correct family of distributions but miss some true predictor. In this paper, we propose a family of semi-Bayesian principles for model selection in misspecified models, which combine the strengths of the two well-known principles. We derive asymptotic expansions of the semi-Bayesian principles in misspecified generalized linear models, which give the new semi-Bayesian information criteria (SIC). A specific form of SIC admits a natural decomposition into the negative maximum quasi-log-likelihood, a penalty on model dimensionality, and a penalty on model misspecification directly. Numerical studies demonstrate the advantage of the newly proposed SIC methodology for model selection in both correctly specified and misspecified models.

1. Introduction

High-dimensional model selection seeks sparse, interpretable models, but classical AIC and BIC do not explicitly account for model misspecification. The paper develops selection criteria for misspecified generalized linear models by expanding Bayesian and KL-divergence principles.

  • When predictors are numerous relative to observations, sparse models can improve prediction accuracy and model interpretability.
  • A fundamental problem in high-dimensional modeling is comparing models that contain different sets of predictors.
  • AIC and BIC arise from KL-divergence and Bayesian principles, respectively, when models are correctly specified.
  • Parametric models are often misspecified, yet neither AIC nor BIC explicitly accounts for model misspecification.
  • The paper derives asymptotic expansions of Bayesian and KL-divergence principles in misspecified GLMs, leading to generalized information criteria.

2. Model misspecification and model selection principles

The section formulates misspecified GLMs through a best approximating model and develops Bayesian and KL-divergence model-selection principles around the QMLE. Their classical correctly specified forms motivate criteria that compare fitted models using likelihood and complexity-related terms.

  • Misspecified GLMs: For a candidate predictor subset M, the working GLM may not contain the unknown true distribution, so its parameter is defined through a best KL-divergence approximation.
  • Misspecified GLMs: The QMLE estimates the parameter of the best approximating model and is consistent and asymptotically normal under regularity conditions.
  • Misspecified GLMs: The covariance matrices A_n and B_n represent working-model and true-model covariance structures, and coincide when the model is correctly specified.
  • Bayesian principle: The Bayesian principle selects among candidate models using posterior probabilities based on model priors and parameter priors.
  • Bayesian principle: With equal model priors, the Bayesian procedure is equivalent to minimizing KL divergence from the marginal response distribution to the true distribution.
  • KL-divergence principle: The KL-divergence principle selects the fitted model with the smallest divergence from the true model, while classical asymptotic expansions yield likelihood criteria with dimensionality penalties.

3. Asymptotic expansions of model selection principles in misspecified models

The paper derives asymptotic expansions for Bayesian and KL-based principles in misspecified GLMs, yielding GBIC, GAIC, and GBICp criteria that explicitly incorporate misspecification. GBICp decomposes selection into fit, complexity, and nonnegative misspecification penalties.

  • Asymptotic expansions: The authors derive rigorous asymptotic expansions of Bayesian and KL divergence principles for misspecified generalized linear models.The Bayesian expansion uses Laplace approximation, while the KL expansion addresses misspecified GLMs with deterministic design matrices.
  • Generalized BIC: GBIC extends BIC by adding the covariance-contrast term −log |bH_n| to the likelihood and dimensionality penalty.The added term explicitly reflects model misspecification, although it is not necessarily nonnegative.
  • GBICp: GBICp uses KL-informed model priors and adds tr(bH_n)−log |bH_n| to the negative maximum quasi-log-likelihood and complexity penalty.Its criterion is −2ℓ_n(y,bβ_n)+(log n)|M|+tr(bH_n)−log |bH_n|.
  • GBICp: GBICp decomposes into goodness of fit, model complexity, and model misspecification, with both penalty terms nonnegative.The misspecification component is motivated by the KL divergence between normal distributions with covariance matrices A_n and B_n.
  • Generalized AIC: GAIC incorporates model complexity and misspecification through a single penalty term based on the trace of the covariance contrast matrix.Under correct specification, this trace is expected to approximate the model dimension, paralleling AIC’s complexity penalty.
  • Covariance contrast estimation: Estimating the trace and log-determinant of the covariance contrast matrix is nontrivial, especially in high dimensions, so the paper studies two specific estimators.The paper considers simple and bootstrap estimators and notes that more accurate high-dimensional estimation requires further theory.

4. Technical conditions and results

The paper imposes regularity conditions supporting identification, consistency, asymptotic normality, and asymptotic expansions for QMLE in misspecified GLMs. Under these conditions, the best misspecified model is uniquely identified, QMLE is consistent, and the fixed-dimensional asymptotic theory applies, while diverging dimensionality is outside the paper’s focus.

  • Regularity conditions: Condition 1 ensures strict concavity and a unique best misspecified GLM under KL divergence, while requiring a twice-differentiable convex cumulant function and full-rank design.These requirements support parameter identifiability and uniqueness of the QMLE.
  • Regularity conditions: Condition 2 supports QMLE consistency through diverging design information and covariance regularity in shrinking neighborhoods.When response variances remain bounded, the eigenvalue requirement reduces to λmin(XT X) →∞, the usual linear-regression design condition.
  • Regularity conditions: Conditions 3 and 4 support QMLE asymptotic normality through continuity of matrix-valued functions and a Lyapunov-type moment condition.The conditions control local covariance behavior and higher-order moments needed for the asymptotic-normality proof.
  • Regularity conditions: Conditions 5 and 6 provide the prior-density and regularity assumptions required for asymptotic expansions of Bayes factors, KL principles, and log-prior probabilities.The prior density must be locally bounded away from zero near the pseudo-true parameter, while additional conditions control the KL expansion.
  • Asymptotic results: Under the stated conditions, the KL divergence has a unique minimizer βn,0 and the QMLE satisfies bβn − βn,0 = oP(1).Thus the estimator consistently targets the theoretically closest model in the misspecified family.
  • Scope boundary: The asymptotic-normality presentation assumes fixed dimensionality d; extending it to diverging d requires more delicate analysis and is not the paper’s focus.In correctly specified models, An = Bn and the results reduce to conventional MLE asymptotic theory.

5. Numerical examples

Numerical studies compare GAIC, GBIC, and GBICp with AIC and BIC across correctly specified and misspecified linear and nonlinear settings. The generalized criteria generally improve prediction and variable selection, with GBICp often performing best.

  • Linear models: GBIC and GBICp outperformed other information criteria across correctly specified and high-dimensional sparse linear settings, especially with small samples.The studies included moderate-size correctly specified models and high-dimensional sparse models with interaction terms.
  • Best subset linear regression: GBICp outperformed all other criteria in selecting the true sparse model across the tested sample sizes.AIC and BIC tended to select larger models, while GAIC and GBIC improved over their classical counterparts.
  • High-dimensional sparse linear regression with interaction: In misspecified high-dimensional regression, GBICp selected models that closely mimicked the oracle estimate, while prediction and selection performance deteriorated as dimensionality increased.GAIC improved over AIC, and GBIC and GBICp improved over BIC on prediction and variable-selection measures.
  • Nonlinear models: GBIC and GBICp outperformed AIC and BIC in polynomial, single-index, and high-dimensional logistic regression settings.These experiments evaluated nonlinear models with and without interaction effects.
  • Polynomial regression with heteroscedasticity: GAIC, GBIC, and GBICp improved over AIC and BIC in heteroscedastic polynomial regression by capturing differences between estimated error variance and apparent residual variance.The criteria were evaluated when fitting constant-variance polynomial models to data generated with heteroscedastic variance.
  • High-dimensional logistic regression with interaction: GAIC, GBIC, and GBICp improved over AIC and BIC for nonlinear logistic regression in prediction and variable selection.The comparison used prediction error, false positives, false negatives, selection probability, and inclusion probability.

6. Discussions

The paper develops GBIC and GAIC from asymptotic expansions of Bayesian and Kullback–Leibler principles in misspecified GLMs, including GBICp with an explicit misspecification penalty. Numerical studies support the methods in correctly specified and misspecified settings, while the authors identify estimation and dimensionality extensions as open challenges.

  • GBIC and GAIC arise from asymptotic expansions of Bayesian and Kullback–Leibler model-selection principles in misspecified generalized linear models.
  • GBICp decomposes into negative maximum quasi-log-likelihood, a model-dimensionality penalty, and a model-misspecification penalty.
  • Numerical studies demonstrate advantages for the proposed methods in both correctly specified and misspecified models.
  • The misspecified-model framework allows sample-size-varying priors with a local density constraint around the theoretically best model.
  • Estimating the trace and log-determinant of the covariance contrast matrix may require better methods and rigorous theory, especially in high dimensions.
  • The asymptotic expansions use fixed dimensionality, leaving extensions to high-dimensional and more general model settings as future work.

A. Proofs

The technical report contains omitted proofs for Theorem 5 and Propositions 1 and 2.

  • Proofs of Theorem 5 and Propositions 1 and 2 are omitted from the paper and provided in a technical report.
  • The omitted results include one theorem and two propositions from the paper’s technical development.
  • The cited technical report is Lv and Liu (2010).

A.1. Proof of Theorem 1

The proof of Theorem 1 derives an asymptotic expansion by controlling a localized quasi-log-likelihood neighborhood and showing remainder terms vanish rapidly.

  • The proof conditions on the event that the QMLE lies in a shrinking neighborhood around the target parameter.
  • The quasi-log-likelihood is smooth and concave, with its maximum attained at the QMLE, enabling a local Taylor expansion.
  • The neighborhood is convex and contains the QMLE, so intermediate points in the Taylor expansion remain within a controlled larger neighborhood.
  • Regularity conditions bound the smallest eigenvalue and control the local remainder term used in the expansion.
  • Several auxiliary lemmas establish that tail and neighborhood-integral terms converge faster than any polynomial rate in n.
  • Combining the bounds with the auxiliary lemmas yields the expansion, with the final step completed on the conditioning event.

A.2. Proof of Theorem 2

The proof of Theorem 2 combines an earlier expansion with the covariance-trace term to obtain the desired asymptotic expansion of the Bayes factor.

  • The proof invokes the first part of the proof of Theorem 3 to establish the needed intermediate relation.
  • The resulting expression includes a covariance-contrast contribution proportional to tr(H_n).
  • Together with Theorem 1, these results yield the desired asymptotic expansion of the Bayes factor.

A.3. Proof of Theorem 3

The proof establishes the theorem through second-order Taylor expansions of the relevant quasi-log-likelihood functions, controlling remainder terms with moment, Lipschitz, and asymptotic normality conditions.

  • Proof strategy: Second-order Taylor expansions around 0 and the quasi-likelihood maximizer provide the proof’s main approximation steps.The expansions retain Lagrange remainder terms and use the maximizer βn,0 or bβn as the expansion reference.
  • Proof strategy: The normalized estimator difference is represented using v = n−1/2Cn(bβn −βn,0), with intermediate points on the segment joining v and 0.This representation supports the remainder control in the expansion around βn,0.
  • Remainder control: The difference between ηn(β) and ℓn(y, β) is linear in β, so their second- and higher-order partial derivatives agree.This lets the proof transfer the second-order expansion between the two functions around bβn.
  • Remainder control: Asymptotic normality and bounded higher moments justify the required Taylor-expansion approximations.The proof invokes D −→N(0, Id), bounded 3+α moments, and moment relations involving the normalized components.

A.4. Proof of Theorem 4

The proof shows consistency of the estimated matrix quantities by comparing their normalized forms with the corresponding population matrices and controlling local deviations around βn,0.

  • Matrix-estimator consistency: The proof relies on bounded eigenvalues of n−1Bn and Lipschitz continuity of n−1An(β) in an asymptotically shrinking neighborhood.These conditions control matrix behavior near βn,0 and support replacing population quantities by evaluations at bβn.
  • Matrix-estimator consistency: n−1 bAn = n−1Bn + oP(1) and n−1 bBn = n−1Bn + oP(1), establishing consistency of both normalized matrix estimators.The two approximations are combined to support consistency of the estimator bHn.
  • Matrix-estimator consistency: The estimator bβn is shown to satisfy bβn −βn,0 = OP(n−1/2), enabling the local matrix approximations.This rate is combined with eigenvalue bounds and Lipschitz conditions in the consistency argument.
  • Matrix-estimator consistency: The proof decomposes n−1 bBn into three terms and shows the additional terms vanish in probability.Uniform boundedness of Yi −EYi and Lipschitz control of the mean-function terms yield G2 = oP(1) and G3 = oP(1).

A.5. Proof of Theorem 6

The proof establishes consistency of bβn by showing that the quasi-log-likelihood is maximized inside a shrinking neighborhood of βn,0 with probability tending to one.

  • Neighborhood argument: The event Qn ensures that the local maximum lies in the interior of Nn(δ), where strict concavity identifies it with bβn.Qn compares the likelihood at βn,0 with its maximum on the neighborhood boundary.
  • Neighborhood argument: Taylor’s theorem and a transformation to the boundary of Nn(δ) provide a lower bound on P(Qn).The proof uses the convexity of Nn(δ), Markov’s inequality, and a suitable choice of δ to control the boundary likelihood.
  • Consistency conclusion: bβn −βn,0 = oP(1), so the quasi-likelihood estimator is consistent for the target parameter.This follows from Qn ⊂ {bβn ∈Nn(δ)} together with the probability control established for Qn.

A.6. Proof of Theorem 7

The proof derives asymptotic normality for the quasi-maximum-likelihood estimator and gives the estimated matrix formulas for common generalized linear models, including a linear-regression misspecification consequence.

  • Asymptotic normality: The normalized estimator difference is decomposed as Cn(bβn −βn,0) = un + wn, with wn = oP(1).Slutsky’s lemma then transfers the limiting distribution from un to the estimator difference.
  • Asymptotic normality: D −→N(0,1) for every unit-vector linear combination implies the stated multivariate asymptotic normality.The argument analyzes vn = aT un for arbitrary unit vectors a and applies Lyapunov’s theorem to independent centered terms.
  • GLM formulas: For linear regression, maximizing the quasi-log-likelihood yields bγ = (XT X)−1Xy, with error variance estimated from the residual sum of squares.The resulting estimator is bβn = bγ/bσ2.
  • GLM formulas: The estimated matrices for linear regression are bAn = bσ2XT X and bBn = XT diag{[y −Xbγ] ◦[y −Xbγ]}X.Here ◦ denotes the Hadamard, or componentwise, product.
  • Misspecified linear regression: When the true model is linear but uses different covariates, GBICp adds a penalty for inflation of the error variance caused by model misspecification.In this case Bn = τ2XT X, and the additional term directly penalizes the variance inflation.
Loading 1005.5483v2…