Source-linked AI summary
Mean and median bias reduction in generalized linear models
Ioannis Kosmidis, Euloge Clovis Kenne Pagui, Nicola Sartori
TL;DR
Finite-sample, boundary-estimate, and high-dimensional-nuisance problems can weaken maximum-likelihood inference in generalized linear models. This paper unifies mean and median bias reduction through adjusted scores and IWLS-compatible algorithms, finding improved inference and practical estimation in difficult settings.
Problem
Maximum-likelihood inference can deteriorate in small or moderate samples, with many parameters, or when ML estimates do not exist for some generalized linear models.
Method
The paper derives general adjusted score equations for mean and median bias reduction and solves them with IWLS using appropriately adjusted working variates.
Results
Mean and median bias reduction improve upon maximum likelihood, including practical handling of infinite estimates and inference for low-dimensional targets with high-dimensional nuisance parameters.
Takeaways & Limitations
The framework provides practical bias-reduced estimation and inference for generalized linear models, including categorical-response settings where standard ML can fail.
Takeaways & Limitations
The mean-bias-reducing iterative scheme has no theoretical convergence guarantee for general generalized linear models, despite substantial empirical evidence of stability.
Abstract
from arXiv · showhide
This paper presents an integrated framework for estimation and inference from generalized linear models using adjusted score equations that result in mean and median bias reduction. The framework unifies theoretical and methodological aspects of past research on mean bias reduction and accommodates, in a natural way, new advances on median bias reduction. General expressions for the adjusted score functions are derived in terms of quantities that are readily available in standard software for fitting generalized linear models. The resulting estimating equations are solved using a unifying quasi-Fisher scoring algorithm that is shown to be equivalent to iteratively re-weighted least squares with appropriately adjusted working variates. Formal links between the iterations for mean and median bias reduction are established. Core model invariance properties are used to develop a novel mixed adjustment strategy when the estimation of a dispersion parameter is necessary. It is also shown how median bias reduction in multinomial logistic regression can be done using the equivalent Poisson log-linear model. The estimates coming out from mean and median bias reduction are found to overcome practical issues related to infinite estimates that can occur with positive probability in generalized linear models with multinomial or discrete responses, and can result in valid inferences even in the presence of a high-dimensional nuisance parameter
1 Introduction
Generalized linear models provide a common framework for diverse responses, but finite-sample and high-dimensional settings can undermine maximum-likelihood inference. The paper develops mean and median bias reduction through adjusted scores, practical IWLS algorithms, and applications addressing boundary estimates and inference quality.
- Maximum-likelihood properties can deteriorate in small or moderate samples or when the parameter count is large relative to observations.
- 54.13% relative mean-squared-error inflation affects the maximum-likelihood dispersion estimator, whereas the moment estimator has smaller bias and improves confidence-interval coverage.
- Bootstrap and analytical higher-order methods improve inference but typically require intensive computation or model-specific algebra and depend on the existence of the ML estimate.
- The paper unifies mean and median bias reduction for generalized linear models using adjusted score equations that preserve computational simplicity and address boundary estimates.
- Explicit adjusted-score formulae allow both procedures to use iteratively re-weighted least squares with adjusted working variates, including settings where standard ML IWLS diverges.
- Mean bias reduction is exactly invariant under linear parameter transformations for transformed-estimator mean bias, supporting inference on arbitrary regression-parameter contrasts.
2 Bias reduction and iteratively reweighted least squares
The paper expresses mean and median bias reduction in GLMs through adjusted scores solved by quasi-Fisher scoring or equivalent IWLS updates with modified working variates.
- 2.3 Mean bias-reducing adjusted score functions: General mean-bias corrections can be implemented through IWLS by replacing the working variate z with z + φξ after ordinary iterations.
- 2.3 Mean bias-reducing adjusted score functions: Mean bias-reducing adjusted scores produce estimators with smaller asymptotic mean bias than maximum likelihood.
- 2.3 Mean bias-reducing adjusted score functions: Quasi-Fisher scoring updates β and, when dispersion is unknown, φ simultaneously because the mean-bias correction generally depends on φ.
- 2.3 Mean bias-reducing adjusted score functions: Mean bias-reduction iterations have no general convergence guarantee, although empirical studies found no divergence, including cases where standard IWLS fails.
- 2.4 Median bias-reducing adjusted score functions: Median bias reduction uses the mean-BR IWLS update followed by translation by φu, equivalently adding φXu to the working variate.
- 2.4 Median bias-reducing adjusted score functions: Median bias-reducing iterations likewise lack a general convergence guarantee, while extensive empirical studies found no evidence of divergence.
3 Inference with mean and median bias reduction
Mean and median bias-reduced estimators retain the ML estimator’s asymptotic distribution, enabling plug-in standard errors and Wald procedures. In normal regression, median bias reduction gives intervals closer to exact Student-t inference, while exact inferential solutions are not generally available for other GLMs with unknown dispersion.
- Mean and median bias-reduced estimators have the same asymptotic distribution as ML, so they are asymptotically unbiased and efficient.Their finite-sample distributions can therefore be approximated using the target parameter and inverse Fisher information.
- Standard errors and Wald tests or confidence intervals are obtained by replacing ML estimates with the corresponding bias-reduced estimates in usual software procedures.
- 3.2 Normal linear regression models: In normal linear regression, mean BR estimates the dispersion parameter without bias, whereas median BR yields Wald intervals closer to exact Student-t intervals.The comparison concerns practically relevant residual degrees of freedom and significance levels.
- 3.2 Normal linear regression models: For n − p = 1, the stated interval inequality holds for every 0 < α < 0.62647.
- Exact inferential solutions are generally unavailable for other GLMs with unknown dispersion, motivating approximate evaluation beyond normal linear regression.The paper investigates whether the desirable median-BR inference behavior persists approximately in Gamma regression.
4 Mixed adjustments for dispersion models
The mixed adjustment combines mean bias reduction for regression parameters with median bias reduction for dispersion, exploiting their complementary invariance properties. In Gamma regression simulations, it preserves the relevant parameterization properties while remaining third-order equivalent to the corresponding separate adjustments.
- The mixed adjustment combines mean bias reduction for β with median bias reduction for φ to exploit their complementary invariance properties.Mean BR is invariant under affine parameter transformations, while median BR is invariant under componentwise monotone transformations.
- For known dispersion, the mixed adjustment reduces to mean BR; for general unknown-dispersion GLMs, its estimators are asymptotically third-order equivalent to the separate mean- and median-BR estimators.
- Gamma regression illustration: The Gamma regression illustration uses 12 independent responses with mean µi = exp(ηi) and variance φµi^2 across alternative parameterizations.
- Gamma regression illustration: In a Gamma regression with three equivalent parameterizations, the mixed adjustment is numerically invariant under linear contrasts of mean parameters and monotone transformations of dispersion.ML and mean BR satisfy the contrast property, whereas ML and median BR satisfy the monotone-dispersion property; the mixed adjustment inherits both.
- Gamma regression illustration: Table 4 evaluates parameterization invariance probabilities from 1000 simulated response vectors under ML, mean BR, median BR, and mixed BR.
5 Illustrations and simulation studies
Case studies and simulations show that mean and median bias reduction achieve their intended bias properties across Gamma and logistic regression settings, including separation and many nuisance parameters.
- 5.1 Case studies and simulation experiments: Mean and median bias reduction provide empirical support for mean and median bias reduction, respectively, across the reported case studies and simulations.The studies cover Gamma regression, logistic regression with separation, and logistic regression with many nuisance parameters.
- 5.2 Gamma regression model for blood clotting times: Mean, median, and mixed bias-reduction methods have similar bias and variance for clotting-model regression coefficients, while improving dispersion estimation and confidence-interval coverage over ML.Mean BR nearly removes ML’s dispersion mean bias; median and mixed BR achieve nearly 50% underestimation probability, with the latter intervals performing best despite liberal coverage overall.
- 5.3 Logistic regression for infant birth weights: Both mean and median BR shrink logistic-regression estimates relative to ML, produce smaller standard errors, and remain finite under data separation.Mean BR estimates are always finite for full-rank binary-regression design matrices; in the simulation, 103 of 10 000 ML fits had infinite components.
- 5.4 Logistic regression with many nuisance parameters: Mean and median BR closely reproduce conditional-likelihood estimates when many nuisance parameters make ML estimation and inference highly biased.Conditional likelihood eliminates matched fixed effects using sufficient statistics, whereas bias reduction remains applicable more generally when such statistics do not exist.
6 Multinomial logistic regression
The paper extends bias reduction to multinomial logistic regression through its equivalent Poisson log-linear representation and examines parameterization invariance using alligator food-choice data.
- Multinomial–Poisson equivalence: Mean and median bias reduction can be implemented for multinomial logistic regression through an equivalent Poisson log-linear model with rescaled Poisson means.Rescaling current Poisson expectations to multinomial totals before the IWLS update yields the corresponding mean or median BR iteration.
- Parameterization invariance: Mean BR is invariant under affine parameter transformations, whereas median BR is not formally invariant under direct contrasts between baseline categories or covariate reference choices.Median BR estimates under two parameterizations can nevertheless be nearly identical in the alligator model.
- Primary food choices of alligators: In the alligator data, mean and median BR estimates are shrunk relative to ML, with corresponding shrinkage in estimated standard errors.Table 10 reports estimates and standard errors for ML, mean BR, median BR, and transformed median BR under the model’s baseline-category parameterization.
- Primary food choices of alligators: When frequencies are halved, two ML estimates become infinite, whereas mean and median BR remain finite for all parameters.The paper attributes such separation-related infinities to sparse data or large covariate effects.
- Simulation study: Across multinomial-total scaling factors r, mean BR is preferable for mean bias, while median BR gives better median centering and both converge toward ML asymptotics as r increases.The simulation uses 10 000 data sets at each r; ML is omitted because infinite-estimate probabilities range from 1.3% at r = 5 to 76.4% at r = 0.5.
7 Discussion
The discussion connects adjusted-score inference, mixed bias-reduction strategies, and the practical value of mean and median bias reduction for nuisance-rich models.
- The mixed approach remains valid for generalized linear models with dispersion covariates because of Fisher orthogonality between mean and dispersion parameters.
- Figure 4 compares underestimation probabilities across r values for mean BR, median BR, and median BRγ′ estimators using 10 000 simulated samples.
- Adjusted scores can be viewed as model-based penalties that provide implicit regularization sufficient to achieve mean or median bias reduction.
- The framework also develops score-type inference that complements Wald-type statistics, although the discussion states that inference and model comparison have used Wald-type statistics.
- Mean and median bias reduction support adjusted-score inference for confidence intervals, regions, hypothesis tests, and analysis-of-deviance-like tables.
- Mean and median bias reduction can aid inference for low-dimensional parameters while improving nuisance-parameter estimates when nuisance dimensions are high.
Proof of Theorem 3.1
The proof establishes nested confidence intervals under stated degrees-of-freedom and significance-level conditions, with additional comparisons of interval lengths.
- The proof shows ˆI_1−α ⊂ I*_1−α ⊂ I†_1−α for n−p ≥ 1 and α ∈ (0,1).
- The inclusion I†_1−α ⊂ I^E_1−α holds when g(ν, α) is positive, which is guaranteed for α below a unique threshold ˜α(ν).
- For fixed natural ν ≥ 1, the threshold ˜α(ν) increases with ν and has minimum ˜α(1) = 0.35562.
- When ν > 1, the absolute difference between the lengths of I†_1−α and I^E_1−α is smaller than the corresponding difference for I*_1−α and I^E_1−α.
- For ν = 1, the function h(ν, α) is positive for α < 0.62647 and negative otherwise.
Supplementary Material for Mean and median bias reduction in generalized linear models
The supplementary material identifies the report’s authors as Ioannis Kosmidis, Euloge Clovis Kenne Pagui, and Nicola Sartori.
- The report lists Ioannis Kosmidis as an author.
- The report lists Euloge Clovis Kenne Pagui as an author.
- The report lists Nicola Sartori as an author.
1 Introduction
The supplementary material reproduces the paper’s numerical results and figures with R and provides code, package setup, model fits, and simulation outputs for reproducibility.
- The report reproduces the numerical results and figures from the main text using R version 3.5.1 and the brglm2 package.
- The setup checks installed versions, installs brglm2 when needed, and loads packages used for reproducing the numerical results.
- The supplementary materials provide scripts for model fits and simulation experiments, with results stored in the glmbias code+results.zip archive.
2 Gamma regression model for blood clotting times
This section reproduces the blood-clotting Gamma regression analyses, including maximum-likelihood, mean bias-reduced, median bias-reduced, and mixed bias-reduced fits. It also reconstructs simulation results used for Tables 1 and 6.
- The blood-clotting data are modeled with a Gamma generalized linear model using lot, log(u), and their interaction.
- Maximum-likelihood estimates are computed first, with coefficient summaries and dispersion-based standard errors extracted for comparison.
- Mean bias reduction is fitted by updating the maximum-likelihood model with the AS_mean adjustment and recomputing coefficient summaries and standard errors.
- Median and mixed bias reduction are analogously fitted with AS_median and AS_mixed adjustments, respectively, and their summaries are recomputed.
- Simulation outputs are combined across maximum likelihood, mean bias reduction, median bias reduction, and mixed bias reduction, with root mean squared error calculated and reported as percentages.
3 Mixed adjustments for dispersion models
This section reproduces applications and simulations comparing maximum likelihood with mean and median bias reduction across binomial, conditional-likelihood, multinomial, and dispersion-related analyses. The code also generates bias and underestimation diagnostics for the alligator example.
- Birth-weight analysis: Birth-weight analyses compare maximum likelihood, mean bias reduction, and median bias reduction in a binomial model, while simulations report coefficient bias and square-root mean squared error.
- Infertility analysis: The infertility example compares a subject-specific binomial model, conditional likelihood, mean bias reduction, and median bias reduction using coefficient estimates and standard errors.
- Alligator analysis: Alligator multinomial analyses fit maximum likelihood, mean bias reduction, median bias reduction, and median-bias-reduced Agresti contrasts, with transformed coefficients and standard errors reported.
- Alligator simulations: The alligator simulation plots empirical relative bias against r and probability of underestimation against r across contrast categories c = 2, 3, 4, and 5.
- Alligator simulations: A second alligator figure uses the same faceting structure to display probability of underestimation, with 50% as the reference line.