Source-linked AI summary
Firth's logistic regression with rare events: accurate effect estimates AND predictions?
Rainer Puhr, Georg Heinze, Mariana Nold, Lara Lusa, Angelika Geroldinger
TL;DR
Firth-type logistic regression reduces coefficient-estimate bias but can introduce non-negligible prediction bias when events are very rare or very common. The paper compares penalization strategies, finding differences in coefficient and prediction performance alongside confidence-interval coverage limitations.
Problem
Firth-type logistic regression reduces coefficient-estimate bias at the cost of bias in predicted probabilities, which may be non-negligible for very rare or very common events.
Method
The comparison considers alternative penalization strategies, including tuning a compromise approach using penalized AIC.
Results
The methods included in the comparison differed in performance, with one method outperforming others for coefficient RMSE in 38 and prediction performance in 44 of 45 simulation scenarios; average predicted-probability bias removal reached 19.4%.
Takeaways & Limitations
RR is suggested when inference is not required, whereas FLAC or CP offer low variability and bias in effect estimates and predictions when valid confidence intervals for effects are needed.
Takeaways & Limitations
Confidence intervals did not reach nominal coverage levels because of combined bias and reduced variance, and naive Wald intervals from the penalized covariance matrix may be inadequate.
Abstract
from arXiv · showhide
Firth-type logistic regression has become a standard approach for the analysis of binary outcomes with small samples. Whereas it reduces the bias in maximum likelihood estimates of coefficients, bias towards 1/2 is introduced in the predicted probabilities. The stronger the imbalance of the outcome, the more severe is the bias in the predicted probabilities. We propose two simple modifications of Firth-type logistic regression resulting in unbiased predicted probabilities. The first corrects the predicted probabilities by a post-hoc adjustment of the intercept. The other is based on an alternative formulation of Firth-types estimation as an iterative data augmentation procedure. Our suggested modification consists in introducing an indicator variable which distinguishes between original and pseudo observations in the augmented data. In a comprehensive simulation study these approaches are compared to other attempts to improve predictions based on Firth-type penalization and to other published penalization strategies intended for routine use. For instance, we consider a recently suggested compromise between maximum likelihood and Firth-type logistic regression. Simulation results are scrutinized both with regard to prediction and regression coefficients. Finally, the methods considered are illustrated and compared for a study on arterial closure devices in minimally invasive cardiac surgery.
1 Introduction
Firth-type penalization reduces small-sample coefficient bias and handles separation, but can bias predicted probabilities toward 1/2, especially with rare events. The paper evaluates this problem and proposes two modifications intended to retain coefficient accuracy while correcting predictions.
- Firth-type penalization reduces small-sample bias in coefficient estimates and provides finite estimates when maximum likelihood fails under separation.
- Firth-type penalization can bias predicted probabilities toward 1/2, making average predictions too high when events are rare.
- The paper investigates the practical size and origin of prediction bias using theoretical considerations and simulations.
- It proposes FLIC, an intercept correction, and FLAC, an added covariate distinguishing original from pseudo-observations, to obtain unbiased predictions.
- The proposed approaches are compared with alternative penalization strategies in simulations and illustrated using arterial closure devices in minimally invasive cardiac surgery.
- In a 2 × 2 example, ML predicts 5% and 20% for x = 0 and x = 1, whereas Firth-type estimates predict 5.44% and 25%, with 11.6% average overestimation.
2 Methods
The methods formulate Firth-type estimation through penalized logistic regression and iterative data augmentation. FLIC corrects only the intercept, while FLAC distinguishes original from pseudo-observations to recalibrate predictions to the observed event rate.
- The logistic model links binary outcomes to covariates through P(y_i = 1|x_i) = (1 + exp(−x_iβ))^−1, with ML estimates obtained by maximizing the log-likelihood.
- Firth-type estimates use Jeffreys invariant penalization and modified score equations involving the hat-matrix diagonal elements h_i.
- Firth-type estimation can be represented as iterative maximum-likelihood estimation on data augmented with pseudo-observations having opposite responses and weights h_i/2.
- The relative bias in average predicted probabilities increases when the number of events k is smaller or the number of parameters p is larger.
- FLIC replaces the Firth-type intercept with an estimated intercept that makes average predicted probability equal the observed event proportion while leaving other coefficients unchanged.
- In the 2 × 2 example, FLIC predicts 4.86% and 22.82%, while FLAC predicts 5.16% and 16.83%.
- FLAC adds an indicator distinguishing original and pseudo-observations, then fits maximum likelihood on the augmented data to recalibrate predictions to the original event proportion.
3 Simulation study
The simulation compared penalized logistic-regression methods across prediction, calibration, discrimination, coefficient estimation, and confidence-interval performance. FLIC and FLAC improved predicted probabilities over ordinary Firth-type estimation, while FL generally gave the least biased covariate coefficients and RR often minimized prediction RMSE.
- Evaluation design: The simulation evaluated bias and RMSE of predictions and coefficients, plus discrimination, calibration slopes, and confidence-interval power, length, and coverage.Prediction results included ML, WF, FL, FLIC, FLAC, LF, CP, AU, AB, and RR; coefficient analyses covered a smaller method set.
- Simulation design: 10 explanatory variables combined continuous, binary, and ordinal covariates, with only linear effects retained because the study focused on prediction.The variables were generated from correlated standard normal variables and transformed into four continuous, four binary, and two ordinal variables.
- Predictions: FLIC, FLAC, LF, AU, and RR produced unbiased predicted probabilities, while RR had considerably lower RMSE than the other methods.The small deviations from zero reported for unbiased methods were attributed to simulation sampling variability.
- Predictions: FLAC or FLIC reduced prediction RMSE versus FL by up to 16.2%, whereas AB performed worst for bias and RMSE in almost all scenarios.FLAC and FLIC also reduced variance; differences between methods diminished with increasing sample size, event rate, and effect size.
- Calibration: RR achieved the calibration slope closest to 1 in some scenarios, while FLAC or RR generally performed best for calibration.All methods except RR had mean calibration slopes below 1 throughout the scenarios, indicating underestimation of small and overestimation of larger event probabilities.
- Prediction profiles: RR overestimated small event probabilities and underestimated large ones, whereas AU was close to unbiased but had considerable variance for larger predictions.FLAC showed the same pattern as RR less strongly; RR nevertheless performed well on both linear predictors and predicted probabilities because its linear-predictor RMSE was smallest.
- Coefficient estimates: FL had the smallest absolute coefficient bias in 28 of 36 non-zero-effect scenarios, with average standardized bias never exceeding 1%.RR generally had lower coefficient RMSE, while ML and WF had the largest and second-largest RMSE across scenarios.
- Confidence intervals: RR confidence intervals were overly conservative without covariate effects but too optimistic with moderate or large effects, and were shorter with less power.Coverage for methods other than RR ranged from 93.9% to 96% across methods and scenarios.
4 Example: arterial closure devices in minimally invasive cardiac surgery
The cardiac-surgery example compared methods for estimating the association between arterial closure devices and complications and for quantifying individual complication risks. All methods supported lower complication rates with arterial closure devices, while their predicted risks and discrimination varied.
- Study and outcome: The retrospective study analyzed 440 eligible patients, including 16 complications in the arterial-closure-device group and 8 among 90 conventional-access surgeries.The reported complication rates were 2.3% for arterial closure devices and 8.9% for conventional surgical access.
- Model specification: Models included surgical access, logistic EuroSCORE, previous cardiac operations, BMI, and diabetes, with estimates obtained by multiple penalized and unpenalized methods.The analysis used four medically selected adjustment variables in addition to surgical access.
- Coefficient estimates: The estimated odds ratio for surgical access ranged from 3.1 with RR to 5.66 with ML, with every 95% confidence-interval lower bound above 1.The authors considered FL estimates potentially preferable to ML because of small-sample bias reduction, although the difference was not substantial.
- Interpretation: The example’s confidence-interval behavior differed from the simulation, where RR intervals were more likely to include zero across effect sizes.The authors attributed this discrepancy partly to the simulation assumption that explanatory variables had similar effect sizes.
- Predicted risks: FLIC and FLAC produced smaller predicted probabilities than FL, AB produced the largest individual probabilities, and RR had the smallest prediction range.AB’s largest individual predictions were 24% for the event group and 33.9% for the non-event group.
- Discrimination: Cross-validated discrimination was best for AB at 68.2% and poorest for AU at 61.4%.Predicted probabilities for average-covariate patients showed the same broad method-specific pattern as the group distributions.
- Main association: All methods estimated a significant association between arterial closure devices and lower complication rates than conventional access.For patients with average covariate values, predicted risks ranged from 1.3% to 2% with arterial closure devices and from 6.1% to 8.2% with conventional access.
5 Discussion
The proposed FLIC and FLAC modifications improve Firth-type predictions, while broader comparisons identify trade-offs among prediction accuracy, effect-estimate accuracy, inference, and implementation. The authors slightly prefer FLAC over FLIC when confidence intervals are needed, and RR when inference is not required.
- Performance of proposed methods: FLIC and FLAC efficiently improve Firth-type predictions, with FLIC preserving Firth-type effect estimates.FLIC leaves bias-corrected effect estimates unchanged, whereas FLAC can alter them.
- Prediction performance: Up to 19.4% bias in average predicted probabilities was removed in the simulations.
- Prediction performance: FLIC and FLAC provided more accurate individual predicted probabilities, while FLIC and FLAC were never worse than FL for prediction improvement.
- Comparison with alternative methods: The weakened Firth compromise improved prediction bias and RMSE over FL but was outperformed by most other methods, including FLIC, FLAC, LF, CP, and RR.
- Comparison with alternative methods: RR outperformed other methods for coefficient RMSE in 38 and prediction RMSE in 44 of 45 simulation scenarios, but its confidence intervals lacked nominal coverage.The authors frame RR as preferable when inference is not required; CP or FLAC may be preferable when valid confidence intervals are needed.
- Practical recommendations: The authors slightly prefer FLAC over FLIC because of lower prediction RMSE, and favor FLAC or CP when valid effect confidence intervals are required.FLAC’s preference is also linked to its invariance properties.
6 Appendix
The appendix documents the simulation-variable structure and acknowledges contributors, data providers, and funding sources.
- Table 5 specifies the structure of explanatory variables used in the simulation study.It follows Binder, Sauerbrei and Royston [11].
- The table notation uses square brackets to remove the non-integer part of an argument.
- The indicator function 1 takes value 1 when its argument is true and 0 otherwise.
- The authors acknowledge Michael Schemper and Mohammad Ali Mansournia for helpful comments.
- Cardiac-surgery complication data were provided by collaborators from the Department of Cardiothoracic Surgery at University Hospital Jena.
- The research received funding from the Austrian Science Fund and the Slovenian Research Agency.