Source-linked AI summary
A general class of zero-or-one inflated beta regression models
Raydonal Ospina, Silvia L. P. Ferrari
TL;DR
Continuous proportions with zeros or ones are not adequately handled by ordinary beta regression because endpoint masses are excluded. The paper proposes an inflated beta regression class with separate regression-linked mixture parameters, develops inference and diagnostics, and demonstrates the approach on real data and simulations.
Problem
Continuous-proportion data may contain zeros or ones, which ordinary beta distributions cannot assign positive probability.
Method
The paper models the response with a beta component and a point mass at zero or one, linking mixture parameters to covariates and providing likelihood, diagnostic, and selection tools.
Results
The model showed reasonable fit in the traffic-accident mortality application, with LR = 0.584.
Takeaways & Limitations
The proposed framework supports modeling unit-interval rates or proportions containing zeros or ones with inference and diagnostics for departures and influential observations.
Takeaways & Limitations
The model focuses on settings where only one endpoint, zero or one, appears in the data and assumes the specified model and regularity conditions for asymptotic inference.
Abstract
from arXiv · showhide
This paper proposes a general class of regression models for continuous proportions when the data contain zeros or ones. The proposed class of models assumes that the response variable has a mixed continuous-discrete distribution with probability mass at zero or one. The beta distribution is used to describe the continuous component of the model, since its density has a wide range of different shapes depending on the values of the two parameters that index the distribution. We use a suitable parameterization of the beta law in terms of its mean and a precision parameter. The parameters of the mixture distribution are modeled as functions of regression parameters. We provide inference, diagnostic, and model selection tools for this class of models. A practical application that employs real data is presented.
1 Introduction
Continuous proportions require regression models suited to bounded outcomes, especially when observations include zeros or ones. The paper develops a general inflated beta regression framework and supplies inference, diagnostics, model selection, and an empirical application.
- Motivation: Standard normal regression models are unsuitable for continuous proportions measured on the unit interval.Examples include retirement contributions, work-related time, food spending, and industrial ammonia loss.
- Motivation: Beta regression is attractive because the beta density can represent many distributional shapes for bounded data.Its flexibility depends on the values of its two indexing parameters.
- Motivation: Zeros and ones create a modeling problem because the beta distribution assigns no positive probability to particular endpoints.A mixed continuous-discrete distribution can instead combine a beta component with a Bernoulli or degenerate endpoint component.
- Contribution: The proposed class models continuous proportions when only zero or only one occurs, using a beta distribution mixed with a degenerate distribution at known point c.The beta mean, precision, and endpoint probability may depend on linear or nonlinear predictors through smooth links.
- Contribution: The paper presents estimation, diagnostic, and model-selection tools, followed by a real-data application.The paper structure covers model definition, maximum likelihood estimation, diagnostics, application, and concluding remarks.
2 Zero-or-one inflated beta regression models
The model represents a continuous proportion as a mixture of a beta distribution and a point mass at one known endpoint. Separate predictors can govern endpoint probability, conditional mean, and precision.
- Beta component: The beta distribution is parameterized by mean µ and precision φ, with E(y)=µ and Var(y)=µ(1−µ)/(φ+1).For fixed µ, larger φ implies smaller variance.
- Mixture distribution: Because beta distributions exclude endpoint masses, the model mixes a beta component with a degenerate distribution at c, where c is zero or one.This targets datasets containing only one observed extreme.
- Mixture distribution: The endpoint probability is α, while the conditional continuous component has mean µ and variance µ(1−µ)/(1+φ).The unconditional mean is a weighted average of the endpoint value c and the beta mean.
- Regression specification: The regression class links α, µ, and φ to covariates through linear or nonlinear predictors and twice-differentiable monotonic link functions.The parameter vectors and predictor derivatives are required to have specified full ranks.
- Model features: For c=0, separate covariates govern zero probability, the conditional mean among nonzero observations, and conditional precision.This permits non-constant response variances and separates consumers from nonconsumers in expenditure proportions.
- Special cases: Zero-inflated and one-inflated beta regressions are special cases, while ordinary beta regression arises as a limiting case.Linear predictors yield corresponding linear models; additional restrictions recover established beta regression forms.
3 Likelihood inference
Likelihood inference separates estimation of the endpoint component from the continuous beta component, enabling numerical maximum likelihood fitting and asymptotic inference. Simulations examine finite-sample behavior under varying endpoint probabilities and sample sizes.
- Likelihood construction: The likelihood factorizes into a discrete-component term for ρ and a continuous-component term for (β,γ), so the parameters are separable.The two parts can be fitted independently using endpoint indicators and observations in (0,1).
- Estimation: Maximum likelihood estimators have no closed-form expressions and are computed numerically with Newton-type, quasi-Newton, or related optimization algorithms.An iterative estimation procedure is also provided, and GAMLSS supports flexible parametric, smooth, and random-effect specifications.
- Large-sample inference: Under regularity conditions, the maximum likelihood estimators are consistent and have asymptotic distributions supporting standard errors and confidence intervals.The inverse Fisher information matrix supplies estimated asymptotic variances for regression parameters and mean responses.
- Hypothesis testing: Likelihood-ratio, score, and Wald tests are available, with large-sample chi-squared approximations for their null distributions.A signed Wald statistic can test individual parameters modeling the mean.
- Finite-sample performance: Increasing endpoint probability raises bias and MSE for continuous-component estimators because fewer observations remain in (0,1).Dispersion-covariate estimators show more pronounced bias and MSE than mean-response estimators.
- Finite-sample performance: At n=50, the estimation algorithm failed to converge in 1.3% of samples, whereas it converged for every sample at larger sizes.Bias in discrete-component parameters was high for small samples but negligible for large samples.
- Finite-sample performance: For all studied sample sizes, mean estimates of β and γ were close to their true values, and root mean square errors decreased as sample size increased.In the simulation, β and γ were essentially estimated from roughly 70% of observations lying in (0,1).
4 Diagnostics
The paper develops diagnostics for assessing model adequacy, detecting outliers and misspecification, measuring influence, and comparing zero-or-one inflated beta regression models. These tools separately examine the discrete and continuous components while also providing global assessments.
- Motivation: Likelihood-based inference can be impaired by severe model misspecification or outliers, motivating residual and goodness-of-fit diagnostics.The paper proposes residuals for detecting departures from the postulated model and measures for assessing goodness-of-fit.
- Residuals: Standardized Pearson residuals assess the discrete and continuous components separately, with leverage-adjusted versions for identifying unusual observations.Discrete-component residuals account for observation leverage, while continuous-component residuals are based on the conditional beta distribution.
- Global diagnostics: Randomized quantile residuals provide a global residual using information from both the discrete and continuous components simultaneously.Apart from parameter-estimation variability, the randomized residuals are standard normal over their randomized intervals and are intended to be continuous.
- Residuals: Residual plots should show no detectable pattern; trends against predictors may indicate link-function misspecification, while simulated-envelope normal probability plots aid diagnosis.Simulation results reported in the paper indicate that randomized quantile residuals perform well at detecting incorrect distributional assumptions.
- Influence measures: Likelihood displacement measures the effect of removing an observation, and the paper recommends separate influence statistics for the discrete parameters and the continuous parameters.This separation reflects the model’s split between the discrete component and the continuous component.
- Model selection: Nested models can be compared with likelihood ratio tests, whereas non-nested models can be compared using GAIC, selecting the model with the smallest value.GAIC combines global fitted deviance with a penalty for the number of parameters; AIC, SBC, and CAIC are special cases.
5 An application
The application models traffic-accident mortality proportions in 200 southeastern Brazilian municipalities using a zero-inflated beta regression, then evaluates fit and influential observations. The analysis finds a positive association of the young-population proportion with both the probability of any death and the conditional mortality proportion.
- Data and objective: The data comprise traffic-accident death proportions for 200 randomly selected southeastern Brazilian municipalities in 2002, with young-population share as the main explanatory variable.Covariates also include population size, urbanization, male share, and education-development index.
- Data and objective: 39% of observations are at zero, while the positive proportions are asymmetric with an inverted-J shape and include outliers.Visual inspection suggested a zero-inflated beta distribution was suitable.
- Model specification: A zero-inflated beta regression was specified with separate logit models for α and µ and a log model for φ, using the municipality covariates.The model-selection procedure used stepGAICAll.B() with AIC.
- Model specification: Λ = 4.89 (p-value = 0.56) for jointly testing urbanization and male-share effects, so those covariates were excluded from the final model.The null hypothesis was not rejected at usual significance levels.
- Results: prop2029 has a positive effect on both the probability of at least one traffic-accident death (1 −α) and the conditional mean proportion of deaths (µ).The likelihood-ratio statistic was LR = 0.584, suggesting a reasonable fit.
- Diagnostics: Observation 138 was potentially atypical and influential for the continuous component, while observation 196 influenced the discrete component.Cases 2, 9, 85, and 138 jointly produced substantial changes in continuous-component estimates and inference.
- Diagnostics: The discrete-component residuals appeared randomly scattered without atypical observations, and their fitted-probability plot was not suggestive of lack of fit.Observation 196 was nevertheless highlighted by the discrete-component Cook statistics.
- Diagnostics: Continuous-component residuals were randomly scattered with no atypical observation, but the fitted-mean plot was not suggestive of lack of fit.Observations 2, 9, 85, and 138 were identified as influential for the continuous component.
6 Concluding remarks
The paper develops a general zero-or-one inflated beta regression class for proportions containing zeros or ones. It supplies inferential, diagnostic, estimation, and model-selection tools and illustrates the approach with real data.
- Concluding remarks: The proposed models target rates or proportions in the standard unit interval when zeros or ones are present.The class is intended for practitioners modeling such responses.
- Concluding remarks: The paper gives explicit score, Fisher-information, and inverse-information formulas, discusses iterative estimation and computation, and presents interval estimation.A real-data application is also presented and discussed.
- Concluding remarks: Diagnostic plots distinguish the discrete and continuous components of the model.Figure 3 contains panels (a)–(c) for the discrete component and (d)–(f) for the continuous component.
A Appendix A: Score vector and observed information matrix
Appendix A defines the score and information components used for inference in the model, including observed Fisher information blocks for the regression parameters.
- The score vector provides the derivative-based quantities used in likelihood inference for the model parameters.
- Table 5 reports estimates, standard errors, p-values, and relative changes caused by excluding observations.
- The observed Fisher information matrix is assembled from parameter-specific blocks for ρ, β, and γ.
- The information expressions incorporate covariate matrices and derivatives of the model components, including second-derivative array terms.
B Appendix B: Iterative algorithm for maximum likelihood estimation
Appendix B describes iterative re-weighted least-squares and Fisher-scoring procedures for obtaining maximum-likelihood estimates and extending regression diagnostics to the model components.
- Maximum-likelihood estimates for ρ and ϑ are obtained with a re-weighted least-squares algorithm.
- The iterative process for ρ interprets the update as fitting a generalized linear model and extends ordinary regression diagnostics to the discrete component.
- The iterative updates continue until successive parameter estimates differ by less than a specified small constant.
- For β, the converged update can be viewed as a weighted least-squares regression of the local modified dependent variable on X.
- The matrix P acts as a generalized leverage matrix, while I_n − P spans the residual space and identifies extreme design points through small 1 − P_tt values.