Source-linked AI summary
New Important Developments in Small Area Estimation
Danny Pfeffermann
TL;DR
Small area estimation seeks reliable estimates and precision assessment when areas have small or no samples. This review synthesizes design-based and model-dependent developments, including frequentist and Bayesian approaches, and discusses reported advances and limitations.
Problem
Small area estimation must produce reliable estimates of means, counts, and quantiles, while assessing prediction error, for areas with small or no samples.
Method
The paper reviews design-based and model-dependent SAE methods, classifying the latter into frequentist and Bayesian approaches and discussing proposed solutions.
Results
The review reports that restricted models yield their largest precision gain when the prior is completely specified, while a proposed PEB predictor outperforms another PEB predictor in PMSE.
Takeaways & Limitations
Bayesian methods offer flexible inference and PMSE or credibility-interval computation without asymptotic properties, whereas M-quantile methods can be more robust to model misspecification.
Takeaways & Limitations
Bayesian approaches require prior-distribution specification, and M-quantile methods have no obvious way to predict means or other targets for nonsampled areas.
Abstract
from arXiv · showhide
The problem of small area estimation (SAE) is how to produce reliable estimates of characteristics of interest such as means, counts, quantiles, etc., for areas or domains for which only small samples or no samples are available, and how to assess their precision. The purpose of this paper is to review and discuss some of the new important developments in small area estimation methods. Rao [Small Area Estimation (2003)] wrote a very comprehensive book, which covers all the main developments in this topic until that time. A few review papers have been written after 2003, but they are limited in scope. Hence, the focus of this review is on new developments in the last 7-8 years, but to make the review more self-contained, I also mention shortly some of the older developments. The review covers both design-based and model-dependent methods, with the latter methods further classified into frequentist and Bayesian methods. The style of the paper is similar to the style of my previous review on SAE published in 2002, explaining the new problems investigated and describing the proposed solutions, but without dwelling on theoretical details, which can be found in the original articles. I hope that this paper will be useful both to researchers who like to learn more on the research carried out in SAE and to practitioners who might be interested in the application of the new methods.
1. PREFACE
Small area estimation addresses reliable estimation and error assessment when areas have small or no samples. This review surveys recent developments while briefly providing older background and emphasizing practical solutions.
- SAE seeks reliable estimates of means, counts, quantiles, and other characteristics for areas with small or no samples, together with estimation-error assessment.
- SAE estimates support fund allocation, educational and health programs, environmental planning, and census-count testing or adjustment.
- Research in SAE has accelerated, with diverse problems and solutions described as both innovative and practical.
- The review focuses on developments from the preceding 7–8 years, while briefly covering older developments for self-containment.
- It explains investigated problems and proposed solutions without dwelling on theoretical details developed in the original articles.
2. SOME BACKGROUND
SAE concerns area- or domain-level estimation, where sample size rather than geographic size creates difficulty. Its methods include design-based and model-based approaches, and auxiliary covariates are vital when samples are small.
- In SAE, “areas” may be geographic districts or other groupings such as sociodemographic groups or industries, often called domains.
- Small area problems require separate point estimates and error measures for each area, including poverty or disease estimates displayed through maps.
- SAE methods divide broadly into design-based and model-based approaches, with model-based methods using frequentist, Bayesian, or empirical Bayes approaches.
- Auxiliary covariate information from surveys, censuses, registers, and administrative records is vital because small samples can limit even elaborate models.
3. NOTATION
The notation defines a finite population partitioned into areas, with samples observed in some areas and covariates recorded for units. The area target θ_i can be a mean, proportion, count, or quantile.
- The population U has N units divided into M exclusive and exhaustive areas, with area i containing N_i units.
- Samples are available in m ≤ M areas, and the overall sample is the union of sampled-area samples s_i with sizes n_i.
- The response y_ij denotes unit-level characteristics, and ȳ_i is the sample mean for sampled area i.
- Each unit has covariates x_ij, with x̄_i denoting sampled-area covariate means and X̄_i the corresponding true area means.
- The area target θ_i may be an area response mean, a proportion when y_ij is binary, a count, or a quantile.
4. DESIGN-BASED METHODS
Design-based SAE methods use sampling-design randomization to estimate small-area characteristics and their errors, often assisted by auxiliary covariates. The reviewed developments address estimator construction, calibration, robustness, sampling design, and the trade-offs among bias, variance, and inference scope.
- Design-based estimators in common use: Direct estimators are design-unbiased but can have large variance when area sample sizes are small.Their conditional design variance is O(1/ni), unless within-area variation is sufficiently small.
- Design-based estimators in common use: Synthetic estimators use common regression coefficients and auxiliary covariates, yielding low variance but potentially large bias when coefficients differ across areas.Under SRSWOR, the pooled OLS coefficient is approximately design-unbiased and consistent, but area-specific coefficients can induce bias.
- Design-based estimators in common use: Survey regression estimators correct synthetic estimates using Horvitz–Thompson estimates and are approximately design-unbiased, but their variance returns to order O(1/ni).They perform well when covariates have good predictive power, and variance may be reduced by scaling the bias correction.
- Design-based estimators in common use: Composite estimators combine synthetic and survey regression estimators to compromise between synthetic bias and survey-regression variance, with weights often depending on area sample size.Choosing the combination weights is not straightforward because the synthetic estimator’s bias is usually difficult to assess accurately.
- Some new developments in design-based small area estimation: New design-based developments include calibration, generalized linear or mixed assisting models, design-consistent model-dependent estimators, model-based direct estimators, and controlled multi-domain sampling.These approaches seek improved efficiency, robustness to model misspecification, or guaranteed bounds on domain sampling errors, though some require strong predictions or may produce overly large samples.
- Pros and cons of design-based small area estimation: Design-based methods face practical limits because confidence intervals may require unjustified large-sample normality, and randomization theory does not support prediction for nonsampled areas.Design-based inference also does not readily permit conditional inference on sampled covariates or clusters.
5. MODEL-BASED METHODS
Model-based small area estimation assumes a model for observed data and predicts area characteristics under that model. The review covers general formulations and frequentist, empirical Bayes, and hierarchical Bayes predictors for continuous and binary outcomes.
- Model-based methods use optimal or approximately optimal predictors of area characteristics under an assumed sample-data model.Prediction-error MSE is also defined and estimated under the model.
- General formulation: A typical small area model links the distribution of observed responses given the target with a model for targets given covariates, using random effects for unexplained area variability.Unknown hyper-parameters are estimated frequently or handled through Bayesian priors.
- Area-level model: The area-level Fay–Herriot model is used when covariates are available only at the area level and combines direct estimates with synthetic covariate-based estimates.Its shrinkage coefficient optimally determines the weights assigned to the two components, unlike ad hoc design-based weighting.
- Frequentist and Bayesian predictors: Empirical BLUP and empirical Bayes predictors replace unknown variance components with estimates, while hierarchical Bayes predictors use priors and posterior distributions for point prediction and intervals.For transformed targets, empirical Bayes and hierarchical Bayes can produce optimal predictors of the original quantity more flexibly than transformed BLUPs.
- Area-level model: BLUPs are unbiased under the joint model distribution but can be biased when conditioning on area random effects or under the randomization distribution.This model-based unbiasedness treats area intercepts as random rather than fixed.
- Generalized linear mixed models: For binary responses, generalized linear mixed models target area proportions or counts; best predictors generally require numerical integration, while empirical and hierarchical Bayes approaches estimate parameters or posterior distributions.Hierarchical Bayes implementation can use MCMC simulations to generate posterior realizations.
6. NEW DEVELOPMENTS IN MODEL-BASED SAE
Recent model-based SAE developments improve prediction-error assessment, prediction intervals, benchmarking, and robustness under measurement errors or outliers.
- Prediction MSE: Conditional PMSE estimation is addressed through a modified jackknife that replaces the usual component and retains controlled bias orders.The modification estimates conditional PMSE with bias of order op(1/m) and unconditional PMSE with bias of order o(1/m).
- Prediction MSE: Bootstrap PMSE estimation achieves bias of order o(1/m) under regularity conditions.
- Prediction MSE: A general cross-validated bias-correction method outperforms double-bootstrap and jackknife procedures in an extensive PMSE simulation study.Under certain conditions, it also estimates the MSE of the bias-corrected estimator.
- Prediction Intervals: Frequentist prediction intervals based on estimated prediction-error variance have coverage error O(1/m), motivating parametric-bootstrap refinements.Bootstrap calibration reduces the coverage error to order O(m^-2), and appropriate choices can produce area-specific intervals.
- Benchmarking: Benchmarking methods impose constraints on predictors, including in time-series models and Bayesian formulations, while restricted models yield precision gains only under informative priors.With a noninformative prior, the restricted model produces no gain in precision.
- Measurement Errors and Outliers: New predictors address measurement errors in covariates and outliers, including asymptotically optimal PEB or HB procedures and robust empirical-Bayes predictors.Simulation results report improved PMSE for one PEB predictor, while robust approaches model outlying direct estimates or random effects.
M-quantile estimation
M-quantile SAE models conditional quantiles rather than expectations, providing a flexible and potentially robust alternative for area estimation. It supports estimation of means and distributions but faces challenges for nonsampled areas and PMSE assessment.
- M-quantile SAE models quantiles of f(y_i|x_i) instead of conditional expectations, using a family indexed by q ∈ (0,1).
- Influence functions and iterative reweighted least squares estimate M-quantile coefficients while allowing robust treatment of outliers.
- For sampled units, the method identifies q_ij values matching observed responses and averages them to predict each area mean.
- The approach is not restricted to means and can estimate an area distribution function, although it assumes continuous outcomes.
- M-quantile estimators can be more robust to model misspecification, but prediction for nonsampled areas and corresponding PMSE estimation remain unresolved without a model.
Use of penalized spline regression
Penalized spline regression extends SAE by avoiding a fixed functional form for the response mean and treating spline coefficients as additional random effects. The resulting EBLUP can use standard mixed-model results, but requires population-wide covariate information.
- The broader SAE review presents penalization as another route to robustifying inference beyond the spline approach.
- P-spline regression approximates smooth response functions with polynomial basis terms and fixed knots rather than assuming a specific functional form.
- Opsomer et al. incorporate spline coefficients as additional random effects alongside area random effects in a unit-level SAE model.
- The resulting responses are dependent across areas because of common spline random effects, but BLUP and EBLUP estimation uses standard results.
- The approach requires covariates for every population element and supports second-order PMSE estimation, bias correction, bootstrap estimation, and hypothesis testing.
Use of empirical likelihood in Bayesian inference
Empirical likelihood combined with proper priors yields a semiparametric Bayesian SAE approach that avoids specifying outcome distributions parametrically. An application to state-wise median income produced better predictions than direct estimates and normality-based HB predictors.
- Empirical likelihood with proper priors defines a semiparametric Bayesian SAE method for continuous and discrete outcomes without specifying their distributions.
- The method represents cumulative distributions through jumps constrained by moment conditions and solves a constrained maximization problem.
- Posterior samples are obtained by combining empirical likelihood with prior distributions and using MCMC simulations.
- In estimating state-wise median income for four-person families, the approach produced much better predictions than direct survey estimates and HB predictors under normality.
Best predictive SAE
New predictive SAE methods target robustness to model misspecification and the distinct difficulty of predicting ordered area means. The reviewed results include strong gains for an observed best predictor under misspecification and less shrinkage for ordered-mean prediction.
- Observed best predictor: The observed best predictor estimates fixed model parameters to make resulting predictors optimal under a loss function, rather than treating parameter estimation as merely intermediate.
- Observed best predictor: The observed best predictor can significantly outperform the EBLUP in PMSE under model misspecification, while the two have similar PMSE under correct specification.
- Ordered means: Predicting ordered means differs from ranking areas because ordering changes the prediction target and requires evaluating predictors jointly under an ordered-means loss.
- Ordered means: For ordered means, constrained predictors replace the usual shrinkage coefficient with approximately γ^-1/2, producing less shrinkage of direct estimators.
- Ordered means: For m = 2, one compared predictor has no greater expected loss than the alternative for all δ, while simulations support extensions to larger m.
- Ordered means: The ordered-mean results also hold empirically when the variance is unknown and replaced by a method-of-moments estimator.
Assessment of literacy
The literacy assessment addresses outcomes with structural zeros and positive continuous scores, requiring a two-part model for area-level averages and positive-score proportions. A Bayesian multilevel analysis accounts for correlations between the two parts’ random effects.
- Outcome structure: Literacy-test outcomes are either structural zeroes indicating illiteracy or positive continuous scores measuring literacy level.Common SAE models are not applicable when the proportion of zeroes is high.
- Two-part model: The model separates the conditional mean among positive scores from the probability of having a positive score.The mean component uses xijk covariates, while the positivity probability uses zijk covariates in a logistic model.
- Two-part model: District and nested-village random effects are included in both model components and allowed to be correlated.The covariate sets may differ between the score and positivity models.
- Estimation: Bayesian estimation with noninformative priors and MCMC produces village- and district-level predictors for average scores and positive-score proportions.The Bayesian approach accounts for cross-component random-effect correlations that separate fitting cannot handle.
- Related approach: A related U.S. literacy study models county and state proportions at the lowest literacy level but does not use a two-part model.Its area proportions are modeled through a logistic transformation with state and county random effects.
Poverty mapping
Poverty mapping estimates nonlinear Foster–Greer–Thorbecke indicators for sampled and nonsampled areas. The reviewed approach transforms welfare outcomes, imputes nonsampled-unit measures through conditional simulation, and evaluates predictors and prediction-error estimators.
- Poverty indicators: Foster–Greer–Thorbecke indicators quantify poverty incidence, the poverty gap, and poverty severity through α = 0, 1, and 2.They are defined from welfare measures such as income or expenditure relative to a poverty threshold.
- Modeling strategy: For α = 1,2, Molina and Rao assume a one-to-one transformation that makes the outcomes satisfy a normal unit-level small-area model.This avoids directly assigning distributions to the nonlinear poverty measures.
- Prediction: Known sampled-unit poverty measures are retained, while missing measures for nonsampled units are imputed from conditional predictions under the transformed-outcome model.The procedure estimates area means by combining observed and imputed unit-level measures.
- Prediction: Conditional normal simulations generate predictions of unobserved outcomes given observed outcomes and estimated model parameters.Monte Carlo averaging is used to form empirical best predictors.
- Evaluation: Simulations and a Spanish real-data application demonstrate good performance of the area predictors and PMSE estimators.The application uses the transformation yij = log(Eij).
- Related approach: The World Bank uses a different procedure that simulates all population values, including sampled units, with random effects tied to design clusters.Molina and Rao argue this effectively treats all areas as nonsampled.
7. SAE UNDER INFORMATIVE SAMPLING AND NONRESPONSE
Informative area or within-area sampling and nonresponse can bias small-area inference when ignored. The reviewed methods adjust predictors using sample, population, and sample-complement relationships, with applications to county BMI and related compositions.
- Scope and motivation: The review notes that noninformative sampling is commonly assumed, although informative sampling and NMAR nonresponse can severely bias predictions.Both frequentist and Bayesian methods address these issues.
- Informative sampling: Pfeffermann and Sverchkov fit a sample model to observed data and use its relationship with population and sample-complement models to obtain unbiased predictors.The approach covers means in both sampled and nonsampled areas.
- Informative sampling: Under two-stage sampling, sample-complement distributions describe random effects and outcomes for areas or units excluded from the sample.When area selection is noninformative, population, sample, and sample-complement distributions coincide.
- Informative sampling: Within-area sampling weights are modeled through their expectations under the sample model, while no model is assumed for the relationship between area selection probabilities and area means.The resulting predictors use estimated model parameters and account for within-area selection.
- Prediction adjustments: A correction term adjusts sampled-area prediction for the difference between sample-complement and sample expectations.For nonsampled areas, another term accounts for nonzero mean random effects induced by informative area selection.
- Applications and extensions: The methods are applied to predicting county mean BMI from NHANES III, while related work also addresses binary outcomes and informative nonresponse.Nandram and Choi additionally account for informative nonresponse, whereas MDC and NC do not consider informative area selection.
- Applications and extensions: Zhang estimates area compositions under NMAR nonresponse using a generalized SPREE model.Compositions are counts or proportions across categories such as household types.
8. MODEL SELECTION AND CHECKING
Model selection and checking are central in SAE because mixed models contain unobservable random effects and standard criteria do not transfer straightforwardly. The review covers frequentist diagnostics, random-effect testing, fence selection, Bayesian checking, and conditional AIC.
- Motivation: SAE model selection is difficult because random effects are often unobservable and their distributions have limited or no direct information.Standard likelihood-based criteria such as AIC do not apply straightforwardly to mixed models.
- Frequentist criteria: Conditional AIC evaluates models operating in small areas by treating predicted random effects as part of the fitted structure.It can select fixed- and random-effects design matrices and performs well in theoretical and empirical evaluations.
- Frequentist diagnostics: Cumulative-sum residual tests assess covariate functional form and link-function adequacy in generalized linear mixed models.Visual comparisons with null realizations accompany a formal supremum test for covariate functional form.
- Random-effect testing: A bootstrap test evaluates whether random effects are needed by comparing the observed statistic with statistics from samples generated under the reduced model.Empirical results indicate good power for the proposed test.
- Model selection: Fence methods isolate a subgroup of correct mixed models and select an optimal model within that subgroup using criteria such as dimension or PMSE.An adaptive tuning coefficient is chosen by parametric bootstrap to maximize the highest model-selection frequency.
- Bayesian checking: Bayesian model checking compares discrepancy measures for observed data with replicated datasets generated under the fitted model.In balanced examples, ensembles of posterior predictive p-values can distinguish covariance misspecification when intra-cluster correlation is sufficiently high and the number of areas sufficiently small.
9. CONCLUDING REMARKS
The review finds model-based predictors generally more accurate and applicable to nonsampled areas, while design-based tools remain useful for model assessment, benchmarking, and informative sampling. Bayesian methods offer greater flexibility, but computational demands and expertise requirements remain important considerations; integrating developments across Bayesian and frequentist approaches is proposed as a future direction.
- Model-based predictors are generally more accurate and can predict for nonsampled areas where design-based theory does not exist.
- Design-based estimators support model-based prediction by supplying inputs, assessing predictors, benchmarking them, and accounting for informative sampling through weights.
- Bayesian methods flexibly generate posterior observations and support PMSE or credibility intervals without relying on asymptotic properties.
- Bayesian analysis can require intensive computation, expert knowledge, and computing skills, although frequentist GLMM fitting is also computation intensive.
- Future research could transfer developments between Bayesian and frequentist approaches, although some extensions may require extensive research or prove infeasible.