Source-linked AI summary

Online Updating of Statistical Inference in the Big Data Setting

Elizabeth D. Schifano, Jing Wu, Chun Wang, Jun Yan, Ming-Hui Chen

arXiv:1505.06354v1stat.COstat.OT

TL;DR

The paper addresses statistical inference for large data streams when historical observations cannot be stored or accessed efficiently. It develops online-updating inference for linear models and estimating equations, including predictive residual tests and a new estimator. Simulations and real-data analysis support favorable performance for the proposed estimating-equation estimator and greater power for predictive residual tests.

  • Problem

    Big-data streams require fast inference without storing historical data, while existing divide-and-conquer work has given limited attention to standard errors and inference.

  • Method

    The paper develops computationally efficient, minimally storage-intensive online-updating algorithms and inference for linear models and estimating equations, including predictive residual tests and a new estimator.

  • Results

    Simulations suggest predictive residual tests are more powerful, while the proposed CUEE estimator is less biased than Lin and Xi’s AEE/CEE estimator in finite samples.

  • Takeaways & Limitations

    The framework supports sequential estimation and inference for linear models and estimating equations when data arrive in streams and subset design matrices may be rank-deficient.

  • Takeaways & Limitations

    How to handle detected outliers in both settings remains an open question for future research.

Abstract

from arXiv · show

We present statistical methods for big data arising from online analytical processing, where large amounts of data arrive in streams and require fast analysis without storage/access to the historical data. In particular, we develop iterative estimating algorithms and statistical inferences for linear models and estimating equations that update as new data arrive. These algorithms are computationally efficient, minimally storage-intensive, and allow for possible rank deficiencies in the subset design matrices due to rare-event covariates. Within the linear model setting, the proposed online-updating framework leads to predictive residual tests that can be used to assess the goodness-of-fit of the hypothesized model. We also propose a new online-updating estimator under the estimating equation setting. Theoretical properties of the goodness-of-fit tests and proposed estimators are examined in detail. In simulation studies and real data applications, our estimator compares favorably with competing approaches under the estimating equation setting.

1 Introduction

The paper addresses inference for large, streaming datasets when historical data cannot be stored, developing online-updating methods for linear models and estimating equations. It provides efficient algorithms, rank-deficiency handling, predictive residual tests, and a new estimating-equation estimator.

  • Motivation: The authors identify limited attention to standard error estimation and inference in divide-and-conquer aggregation.
  • Motivation: Online-updating inference is motivated by data streams that require sequential analysis without storing historical observations.The framework targets computationally efficient, minimally storage-intensive analysis as new data arrive.
  • Contributions: The framework develops iterative estimation and inference procedures for linear models and estimating equations.These procedures update as new data arrive and accommodate possible rank deficiencies in subset design matrices.
  • Contributions: For online linear models, predictive residual tests are proposed for outlier detection and goodness-of-fit assessment.The paper derives exact test-statistic distributions under normal errors and asymptotic distributions for non-normal cases.
  • Contributions: For estimating equations, the paper introduces a new online-updated estimator, establishes consistency and uniqueness results, and evaluates it in simulations and real data.The proposed estimator outperforms competing divide-and-conquer or online-updated estimators in bias and mean squared error.

2 Normal Linear Regression Model

This section formulates online least-squares estimation and inference for data arriving in chunks, using cumulative summaries rather than retaining all observations. It addresses rank-deficient subsets and develops predictive residual diagnostics for online model assessment.

  • Online updating: Online analysis treats observations as arriving in chunks, requiring estimation and inference from cumulative data rather than all historical records.The proposed procedure is computationally efficient and minimally storage-intensive.
  • Online updating: Sequential formulas update regression estimates, residual sums of squares, mean squared error, t-tests, and ANOVA quantities as new subsets arrive.At the terminal accumulation point, the online formulas coincide with the corresponding batch divide-and-conquer expressions.
  • Rank deficiencies: Although subset coefficient estimates may be nonunique under rank deficiency, the cumulative coefficient and mean-square-error quantities can remain unique.The paper establishes this invariance through generalized-inverse arguments.
  • Rank deficiencies: The online scheme requires cumulative cross-product summaries and does not require each subset cross-product matrix to be invertible.This permits online updating when rare-event covariates produce rank-deficient subsets.
  • Model fit diagnostics: Predictive residuals use the previous cumulative estimate to evaluate newly arriving observations for online model-fit diagnostics.They are formed from current responses minus predictions based on the preceding cumulative data.

3 Online Updating for Estimating Equations

The paper develops online-updating estimators for estimating equations, addressing the finite-sample gap between cumulative approximations and full-data estimating-equation estimators. The cumulative estimator uses prior summaries and current subset information, supports rank-deficient subset matrices, and is shown to be consistent under regularity conditions.

  • Motivation: Nonlinear estimating functions make divide-and-conquer estimators generally differ from the estimator obtained by fitting all observations simultaneously.This motivates an online-updating approach for estimating equations rather than relying only on linear approximations.
  • Rank deficiency: The framework does not require each subset matrix A_nk,k to be invertible, but requires their cumulative sum to be invertible.The asymptotic theory additionally assumes subset matrices are invertible for sufficiently large subset sizes.
  • Cumulative estimating equation estimator: The cumulative estimating equation estimator uses previous intermediary estimators, the current subset estimator, and a bias-correction term for omitted earlier contributions.The update is expressed recursively through the intermediary and cumulatively updated estimators.
  • Relation to aggregated estimation: When the intermediary expansion point equals the subset estimator, the cumulative estimating equation estimator reduces to Lin and Xi’s aggregated estimating equation estimator.The paper introduces a different choice that uses information from previous accumulation points.
  • Exactness and finite-sample differences: Under the normal linear regression model, the full-data, cumulative, aggregated, and online estimators are exactly identical.The recursive update is also equivalent to the aggregated combination at the terminal update, whereas estimating-equation estimators generally are not identical in finite samples.
  • Asymptotic results: Under regularity conditions and a partition number K = O(n^γ) with an appropriately bounded growth rate, the cumulative estimator is consistent.The result also holds for unequal subset sizes when consecutive subset-size ratios remain bounded.

4 Simulations

The simulations evaluate individual and global outlier tests under normal and skew-t errors using data with a single potentially contaminated subset. Outlier strength is varied from none to large.

  • Simulation design: The simulations evaluate individual and global outlier tests for data streams containing a single subset that may include outliers.The individual test is assessed separately from the global tests.
  • Simulation design: Outlier magnitude is controlled by δ ∈ {0, 2, 4, 6}, representing no, small, medium, and large outliers.The baseline regression coefficient vector is β = (1, 2, 3, 4, 5)′ with normally distributed predictor variables.
  • Simulation design: The error distributions include independent N(0, 1) errors and standardized independent skew-t errors with ν = 3 and γ = 1.5.The skew-t setting is used for evaluating the global outlier tests.

1. To be precise, we use the skew t density

The simulations evaluate online outlier tests and estimating-equation estimators under varying data-stream configurations. Predictive residual tests improve power while maintaining low false-positive counts, and CUEE is more stable than CEE as blocks increase.

  • Outlier tests: Under normal errors, both the F test and asymptotic F test control Type I error across scenarios.
  • Outlier tests: Power increases with outlier strength, number of outliers, and the proportion of preceding outlier-free data, with diminishing gains at large denominator degrees of freedom.
  • Outlier tests: The normality-based F test is more powerful than the asymptotic F test when errors are normally distributed.
  • Outlier tests: Under nonnormal errors, the normality-based F test has severely inflated Type I error, whereas the asymptotic F test maintains appropriate size.
  • Outlier tests: The predictive residual test has higher power and low false-positive counts than the externally studentized residual test, whose false positives are lower but false negatives higher.
  • Estimating-equation estimators: As the number of blocks increases and block size decreases, CEE RMSE increases rapidly while CUEE RMSE remains relatively stable.

5 Data Analysis

The airline analysis applies online-updating estimators and predictive residual tests to large flight datasets. CUEE estimates most closely match full-data EE estimates, while the outlier tests largely agree but differ in rejection frequency.

  • Data and setup: The analysis uses 120,748,239 complete airline observations and online-updating subsets of 50,000 observations, with a smaller final subset.
  • Data and setup: A blockwise detection mechanism combines subsequent blocks when rare extreme-distance observations create potential data-separation problems.
  • Logistic regression: All three methods find that covariates other than extreme distance are highly associated with late arrival, while extreme distance is not associated (p = 0.613).
  • Estimator comparison: CUEE regression coefficients tend to be most similar to Revolution R's full-data EE estimates, whereas CEE differs most for the intercept and binary-covariate coefficients.
  • Outlier detection: The normality-based and asymptotic F tests agree on outlier presence in 96% of 12,803 subsets.
  • Outlier detection: Among disagreements, the normality-based F test identifies 488 additional outlier-containing subsets, compared with 23 identified only by the asymptotic F test.

6 Discussion

The paper develops online-updating inference for linear models and estimating equations, including predictive-residual outlier tests and a new estimating-equation estimator. Its methods target sequential data without historical-data access, while leaving information-retention and high-dimensional inference questions open.

  • Online-updating framework: Online-updating algorithms and inferences apply to linear models and estimating equations for sequential data without access to historical data.The framework includes updated coefficient and variance inferences.
  • Linear-model inference: Predictive-residual tests for linear models were more powerful than tests using only the current dataset in simulations.The tests support outlier detection and model assessment within the linear-model setting.
  • Estimating-equation inference: The proposed CUEE estimator borrows information from previous stream datasets and was less biased than the AEE/CEE estimator in finite-sample simulations.Both estimators were shown to be asymptotically consistent, and the proposed estimator outperformed competing approaches in bias and mean squared error.
  • Scope and open questions: The methods are designed for small to moderate covariate dimensionality p and large N; penalized-parameter inference remains challenging.Variance estimation for penalized parameters is identified as an area of future research.
  • Information retention: Under the normal linear regression model, the proposed scheme does not lose information for inference on β when the full design matrix has full rank.The stated sufficient and complete statistics support inference for β and σ2.
  • Scope and open questions: In the estimating-equation setting, some information is lost, and the amount required for specific inferences remains an open question.The paper also leaves handling detected outliers for future research and notes that sequential testing requires α-level adjustment for multiple testing.

Statistical Inference in the Big Data Setting

The paper develops Bayesian-motivated online updates for linear-model estimates, variances, tests, ANOVA quantities, and goodness-of-fit diagnostics while retaining only low-dimensional summaries.

  • Online updating: Online updates for regression coefficients and error variance require storing only a few low-dimensional quantities computed within the current subset.The stored quantities need not include historical subset-level values.
  • Online updating: The Bayesian formulation yields posterior distributions with the same form after each cumulative data update.The posterior is parameterized by updated quantities such as µ_k, V_k, ν_k, and τ_k.
  • Statistical inference: At the final update, the Bayesian-derived coefficient and mean-square-error updates coincide with the paper’s earlier online formulas.This establishes equivalence between the Bayesian online-update derivation and the proposed linear-model updates.
  • Statistical inference: Online-updated inference includes coefficient t-tests, ANOVA tables, general linear hypothesis F-tests, and coefficients of determination.For coefficient t-tests, only current V_k, β̂_k, N_k, and MSE_k are needed.

C: Proof of Proposition 2.4

The proposition’s proof uses bounded-design and independent-error conditions, then applies matrix decompositions to address computation when subset sizes are large.

  • Proof strategy: Bounded design-matrix elements are combined with Chebyshev’s inequality to control the relevant quantities.The argument treats each covariate column separately under the stated boundedness condition.
  • Assumptions: The proof assumes independent and identically distributed errors and independence across subgroups.These assumptions support the probabilistic bounds and central-limit arguments used in the proposition.
  • Computation: For large subset sizes, a Cholesky decomposition followed by singular value decomposition provides an alternative computational route.The SVD diagonalizes the relevant expression, making the remaining inverse calculation diagonal.

E: Proof of Theorem 3.2

The theorem proof establishes control of the online estimator through matrix-norm bounds, induction, and remainder-term control under the stated regularity conditions.

  • Matrix bounds: The proof begins with a matrix-norm definition and eigenvalue inequalities for positive definite matrices.These facts provide lower bounds used in subsequent probability and norm estimates.
  • Inductive step: Induction combines the previous online estimate with the current subset estimate through bounded matrix operators.The proof explicitly bounds the operators multiplying the previous and current contributions.
  • Conclusion: Under the stated growth condition K = O(n^γ), the proof controls the cumulative estimator and its remainder term.The remainder is bounded quadratically in the difference between the batch and online estimates.

F: Proof of Proposition 3.5

The proposition analyzes online IRLS when subset design matrices are rank deficient, showing convergence and estimator invariance despite nonunique generalized-inverse representations.

  • Rank deficiency: When X_k lacks full column rank, the IRLS coefficient vector is not unique because it uses a generalized inverse.The proof isolates a full-rank column submatrix to analyze convergence.
  • Convergence: Under the stated conditions, the IRLS iterates converge for a special generalized inverse and then for any generalized inverse.The argument transfers convergence using the invariance properties established in the proof.
  • Invariance: The combined estimator and the relevant fitted quantities are invariant to the choice of generalized inverse.This includes the accumulated estimator and the product involving the current weighted design matrix.
Loading 1505.06354v1…