Source-linked AI summary

Propagation of outliers in multivariate data

Fatemah Alqallaf, Stefan Van Aelst, Victor J. Yohai, Ruben H. Zamar

arXiv:0903.0447v1math.ST

TL;DR

The paper asks how robust multivariate location estimates behave when contamination occurs independently across variables, rather than by whole cases. It develops influence functions for flexible contamination models and finds that standard affine equivariant high-breakdown estimators can propagate outliers and have poor breakdown behavior under componentwise contamination in high dimensions.

  • Problem

    High-dimensional data challenge contamination models that assume most whole cases are clean, while componentwise contamination can make perfectly observed cases increasingly rare.

  • Method

    The paper defines and derives generalized influence functions for robust multivariate location under contamination models including FDCM and FICM.

  • Results

    Theorem 1 shows that under FICM the breakdown point of equivariant location estimates that are δ-consistent at infinity is at most 1−(1/2−δ)1/d, tending to 0 when δ is independent of d.

  • Takeaways & Limitations

    Coordinatewise procedures can protect against outlier propagation, whereas affine equivariant procedures are needed for structural outliers, creating a robustness trade-off.

  • Takeaways & Limitations

    The paper concludes that sufficiently robust estimators against all outlier types are intrinsically difficult in high dimensions, requiring trade-offs among robustness properties.

Abstract

from arXiv · show

We investigate the performance of robust estimates of multivariate location under nonstandard data contamination models such as componentwise outliers (i.e., contamination in each variable is independent from the other variables). This model brings up a possible new source of statistical error that we call "propagation of outliers." This source of error is unusual in the sense that it is generated by the data processing itself and takes place after the data has been collected. We define and derive the influence function of robust multivariate location estimates under flexible contamination models and use it to investigate the effect of propagation of outliers. Furthermore, we show that standard high-breakdown affine equivariant estimators propagate outliers and therefore show poor breakdown behavior under componentwise contamination when the dimension $d$ is high.

1. Introduction.

Robust multivariate methods are built around contamination models, especially Tukey–Huber’s mixture model. The paper motivates alternatives for high-dimensional data, where whole-case downweighting and clean-case assumptions become problematic.

  • Tukey–Huber models data as a dominant clean component mixed with an unspecified minority contamination component.Robust analysis targets the dominant component by identifying and downweighting outlying cases.
  • Robust procedures and concepts such as influence functions and breakdown points were strongly shaped by the Tukey–Huber model.
  • The paper introduces contamination models, derives influence functions under them, studies outlier propagation, and analyzes breakdown under componentwise contamination.
  • Classical multivariate contamination assumes a majority of cases is uncontaminated, but this assumption becomes restrictive as dimension increases.The fraction of perfectly observed cases can become small in high-dimensional settings.

2. Alternative contamination frameworks.

The paper contrasts fully dependent and fully independent contamination models. Under independent componentwise contamination, the chance of a completely clean case decreases rapidly with dimension, creating potential for outlier propagation and motivating a generalized influence function.

  • FDCM: Under FDCM, contamination indicators are fully dependent, preserving a dimension-independent majority of perfectly observed cases.FDCM also preserves the percentage of contaminated cases under affine equivariant transformations.
  • FICM: Under FICM, independent contamination makes the probability of a perfectly observed case equal to (1−ǫ)^d.With ǫ = 0.05 this falls below 1/2 at d ≥14, and with ǫ = 0.01 at d ≥69.
  • FICM: FICM is not affine equivariant, so linear combinations can contain a lower fraction of clean values than individual columns.This creates potential for outlier propagation during data processing.
  • Generalized influence function: The paper generalizes the influence function because existing multivariate definitions covered only the classical FDCM.

3. The influence function.

The paper defines influence functions for robust multivariate location under flexible contamination configurations and illustrates how contamination structure changes estimator sensitivity. Under FICM, influence functions need not redescend and can depend nonlinearly on correlation.

  • Definition and derivation: The generalized influence function is defined for contamination configuration distributions satisfying exchangeable pattern-probability conditions that include FDCM and FICM.
  • Definition and derivation: The derivation differentiates the estimating functional with respect to contamination at zero, assuming Fisher consistency and differentiability of the scatter functional.
  • Illustrations: Under FDCM, Tukey bisquare influence functions are fully redescending, whereas under FICM they are not.A vanishingly small fraction of large coordinatewise contamination can therefore retain influence on the location M-estimate.
  • Illustrations: Under FICM, changing correlation changes the influence function beyond a linear transformation by Σ1/2.For zero correlation, components are influenced mainly by matching-component contamination; with correlation, both components contribute.
  • Illustrations: The coordinatewise estimator is influenced only by contamination in the corresponding component, regardless of correlation.

4. Propagation of outliers.

FICM is not affine equivariant: linear transformations can combine independently contaminated columns so that contamination propagates into transformed components. A two-dimensional example shows clean-majority marginals becoming majority-contaminated after transformation, while higher-dimensional contamination can exceed 90% per cell.

  • Affine equivariance: Under FDCM, affine transformations preserve the contamination probability, unlike the independent contamination model.FICM is generally not preserved after an invertible linear transformation unless the transformation matrix is diagonal.
  • Affine equivariance: Outlier propagation occurs because affine transformations linearly combine columns and destroy their independent contamination structure.Each original column may contain an average contamination fraction ǫ, but transformed columns need not retain that fraction.
  • Two-dimensional illustration: With n = 20 and 30% independent contamination in each component, transformed components contained majority-contaminated cells and groups of 49%, 42%, and 9%.The transformed medians no longer reflected the location of the clean data.
  • High-dimensional consequence: In a 15-dimensional example with 15% contamination in each column, linear transformation produced more than 90% contamination probability for each cell.The passage notes that affine equivariant robust estimators treat the original and transformed data as equivalent, with devastating performance consequences in this setting.
  • Scope: The analysis assumes equal marginal contamination probabilities across components, although the results extend with modifications to unequal probabilities.The maintained assumption is P(B1 = 1) = ··· = P(Bd = 1) = ǫ.

5. Affine equivariance and independent contamination.

Under independent componentwise contamination, a small contamination fraction in each variable can make affine-equivariant robust location estimators break down as dimension increases. The section formalizes this behavior through δ-consistency at infinity and applies it to several high-breakdown estimators.

  • Affine equivariance and independent contamination.: An estimator is δ-consistent at infinity when at least 1/2 + δ of the mass moves to infinity in every coordinate and all estimated coordinates then diverge.This property is used to establish upper bounds on breakdown under componentwise contamination.
  • Affine equivariance and independent contamination.: The FICM contamination neighborhood consists of distributions formed by independently replacing components with arbitrary contaminated values.The contamination indicators are independent Bernoulli variables with probability ǫ in each component.
  • Affine equivariance and independent contamination.: Scatter estimates also break down whenever the multivariate location estimates used to center the data break down.The theorem therefore has direct implications for companion scatter procedures.
  • Affine equivariance and independent contamination.: Independent contamination in each component can cause breakdown when the dimension is moderately large, despite a small per-variable contamination fraction.The breakdown mechanism arises because contamination can accumulate across variables and affect many cases.
  • Affine equivariance and independent contamination.: For d ≥10, a small amount of contamination in each variable suffices to break down several affine-equivariant high-breakdown estimators, including MCD, S-estimators, τ-estimators, and minimum volume ellipsoid estimators.The MCD and S-estimator arguments establish δ-consistency at infinity, which places them under the theorem’s breakdown bound.

6. Concluding remarks.

The paper distinguishes contamination models by how clean cases and cells are distributed, showing that independent contamination creates a trade-off between protection from propagated and structural outliers.

  • Under PSICM, the influence function is the average of the influence functions under FICM and FDCM.
  • PCICM and PSICM both contain independent componentwise contamination, so outlier propagation occurs in both models.The effect is more devastating in PSICM because no clean cases are guaranteed.
  • High-dimensional robustness against all outlier types is intrinsically difficult because affine-equivariant estimators’ breakdown point under FICM decreases with dimension.The paper concludes that desirable robustness properties require a trade-off.
  • Coordinatewise procedures such as the median protect against outlier propagation, whereas affine-equivariant methods handle structural outliers.Applying affine-equivariant methods to column subsets trades protection against structural outliers against protection from propagation.

APPENDIX

For elliptically symmetric distributions, the estimating function vanishes at the target location for every positive definite scatter matrix.

  • g(H0,µ0,Σ) = 0 for every positive definite Σ when H0 is elliptically symmetric.

A.1. Derivation of (9).

The appendix derives the target influence-function expression using spherical symmetry and a constant independent of the location and scatter parameters.

  • The transformed variable w has a spherical distribution, which is used in deriving the influence-function expression.
  • The derivation obtains the result from equations (19) and (20).
  • Aψ is independent of µ0 and Σ0, allowing equation (9) to follow from equation (18).
  • The proof assumes a distribution G0 with finite first moment and constructs Hh by mixing H0 and Gh with weights 1/2 − δ and 1/2 + δ.

A.2. Proof of Lemma 1.

The proof establishes divergence by showing one term tends to positive infinity while the remaining term stays uniformly bounded.

  • Scale equivariance of T(H) is used in the proof.
  • The transformed distribution converges weakly to a point mass at zero as h tends to infinity.
  • The proof analyzes the set Ah of vectors whose components are all at least h.
  • The proof separately bounds the first and second terms on the right-hand side of equation (22).
  • The second term tends to zero because G0 has finite first moments.
  • The first term tends to +∞ while the second term remains uniformly bounded.
  • The weights w(x,H) are nonnegative, and u(t) is bounded under assumption (A2).

A.3. Proof of Lemma 2.

The proof selects a threshold t0 and corresponding radius r0 to establish positive lower bounds for the weighting function and probability of Bt0. These bounds verify the assumptions needed for Lemma 1.

  • The proof introduces Bt and uses preceding equations (13) and (14) to continue the bound construction.
  • The threshold satisfies t0 = 1/(1 + 2δ0), with 1/2 ≤ t0 < 1 when 0 < δ0 ≤ 1/2.
  • Setting r0 = ρ−1(t0) and ζ = u(r0) gives ζ > 0 and establishes κ ≥ ζP(Bt0) ≥ ζ/4.
  • The resulting bounds on w(x,H) verify assumptions (i)–(iii) of Lemma 1 together with (29).
Loading 0903.0447v1…