Source-linked AI summary

On a minimum distance procedure for threshold selection in tail analysis

Holger Drees, Anja Janßen, Sidney I. Resnick, Tiandong Wang

arXiv:1811.06433v2math.STstat.MEstat.TH

TL;DR

The paper examines how the widely used MDSP selects power-law thresholds and estimates tail indices, addressing limited theoretical justification for the procedure. Through asymptotic analysis and simulations, it finds that MDSP selections can produce non-normal estimators and inflated Hill-estimator variance and RMSE. The method remains convenient but should be applied cautiously because its efficiency and behavior depend on the data-generating setting.

  • Problem

    The MDSP is widely used to select power-law thresholds, but its theoretical performance and inferential properties have not been sufficiently analyzed.

  • Method

    The paper analyzes MDSP asymptotics for iid Pareto-like data and deviations from Pareto behavior, and studies dependent preferential-attachment network simulations.

  • Results

    The MDSP often selects unsuitable k*_n values, yielding a non-normal tail-index limit and increased Hill-estimator variance and RMSE; network performance depends on model parameters.

  • Takeaways & Limitations

    MDSP estimates can be reasonable in some network settings, but the procedure does not generally achieve minimal asymptotic RMSE and should be applied with caution.

  • Takeaways & Limitations

    For preferential-attachment network data, theoretical MDSP analysis is unavailable because iid Brownian-motion embedding techniques do not extend to the network case.

Abstract

from arXiv · show

Power-law distributions have been widely observed in different areas of scientific research. Practical estimation issues include how to select a threshold above which observations follow a power-law distribution and then how to estimate the power-law tail index. A minimum distance selection procedure (MDSP) is proposed in Clauset et al. (2009) and has been widely adopted in practice, especially in the analyses of social networks. However, theoretical justifications for this selection procedure remain scant. In this paper, we study the asymptotic behavior of the selected threshold and the corresponding power-law index given by the MDSP. We find that the MDSP tends to choose too high a threshold level and leads to Hill estimates with large variances and root mean squared errors for simulated data with Pareto-like tails.

1. Introduction.

Power-law tail-index estimation depends critically on selecting a threshold, but the widely used MDSP lacks sufficient theoretical analysis. This paper studies its asymptotic behavior and finds that it can select unsuitable thresholds, increasing Hill-estimator uncertainty and error.

  • 1. Introduction.: Threshold choice governs a bias–variance trade-off: thresholds that are too low create bias from nonlinearity, whereas thresholds that are too high use too little data and inflate variance.The selected threshold also determines the order statistic used by the Hill estimator.
  • 1. Introduction.: The MDSP chooses the cutoff minimizing the Kolmogorov-Smirnov distance between empirical exceedances and a fitted Pareto tail, then applies the Hill estimator to estimate the tail index.The procedure was proposed by Clauset et al. and is widely used in social-network and income-distribution analyses.
  • 1. Introduction.: The paper addresses a theoretical gap because the MDSP is widely applied, yet its performance had not been mathematically analyzed even for iid repeated-sampling models.The analysis considers exact Pareto data, deviations below a Pareto threshold, and dependent observations from preferential-attachment networks.
  • 1.2. Summary.: The MDSP tail-index estimator has a non-normal, difficult-to-calculate limit distribution, complicating the construction of reliable confidence intervals.This result directly limits straightforward inferential procedures for MDSP-based estimates.
  • 1.2. Summary.: The procedure often selects too small a k*_n, increasing the Hill estimator’s variance and RMSE relative to a choice minimizing asymptotic mean squared error.For preferential-attachment simulations, performance depends strongly on model parameters; some settings perform well, whereas others select k*_n that is too large.
  • 1.2. Summary.: MDSP is convenient and often reasonable in network simulations, but it does not achieve minimal asymptotic RMSE and lacks theoretical analysis for node-based random-graph data.The authors therefore recommend applying it with caution, especially when confidence intervals or efficiency are important.

2. The Pareto case.

In the Pareto case, the MDSP’s selected threshold has a nondegenerate asymptotic distribution, making the resulting Hill estimator non-normal and often suboptimal. Simulations show that threshold selection can substantially increase estimator variance, RMSE, and confidence-interval error.

  • Asymptotic theory: The MDSP-selected threshold k*_n/n converges to a random minimizer T of a Gaussian-process criterion, under a unique-minimum condition.Theorem 2.1 establishes the underlying process convergence, while Corollary 2.2 gives the joint asymptotic distribution of the selected k and Hill estimator.
  • Asymptotic theory: Because T is random, the MDSP Hill estimator has a non-normal limiting distribution rather than the normal limit obtained with deterministic k.The resulting limit distribution exhibits strong deviation from normality and heavier tails than the estimator using the top n order statistics.
  • Simulation evidence: In the limit, k*_n is at least 0.9n in about 3/5 of cases, but falls below 3n/4 with probability about 1/4 and below n/2 with probability 8.4%.The probability that k*_n is below n/3 is about 2%, so substantially smaller selections remain possible.
  • Simulation evidence: Relative to the minimal-RMSE deterministic choice, the MDSP estimator has about 88% higher limiting variance and 37% higher limiting RMSE.For finite samples, RMSE increases are 33% for n = 100, 36% for n = 1,000, and 37% for n = 10,000.

3. Deviations from the Pareto model.

The paper extends its MDSP analysis beyond pure Pareto tails to distributions with structural deviations below or near a threshold. These deviations alter threshold selection behavior, but the MDSP still often underestimates the break point and increases Hill-estimator error.

  • Asymptotic analysis: The asymptotic theory assumes a continuous deviation function with a non-vanishing right derivative at t0, creating a smooth but detectable tail change.Theorem 3.1 gives uniform approximations below t0, divergence of the distance above t0, and a limiting characterization near the break point.
  • Asymptotic analysis: Near t0, the Kolmogorov-Smirnov distance behaves on a finer n^1/2 scale, complicating conclusions about the exact threshold selected by the MDSP.The limiting process may have multiple minimizers, and finite-sample approximations are ambiguous near t0.
  • Piecewise Pareto model: For a piecewise Pareto model, the limiting threshold distribution has positive mass at t0, shifting MDSP selections toward the true break point and improving performance.Finite samples smear this mass around t0; the approximation is poor for n = 200, improves for n = 2,000, and is almost perfect for n = 20,000.
  • Different deviation models: The MDSP’s threshold distribution depends strongly on how the tail departs from Pareto behavior near t0: smooth, continuous, and discontinuous deviations yield different detection accuracy.The smooth model detects the structural change slowly, while the discontinuous model is less accurate than the continuous but non-differentiable model Hc.
  • Beyond exact Pareto tails: When the MDSP is restricted to an intermediate sequence under regularly varying tails, it still fails to select a threshold that asymptotically minimizes Hill-estimator RMSE.Simulations indicate that sequential and bootstrap procedures usually outperform this restricted MDSP in RMSE.

4. Linear preferential attachment (PA) networks.

The paper models directed linear preferential-attachment networks and establishes their limiting degree-count and power-law behavior. Simulations show that MDSP performance varies substantially across parameter settings, depending on how sensitive Hill estimation is to threshold choice.

  • 4.1. The linear PA model.: The directed linear PA model grows by adding edges through α-, β-, and γ-schemes, with degree-dependent attachment controlled by δin and δout.The construction permits self-loops in the β-scheme and multiple edges, while the proportion of self-loops vanishes asymptotically.
  • 4.2. Power law of degree distributions.: Empirical degree frequencies converge to a limiting distribution, whose marginal in- and out-degree tails follow power laws under the stated parameter conditions.The in-degree tail satisfies fin_i ∼ Cin i^−(1+αin) when αδin + γ > 0, while the out-degree tail has the analogous condition γδout + α > 0.
  • 4.3. Simulations.: For the Baidu parameter setting, MDSP’s Hill-estimator RMSE is only 6.8% above the minimum, despite selected k values spread across [500, 3500].The minimum RMSE occurs at k = 1187, and a broad near-optimal range makes the estimator relatively insensitive to threshold choice.
  • 4.3. Simulations.: The Facebook parameter setting yields an RMSE about 50.0% above the minimum, because MDSP selects k values far above the RMSE-minimizing k = 523.The selected k distribution has estimated mean 1857, while RMSE rises rapidly away from the optimum.
  • 4.3. Simulations.: Across simulations, MDSP can perform well for linear PA models when Hill estimates are insensitive to threshold choice, a behavior reported as more common for network data than many iid models.The paper therefore concludes that MDSP often works well for linear PA models under suitable parameter choices, although performance is mixed.

5. Conclusions.

The conclusions characterize MDSP as asymptotically and finitely sample-dependent, with threshold selection and Hill-estimator accuracy varying by tail-transition structure and data dependence. For preferential-attachment simulations, MDSP can select unsuitable thresholds, although estimator efficiency may remain acceptable when RMSE is flat across thresholds.

  • 5. Conclusions.: MDSP often selects too small a k at structural breaks or too late a change point under smooth transitions, increasing Hill-estimator variance and RMSE.The selected threshold need not concentrate at one point, and both abrupt and smooth departures from the Pareto tail can impair selection.
  • 5. Conclusions.: For dependent preferential-attachment data, a spread-out selected k does not necessarily reduce efficiency when bias and variance balance over a broad k range.The Hill estimator can be relatively insensitive to k when increasing bias is offset by decreasing variance.
  • 5. Conclusions.: For fixed-size linear PA networks, the degree distribution changes with sample size and has a power tail only asymptotically, so the Hill estimator’s RMSE interpretation is not straightforward.Unlike iid regularly varying observations, finite networks have bounded degrees and no strict finite-sample tail index.

Appendix A. Proofs.

The appendix proves asymptotic results for the MDSP by approximating empirical processes with exponential representations and Brownian motion. It derives threshold-selection behavior and corresponding Hill-estimator asymptotics under Pareto and non-Pareto settings.

  • Appendix A. Proofs.: The proofs represent uniform order statistics through normalized exponential sums and then use quantile transformations to analyze the empirical tail process.The exponential representation provides the probabilistic starting point for the asymptotic arguments.
  • Appendix A. Proofs.: The non-Pareto proof analyzes the transition function H, including its continuity, vanishing region, and behavior near t0, to establish the relevant distance-process limits.These arguments yield the assertions for thresholds above and within an n^1/2 neighborhood of t0.
  • Appendix A. Proofs.: Brownian-motion approximations, moduli of continuity, and the law of the iterated logarithm control suprema and discretization errors uniformly over threshold ranges.The proofs replace discrete maxima with interval suprema and sums with integrals while bounding approximation errors.
  • Appendix A. Proofs.: Under a unique minimum of the limiting distance process, the selected threshold lies beyond arbitrary kn = o(n) with probability tending to one, and its asymptotics determine the Hill estimator’s limit behavior.The proof establishes that n/k*_n is stochastically bounded and then derives the estimator asymptotics from the process approximations.
  • Appendix A. Proofs.: For thresholds below a non-Pareto transition point, the asymptotic analysis differs from the local n^1/2 neighborhood around that point.The appendix separately analyzes k ∼ nt for t < t0 and |k − nt0| = O(n^1/2), reflecting distinct limiting behaviors.
Loading 1811.06433v2…