Source-linked AI summary

Vocabulary Growth Fundamentals: Bernstein Functions and Hausdorff Sequences

Łukasz Dębowski

arXiv:2608.29449v1math.PRcs.CL

TL;DR

The paper addresses how vocabulary growth and lexicon structure should be modeled as linguistic data scale and as simple memoryless models face broader stochastic settings. It surveys Bernstein functions and Hausdorff sequences, connects them with stochastic sources and hapax-rate models, and proves a key result for the logistic model. The paper also examines the framework’s scope through stationary and Weibull renewal processes.

  • Problem

    The paper addresses the need for a mathematical framework for vocabulary growth, including modeling heterogeneous lexicon components and extending beyond memoryless sources.

  • Method

    The paper conducts a theoretical survey integrating Bernstein functions, Hausdorff sequences, stochastic-process models, inverse Laplace methods, and parametric hapax-rate models.

  • Results

    The paper proves that the logistic hapax-rate model has non-negative spectrum elements and thus defines a Bernstein function.

  • Takeaways & Limitations

    Vocabulary growth admits a richer mathematical structure connecting marginal lexical statistics with Bernstein functions, Hausdorff sequences, and stochastic processes.

  • Takeaways & Limitations

    The paper does not analyze empirical data and studies a theoretical framework whose mixture representations are not generally unique.

Abstract

from arXiv · show

We survey the theory of vocabulary growth founded in the setting of stochastic processes. In particular, we model the expected number of types through Bernstein functions and Hausdorff sequences. These classes of mathematical objects, defined by alternating signs of their derivatives or differences, can be related to continuous-time Poisson point processes and discrete-time IID processes, respectively. Building on previous accounts of the vocabulary growth, we integrate the broader theories of Bernstein functions and Hausdorff sequences and connect them with recently developed hapax rate models. In particular, we prove that the logistic hapax rate model has a non-negative spectrum and hence it defines a Bernstein function, thereby solving an earlier posed problem. We also analyze the limitations of the Bernstein--Hausdorff theory of the vocabulary growth by considering its generalizations under stationary and Weibull renewal processes.

1 Introduction

The paper frames vocabulary growth as a mathematical problem involving large-scale linguistic data, heterogeneous lexicon components, and stochastic models. It integrates Bernstein functions, Hausdorff sequences, and hapax-rate models, while testing the framework beyond memoryless sources.

  • Motivation: Trillion-token corpora create an empirical setting for testing theoretical models of word-frequency distributions and raise questions about what quantitative linguistics should model.
  • Motivation: Vocabulary growth supports interpolation, extrapolation, and mixture identification, including decomposing the lexicon into components with distinct statistical behavior.
  • Motivation: Extreme text collections can produce two Zipf regimes and a U-shaped hapax rate, motivating models that combine decreasing and nuisance components.
  • Framework: Bernstein functions and Hausdorff sequences provide generic models for expected type counts through alternating signs of derivatives or differences, linked to Poisson and IID processes.
  • Contributions: The paper proves that the logistic hapax-rate model has non-negative spectrum elements and therefore yields a Bernstein function.
  • Scope: The paper is a theoretical survey that reviews mathematical theory rather than analyzing empirical data.

2 Stochastic setting

The paper models word-token generation with four stochastic sources and derives corresponding counting measures, spectra, and type counts. These sources range from stationary and IID sequences to Weibull renewal and Poisson point processes.

  • Counting measures: The framework counts tokens, types, and spectrum elements, with hapaxes defined as types occurring exactly once.
  • Invariant and multinomial sources: The invariant source is a stationary sequence of word types, while the multinomial source is an IID sequence of words at integer positions.
  • Weibull source: The Weibull source uses independent IID interword intervals, with shape parameter α_w representing burstiness in quantitative linguistics.
  • Poisson source: The Poisson source is a symmetric special case of the Weibull source and consists of independent Poisson point processes.
  • Poisson source: Poisson counting measures have independent counts on disjoint intervals and Poisson-distributed interval counts.
  • Source relationships: For large intervals and small word probabilities, the Poisson distribution approximates the binomial distribution, making the Poisson source useful for long samples.

3 Poisson sources

The Poisson-source framework characterizes vocabulary growth with Bernstein functions, whose derivative structure and measure representations connect expected type counts, spectra, and hapax rates. It also establishes mixture-identification results while marking limits of the model family and simple vocabulary-growth models.

  • 3 Poisson sources: Poisson sources allow exact computation of expected vocabulary-growth quantities and connect expected spectra to derivatives of the expected number of types.The framework treats non-negativity and alternating derivative signs as the defining structure behind Bernstein-function models.
  • 3 Poisson sources: A Bernstein function is non-negative with derivatives having alternating signs, and its derivative admits a Laplace-transform representation through a unique non-negative Bernstein measure.Standard Bernstein functions additionally satisfy v(0+)=0 and have Bernstein measure of total mass one.
  • 3 Poisson sources: Bernstein functions are mixtures of elementary functions t ↦ p^-1(1 − e^-pt), providing a measure-based representation of vocabulary-growth models.The associated Bernstein measure may be discrete or continuous, and its support can extend beyond [0,1] for ultra-frequent types.
  • 3 Poisson sources: Hapax rates correspond one-to-one with standard Bernstein functions under the stated regularity condition, while fixing values at lengths t ∈ {0,1} yields a correspondence with Bernstein functions more generally.The relevant condition holds exactly when the right derivative at zero is finite.
  • 3 Poisson sources: Mixtures of power laws are identifiable when the mixture exists, but broader mixtures with varying parameter c remain unresolved and general mixture measures need not be unique.Uniqueness is guaranteed for some convex representations, including Bernstein functions, but not in general.
  • 3 Poisson sources: The constant model is visibly incorrect for real textual data, whereas affine transformation yields a saturation model and the model in Example 12 is bounded and standard for γ ∈ (0,1].The paper also shows that the formulas in (74)–(76) define a singular and expansive Bernstein function for t ≥ 1.

4 Multinomial sources

The multinomial-source analysis characterizes expected vocabulary growth through differences, Hausdorff sequences, their measures, and hapax rates. It shows that only standard Hausdorff sequences satisfy both consistency conditions required for word-frequency models.

  • Expected multinomial-source quantities can be computed exactly, and the expected spectrum can be obtained from the expected number of types by taking differences.
  • Hausdorff sequences: Hausdorff sequences are non-negative sequences whose subsequent differences have alternating signs; standard sequences additionally satisfy v(0)=0 and Δv(1)=1.
  • Hausdorff measures: Hausdorff sequences can be represented as mixtures of elementary sequences indexed by p, including discrete and continuous measure cases.
  • Hausdorff measures: Every Hausdorff sequence corresponds uniquely to a non-negative Hausdorff measure on [0,1], and standard sequences correspond to probability measures.
  • Consistency: Only standard Hausdorff sequences satisfy both consistency conditions and can therefore be applied as models of word-frequency distributions.
  • Hapax rates: Not every sequence-valued hapax rate is valid, but each valid hapax rate determines exactly one standard Hausdorff sequence.

5 Invariant sources

For stationary sources, expected vocabulary growth retains positivity, growth, and concavity, but need not remain completely alternating. Dependent processes can cause higher differences to oscillate, while hapax and dis legomenon counts remain bounded by vocabulary-growth differences.

  • Entropy: The block Shannon entropy is also positive, growing, and concave, with H(0)=0, while H(n)/n decreases.
  • Limitations: For some dependent processes, including periodic processes, the second differences of entropy and expected vocabulary can oscillate with n.
  • Spectrum bounds: Stationary-process formulas provide an upper bound for expected hapaxes and dis legomena in terms of differences of the total expected number of types.

6 Weibull sources

The Weibull-source analysis extends vocabulary-growth theory beyond IID and stationary settings. Exact results are limited, but under specified burstiness conditions the expected vocabulary growth is a Bernstein function and admits an analogue of an earlier formula.

  • Exact results for Weibull sources are limited, and the section is presented as an invitation to future research.
  • In bursty generating processes, hapaxes tend to occur near the edge of the text.
  • For Weibull sources with αw≤1, the expected vocabulary-growth function is a Bernstein function.
  • If burstiness satisfies a uniform lower bound α<αw≤1 across word types, an analogue of formula (20) follows.

7 Conclusion

The paper surveys Bernstein functions and Hausdorff sequences as a richer mathematical framework for marginal lexical statistics. It connects vocabulary-growth models with stochastic processes and develops consequences for word-frequency analysis.

  • The paper presents Bernstein functions and Hausdorff sequences as a mathematical theory for expected vocabulary growth and marginal lexical statistics.
  • Its perspective connects counting distinct words with broader mathematical structures underlying Zipf’s and Heaps’ laws.

Funding

The paper received no funding from funding agencies.

  • The work received no funding from funding agencies.

AI usage

The paper describes a dialogue with ChatGPT as the origin of its general idea and used the chatbot to search for established mathematical notions and references.

  • The paper’s general idea originated from a dialogue with ChatGPT.
  • ChatGPT was used as an intelligent search engine for established mathematical notions.
  • The chatbot drew attention to Bernstein functions, the Hausdorff moment theorem, stable densities, and related literature references.

A Proofs for Poisson sources

The appendix proves propositions using a mixture of cited results, explicit derivations, and arguments from Bernstein-function theory. A central result establishes that the logistic model has a non-negative spectrum.

  • Proof strategy: Proofs of Propositions 3–4 and 8–9 rely on established references, while several other propositions are demonstrated explicitly.The appendix cites Schilling et al., Dębowski (2025a), and Baayen (2001), and reproduces some known arguments for completeness.
  • Bernstein-function consequence: A Bernstein function cannot be both closed and expansive at all points t > 0.The contradiction uses non-negative spectrum elements and shows that some spectrum component has a negative derivative.
  • Logistic model: The logistic model has non-negative spectrum elements because its recursively defined coefficients are non-negative.Induction shows that the coefficients W_kj are non-negative; the associated polynomials are also growing and convex for x ≥ 0.

A.14 Proof of Proposition 19

The proof section derives a density for the Bernstein measure by differentiating an integral representation, interchanging integrals, and invoking uniqueness. It also situates these results within prior proposition proofs.

  • Derivation: The proof computes the derivative of v(t) and uses its integral representation.
  • Bernstein measure: Interchanging the order of integration and using uniqueness of the Bernstein measure yields its density.
  • Proof provenance: The appendix refers to Akhiezer for Proposition 22 and reproduces proofs previously derived in Dębowski (2025c).

C Proofs for invariant sources

This section establishes the stated propositions for invariant sources, noting that Propositions 27–29 follow earlier work while Propositions 30 and 31 appear to be new. The claims are obtained by propagating the relevant pattern inductively and summing contributions over types.

  • Propositions 27–29 are proved in Dębowski (2025c, Appendix A.3), whereas Proposition 30 appears to be new.
  • Proposition 31 also appears to be new.
  • The proof introduces notation, suppresses the type subscript for a chosen type, and uses an inductively propagating pattern with nonnegative coefficients w_kj and w_kk = α_k.
  • The propositions follow by summing the contributions ρ(t) and ρ(t|k) over types.
Loading 2608.29449v1…