Source-linked AI summary

Beta-Negative Binomial Process and Poisson Factor Analysis

Mingyuan Zhou, Lauren Hannah, David Dunson, Lawrence Carin

arXiv:1112.3605v4stat.MLstat.ME

TL;DR

The paper addresses limited flexible modeling tools for high-dimensional multivariate count data with latent structure. It proposes the beta-negative binomial process and embeds it in an infinite Poisson factor analysis hierarchy, with efficient inference and encouraging document-factorization results. The beta-gamma-gamma-Poisson model captures topics with diverse characteristics while automatically inferring factor structure.

  • Problem

    High-dimensional multivariate count data need latent-variable models that accommodate discreteness, nonnegativity, limited ranges, and overdispersion without restrictive Gaussian assumptions.

  • Method

    The paper proposes a beta-negative binomial process, augments it into a beta-gamma-gamma-Poisson hierarchy, and uses it as a nonparametric prior for infinite Poisson factor analysis.

  • Results

    βγΓ-PFA produces the smallest held-out perplexity among the compared PFA models and automatically infers the number of factors up to Kmax = 400.

  • Takeaways & Limitations

    The model accurately captures document topics with diverse characteristics in document count matrix factorization.

  • Takeaways & Limitations

    For a general measure R0, the relevant distribution does not have a closed-form solution because R0 is not conjugate to the negative binomial distribution.

Abstract

from arXiv · show

A beta-negative binomial (BNB) process is proposed, leading to a beta-gamma-Poisson process, which may be viewed as a "multi-scoop" generalization of the beta-Bernoulli process. The BNB process is augmented into a beta-gamma-gamma-Poisson hierarchical structure, and applied as a nonparametric Bayesian prior for an infinite Poisson factor analysis model. A finite approximation for the beta process Levy random measure is constructed for convenient implementation. Efficient MCMC computations are performed with data augmentation and marginalization techniques. Encouraging results are shown on document count matrix factorization.

1 Introduction

The paper introduces the beta-negative binomial process as a flexible nonparametric Bayesian prior for multivariate count data and uses it to build an infinite Poisson factor analysis model.

  • Multivariate count models with latent variables remain underdeveloped, while Gaussian latent-factor approaches can be computationally prohibitive and restrictive for count data.Count data are discrete, nonnegative, bounded in range, and often overdispersed.
  • The beta-negative binomial process extends the beta process to a marked space and models count variables through a negative binomial process.Its beta-gamma-Poisson augmentation provides the process construction.
  • The process is interpreted as a multi-scoop Indian buffet process, where customers may select multiple scoops of each dish.The number of scoops is controlled by a dish-dependent negative binomial distribution.
  • The resulting beta-gamma-gamma-Poisson hierarchy produces an infinite Poisson factor analysis model for count matrix factorization.A gamma prior on the Poisson rate yields a negative binomial distribution, and latent counts represent factor occurrences in documents.
  • The paper contributes efficient inference and a flexible count-matrix factorization model that captures topics with diverse characteristics.The contributions include extending the beta process, developing efficient inference, and applying the model to document corpora.

2 Preliminaries

The preliminaries define the negative binomial and Lévy random-measure foundations used to construct the beta-negative binomial process and its beta-process components.

  • 2.1 Negative Binomial Distribution: A gamma prior on a Poisson rate produces a negative binomial distribution that can model overdispersed count data.Its variance exceeds its mean when parameterized by r > 0 and p ∈ (0, 1).
  • 2.2 L´evy Random Measures: A Lévy random measure assigns independent infinitely divisible random variables to disjoint Borel sets through a Poisson construction.A nonnegative Lévy random measure satisfying the stated integrability condition is called a completely random measure.
  • 2.3 Beta Process: The beta process is a completely random measure on [0, 1] × Ω with concentration parameter c and finite base measure B0.Its mass parameter is α = B0(Ω), and its draw has countably infinitely many atoms with weights in [0, 1].

3 The Beta Process and the Negative Binomial Process

The beta process is extended with marked negative-binomial counts, producing a beta-negative binomial process whose shared atoms support multiple counts per observation. The construction yields an infinite-process prior with a finite approximation for computation, while general predictive probabilities may lack closed forms.

  • 3.1 The Beta-Negative Binomial Process: The beta-negative binomial process marks beta-process atoms with r_k and uses negative-binomial counts κ_ki at shared locations ω_k.The parameters are shared across draws, while counts vary by observation.
  • 3.2 Model Properties: Unlike the beta-Bernoulli process, the BNB process allows each selected atom to receive multiple counts rather than only zero or one.This motivates its interpretation as a multiple-scoop generalization of the Indian buffet process.
  • 3.2 Model Properties: For a general base measure R0, the multiple-scoop IBP predictive posterior has no closed-form solution, although the number of new dishes can be computed analytically.The difficulty arises because R0 is not conjugate to the negative binomial distribution.
  • 3.2 Model Properties: When R0 = δ1, the BNB process reduces to the geometric process, but this special case is excluded from experiments because it is often overly restrictive.Under this choice, the number of new dishes has a Poisson distribution.
  • 3.3 Finite Approximations for Beta Process: A finite Levy-measure approximation retains nonnegligible beta-process weights and converges to the infinite Levy measure as ϵ approaches zero.The approximation restricts p to a finite range and supports practical inference with finite or variable K.

4 Poisson Factor Analysis

Poisson factor analysis represents count matrices as sums of latent factor-specific Poisson counts, with gamma augmentation yielding negative-binomial structure. The beta-gamma-gamma-Poisson hierarchy supports infinite factors and efficient inference through count allocation and conjugate updates.

  • 4 Poisson Factor Analysis: Poisson factor analysis explains each observed count as a sum of smaller counts generated by hidden factors.The factor loading matrix Φ encodes term importance, while the factor score matrix Θ encodes atom importance in each sample.
  • 4.1 Beta-Gamma-Gamma-Poisson Model: The Poisson and multinomial augmentations are distributionally equivalent, as established by matching characteristic functions.This equivalence underlies the latent-count representation used for inference.
  • 4.1 Beta-Gamma-Gamma-Poisson Model: Observed total counts are allocated to latent factors using an equivalent multinomial augmentation.These augmentations enable inference for factor loadings and scores.
  • 4.1 Beta-Gamma-Gamma-Poisson Model: Gamma priors on Poisson rates produce negative-binomial distributions within the beta-gamma-gamma-Poisson hierarchy.The hierarchy is used to construct an infinite PFA model for count-matrix factorization.

5 Related Discrete Latent Variable Models

The beta-gamma-gamma-Poisson PFA framework contains several related discrete latent-variable models as special cases or close relatives. Its varying factor-specific gamma scales distinguish it from models whose normalized factor scores follow a Dirichlet distribution.

  • 5 Related Discrete Latent Variable Models: The shared PFA framework enables comparison of NMF, LDA, GaP, FTM, and βγΓ-PFA using MCMC inference.These models are connected through the prior settings summarized in Table 1.
  • 5 Related Discrete Latent Variable Models: βγΓ-PFA permits different gamma scale parameters across factors, so normalized factor scores do not generally follow a Dirichlet distribution.The table organizes related algorithms by gamma scale, shape, and factor-score sparsity parameters.
  • 5 Related Discrete Latent Variable Models: NMF is a special case of Γ-PFA, while GaP is a special case of βΓ-PFA within the broader βγΓ-PFA framework.These relationships arise by fixing or simplifying selected priors and parameters.
  • 5 Related Discrete Latent Variable Models: Dir-PFA has the same block Gibbs sampling and variational Bayes inference equations as LDA.The equivalence follows from imposing Dirichlet priors on both factor loadings and scores.
  • 5 Related Discrete Latent Variable Models: SγΓ-PFA is closely related to the focused topic model, and its inference equations arise naturally when p_k = 0.5.This provides justification for corresponding inference steps used in the focused topic model.

6 Example Results and Discussions

Experiments on JACM and PsyRev evaluate βγΓ-PFA for document count matrix factorization, showing strong held-out prediction and automatically inferred factor structure. The model also represents topics with varied mean and variance-to-mean characteristics.

  • Experimental setup: The evaluation uses 80% of each document's words for training and 20% for testing, averages five random partitions, and reports held-out per-word perplexity.PFA factorization is performed on the training matrix, with 2,500 MCMC iterations and the first 1,000 discarded.
  • Topic characteristics: The first dominant factors capture common corpus topics, whereas remaining factors span diverse means and variance-to-mean ratios.Examples include topics with large mean and large VMR and topics with small mean and large VMR, indicating distinct negative-binomial parameter settings.
  • Topic characteristics: With stopwords present, βγΓ-PFA usually absorbs them into a few dominant topics with large mean and small VMR, leaving remaining topics more interpretable.On PsyRev, examples of dominant topics concern general research language; on JACM, stopword-heavy topics dominate.
  • Perplexity comparison: βγΓ-PFA produces the smallest held-out perplexity on both JACM and PsyRev, followed by SγΓ-PFA and βΓ-PFA.Γ-PFA overfits quickly as K increases, while Dir-PFA shows overfitting around K = 100; the nonparametric models infer K automatically up to Kmax = 400.
  • Prior sensitivity: βγΓ-PFA yields the best results under each tested aφ and automatically infers the number of active factors as aφ varies.Smaller aφ generally supports larger K and better held-out prediction, but values that are too small produce overly specialized topics concentrated on few terms.

7 Conclusions

The paper proposes the BNB process and embeds it in a beta-gamma-gamma-Poisson hierarchy for infinite PFA. Results on document count matrices indicate that learning both negative-binomial mean and variance supports topic modeling with quantitative and qualitative structure.

  • Conclusions: The BNB process extends beta-process modeling to multivariate count data and yields a beta-gamma-Poisson process for an infinite PFA model.The hierarchy is used as a nonparametric Bayesian prior for count matrix factorization.
  • Conclusions: A finite beta-process Lévy random-measure approximation and efficient MCMC inference make the proposed model convenient to implement.The inference exploits relationships among beta, gamma, Poisson, negative binomial, multinomial and Dirichlet distributions.
  • Conclusions: Learning both the mean and variance of latent negative-binomial factors supports topic modeling that captures common and specific aspects of document corpora.The paper evaluates this through perplexity and through the inferred characteristics of document topics.
Loading 1112.3605v4…