Source-linked AI summary
Negative Binomial Process Count and Mixture Modeling
Mingyuan Zhou, Lawrence Carin
TL;DR
Count and mixture modeling are treated jointly through negative binomial processes built from gamma and Poisson processes. The paper develops distributional connections and Bayesian inference, showing relationships to Dirichlet-process models and topic-modeling applications.
Problem
Mixture modeling assigns samples to clusters but is typically separated from count modeling, while existing inference can be difficult and model flexibility limited.
Method
The paper models mixture-component counts with NB processes, using gamma-process rate measures, Poisson-logarithmic augmentation, and normalization or marginalization.
Results
The paper derives efficient Bayesian inference, connects NB processes to Dirichlet-process models, and applies a family of NB processes to topic modeling.
Takeaways & Limitations
The relationships among NB, Poisson, logarithmic, CRT, Dirichlet-process, and HDP constructions provide theoretical, structural, and computational advantages for count and mixture modeling.
Takeaways & Limitations
HDP inference remains challenging, and truncation-based inference risks wasted computation as the truncation level increases.
Abstract
from arXiv · showhide
The seemingly disjoint problems of count and mixture modeling are united under the negative binomial (NB) process. A gamma process is employed to model the rate measure of a Poisson process, whose normalization provides a random probability measure for mixture modeling and whose marginalization leads to an NB process for count modeling. A draw from the NB process consists of a Poisson distributed finite number of distinct atoms, each of which is associated with a logarithmic distributed number of data samples. We reveal relationships between various count- and mixture-modeling distributions and construct a Poisson-logarithmic bivariate distribution that connects the NB and Chinese restaurant table distributions. Fundamental properties of the models are developed, and we derive efficient Bayesian inference. It is shown that with augmentation and normalization, the NB process and gamma-NB process can be reduced to the Dirichlet process and hierarchical Dirichlet process, respectively. These relationships highlight theoretical, structural and computational advantages of the NB process. A variety of NB processes, including the beta-geometric, beta-NB, marked-beta-NB, marked-gamma-NB and zero-inflated-NB processes, with distinct sharing mechanisms, are also constructed. These models are applied to topic modeling, with connections made to existing algorithms under Poisson factor analysis. Example results show the importance of inferring both the NB dispersion and probability parameters.
1 INTRODUCTION
The paper reframes mixture modeling as count modeling through negative binomial processes, unifying count and mixture problems with tractable completely random measures and Bayesian inference.
- Mixture modeling assigns samples to clusters, but is typically formulated through the Dirichlet-multinomial framework rather than as count modeling.
- The paper directly models counts assigned to mixture components as negative binomial variables, enabling joint count and mixture modeling through completely random measures.
- A Poisson-logarithmic bivariate distribution connects Poisson, logarithmic, NB, and Chinese restaurant table distributions for augmentation, marginalization, and inference.
- A gamma process models the Poisson rate measure; normalization supports mixture modeling, while marginalization yields an NB process for count modeling.
- The paper unifies earlier conference materials, extends the Chinese restaurant process to random customer and table counts, and gives conditions recovering NB processes from Dirichlet-process models.
- The authors construct a broad family of NB processes, including beta-NB and marked variants, with distinct sharing mechanisms.
2 PRELIMINARIES
The preliminaries introduce completely random measures, Poisson and gamma processes, their normalized forms, and the Chinese restaurant process representation of exchangeable assignments.
- Completely random measures assign independent infinitely divisible random variables to disjoint Borel sets.
- A Poisson process places Poisson counts on subsets, remaining valid for continuous, discrete, or mixed discrete-continuous base measures.
- A gamma process is a completely random measure whose mass on each subset follows a gamma distribution with base measure G0 and scale parameter 1/c.
- Normalizing a gamma process with space-invariant scale produces a Dirichlet process, but the normalized measures on disjoint sets become negatively correlated.
- Conversely, an independent gamma-distributed total mass multiplied by a Dirichlet process recovers a gamma process.
- The Chinese restaurant process describes exchangeable assignments, with the number of nonempty tables represented by a CRT random variable.
3 NEGATIVE BINOMIAL DISTRIBUTION
The section develops the negative binomial distribution as an overdispersed count model, relates it to Poisson, logarithmic, and CRT distributions, and introduces a useful bivariate construction.
- Poisson counts have equal mean and variance, making the Poisson assumption restrictive for overdispersed data.
- Gamma-Poisson marginalization produces an NB distribution with mean µ = rp/(1 −p) and variance σ2 = rp/(1 −p)2.
- The NB distribution supports both a compound-Poisson representation with logarithmic variables and limiting connections to Poisson and logarithmic distributions.
- Inference for the NB dispersion parameter r is difficult because its conjugate prior is unknown and maximum-likelihood estimates can be non-robust or fail to converge, especially with small samples.
- These distributional connections provide augmentation and marginalization tools for efficient NB inference and reveal links to Dirichlet-process, HDP, and beta-NB models.
- The Poisson-logarithmic bivariate distribution has equivalent CRT-NB and sum-logarithmic-Poisson representations, linking customer and table counts.
4 JOINT COUNT AND MIXTURE MODELING
The paper unifies count and mixture modeling through Poisson and gamma-process constructions, then uses NB-process representations to model overdispersion and connect to Dirichlet processes.
- Normalizing the shared rate measure turns the same Poisson construction into mixture modeling that allocates observations across partition cells.
- A Poisson process with a shared completely random measure jointly models counts across measurable partitions of grouped data.
- A gamma-Poisson construction is equivalent in distribution to an NB process after marginalizing the gamma rate measure.
- An NB-process draw contains a finite Poisson-distributed number of distinct atoms, with logarithmically distributed sample counts at each atom.
- The NB probability parameter p controls the prior distributions of distinct-atom counts, per-atom sample counts, and total sample counts.
- Normalizing the gamma process recovers a Dirichlet process, whereas the gamma-Poisson NB process can be more restrictive because it imposes shared mixture proportions and atom-level count distributions across groups.
5 JOINT COUNT AND MIXTURE MODELING OF GROUPED DATA
The gamma-NB process jointly models grouped-data counts and mixtures, with equivalent augmentations that support analytic inference and connect to the HDP under normalization.
- Gamma-Negative Binomial Process: The gamma-NB process couples an NB process with a gamma process for grouped-data modeling and has analytic conditional posteriors.The construction shares NB dispersion across groups while allowing group-dependent probability parameters.
- Equivalent Constructions: The gamma-gamma-Poisson augmentation introduces group-specific gamma processes whose normalizations provide random probability measures for mixture modeling.The same construction yields Poisson counts conditional on gamma-distributed rates.
- Equivalent Constructions: Gamma-gamma-Poisson, gamma-compound Poisson, and gamma-NB-CRT constructions provide equivalent representations of the gamma-NB process.The center and right constructions are equivalent in distribution.
- Relationship with Hierarchical Dirichlet Process: Augmentation and normalization reduce the gamma-NB process to an HDP when total-count and scale variables are not modeled as random.The gamma-NB process remains a completely random measure, whereas normalization removes that property in the HDP.
- Relationship with Hierarchical Dirichlet Process: Unlike standard HDP inference, the gamma-NB process supports analytic updates for shared mass and total base-measure mass, including with discrete base measures.HDP concentration parameters are nontrivial to infer, and approximate sampling of γ0 can be biased when the truncation level is insufficiently large.
6 THE NEGATIVE BINOMIAL PROCESS FAMILY
The NB process family varies which dispersion, probability, and sparsity parameters are shared or group dependent. These constructions support different count-sharing mechanisms and connections to existing process models.
- Sharing Mechanisms: NB process variants explore sharing the NB probability measure across groups while allowing dispersion parameters to be group specific or atom dependent.This contrasts with the gamma-NB process, which shares dispersion across groups and makes probabilities group dependent.
- Beta-Based Processes: The beta-NB process uses a shared probability measure and group-dependent dispersion, while fixed dispersion one yields the beta-geometric process.The beta-NB formulation also provides analytic conditional posteriors for group-specific dispersion parameters.
- Connections to Mixture Modeling: Beta-NB, gamma-NB, and HDP constructions can all support mixed-membership modeling through normalized group-specific random measures.Their hierarchical structures differ in how process parameters are shared and inferred.
- Marked Processes: Marked-beta-NB and marked-gamma-NB processes introduce gamma or beta marks to construct alternative sharing patterns for NB probability and dispersion measures.The marked-beta-NB process shares both measures before applying independent gamma marks.
- Zero Inflation: The zero-inflated-NB process adds Bernoulli-process indicators to model excessive zeros beyond those governed by the NB process.Its construction connects to focused topic models and spike-and-slab process priors, while retaining tractable inference and learnable probability parameters.
7 NEGATIVE BINOMIAL PROCESS TOPIC MODELING AND POISSON FACTOR ANALYSIS
NB process topic models factorize document term counts while using different NB parameter-sharing schemes for topic weights. They connect to Poisson factor analysis and existing topic models.
- Topic Modeling: NB process topic modeling treats document words as grouped exchangeable observations assigned to topics with categorical word likelihoods.Each word is drawn from the topic-specific distribution associated with its latent topic index.
- Poisson Factor Analysis: Poisson factor analysis represents the term-document count matrix as M ∼Pois(ΦΘ), with factors encoding term importance and scores encoding sample importance.Under the NB process formulation, factor scores are gamma distributed and topic-assigned counts follow NB distributions.
- Model Family: Different NB process topic models arise by choosing which dispersion, probability, and sparsity parameters are shared or inferred.Table 1 summarizes these mechanisms through inferred parameters, variance-mean ratios, overdispersion levels, and related algorithms.
- Model Comparisons: The gamma-NB topic model learns document-dependent NB probabilities and improves over NB-HDP by allowing those probabilities to be inferred.Without modeling total-count and scale variables, the gamma-NB process reduces to the HDP.
- Model Comparisons: Beta-NB and marked-beta-NB models provide analytic conditional posteriors for their group- or atom-specific dispersion parameters.The beta-geometric process is more restrictive because it fixes dispersion parameters at one.
- Approximate and Exact Inference: Infinite atom representations require finite truncation for block Gibbs sampling, while slice sampling offers adaptive truncation as an alternative.Larger truncation levels can improve approximation but increase the risk of wasted computation.
8 EXAMPLE RESULTS AND DISCUSSIONS
Experiments on the Psychological Review corpus compare NB process topic models with LDA and CRF-HDP using held-out-word perplexity and inferred parameter-sharing behavior. Results show that learning document- and topic-specific NB dispersion and probability parameters can improve flexibility and data fitting, while restrictive parameter choices can degrade performance.
- Experimental Setup: The study evaluates topic models on 1281 abstracts using held-out-word perplexity across multiple training percentages and Gibbs-sampling implementations.The corpus contains 71,279 word counts and a vocabulary of 2,566 terms; results average five random training/testing partitions.
- Experimental Setup: NB process topic models automatically infer active topics, whereas LDA and NB-LDA require tuning the topic number K.All models use K = 400 as an upper bound for fair comparison.
- Parameter Sharing: Learning document-dependent rj and pj allows NB-LDA to model highly overdispersed topic usage and, with appropriate K, outperform NB-HDP and NB-FTM as training data increases.NB-LDA needs tuning only K, while LDA tunes both K and α.
- Model Restrictions: Restrictive parameterizations can hurt prediction: Beta-Geometric fixes rj = 1, while NB-HDP fixes pj = 0.5 and thereby imposes the same VMR of 2 on every count vector.The Beta-Geometric process required substantially underestimated pk values and showed degraded performance relative to Beta-NB.
- Parameter Sharing: Using pk creates a sharp active-topic transition and positively couples topic mean with overdispersion, whereas using rk produces smoother transitions and a negative coupling.The Gamma-NB process learned 177 active topics versus 107 for the Beta-NB process.
- Parameter Sharing: The Marked-Beta-NB process permits more diverse combinations of topic mean and overdispersion and showed superior performance in the reported experiments.Both rk and pk contribute to the topic mean, allowing large mean with either small or large overdispersion.
9 CONCLUSIONS
The paper proposes negative binomial processes that unify count and mixture modeling through shared distributional relationships, augmentations, and completely random measures. These constructions support model analysis, efficient inference, and topic-modeling applications.
- Negative binomial processes naturally apply count modeling to mixture modeling while using completely random measures.
- Distributional connections and specialized augmentation methods unify count and mixture modeling, analyze model properties, and enable efficient Bayesian inference.
- The NB process and gamma-NB process can be recovered from the Dirichlet process and hierarchical Dirichlet process, respectively.
- The proposed family includes beta-geometric, beta-NB, marked-beta-NB, marked-gamma-NB, and zero-inflated-NB processes with distinct sharing mechanisms.
- Topic-modeling experiments show the importance of modeling both NB dispersion and probability parameters.
PROOF OF THEOREM 1
The proof establishes a bivariate count distribution by expressing negative binomial counts through Chinese restaurant table counts and logarithmic summands. Equivalent Poisson-logarithmic constructions yield the same joint distribution.
- The joint distribution of negative binomial counts and Chinese restaurant table counts factors into a CRT conditional distribution and an NB marginal.
- A count obtained by summing l independent logarithmic random variables has a probability generating function derived from the logarithmic PGF.
- The logarithmic-sum probability mass function is expressed using signed Stirling numbers of the first kind.
- Letting l follow a Poisson distribution with rate −r ln(1−p) produces the same joint distribution as the NB–CRT factorization.
APPENDIX B BLOCK GIBBS SAMPLING FOR THE GAMMA-NEGATIVE BINOMIAL PROCESS
The appendix specifies block Gibbs sampling for the gamma-negative binomial process under beta and gamma priors with a discrete base measure. The sampler uses topic-activation counts and normalized base-measure parameters.
- The model assigns beta priors to p_j and a gamma prior to γ_0, while using a discrete base measure over K atoms.
- The block Gibbs sampler updates the gamma-negative binomial process using the discrete base measure and topic-activation indicators.