Source-linked AI summary
Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions
Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
TL;DR
Robust clustering must accommodate heavy-tailed behavior and atypical observations without relying only on variance inflation or exclusion. The paper proposes MPVM, combining Gaussian and Cauchy components with shared structure and EM estimation; results report more stable and accurate performance under non-Gaussian conditions.
Problem
Existing robust clustering methods use diffuse noise or variance-inflated Gaussian components, limiting direct representation of continuously decaying heavy-tailed behavior.
Method
MPVM models each cluster with a Gaussian–Cauchy pseudo-Voigt mixture, shared location and scale structure, and EM-based estimation.
Results
MPVM provides more stable and accurate performance than competing approaches when the underlying distribution deviates from Gaussian assumptions.
Takeaways & Limitations
MPVM can represent heavy-tailed distributions within a single cluster while supporting clustering and outlier detection.
Takeaways & Limitations
When contamination has a variance-inflated Gaussian structure, MPVM may interpret increased variability as cluster separation rather than within-cluster dispersion.
Abstract
from arXiv · showhide
We propose a multivariate extension of the pseudo-Voigt profile-a weighted convex combination of Gaussian and Cauchy distributions-within a finite mixture modeling framework for robust model-based clustering and outlier detection. To ensure parsimony and coherence within clusters, shared location and scale parameters are imposed between the Gaussian and Cauchy components. Parameter estimation is carried out via an Expectation Maximization algorithm, with latent variables facilitating efficient likelihood-based inference. The performance of the proposed model is evaluated through simulation studies and applications to real-world data. Comparisons with established robust models, including mixtures of contaminated normal distributions, are provided to illustrate the model's clustering accuracy and outlier detection capabilities. The framework is shown to be particularly effective for data characterized by heavy-tailed behavior.
1 Introduction
The section motivates robust clustering for heavy-tailed data and atypical observations, then proposes MPVM as a heavy-tailed alternative for joint clustering and outlier detection.
- 1 Introduction: Gaussian mixtures can misrepresent heavy-tailed data, inflating covariance estimates or adding components to accommodate extreme observations.Mixtures of multivariate t-distributions address heavier tails through a degrees-of-freedom parameter.
- 1 Introduction: Mild outliers may reflect insufficiently flexible component tails, favoring heavier-tailed cluster distributions over removing observations.Gross outliers may instead require trimming or suppression when their mechanisms are difficult to model.
- 1 Introduction: Existing robust approaches use diffuse uniform noise or variance-inflated Gaussian components, without directly modeling continuously decaying heavy tails.This limitation may restrict their ability to capture genuinely heavy-tailed structures.
- 1 Introduction: The MPVM represents each cluster with Gaussian and Cauchy components, combining central modeling with heavy-tailed behavior for clustering and outlier detection.The Cauchy component supplies polynomial tail decay and supports identifying atypical observations within clusters.
- 1 Introduction: The paper evaluates MPVM through simulations and real-data applications, comparing clustering accuracy and outlier detection with established robust models.The stated comparisons include mixtures of contaminated normal distributions.
2 Background
This section introduces mixture-model notation, contaminated normal distributions, latent memberships, EM convergence, and BIC-based model selection as background for robust clustering.
- 2 Background: Finite mixture models represent multivariate observations as arising from weighted latent component distributions.Each component corresponds to a distinct subpopulation within the overall population.
- 2.2 Contaminated Normal Distribution: Contaminated normal models combine a reference Gaussian density with a covariance-inflated Gaussian density sharing a common mean.The mixing proportion denotes good observations, while the inflation factor exceeds one for contaminated observations.
- 2.2 Contaminated Normal Distribution: Identifiability constraints assign most probability mass to the smaller-covariance good component and use the inflated component for tails and potential outliers.The stated constraints are αg ≥ 0.5 and ηg > 1.
- 2.3 Model-Based Clustering: Latent component memberships support posterior probabilities, MAP hard classification, and a measure of classification uncertainty.The MAP rule assigns each observation to the cluster with the largest posterior probability.
- 2.4 Convergence Criteria: The EM algorithm increases observed-data log-likelihood monotonically, with convergence assessed using Aitken acceleration and a tolerance threshold.The criterion declares convergence when the estimated remaining log-likelihood increase falls below ε.
- 2.5 Model Selection: BIC balances model fit against the number of free parameters, preferring models with lower values.The criterion is used to compare alternative model complexities and cluster counts.
3.1 Mixture of Multivariate Pseudo-Voigt Distributions
The multivariate pseudo-Voigt distribution combines Gaussian and Cauchy components with shared location and scale structure, yielding heavy tails and a parsimonious cluster geometry.
- 3.1 Mixture of Multivariate Pseudo-Voigt Distributions: A multivariate pseudo-Voigt distribution is a convex combination of Gaussian and Cauchy distributions sharing a common location parameter.The within-cluster mixing proportion controls the combination of the two densities.
- 3.1 Mixture of Multivariate Pseudo-Voigt Distributions: The Gaussian and Cauchy components use distinct density forms while sharing the cluster location parameter.The Gaussian uses covariance matrix Σg, whereas the Cauchy uses scale matrix Γg.
- 3.1 Mixture of Multivariate Pseudo-Voigt Distributions: Imposing Γg = Σg gives both components a common geometric structure while allowing the Cauchy component to introduce heavier tails.The constraint preserves the cluster orientation and reduces the number of parameters.
- 3.1 Mixture of Multivariate Pseudo-Voigt Distributions: Compared with contaminated normal mixtures, the proposed formulation requires G fewer parameters and is therefore more parsimonious.The parameter count includes mixing proportions, mean vectors, scale matrices, and within-cluster Gaussian proportions.
3.2 Parameter Estimation
The MPVM is estimated with EM using cluster and within-cluster latent indicators, plus a scale-mixture representation for the Cauchy component. This augmentation enables closed-form updates despite the absence of direct closed-form estimates for shared location and scale parameters.
- EM framework: EM replaces direct observed-likelihood maximization with iterative maximization of the conditional expected complete-data log-likelihood.The observed likelihood is analytically challenging because of its log-sum structure.
- Latent variable structure: Three latent-variable layers represent cluster membership, Gaussian-versus-Cauchy membership, and Cauchy-component scaling.Z_ig identifies clusters, V_ig identifies the within-cluster component, and U_ig supports the Cauchy scale-mixture representation.
- Cauchy augmentation: The Cauchy distribution is represented as a scale mixture of Gaussian distributions, with U_ig conditionally distributed through a Gamma density.This representation supplies the additional latent scaling variable needed for EM estimation.
- Parameter updates: The scale-mixture formulation enables closed-form parameter estimation within the EM framework.Without this augmentation, the latent-indicator formulation alone does not provide closed-form estimates for µ and Σ.
- EM iterations: At each E-step, conditional expectations and posterior probabilities are computed for cluster membership, Gaussian membership, Cauchy membership, and the scaling variable.The resulting expected complete-data likelihood is maximized sequentially over additive parameter blocks, including mixing proportions and shared location and scale parameters.
3.3 Outlier Detection
The MPVM identifies outliers using both Cauchy-component evidence and geometric tail evidence, rather than relying on latent sub-cluster membership alone. Its shared Gaussian–Cauchy structure supports heavy-tailed clusters without requiring the Gaussian component to dominate.
- Motivation: Contaminated-mixture constraints can restrict cluster structures by assuming atypical observations are a minority and variance inflation adequately describes contamination.These assumptions become restrictive for genuinely heavy-tailed clusters or clusters with substantial tail-associated mass.
- Model distinction: The MPVM imposes no restriction requiring α_g to exceed 0.5, allowing the Cauchy component to contribute substantially at the center or even dominate near µ_g.Thus, Cauchy membership does not by itself identify an outlier.
- Dual criterion: An observation is classified as an outlier only when it is more likely from the Cauchy sub-component and lies sufficiently far into the cluster’s tail.The criterion combines generative posterior evidence with geometric evidence based on relative density.
- Geometric criterion: The relative density ratio R(x) is compared with its cluster-center value R(µ_g), which defines the baseline for acceptable Cauchy dominance in the core.The relative log-ratio depends only on squared Mahalanobis distance from the cluster center.
- Tail boundary: A unique dominance threshold defines an ellipsoidal central-to-tail boundary, and observations beyond it remain in the tail as Mahalanobis distance increases.Central observations are not outliers even when the Cauchy component predominates at the cluster peak.
- Advantages: The shared-mean Gaussian core and heavy-tailed Cauchy component separate core and tail behavior while reducing extreme observations’ influence on location estimation.Removing clean-component dominance constraints also accommodates clusters with substantial heavy-tailed behavior or contamination mass.
3.4 Initialization
MPVM initialization uses a robust trimmed k-means partition to assign retained observations to Gaussian components and trimmed observations softly to Cauchy components.
- Partitioning: A trimmed k-means partition divides observations into retained set K and trimmed set T.The trimming level is controlled by α_trim.
- Retained observations: Retained observations receive hard cluster memberships and initialize in the Gaussian sub-component.This uses the partition obtained from trimmed k-means.
- Trimmed observations: Trimmed observations receive soft memberships based on normalized inverse Euclidean distances and initialize in the Cauchy sub-component.The soft assignment distributes trimmed observations across cluster centers.
4 Simulation Studies and Real Data Analysis
The study evaluates MPVM for cluster recovery, component selection, and atypical-observation detection using simulations and real datasets. Results indicate that robust trimmed-k-means initialization supports stable estimation, while MPVM represents diverse tail behaviors without the extra clusters or variance inflation sometimes used by CNM.
- Study design: MPVM is evaluated on cluster recovery, mixture-component selection, and atypical-observation detection in simulations and real-data applications.The experiments include controlled edge cases and datasets with known cluster structures and varying contamination.
- Initialization: The modified trimmed-k-means initialization generally provides better MPVM starting values than random and GMM-based alternatives across separation scenarios.Positive mean ΔBIC values indicate lower BIC for trimmed initialization relative to the alternatives.
- Initialization: The trimmed initialization reduces heavy-tailed observations’ influence during early estimation by producing more stable initial partitions.The procedure uses robust partitioning before the first EM iteration, with retained observations initialized in the Gaussian sub-component and trimmed observations in the Cauchy sub-component.
- Edge-case analysis: In normal-data Case 1, both models identify one cluster without outliers, while MPVM uses one fewer free parameter and is expected to reduce BIC by approximately log(1000) ≈6.91 units.The models produce nearly identical cluster-center estimates, but MPVM incurs a smaller complexity penalty.
- Edge-case analysis: For pseudo-Voigt and multivariate Cauchy cases, MPVM represents heavy-tailed observations within one cluster, whereas CNM introduces additional cluster structure.For contaminated-normal data, the pattern reverses: CNM uses one cluster with outliers, while MPVM uses two Gaussian clusters without outliers.
- Edge-case analysis: Across edge cases, MPVM provides a stable representation from normal through fully heavy-tailed settings, while CNM may add clusters or inflate variance in extreme cases.This difference reflects MPVM’s Gaussian–Cauchy representation versus CNM’s Gaussian reference structure.
4.4 Simulation Study: Cluster Recovery and Outlier Detection
The simulation study compares MPVM and CNM using BIC and F1-score under Gaussian and heavy-tailed t-distribution settings, then evaluates clustering and outlier detection on Raman spectroscopy data. MPVM is comparable under Gaussian data and performs more stably and accurately when distributions are heavy-tailed or deviate from Gaussian assumptions.
- Simulation design: The study compares MPVM and CNM using BIC for model fit and F1-score for outlier detection across Gaussian and t-distribution scenarios.The simulations use known outlier labels and average performance over 10 independent repetitions.
- Simulation results: In the Gaussian scenario, CNM achieves lower average BIC while MPVM attains a marginally higher F1-score, indicating comparable performance.
- Simulation results: Under t-distributed data, MPVM achieves lower average BIC and higher average F1-score than CNM, with lower F1-score variability across repetitions.The results indicate more stable outlier detection and greater CNM sensitivity to heavier tails.
- Simulation results: Overall, both models suit Gaussian data, whereas MPVM provides more stable and accurate performance when the distribution departs from Gaussian assumptions.
- Real-data evaluation: MPVM achieves the lowest BIC of −127736.4 and the highest ARI of 0.975 among the evaluated models.The corresponding BIC differences versus MPVM are 20.5 for CNM, 84.6 for GMM, and 6.3 for the t-mixture.
- Real-data evaluation: In Raman spectroscopy data, MPVM, CNM, and the multivariate t-mixture recover three biological groups with minimal misclassification, while GMM splits H460 into two clusters.
- Real-data evaluation: MPVM identifies 11 observations as outliers compared with 78 identified by CNM, consistent with different estimated component weights.CNM estimates α = (0.84, 0.85, 0.85), whereas MPVM estimates α = (0.97, 0.96, 0.95).
5 Discussion
The discussion presents MPVM as a Gaussian–Cauchy mixture that models light- and heavy-tailed behavior while combining generative and geometric evidence for outlier identification. Simulations and Raman data show robust performance, but variance-inflated Gaussian contamination can be interpreted as cluster separation.
- Model contribution: MPVM combines Gaussian and heavy-tailed Cauchy components within each cluster to represent both light-tailed and heavy-tailed behavior.
- Model contribution: MPVM identifies outliers using both component-based and positional evidence, allowing heavy-tailed observations to remain within cluster structure.
- Limitations: When contamination has a variance-inflated Gaussian structure, MPVM may interpret increased variability as cluster separation rather than within-cluster dispersion.
- Simulation findings: In Gaussian simulations, MPVM performs comparably to contaminated normal mixtures, while under heavy-tailed scenarios it shows greater stability and robustness.
- Simulation findings: MPVM reduces oversegmentation under heavy-tailed conditions, where Gaussian-based frameworks commonly introduce additional clusters for extreme observations.
- Real-data findings: In Raman data, MPVM recovers the three biological groups with near-perfect agreement and attains the lowest BIC and highest adjusted Rand index.Allowing departures from Gaussian assumptions provides a more flexible representation of the observed cluster structure in this application.
- Future work: Future extensions may include parsimonious covariance structures and asymmetric component distributions.
Appendix A Existence and Uniqueness of the Dominance Threshold
The appendix proves that, under shared Gaussian and Cauchy location and scale parameters, the relative dominance threshold is unique and depends only on squared Mahalanobis distance. This yields a single ellipsoidal boundary beyond which observations remain in the tail region.
- Threshold formulation: The normalized Cauchy-to-Gaussian log-ratio depends solely on squared Mahalanobis distance from the common location parameter.
- Threshold formulation: With common location µ and scale matrix Σ, the dominance boundary is defined by the unique threshold satisfying (x −µ)⊤Σ−1(x −µ) = δ⋆.
- Existence and uniqueness: The scalar function m(δ) decreases on (0, p), increases on (p, ∞), and therefore has a unique global minimum at δ = p.
- Existence and uniqueness: Because m(p) < 0 and m(δ) tends to +∞ as δ grows, the Intermediate Value Theorem guarantees a unique dominance threshold.
- Geometric implication: Observations with δ(x) > δ⋆ lie beyond the ellipsoidal boundary and remain in the tail region for all larger Mahalanobis distances.
Appendix B Detailed Parameter Estimates for Real Data Analysis
Appendix B reports estimated mixing proportions and component weights for MPVM and CNM, alongside estimated mean vectors used in the real-data analysis.
- Parameter estimates: Table B1 reports estimated mixing proportions π and component weights α for the MPVM and CNM models.
- Parameter estimates: Table B2 reports estimated mean vectors µ for the MPVM and CNM models.