Source-linked AI summary
On Mixtures of Skew Normal and Skew t-Distributions
Sharon X. Lee, Geoffrey J. McLachlan
TL;DR
Existing skew-mixture proposals are numerous and difficult to relate, motivating a systematic account of their forms and performance. The paper classifies skew-normal and skew-t distributions, explains links among EM algorithms, and compares mixture models on a real clustering dataset. It concludes that the four-form classification clarifies relationships, while mixtures based on more general forms remain a direction for further research.
Problem
Numerous similar but nonidentical skew-mixture proposals make their relationships and relative performance unclear.
Method
The paper classifies skew-normal and skew-t distributions into four forms, relates EM fitting algorithms, and illustrates model performance on a real dataset.
Results
The paper presents a four-form classification that clarifies connections among skew distributions and compares their mixture-model clustering performance on a real dataset.
Takeaways & Limitations
Restricted and unrestricted forms are both special cases of the extended form, which is itself a special case of the generalized form.
Takeaways & Limitations
Existing finite-mixture work has investigated only restricted and unrestricted multivariate skew distributions; mixtures using more general forms remain for future research.
Abstract
from arXiv · showhide
Finite mixture of skew distributions have emerged as an effective tool in modelling heterogeneous data with asymmetric features. With various proposals appearing rapidly in the recent years, which are similar but not identical, the connections between them and their relative performance becomes rather unclear. This paper aims to provide a concise overview of these developments by presenting a systematic classification of the existing skew distributions into four types, thereby clarifying their close relationships. This also aids in understanding the link between some of the proposed expectation-maximization (EM) based algorithms for the computation of the maximum likelihood estimates of the parameters of the models. The final part of this paper presents an illustration of the performance of these mixture models in clustering a real dataset, relative to other non-elliptically contoured clustering methods and associated algorithms for their implementation.
1 Introduction
Finite mixtures of skew distributions address heterogeneous, asymmetric, multimodal, and heavy-tailed data, but many similar proposals have obscured their relationships and relative performance. This paper classifies these distributions systematically, connects fitting algorithms, and compares models on a real clustering dataset.
- Motivation: Finite skew-distribution mixtures offer an alternative to Gaussian mixtures for datasets with asymmetry, multimodality, or heavy tails.Applications span medical sciences, bioinformatics, environmetrics, engineering, economics, and financial sciences.
- Research gap: The rapid growth of skew-normal and skew-t mixture proposals has made their relationships and relative performance unclear.The literature includes multiple characterizations of both distribution families.
- Contributions: The paper presents a systematic classification of multivariate skew distributions into restricted, unrestricted, extended, and generalized forms.The terminology expands earlier restricted and unrestricted classifications.
- Empirical illustration: The paper illustrates relative performance by applying skew-mixture models and related clustering algorithms to a real dataset.The application compares these models with other related model-based clustering methods.
- Contributions: The classification clarifies connections among existing proposals and supports understanding equivalences among some EM-based mixture-fitting algorithms.The paper organizes these connections across its classification and algorithmic discussions.
2.1 Multivariate skew normal distributions
The multivariate skew-normal family can be organized through conditioning constructions into restricted, unrestricted, extended, and generalized forms. These forms differ in latent-variable dimension, extension parameters, and latent-variable distribution, while several named proposals are equivalent after reparameterization.
- General framework: The fundamental skew-normal distribution encompasses the restricted, unrestricted, and extended multivariate skew-normal forms.It is generated by conditioning a multivariate normal variable on another random variable.
- Restricted and unrestricted forms: The restricted form uses a univariate latent normal variable with q = 1, τ = 0, and Γ = 1.The unrestricted form instead uses p-dimensional normal latent variables with q = p.
- Restricted and unrestricted forms: Restricted and unrestricted forms coincide in the univariate case but are not generally nested because their restrictions concern the latent stochastic construction.Thus, “restricted” does not mean a restriction on the parameter space.
- Extended and generalized forms: The extended form allows unrestricted latent-variable dimension and a non-zero extension parameter τ.When the latent variable is non-normal, the distribution has the generalized form.
- Classification tables: The section summarizes representative multivariate skew-normal distributions using the abbreviations rMSN, uMSN, eMSN, and gMSN.The cited tables provide the classification and abbreviation summary, while the listed examples are not exhaustive.
- Restricted multivariate skew normal distributions: Azzalini–Dalla Valle, Branco–Dey, Lachos, and Pyne skew-normal distributions are identical after reparameterization and can be represented as restricted forms.The convolution representation uses independent latent variables and supports a hierarchical form for EM implementation.
The skew normal distribution of Branco and Dey (B-MSN)
The B-MSN is a restricted skew normal parameterization that is algebraically simpler than A-MSN and equivalent to it through reparameterization. Related restricted formulations likewise share equivalent densities or stochastic representations, supporting EM implementation.
- The skew normal distribution of Branco and Dey (B-MSN): B-MSN removes the scaling matrix D from A-MSN, simplifying the algebra but making the skewness parameter scale-dependent.The parameterization is denoted B-MSN and uses density (7).
- The skew normal distribution of Branco and Dey (B-MSN): The B-MSN conditioning representation conditions Y 1 on a scalar latent normal variable Y0 exceeding zero.Its convolution representation uses independent normal variables and a residual covariance ˜Σ = Σ −δδT.
- The skew normal distribution of Branco and Dey (B-MSN): B-MSN is a reparameterization of A-MSN, recovered by replacing δ with DδA.The two parameterizations therefore describe the same restricted skew normal form under transformed parameters.
- The skew normal distribution of Branco and Dey (B-MSN): SNI-SN is a scale-mixture family whose basic degenerate case includes the multivariate skew normal distribution and is itself a reparameterization of restricted forms.Its density becomes identical to B-MSN after replacing δ in the SNI-SN form with Σ^2δS.
2.1.2 Unrestricted multivariate skew normal distributions
Unrestricted multivariate skew normal distributions replace the scalar latent skewing variable with a p-dimensional vector and impose positivity componentwise. Their parameterization uses a p × p skewness matrix, with Sahu et al.’s form restricting that matrix to be diagonal.
- 2.1.2 Unrestricted multivariate skew normal distributions: The unrestricted MSN replaces the scalar latent variable of the restricted case with a p-dimensional normal vector.The conditioning constraint becomes Y 0 > 0 componentwise, requiring every latent-vector element to be positive.
- 2.1.2 Unrestricted multivariate skew normal distributions: The unrestricted model uses a p × p skewness matrix ∆ rather than a p-dimensional skewness vector.Its convolution representation uses independent ˜Y 0 and ˜Y 1 normal vectors.
- 2.1.2 Unrestricted multivariate skew normal distributions: The Sahu et al. skew normal is an unrestricted MSN whose skewness matrix ∆ is restricted to be diagonal.The associated parameter relationship is ˜Σ = Σ −∆∆T.
- 2.1.2 Unrestricted multivariate skew normal distributions: Sahu et al.’s unrestricted MSN density involves a multivariate normal distribution function, unlike restricted forms defined through a univariate distribution function.The distinction follows from the multivariate latent-variable conditioning construction.
2.1.3 Extended multivariate skew normal distributions
Extended and generalized skew normal distributions broaden the latent-variable construction through nonzero thresholds, higher-dimensional latent variables, and relaxed latent distributions. The SUN distribution contains restricted, unrestricted, and extended MSN forms as special cases.
- 2.1.3 Extended multivariate skew normal distributions: The ESN distribution conditions on a latent variable shifted by an arbitrary threshold τ, called the extension parameter.Its normalizing constant depends on τ rather than being a fixed value.
- 2.1.3 Extended multivariate skew normal distributions: The SUN distribution replaces the scalar latent variable with a q-dimensional one to unify the previously described skew normal distributions.Its convolution construction uses a q-dimensional truncated normal latent variable and a p × q skewness matrix.
- 2.1.3 Extended multivariate skew normal distributions: SUN includes restricted MSN, unrestricted MSN, and ESN distributions as special cases, alongside equivalent variants such as HSN and CSN.These equivalences organize several apparently distinct skew normal proposals within one framework.
- 2.1.3 Extended multivariate skew normal distributions: The generalized MSN form retains only a multivariate normal symmetric component while allowing broader latent-variable distributions and skewing functions.The fundamental skew normal is identified as a prominent example of this generalized form.
- 2.1.3 Extended multivariate skew normal distributions: For CFUSN, q = p with diagonal ∆ yields the unrestricted skew normal, whereas q = 1 yields the restricted B-MSN.Its density is f(y; µ, Σ, ∆) = 2qφp(y; µ, Σ)Φq(∆TΣ−1(y −µ); 0, Λ).
2.2 Multivariate skew t-distributions
Multivariate skew t-distributions extend skew-normal models by combining skewness and kurtosis, but have been proposed in several nonidentical forms. The paper organizes these variants and relates their stochastic representations and parametrizations.
- The multivariate skew t-distribution offers greater flexibility than the normal distribution by combining skewness and kurtosis while retaining algebraic tractability.
- Many variants have been proposed, including skew elliptical, Azzalini–Capitanio, Gupta, Sahu et al., SNI-ST, and closed skew t distributions.
- Restricted skew t-distributions condition on a univariate latent variable being positive, with dependence between the latent and observed variables described by δ.
- The restricted MST also has a convolution-type representation involving jointly distributed central t variables and ν degrees of freedom.
- The skew t-distributions of Branco and Dey, Azzalini and Capitanio, Gupta, SNI-ST, and Pyne et al. are equivalent to the restricted MST up to reparametrization.
The skew t-distribution of Branco and Dey (B-MST)
The Branco–Dey skew t-distribution is a scale-mixture extension of a skew-normal model. Its density and stochastic representations connect it to the restricted MST formulation.
- The Branco–Dey skew t-distribution is a special case of a scale mixture of the B-MSN distribution within the skew elliptical class.
- Its density uses the squared Mahalanobis distance d(y) and multivariate and univariate t-distribution terms.
- The distribution has both conditioning-type and convolution-type stochastic representations.
- The conditioning representation is identical to the restricted MST representation in equation (29).
The skew t-distribution of Azzalini and Capitanio (A-MST)
The Azzalini–Capitanio skew t-distribution and several other restricted skew t formulations are linked through reparametrization and alternative scale parameterizations. The section also distinguishes restricted and unrestricted models.
- The Azzalini–Capitanio skew t-distribution extends the Azzalini–Dalla Valle multivariate skew normal distribution to the skew-t case.
- A-MST: The A-MST uses a conditioning representation involving D(Y 1 | Y0 > 0), while its convolution representation provides a parallel formulation.
- A-MST: Factoring the scale matrix as DRD makes the skewness parameter invariant to a change of scale, and setting δ to DδA yields the B-MST distribution.
- G-MST: The Gupta formulation replaces δA with D−1ΣδG, and its density is identical to the B-MST density after rewriting δ as ΣδG.
- Restricted MST mixtures: The restricted MST is identical to the general restricted representation, and mixtures have been fitted using EM, alternative exact EM, and Bayesian algorithms.
- Unrestricted MST: In unrestricted models, the latent variable is p-dimensional and positivity applies elementwise; maximum-likelihood estimation is computationally difficult.
- Restricted versus unrestricted: The restricted and unrestricted MST densities are equivalent only in the univariate case, so the unrestricted density does not generally incorporate the restricted one.
3 Mixtures of multivariate skew normal and skew t-distributions
Finite skew-mixture models represent heterogeneous populations through component densities and mixing proportions, with EM algorithms using latent component labels and distribution-specific latent variables. Existing algorithms differ in computational strategy and cost.
- A finite mixture models a population as g subpopulations, combining component densities f(y; θh) with nonnegative weights πh that sum to one.
- The EM algorithm treats observed data as incomplete by introducing latent component labels, with the E-step computing the Q-function.
- Skew normal mixtures: Restricted MSN mixtures admit closed-form E- and M-steps, with implementations developed for FM-rMSN, FM-SNI-SN, and FM-A-MSN models.
- Skew normal mixtures: Unrestricted MSN mixtures also have closed-form expressions, but replacing a scalar latent variable with a multivariate one increases computational cost.
- Skew t mixtures: Restricted MST mixtures use half-normal and gamma latent variables, and have been fitted with EM, ECME, and Bayesian approaches.
- Algorithm equivalence: Pyne et al. and Vrbik and McNicholas used equivalent EM formulations for FM-rMSN, expressing required moments through truncated-t moments or hypergeometric functions.
- Algorithm equivalence: For FM-uMST, implementations use either a Monte Carlo E-step or an OSL and multivariate truncated-t moments approach.
- Computational cost: Replacing a univariate t distribution function with a normal distribution function considerably reduces computation time by limiting intensive E-step calculations to univariate truncated-normal moments.
4 Clustering DLBCL samples
The paper evaluates skew mixture models by clustering a trivariate DLBCL cell dataset into three groups. Multivariate skew t-mixtures outperform the other evaluated methods, with FM-uMST and FM-rMST contours resembling manually identified clusters.
- Dataset and evaluation: The DLBCL dataset contains over 3000 cells measured using CD3, CD5, and CD19 markers, with cells clustered into three groups.The sample population comprises 3290 cells, and dead cells were excluded before evaluating misclassification rates.
- Dataset and evaluation: Performance is assessed using misclassification rates against human-expert cluster labels, minimizing error over permutations of cluster labels.Lower misclassification rates indicate closer agreement with the expert-defined labels.
- Results: Multivariate skew t-mixture models outperform the other methods on the DLBCL dataset.The comparison includes FM-uMST, FM-rMST, FM-MNIG, FM-MSAL, and related clustering methods.
- Results: The FM-uMST and FM-rMST component contours resemble the shapes of clusters identified by manual gating.Figure 1 compares expert labels with fitted component contours for the evaluated mixture models.
5 Concluding Remarks
The paper classifies multivariate skew distributions into four forms and clarifies their nesting relationships. Existing finite-mixture work has focused on restricted and unrestricted forms, leaving more general mixture models as a direction for further research.
- Classification: The proposed classification distinguishes restricted, unrestricted, extended, and generalized forms of multivariate skew distributions.The classification provides a schematic framework for organizing these distributional variants.
- Classification: Restricted and unrestricted skew forms are not nested, coincide in the univariate case, and are both special cases of the extended form.The extended form is itself a special case of the generalized form.
- Future research: Existing finite-mixture research has investigated only restricted and unrestricted multivariate skew distributions.Mixtures based on more general forms are identified as an area for further research.