Source-linked AI summary
Statistical power for cluster analysis
E. S. Dalmaijer, C. L. Nord, D. E. Astle
TL;DR
Biomedical clustering lacks clear guidance on power and appropriate data conditions. This study evaluates common clustering pipelines through simulation and finds higher power for c-means and finite Gaussian approaches than for traditional discrete methods.
Problem
A major concern is that clustering techniques in medical research are not necessarily used on data for which they were developed or validated.
Method
The study investigates the statistical power and accuracy of popular clustering approaches and offers guidance on applicability, sample size, and technique choice.
Results
Better power was observed for c-means and finite Gaussian approaches than for traditional discrete methods.
Takeaways & Limitations
The findings support considering c-means and finite Gaussian approaches when seeking higher power than traditional discrete methods.
Takeaways & Limitations
Common evaluation metrics may assume unrealistically large, well-separated, and non-overlapping clusters.
Abstract
from arXiv · showhide
Cluster algorithms are increasingly popular in biomedical research due to their compelling ability to identify discrete subgroups in data, and their increasing accessibility in mainstream software. While guidelines exist for algorithm selection and outcome evaluation, there are no firmly established ways of computing a priori statistical power for cluster analysis. Here, we estimated power and accuracy for common analysis pipelines through simulation. We varied subgroup size, number, separation (effect size), and covariance structure. We then subjected generated datasets to dimensionality reduction (none, multidimensional scaling, or UMAP) and cluster algorithms (k-means, agglomerative hierarchical clustering with Ward or average linkage and Euclidean or cosine distance, HDBSCAN). Finally, we compared the statistical power of discrete (k-means), "fuzzy" (c-means), and finite mixture modelling approaches (which include latent profile and latent class analysis). We found that outcomes were driven by large effect sizes or the accumulation of many smaller effects across features, and were unaffected by differences in covariance structure. Sufficient statistical power was achieved with relatively small samples (N=20 per subgroup), provided cluster separation is large (Δ=4). Fuzzy clustering provided a more parsimonious and powerful alternative for identifying separable multivariate normal distributions, particularly those with slightly lower centroid separation (Δ=3). Overall, we recommend that researchers 1) only apply cluster analysis when large subgroup separation is expected, 2) aim for sample sizes of N=20 to N=30 per expected subgroup, 3) use multidimensional scaling to improve cluster separation, and 4) use fuzzy clustering or finite mixture modelling approaches that are more powerful and more parsimonious with partially overlapping multivariate normal distributions.
Background
Cluster analysis enables data-driven identification of clinically relevant subgroups without a priori labels, but biomedical data often contain partially overlapping distributions that challenge standard clustering assumptions. This study therefore investigated the power, accuracy, applicability, sample-size requirements, and technique choices of common cluster-analysis approaches.
- Background: Cluster algorithms identify discrete subgroups in data without requiring a priori labelling, contributing to their growing use in biomedical research.Their increasing popularity was supported by implementation in mainstream software and applications including disease subtyping, treatment response, and behavioural phenotyping.
- Background: Biomedical datasets frequently contain partially overlapping multivariate normal subgroups, whereas clustering tutorials and evaluation metrics often assume strongly separated clusters.This mismatch raises uncertainty about when clustering is suitable for realistic biomedical data.
- Background: A typical cluster-analysis pipeline combines dimensionality reduction, subgroup identification, and outcome evaluation across methods based on centroids, linkage, density, fuzzy assignments, or mixture models.Dimensionality reduction addresses visualization and the curse of dimensionality, while mixture models estimate the probability that observations belong to each distribution.
- Background: The study simulated datasets to determine how subgroup size, separation, covariance structure, dimensionality reduction, and clustering approach affect statistical power and accuracy.It focused on practical guidance about applicability, required observations, and techniques providing stronger statistical power.
- Background: Dimensionality reduction predictably altered cluster-centroid differences, which ultimately drove outcomes, while agglomerative clustering produced highly similar results to k-means.Covariance structure did not affect outcomes in the reported simulations.
1) What drives … Results
Cluster separation, especially when driven by larger within-feature effects across more features, determined clustering performance more than covariance structure or sample size. MDS and UMAP improved separation and detection, while k-means achieved adequate power with about 20 observations per subgroup when separation was sufficiently large.
- Cluster centroid separation: Cluster separation increased with larger within-feature effects and more differing features, while covariance structure and cluster number or size had negligible effects after MDS.UMAP showed the same general pattern but was nonlinear and more variable, improving separation mainly at large differences or feature proportions.
- Membership classification (adjusted Rand index): K-means and both agglomerative approaches performed similarly across dimensionality-reduction methods, whereas HDBSCAN performed well primarily after UMAP.HDBSCAN likely left many observations unassigned while identifying denser centers.
- Data subgrouping (silhouette coefficient): Only a minority of datasets exceeded the silhouette threshold of 0.5, although dimensionality reduction—especially UMAP—raised scores and large effects across many features enabled detection.Without reduction, none were correctly identified as clustered; after UMAP, k-means and HDBSCAN detected weaker separations than after MDS.
- K-means: At Δ=4, k-means achieved at least 80% accuracy for the true cluster number with samples of about 20 per subgroup when clusters were equally sized.For equally sized clusters, that accuracy was also reached at Δ=3 with approximately 20 observations per subgroup.
- K-means: 20 observations per subgroup provided sufficient k-means power when cluster separation was Δ=4 or higher, with smaller subgroups requiring Δ=5 or higher.These conditions also yielded near-perfect identification of the true cluster number and 90–100% individual classification accuracy.
2 Clusters … 4 Clusters
Across two-, three-, and four-cluster simulations, HDBSCAN achieved sufficient detection power at 20–30 observations per subgroup when separation was Δ=4 or higher, whereas c-means reached this at about 20 observations per subgroup and Δ=3 or higher. C-means also produced high subgroup-count and membership accuracy, while HDBSCAN required stronger separation for comparable performance, especially with unequal subgroup sizes.
- HDBSCAN: For HDBSCAN, power depended primarily on separation after sample size exceeded a threshold, reaching 84% for unequal clusters at N=80 and Δ=6.For equal-sized clusters, power was 66% at N=40 and 83% at N=80 for two clusters with Δ=3, 66% at N=160 and 84% at N=80 for three clusters with Δ=3 and Δ=4, and 75% at N=80 for four clusters with Δ=4.
- HDBSCAN: HDBSCAN identified the true number of clusters most accurately at separations Δ=5 or higher, with roughly 70–80% accuracy at Δ=4 for equal-sized groups.Equal-sized groups reached this accuracy from N=40, whereas unequal 10%/90% groups required N=80 or more.
- HDBSCAN: HDBSCAN’s individual-membership classification exceeded chance from Δ=4 for equal-sized groups, but required Δ=8 and N=80 for 10%/90% groups.For equal-sized populations, this threshold held at N=40 for two and three groups and at Δ=3 for four groups; chance levels were 50%, 33%, and 25%, respectively.
- HDBSCAN / C-means: The simulations used silhouette scores, correct-clustering proportions, true-cluster-number proportions, and subgroup-assignment proportions across 100 iterations for each method.Figures 9 and 10 summarize these outcomes across sample sizes and simulated cluster separations.
- C-means: For c-means, detection power was driven mainly by separation: power was 81% at N=20 and Δ=5 for unequal clusters, and 77–82% at Δ=3 for two equal clusters.For three and four equal clusters, power was 76–77% at selected N=10–40 and Δ=3–4, increasing to 89–100% or higher with larger effects and samples.
- C-means: C-means achieved at least 80% accuracy for detecting the true number of clusters at Δ=4 with N=40, or at Δ=3 with about 20 observations per equal-sized subgroup.The reported summary describes near-perfect true-cluster-number accuracy under these conditions.
- C-means: C-means classified equal-sized subgroup membership above chance across tested conditions, reaching 80–90% or higher from Δ=3 and exceeding chance for unequal groups at N=40 and Δ=5.The summary reports 90–100% classification accuracy under the recommended conditions.
2 Clusters · 3 Clusters · 4 Clusters
Across two-, three-, and four-cluster simulations, c-means and Gaussian mixture modelling detected clustering with greater power and identified the correct cluster count at lower separations than k-means. Their higher fuzzy silhouette scores did not increase false positives, and silhouette-based cluster-count selection outperformed BIC for mixture modelling.
- Direct comparison of discrete and fuzzy clustering: For equally sized clusters, c-means showed 80-100% power at Δ=3, whereas k-means showed 70-100% power at Δ=4.The methods used fuzzy and traditional silhouette scores, respectively.
- Direct comparison of discrete and fuzzy clustering: The higher c-means sensitivity may partly reflect inflated fuzzy silhouette scores, although such inflation could also have increased false positives.The study tested this possibility using data with no clustering and with two to four equidistant clusters.
- Direct comparison of discrete and fuzzy clustering: Gaussian mixture modelling provided membership-confidence measures that enabled computation of a fuzzy silhouette coefficient comparable to c-means.This allowed the two non-discrete approaches to be assessed with the same fuzzy silhouette framework.
- Direct comparison of discrete and fuzzy clustering: The false positive rate was 0% for k-means, c-means, and Gaussian mixture modelling when no clustering was present.This rate was based on 100 iterations using the same simulated data for each algorithm.
- Direct comparison of discrete and fuzzy clustering: Fuzzy silhouette scores were higher than traditional silhouette scores, increasing cluster-detection likelihood without increasing false positives.The comparison used traditional k-means silhouettes versus fuzzy silhouettes from c-means and Gaussian mixture modelling.
- Direct comparison of discrete and fuzzy clustering: At Δ=3, c-means and Gaussian mixture modelling achieved 100% power, compared with 100% for k-means at Δ=4.Both non-discrete approaches were therefore more powerful than k-means at lower centroid separation.
- Direct comparison of discrete and fuzzy clustering: C-means and mixture modelling identified the correct number of clusters at lower separations than k-means.This comparison covered simulated datasets with two to four equally sized, equidistant clusters and separations from Δ=1 to 10.
- Direct comparison of discrete and fuzzy clustering: Silhouette-based cluster-count selection outperformed lowest-BIC selection for Gaussian mixture modelling.All methods were evaluated across guesses k=2 to k=7.
Discussion
The simulations show that cluster-analysis power depends primarily on subgroup separation, with c-means and finite Gaussian mixture modelling detecting weaker separation than k-means and HDBSCAN. Adequate power generally requires 20–30 observations per expected subgroup, while covariance structure has little effect.
- Dimensionality reduction: MDS increased subgroup separation by about Δ=1, whereas UMAP decreased separation below Δ=4 but increased it when original separation exceeded Δ=5.The authors therefore recommend MDS for multivariate normal distributions.
- Covariance structure: Cluster-analysis outcomes were unaffected by covariance structure or differences in covariance structure between subgroups.Centroid separation, rather than covariance structure, was the main driver of statistical power.
- Practical recommendations: The authors recommend using fuzzy clustering or finite mixture modelling when expected centroid separation is Δ=3 to 4, where k-means is less reliable.These methods can improve power without elevating the false positive rate.
- Main findings: 80% power required Δ=4 for k-means and HDBSCAN, whereas c-means and finite Gaussian mixture modelling achieved 80% power at Δ=3 without inflating false positives.Fuzzy and finite mixture approaches were more powerful than discrete clustering methods.
- Sample size: Sampling 20–30 observations per expected subgroup provided satisfactory power, particularly for k-means and HDBSCAN, with larger separation needed for small subgroups.Most algorithms performed optimally at lower separations only with at least 20–30 observations per subgroup.
Methods
The study used simulated multivariate-normal datasets to test statistical power and accuracy across subgroup structures, separations, covariance patterns, dimensionality-reduction strategies, and clustering algorithms. Analyses included multiple simulation designs comparing conventional, density-based, fuzzy, and Gaussian-mixture approaches.
- Simulation design: The primary datasets were multivariate normal distributions with standard deviations of 1, controllable mean vectors, and covariance structures ranging from uncorrelated to random or factor-based.Subgroup distributions could have identical or different covariance structures, including 3- or 4-factor patterns.
- Model comparison: Additional simulations directly compared k-means, c-means, and Gaussian mixture modelling for power and accuracy across one to four equally sized distributions with Δ=1–10.These comparisons used 120 observations and 2 uncorrelated features.
- Dimensionality reduction: Three dimensionality-reduction strategies—None, MDS, and UMAP—were evaluated before clustering.MDS was described as preserving inter-observation distances, whereas UMAP performs nonlinear dimensionality reduction.
- Clustering algorithms: Five clustering algorithms were tested: k-means, two agglomerative variants, HDBSCAN, and c-means fuzzy clustering.The hierarchical variants used Ward linkage with Euclidean distance or average linkage with cosine distance.
Declarations
This simulation-only study involved no participants, and the authors report public availability of data and software, no competing interests, funding support, and author contributions.
- No participants were tested because the study relied exclusively on simulations.
- The authors made their data and analysis software publicly available.
- The authors declared no competing interests and acknowledged support from the Templeton World Charity Foundation, AXA Fellowship, and UK Medical Research Council.
- ESD and CLN initiated the study, while ESD led its conceptualisation, methods, simulation, analysis, and manuscript drafting; all authors interpreted results, provided feedback, and approved the final version.