Source-linked AI summary
What are the true clusters?
Christian Hennig
TL;DR
The paper addresses how to define and assess “true clusters” when clustering aims and contexts differ and no unique truth is generally available. It combines constructivist philosophy with Hasok Chang’s active scientific realism, reviews context-dependent cluster concepts and formal definitions, and discusses implications for method comparison and practical choices. It concludes that clustering can remain scientific through explicit requirements, transparent comparison, and open communication rather than uniqueness.
Problem
Cluster analysis lacks a generally agreed definition of “true clusters”, while researchers use it for different aims and often leave their target cluster concept unspecified.
Method
The paper combines constructivism and active scientific realism, then examines context-dependent aims, cluster characteristics, truth definitions, method comparison, and practical decisions.
Results
The paper concludes that researchers should specify the intended kind of truth and real-cluster requirements, using formal definitions as clarifying constructs rather than ultimate truths.
Takeaways & Limitations
Clustering becomes scientific through explicit concepts, transparent comparison, and communication that remains open to constraints from reality outside researchers’ control.
Takeaways & Limitations
All proposed definitions have shortcomings: data-only definitions do not capture unobserved truth or generalization, external classes may not represent target clusters, and probability-model definitions can be unstable.
Abstract
from arXiv · showhide
Constructivist philosophy and Hasok Chang's active scientific realism are used to argue that the idea of "truth" in cluster analysis depends on the context and the clustering aims. Different characteristics of clusterings are required in different situations. Researchers should be explicit about on what requirements and what idea of "true clusters" their research is based, because clustering becomes scientific not through uniqueness but through transparent and open communication. The idea of "natural kinds" is a human construct, but it highlights the human experience that the reality outside the observer's control seems to make certain distinctions between categories inevitable. Various desirable characteristics of clusterings and various approaches to define a context-dependent truth are listed, and I discuss what impact these ideas can have on the comparison of clustering methods, and the choice of a clustering methods and related decisions in practice.
1 Introduction
Cluster analysis lacks a universally agreed definition of “true clusters” because applications pursue different aims and cluster concepts. The paper proposes a philosophically informed, context-dependent approach to choosing, assessing, and interpreting methods.
- Cluster analysis groups objects represented by variables, dissimilarities, or graph edges, with memberships that may be partitioned, overlapping, non-exhaustive, crisp, or fuzzy.
- There is no agreed definition of a cluster, and many proposed methods do not formally specify the “true clusters” they aim to find.
- Different applications use clustering for different aims, so the meaning of “cluster” varies across situations.
- The paper argues that researchers should identify the problem-specific truth they seek and the cluster characteristics each method is suited to finding.
- Its structure develops a philosophical basis, context-dependent clustering aims, definitions for comparing methods, and practical consequences for method selection and interpretation.
2 Philosophical background
The paper combines constructivism with Hasok Chang’s active scientific realism to explain how clustering concepts are constructed yet constrained by realities experienced as outside human control. Natural kinds and categorization therefore support plural, context-sensitive accounts rather than unique clustering truths.
- Constructivism and science: Constructivism treats personal and social worldviews as constructed through bodily, cognitive, and communicative activity, while recognizing that construction is constrained rather than arbitrary.
- Constructivism and science: Science is described as a social effort to build a stable, shared, criticizable view of the world, assessed through stability, agreement, and pragmatic use rather than direct access to objective truth.
- Active scientific realism: Active scientific realism defines reality through resistance to human will and emphasizes continual, pluralistic knowledge-seeking through confrontation with observed realities outside researchers’ control.
- Natural kinds: The concept of natural kinds is valuable not as proof that categories uniquely match reality, but as a description of categorizations that seem difficult to escape under such constraints.
- Natural kinds: Scientific classification should therefore be guided by observation toward stable agreement about legitimacy and use, rather than assumed uniqueness.
- Categorization: Human categorization is pluralist and context-dependent, making cognitive theories useful methodological inspiration but limited as definitions of true clusters for data analysis.
3 Clustering aims and cluster concepts
Clustering aims range from discovering meaningful real structures to constructively organizing data for practical purposes, and each aim requires a corresponding cluster concept. The paper therefore links method choice to the desired characteristics, data connection, and intended use of the clusters.
- A list of aims of clustering: Cluster definition and methodology must adapt to the specific application because different aims imply different meanings of “cluster”.
- A list of aims of clustering: Applications include species delimitation, medical classification, archaeology, image segmentation, object recognition, database organization, exploratory analysis, and information reduction.
- Realist and constructive aims of clustering: Realist aims seek meaningful structures experienced as real, whereas constructive aims divide data for pragmatic reasons without requiring essential differences between groups.
- Realist and constructive aims of clustering: Even realist clustering requires researchers to connect the intended real structure to available data using subject-matter knowledge and methodological decisions.
- Realist and constructive aims of clustering: Depending on the data connection, suitable methods may range from graph-theoretic approaches for interbreeding information to methods tolerant of heterogeneous species distributions; k-means or complete linkage may be inappropriate when large within-cluster distances are expected.
- Desirable characteristics of clusters: Desirable cluster characteristics include small within-cluster dissimilarities, large between-cluster dissimilarities, model fit, centroid representation, stability, connected high-density regions, and specific shapes.
- Desirable characteristics of clusters: Other requirements may include few variables, correspondence with external information, within-cluster feature independence, similar cluster sizes, and a low number of clusters.
- Desirable characteristics of clusters: These characteristics require precise operational definitions because choices such as maximum versus average within-cluster dissimilarity affect how outliers and separation are treated.
4 Definitions of true clusters
True clusters can be defined through data-based objectives, external classifications, or probability models, but each definition is context-dependent, idealized, and limited. Explicit formalizations nevertheless clarify cluster concepts and enable transparent method comparisons, provided they are not treated as ultimate truth.
- General principles: Mathematical definitions of true clusters are necessarily idealized, because different situations require different clustering concepts and formal objects cannot capture all informal ideas.Discrepancies between formal definitions and complex ideas about reality should remain visible.
- Definitions based on the data alone: A data-based definition is appropriate when the practical aim exactly matches its loss function, but it can be unsuitable for other aims such as identifying high-density regions.For k-means, the number of clusters must also be specified in advance.
- Definitions based on the data alone: Defining truth by the k-means objective makes k-means optimal by construction, creating a potentially tautological comparison while still permitting algorithmic optimization studies.Quality indices can aggregate desirable characteristics, but direct optimization may be infeasible or inappropriate when constraints are needed.
- Definitions based on external information: External true classes often provide weak clustering benchmarks because supervised classes may lack the desired cluster characteristics, omit alternative categorical groupings, or be unavailable in real applications.Comparisons are informative only when the external classes model the clusters sought in the new dataset.
- Definitions based on probability models: Probability-model definitions, such as treating mixture components as clusters, offer clear prototypes but depend on identifiable models and may not justify component-wise interpretations.Mixture components can be close, produce fewer modes than components, contain large within-component distances, or be unstable under slight model misspecification.
- Limitations of formal definitions: The same distribution can support genuinely different true clusterings, so definitions should be used as clarifying constructs for transparent comparison rather than as ultimate clustering truth.A five-component Gaussian mixture can have four modes, three high-density level sets, and a distinct 5-means partition.
5 Implications for cluster analysis research and practice
Choosing and assessing clustering methods requires specifying the desired cluster concept and recognizing that several characteristics, purposes, and preprocessing decisions may be relevant. Comparisons are useful when they clarify method characteristics rather than produce a universal ranking.
- Researchers should specify what kind of truth and what characteristics define a real cluster before choosing a clustering method.
- Method choice is harder when clustering serves several purposes or when desired characteristics are poorly defined, especially in exploratory analyses.
- Cluster concepts often require revision after examining the resulting clustering, such as enforcing links to external variables or avoiding unusably small clusters.
- Application-independent method comparisons remain useful when they reveal method characteristics, whereas aggregating varied simulations into one ranking can obscure those differences.
- Standardization, transformations, dissimilarity measures, variable selection, and dimension reduction should likewise be chosen according to context and clustering aims.
6 Conclusion
The paper rejects a universal, context-independent clustering truth while arguing that context-sensitive analyses can remain scientific through transparent concepts, decisions, and validation. Its philosophical perspective treats clustering as open both to researcher judgment and to reality outside the researcher’s control.
- A desire for uniqueness and context-independent objectivity can discourage researchers from specifying desired cluster characteristics and selecting methods accordingly.
- Scientific clustering requires transparent decisions and rationales while allowing reality outside the researcher’s control to challenge initial choices.
- The proposed perspective combines contextual dependence on aims and researcher decisions with scientific transparency and openness to external reality.
- The plurality of clustering truths is especially visible in cluster analysis, while belief in a unique natural truth can also create problems elsewhere in data analysis.