Source-linked AI summary

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao, Jieran Li, Dongbiao Sun, José Miguel Hernández-Lobato, Hao Zhang, Xue Liu

arXiv:2608.17731v1cs.AI

TL;DR

Single-number diversity scores can produce conflicting, resolution-dependent comparisons of AI-generated content. The paper introduces diversity profiles, whose curves expose whether comparisons remain robust across parameter choices, yielding a more transparent evaluation framework.

  • Problem

    Diversity spans modalities, representations, and resolutions, making it difficult for a single scalar to summarize generated-content diversity reliably.

  • Method

    The paper evaluates parameterized diversity metric families as condition-aware curves across parameter domains, supported by axiomatic and high-dimensional geometric analyses.

  • Results

    Diversity profiles reveal robust comparisons when curves dominate across a parameter domain and scale-dependent trade-offs when curves cross.

  • Takeaways & Limitations

    Diversity evaluation should characterize comparisons across resolutions and metric families rather than select a single supposedly best scalar metric.

  • Takeaways & Limitations

    Profile interpretation depends on the embedding representation, distance or similarity function, metric family, and parameter domain.

Abstract

from arXiv · show

Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarities, and aggregate them into a single scalar score. Such scalar summaries are convenient, but they often encode different inductive biases and may yield contradictory rankings of the same sample sets. In this paper, we argue that diversity evaluation for AI-generated content is intrinsically under-specified when reduced to a single number. We first review representative diversity metrics, and then diagnose their limitations from two complementary perspectives: an axiomatic analysis showing that no representative scalar metric satisfies all desirable properties simultaneously, and an empirical analysis showing that high-dimensional representation spaces can induce concentrated, modality-dependent distance distributions. To address these issues, we propose diversity profiles: curve-valued, condition-aware summaries that evaluate a parameterized diversity family across a range of thresholds, scales, exponents, or orders under a specified representation and distance or kernel function. Diversity profiles reveal whether a comparison is robust across resolutions or instead depends on an arbitrary parameter choice. We instantiate profiles for several representative metric families and demonstrate their practical use in generative AI evaluation. Overall, diversity profiles provide a more transparent and resolution-aware framework for comparing the diversity of AI-generated content.

1 Introduction

Diversity evaluation for generative AI is important but intrinsically under-specified by a single scalar, because diversity spans modalities, representations, resolutions, and competing metric assumptions. The paper proposes diversity profiles—curve-valued summaries that expose parameter-dependent and metric-family-specific comparisons.

  • Motivation: Generative AI diversity evaluation seeks to prevent mode collapse, repetition, and narrow coverage across images, language, graphs, and molecules.In language generation, diversity can concern topics, reasoning paths, lexical choices, syntax, or semantics.
  • Illustrative example: A minimal example shows richness prefers A, average distance prefers B, and Circles, Vendi, and Magnitude rankings can change as threshold, order, or scale varies.Thus, a single scalar can obscure trade-offs and make conclusions depend on an arbitrary resolution.
  • Existing evaluation: Common metrics embed samples, compute pairwise distances or similarities, and aggregate them into one score, including average distance, energy, Circles, Vendi, and Magnitude.These metrics are useful and often interpretable, but their different constructions encode different aspects of diversity.
  • Limitations: No representative scalar metric satisfies all desirable axioms simultaneously, while high-dimensional representations can make diversity comparisons scale- and modality-dependent.The paper identifies size monotonicity, duplicate invariance, distance monotonicity, and continuity as examples of competing desirable properties.
  • Proposed framework: Diversity profiles evaluate parameterized metric families across meaningful ranges under specified representations and distance or similarity functions.They support comparisons across resolutions and metric families, helping distinguish robust conclusions from metric-family-specific ones at modest additional computational cost.

2 Existing Diversity Metrics for AI-Generated Content

Existing diversity metrics commonly embed generated items, compute pairwise relations, and reduce them to a scalar, while also including similarity-based and count-based alternatives. Representative families capture average separation, penalties for close pairs, thresholded subset size, kernel-spectrum diversity, and effective metric-space size, each with distinct assumptions and interpretations.

  • Distance-based metrics: Distance-based metrics represent generated items in an embedding or feature space, compute pairwise distances, and summarize the distance matrix with a permutation-invariant scalar.Their geometry depends on the chosen embedding model and distance function, and they use only pairwise relations among generated items.
  • Distance-based metrics: Representative distance-based metrics include AvgDist, Energy, and packing number, which respectively summarize average separation, penalize small distances, and count mutually separated subsets.Energy uses exponent s to control the penalty for small pairwise distances and may require ε > 0 to avoid singularities.
  • Similarity-based metrics: The Vendi score is similarity-based and derives diversity from the eigenvalues of a normalized positive-semidefinite similarity matrix.Order q yields Rényi-entropy variants, while q = 1 gives the exponential of Shannon entropy; the order-2 case relates to Rényi kernel entropy.
  • Other metric constructions: Magnitude provides an effective size of a finite metric space at scale t through a scale-dependent construction.The construction is the finite-space version of magnitude in metric geometry and requires nonsingular Z_t(D).
  • Count-based metrics: Count-based metrics such as Richness count unique generated elements and are interpretable, but they do not capture graded distances among distinct items.Domain-specific variants count unique words or n-grams in text and unique molecular scaffolds in cheminformatics.

3 Diagnosis of Existing Metrics

Existing diversity metrics face two complementary problems: no representative scalar metric satisfies all four desirable axioms, and high-dimensional representations concentrate distances in modality- and representation-dependent ways. These limitations make single-parameter diversity judgments potentially fragile and unreliable.

  • Axiomatic analysis: Four desiderata structure the axiomatic diagnosis: size monotonicity, the twin property, distance monotonicity, and continuity.Together, these axioms require sensible responses to added samples, duplicates, increased separations, and small distance perturbations.
  • Axiomatic analysis: No single representative metric satisfies all four axioms in its natural domain.The incompatibility suggests that one scalar score cannot reliably describe diversity in every setting.
  • High-dimensional geometry: As dimension grows, pairwise distances among random points can become nearly indistinguishable, weakening relative distance contrast.Nearest and farthest pairwise distances become similar in relative terms under standard concentration conditions.
  • High-dimensional geometry: Distance concentration weakens scalar diversity notions based on nearest neighbors, density, locality, and separation.These effects undermine distance-based quantities that many existing diversity metrics rely on.
  • High-dimensional geometry: Concentrated distance distributions differ substantially across modalities and distance definitions, making thresholds, scales, exponents, or orders fragile and representation-dependent.When distances occupy a narrow empirical range, large parameter regions may become saturated or uninformative.

4 Diversity Profiles

Diversity profiles represent diversity as condition-aware curves over parameter domains rather than single scalar scores. They expose whether comparisons remain robust across resolutions or depend on the chosen metric family and parameter.

  • Definition: A diversity profile specifies a representation, distance function, parameterized diversity family, and admissible parameter set.It is defined as P = (ϕ, d, µ, Θ), with the associated profile evaluated over scales, thresholds, or orders.
  • Profile properties: Profiles are permutation-invariant, but their parameter-dependent behavior differs substantially across metric families.Energy, Circles, Vendi score, and Magnitude define nontrivial profiles, whereas Richness and Average distance are scale-rigid.
  • Profile properties: Circles is a nonincreasing step function in threshold, interpreting small thresholds as near-duplicate sensitivity and large thresholds as separated-representative coverage.Its jumps occur at observed pairwise distances, making it a resolution-dependent packing number.
  • Interpreting comparisons: Uniform profile dominance yields a robust ordering, whereas curve crossings indicate scale-dependent trade-offs without an unconditional diversity ranking.The relevant parameter domain should be chosen from empirical distance or similarity distributions rather than arbitrary intervals.
  • Empirical illustration: In Figure 3, Energy favors set A, Magnitude favors set B, Circles crosses, and Vendi score shows negligible separation.The discrepancy reflects that metric families emphasize different aspects of diversity, including close-pair repulsion, threshold-based packing, and spectral effects.

5 Conclusion and Discussion

The paper concludes that scalar diversity metrics encode different inductive biases and that none analyzed satisfies all four desiderata simultaneously. Diversity profiles offer a more transparent, resolution-aware alternative, while open theoretical and joint quality–diversity evaluation challenges remain.

  • Conclusion: None of the representative scalar diversity metrics analyzed satisfies all four desiderata simultaneously.The paper reaches this conclusion through an axiomatic analysis of distance-based diversity metrics.
  • Open directions: Diversity profiles reduce dependence on a single arbitrary parameter, but their metric families remain theoretically unsatisfactory and require deeper analysis.The paper identifies relationships among diversity metrics, axioms, and profile-level properties as an open direction.
  • Limitations: Joint evaluation of fidelity, utility, and diversity remains an important challenge because quality and diversity are tightly coupled.Low-quality or off-distribution samples can appear diverse, while high-quality systems can still exhibit mode collapse or limited coverage.
  • Broader impacts: Diversity profiles encourage more transparent, resolution-aware, and less hyperparameter-dependent evaluation of generative AI systems.They can help reveal mode collapse, excessive repetition, narrow coverage, and metric-specific conclusions hidden by a single score.

A Proofs of Table 1

The proofs and counterexamples show that representative diversity metrics satisfy different subsets of the axioms, often only under explicit domain assumptions. Richness and Circle count preserve some desirable properties, whereas average distance, Energy, Vendi, and Magnitude each violate at least one key property.

  • Richness: Richness satisfies size monotonicity and the twin property because duplicates join existing equivalence classes.Adding an element either creates a new equivalence class or joins an existing one; adding an exact duplicate leaves the number of classes unchanged.
  • Average distance: Average distance violates size monotonicity and the twin property: duplicating one of two points changes the score from 1 to 2/3.The counterexample starts with two points at distance 1 and adds a duplicate of the first point.
  • Energy: Energy is distance-monotone and continuous for strictly positive off-diagonal distances, but fails size monotonicity and has no well-defined twin property on exact duplicates.Exact duplicates create zero off-diagonal distances, making D^-s singular; ε-regularization instead introduces a large negative penalty.
  • Vendi and Magnitude: Vendi and Magnitude have conditional continuity or monotonicity results but each has counterexamples involving size, twins, or distance monotonicity.Vendi can decrease after adding an exact duplicate or increasing dissimilarities, while Magnitude requires positive definiteness for size monotonicity, is undefined for duplicates, and can decrease as distances increase.

B Details of Figure 2

Figure 2 estimates empirical pairwise-distance distributions from 100,000 sampled pairs across image, text, and molecular data. It shows that distance concentration recurs across modalities, representations, and distance functions, while distribution shape and location depend on domain and representation.

  • Experimental setup: 100,000 randomly sampled pairwise distances estimate empirical distributions for data instances from each domain in Figure 2.Representations are computed with standard feature extractors before distances are measured.
  • Images: Images use 2048-dimensional ResNet-50 penultimate-layer features from ImageNet and pairwise cosine distances.RGB images are sampled from ImageNet and encoded with a pretrained ResNet-50 recognition model.
  • Natural language text: Natural-language question-answer pairs use 4096-dimensional Qwen3-8B embeddings from 2WikiMultiHopQA with cosine-distance dissimilarity.The embeddings are dense semantic representations produced by Qwen3-8B.
  • Molecular structures: Molecular structures use fixed-length 2048-dimensional ECFP fingerprints from ChEMBL with pairwise Tanimoto distances.The binary fingerprints encode local molecular substructures and are computed with RDKit.
  • Overall finding: Distance concentration recurs across modalities, representation types, and distance functions, but distribution location and shape depend on domain and representation.This demonstrates why high-dimensional representations can produce modality-dependent distance distributions.

C More Examples of Diversity Profiles

The paper extends diversity-profile examples beyond natural language to real image and molecular-structure data, using matched sampling and evaluation setups across modalities.

  • Examples across modalities: Two further real-data examples apply diversity profiles to images and molecular structures, alongside the natural-language example.The examples appear in Figures 4 and 5.
  • Experimental setup: For each modality, sets A and B contain 50 data items sampled uniformly at random under the same source, representation, and distance function.These choices follow the setup described in Appendix B.
  • Image examples: Figure 4 presents diversity profiles for two real image sample sets.
  • Molecular examples: Figure 5 presents diversity profiles for two real molecular sample sets.
Loading 2608.17731v1…