Source-linked AI summary
When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics
Charu Karakkaparambil James
TL;DR
Traditional coherence can misjudge evolving topics when vocabulary changes but meaning persists. This paper evaluates lexical-change-aware traditional and LLM-based metrics against human judgments across two models and three datasets, finding that metric validity varies by model and condition, supporting complementary rather than single-metric evaluation.
Problem
Traditional coherence measures may diverge from human judgments when temporal topics change vocabulary while preserving semantic continuity.
Method
The study compares traditional coherence and LLM-based semantic similarity with human judgments across 120 CoNTM and DLDA topics from NYT, DBLP, and arXiv, stratified by lexical change.
Results
Metric validity varies by condition: traditional coherence ranges from strong positive to negative human agreement, while LLM similarity is strong for CoNTM but inconsistent for DLDA.
Takeaways & Limitations
Lexical change makes reliance on a single evaluation metric unsafe, so traditional coherence and LLM-based semantic similarity should be reported as complementary signals.
Takeaways & Limitations
The study covers two models, three English-language datasets, and one LLM configuration, while category-level correlations use only six or seven topics.
Abstract
from arXiv · showhide
Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv, using three human annotators and Low, Medium, and High lexical-change categories. Traditional temporal coherence shows highly variable agreement with human judgments ($ρ$=-0.256 to 0.614). In contrast, LLM-based semantic similarity agrees strongly with human semantic judgments for CoNTM on NYT ($ρ$=0.609), DBLP ($ρ$=0.721), and arXiv ($ρ$=0.502), but is less consistent for DLDA. Lexical-change stratification reveals variation hidden by aggregate evaluation. We therefore advocate lexical-change-aware evaluation, jointly reporting traditional coherence and LLM-based semantic measures as complementary rather than interchangeable signals.
1 Introduction
Dynamic topic models can undergo substantial vocabulary change while preserving semantic continuity, making lexical coherence an incomplete measure of temporal topic quality. The section motivates lexical-change-aware evaluation that combines traditional coherence with LLM-based semantic similarity as complementary signals.
- Dynamic topics may use substantially different vocabularies over time while describing one continuous underlying phenomenon.This temporal evolution creates a distinction between lexical change and semantic continuity.
- Traditional coherence measures based on statistical relationships between topic words do not always correspond to human interpretability.Prior work has also questioned their validity across newer models and application settings.
- Temporal lexical change can lower co-occurrence-based coherence even when humans recognize a coherent semantic evolution.A topic shifting from desktop and software to smartphone and cloud may still represent continuous digital-technology development.
- LLMs offer a complementary signal by assessing semantic relationships across different surface realizations of a topic.Their effectiveness can depend on the evaluation task, motivating their extension to temporal topic evolution.
- The study compares CoNTM and DLDA across three temporally structured corpora and finds that metric validity varies by model and condition.Lexical-change stratification and a multi-view framework are proposed because aggregate evaluation can obscure these patterns.
2 Related Work
Prior work questions coherence as a universal proxy for human judgment and explores LLMs as semantic evaluators. This section extends that direction to temporal topic evolution while treating traditional coherence and LLM-based semantic similarity as complementary signals.
- Topic model evaluation: Traditional coherence metrics based on statistical associations among topic words do not always align with human interpretability.This mismatch has been reported across neural topic models and different domains.
- LLM-based evaluation: LLM judgments can correlate strongly with human topic evaluations, but performance depends on the evaluation formulation.Existing work has primarily considered static topic representations.
- LLM-based evaluation: This work examines temporal topic evolution using traditional coherence and LLM-based semantic similarity as complementary signals for evaluating dynamic topics.
3 Experimental Setup
The experiments evaluate CoNTM and DLDA across three temporally structured datasets using human ratings, traditional coherence, and LLM-based semantic similarity. Lexical-change stratification and Spearman correlation assess whether metric agreement varies with vocabulary evolution.
- Models and datasets: 120 temporal topic trajectories come from CoNTM and DLDA evaluated on NYT, DBLP, and arXiv, with 20 topics per model–dataset combination.Topics are represented at three dataset-specific time points.
- Human evaluation: 1,800 scalar human judgments result from three annotators rating five topic-quality dimensions on five-point scales, averaged per topic.The dimensions include temporal coherence, smoothness, beginning and ending theme accuracy, and semantic similarity.
- Evaluation metrics: Traditional evaluation uses automatically computed Temporal Topic Coherence, while GPT-5.5 generates temporal semantic descriptions and scores similarity from 1 to 5.GPT-5.5 receives dataset, model, topic, timestamps, and top topic words through the OpenAI Responses API.
- Evaluation metrics: The traditional and LLM-based metrics are compared with corresponding human judgments rather than interpreted as direct comparisons between competing metrics.Traditional coherence is assessed against human temporal coherence, whereas LLM similarity is assessed against human semantic similarity.
- Lexical-change analysis: Spearman’s rank correlation (ρ) is the primary agreement statistic, and lexical-change categories test whether aggregate correlations conceal behavior linked to vocabulary evolution.Each model–dataset combination contains seven Low, seven Medium, and six High lexical-change topics; category results require caution because groups are small.
4 Results
Results show that metric–human agreement varies substantially across models, datasets, and lexical-change regimes. LLM semantic scores complement rather than replace traditional coherence, whose reliability is condition-dependent.
- Overall correlations: Traditional temporal coherence agrees significantly with humans for DLDA–DBLP (ρ = 0.614), DLDA–NYT (ρ = 0.531), and CoNTM–arXiv (ρ = 0.500).Agreement differs across experimental conditions.
- Overall correlations: DLDA–arXiv shows near-zero agreement (ρ = 0.007), while CoNTM–NYT shows a negative overall correlation, preventing universal interpretation of traditional coherence.A high traditional coherence score does not consistently correspond to human temporal interpretability across models and corpora.
- Lexical-change-stratified results: For CoNTM–NYT under Low lexical change, traditional coherence correlates negatively with human coherence (ρ = −0.786), whereas LLM semantic similarity correlates positively with human semantic judgments (ρ = 0.801).This contrast demonstrates metric instability within a single model, dataset, and lexical regime.
- Lexical-change-stratified results: DLDA traditional coherence is strongest for Medium lexical change, reaching ρ = 0.873 on NYT and ρ = 0.898 on DBLP, so lexical change does not uniformly invalidate it.The effect depends on interactions among lexical change, model behavior, and corpus characteristics.
- LLM semantic evaluation: Pooled LLM–human semantic agreement is strongest for Low lexical change (ρ = 0.556), remains significant for Medium change (ρ = 0.405), and weakens for High change (ρ = 0.294).The High lexical-change correlation is not statistically significant (p = .082).
- LLM semantic evaluation: CoNTM shows strong LLM–human semantic agreement across all three datasets, whereas DLDA is less consistent, with NYT at ρ = 0.445 and nonsignificant DBLP and arXiv correlations.These findings support using LLM judgments as a second evaluation view whose reliability requires validation, rather than replacing traditional metrics.
5 Discussion: When Does Lexical Change Mislead?
Lexical change should be treated as an evaluation context variable, not as evidence that traditional coherence is useless. Dynamic topics should instead be evaluated through complementary lexical, semantic, and lexical-change dimensions reported jointly.
- Evaluation context: Lexical change is an evaluation context variable rather than merely another topic-quality score.Its role is to contextualize evaluation rather than replace quality assessment.
- Complementary signals: Traditional coherence measures lexical or statistical compatibility, whereas semantic continuity asks whether changing words preserve the same underlying theme.An LLM can compare beginning and ending topic descriptions without requiring persistent vocabulary.
- Limits of lexical change: High lexical change does not automatically mean traditional metrics fail or LLMs succeed, because highly changed trajectories may be ambiguous between meaningful evolution, topic drift, and unrelated concepts.The LLM–human relationship weakens under several High-change conditions.
- Three dimensions: Dynamic topics should be evaluated along three complementary dimensions: Ctraditional, SLLM, and Lchange.These represent lexical/statistical coherence, semantic continuity, and the amount of vocabulary evolution, respectively.
- Joint reporting: The paper recommends reporting these dimensions jointly rather than collapsing them into fixed weights, because their behavior varies across models and lexical regimes.Low traditional coherence with high semantic similarity and high lexical change may indicate legitimate evolution, unlike poor scores on both coherence and semantic similarity.
6 Conclusion
The conclusion finds that traditional coherence and LLM-based semantic similarity provide variable, model-dependent evidence of dynamic topic quality. It therefore advocates lexical-change-aware, multi-view evaluation validated against corresponding human judgments.
- Conclusion: Traditional coherence shows substantial variation in agreement with human judgments, including strong positive, near-zero, and negative correlations.The evaluation examined lexical change, traditional coherence, LLM-based semantic similarity, and human judgments across CoNTM and DLDA on NYT, DBLP, and arXiv.
- Conclusion: LLM-based semantic similarity offers strong complementary evidence for several CoNTM conditions but remains model dependent.Its usefulness is therefore not uniform across dynamic topic-model settings.
- Conclusion: Dynamic topic evaluation should move beyond seeking one universal automatic metric because temporal topics are both lexical objects and semantic trajectories.Traditional coherence assesses whether words fit together, whereas LLM-based evaluation assesses whether meaning persists or evolves coherently.
- Conclusion: Lexical change identifies when the distinction between lexical coherence and semantic continuity matters.This motivates treating lexical and semantic signals as complementary views rather than interchangeable measures.
- Conclusion: The paper advocates lexical-change-aware, multi-view evaluation that jointly reports traditional and LLM-based metrics and validates them against corresponding human judgments.The proposed approach explicitly preserves both metric perspectives in dynamic topic-model evaluation.
Limitations
The study is limited by its scope, small category-level samples, single-configuration LLM analysis, and differing human validation questions. Future work should broaden evaluation and test robust, combined metrics against holistic human judgments.
- Study scope: The evaluation covers 120 topics from two models and three English-language datasets, limiting its scope.The study’s category-level experiments contain only six or seven topics.
- Sample size: Six or seven topics per category make individual correlations sensitive to outliers.This limitation affects the interpretation of category-level experiments.
- LLM analysis: The LLM analysis uses one configuration, while judgments may vary with prompting, model family, and generated semantic themes.The passage identifies these factors as potential sources of variation in LLM judgments.
- Metric validation: Traditional coherence and LLM semantic similarity are validated against different human questions, so their correlation coefficients should not be interpreted as directly equivalent.The supplied passage ends mid-word after “interpre,” but establishes that the metrics rely on different human validation questions.
- Future work: Future work should expand topics and annotators, evaluate more temporal topic models and LLMs, test prompt robustness, and assess lexical-change-aware metric combinations.It should also test whether such combinations predict holistic human judgments better than any individual metric.