Source-linked AI summary
Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text
Sebastian Gehrmann, Elizabeth Clark, Thibault Sellam
TL;DR
NLG evaluation has persistent problems across datasets, automatic metrics, and human assessment, while improved practices are rarely adopted. This paper surveys those obstacles and proposes evaluation reports and related processes focused on model shortcomings. It concludes that current evaluation is unsustainable for distinguishing improved models and that many proposed practices remain only partially followed.
Problem
NLG evaluation suffers from dataset coverage problems, similarity-focused metrics, high-variance human judgments, and poor replicability, making improved evaluation practices necessary.
Method
The paper surveys critiques of human and automatic evaluations and NLG datasets, synthesizes responses, proposes evaluation reports, and analyzes 66 recent NLG papers against its suggestions.
Results
Current evaluation is not sustainable because model differences are increasingly difficult to detect from surface-level phenomena and careful annotation is required to characterize output quality.
Takeaways & Limitations
Evaluation reports should focus on model limitations and combine complementary metrics, rigorous human evaluations, evaluation suites, and data releases enabling re-analysis.
Takeaways & Limitations
The survey is constrained by subjectivity in evaluation and dataset choices affecting representation, language coverage, communicative goals, and noise.
Abstract
from arXiv · showhide
Evaluation practices in natural language generation (NLG) have many known flaws, but improved evaluation approaches are rarely widely adopted. This issue has become more urgent, since neural NLG models have improved to the point where they can often no longer be distinguished based on the surface-level features that older metrics rely on. This paper surveys the issues with human and automatic model evaluations and with commonly used datasets in NLG that have been pointed out over the past 20 years. We summarize, categorize, and discuss how researchers have been addressing these issues and what their findings mean for the current state of model evaluations. Building on those insights, we lay out a long-term vision for NLG evaluation and propose concrete steps for researchers to improve their evaluation processes. Finally, we analyze 66 NLG papers from recent NLP conferences in how well they already follow these suggestions and identify which areas require more drastic changes to the status quo.
1 Introduction
NLG evaluation practices often miss important quality and robustness problems, while researchers rarely adopt proposed improvements. The paper distinguishes selecting a best-performing model from characterizing its limitations and advocates evaluation reports designed to expose shortcomings.
- Evaluation problems: Datasets often miss robustness tail effects and non-English languages, while automated metrics poorly measure fine-grained quality and human evaluations are rarely replicable.Human ratings also exhibit high variance because evaluation procedures are insufficiently documented.
- Evaluation problems: Newer models make surface-level evaluation less informative because existing methods are poorly suited to detect attribution and content problems.Deeper analyses reveal shortcuts, overfitting, hallucinations, and misalignment with communicative goals.
- Barriers to adoption: Researchers propose evaluation improvements, but incentive mismatches, resource demands, and collaborative dataset refinement create a circular dependency that limits adoption.Better evaluations require resources such as model outputs and large quantities of human assessments, while better models are expected to use better evaluations.
- Paper approach: The paper surveys critiques of evaluation approaches and NLG datasets, then describes how model developers can improve evaluations using methods available today.Its central distinction is between selecting the best model and characterizing how good that model really is.
- Proposed direction: The authors propose evaluation reports centered on model shortcomings, combining complementary automatic metrics, rigorous human evaluations, and data releases for re-analysis.These reports are intended to extend existing work on model documentation and formal evaluation processes.
- Evidence from recent papers: In 66 recent papers, improved-evaluation practices appeared at an average rate of 27%, with 84% reporting multiple datasets but only one contributing to dataset documentation.More than 28% of papers pointed out dataset issues, yet future researchers were generally left to re-identify them.
2 Background
NLG broadly covers language production tasks, but this survey focuses on conditional text generation with explicit communicative goals. Evaluations may be intrinsic or extrinsic, and evidence from common tasks cannot automatically be assumed to generalize across all NLG tasks.
- Scope: NLG now broadly includes summarization, machine translation, paraphrasing, and story generation, although this survey focuses on conditional generation tasks.Conditional NLG models generate natural language y from structured or natural-language input x by maximizing p(y|x).
- Scope: The survey excludes multimodal tasks and tasks with non-textual outputs because they require substantially different evaluation processes.Examples include image captioning, speech-to-text, sign language, and audio output.
- Scope: In-scope NLG tasks have explicit communicative goals requiring content and structure planning plus fluent, error-free realization.Evaluations must capture all these aspects, making NLG evaluation especially challenging.
- Evaluation categories: Intrinsic evaluation assesses text itself, whereas extrinsic evaluation measures how text affects people performing a task.Intrinsic approaches include human ratings and automatic metrics.
- Evidence boundaries: Meta-evaluations most often study summarization and machine translation, so their findings may not transfer directly to every NLG task.The survey notes this limitation and adopts a cautious worst-case assumption that failure modes may transfer across tasks.
3 Challenges of Automatic Evaluation
Automatic NLG metrics often treat similarity to references as a proxy for quality, but this can reward surface overlap while missing meaning, factuality, and task-specific dimensions. Their validity is further constrained by language coverage, implementation choices, reference artifacts, and evaluation designs, motivating multidimensional and self-critical evaluation.
- 3.1 The Status Quo: Most metrics are designed only for English, limiting their applicability as non-English NLG models become more prevalent.Machine translation metrics are a notable exception, while multilingual metric development remains difficult.
- 3.2 Similarity to References is a Red Herring: Similarity-based metrics can reward surface overlap while overlooking semantic differences, unsupported information, and the broad output space of open-ended tasks.These problems are especially evident when references are flawed or do not match the information in the input.
- 3.2 Similarity to References is a Red Herring: Reference artifacts such as translationese can cause BLEU, METEOR, and BERTSCORE to fail to reward good translations.Low-quality references favor systems that produce similarly low-quality outputs.
- 3.2 Similarity to References is a Red Herring: Learned and embedding-based metrics also break down on simple adversarial examples, showing that greater flexibility does not eliminate dependence on reference coverage.Studies report that current metrics can still rely on surface-level features despite their distributional representations.
- 3.3 Do Benchmarks Help?: Evidence on metric validity is mixed: ROUGE does not significantly correlate with several content and linguistic quality measures, while metric correlations can change with human-evaluation methodology.In machine translation, DA was later found insufficient for identifying a best metric, whereas metrics correlated better with MQM, including metrics trained on DA annotations.
- 3.2 Similarity to References is a Red Herring: Implementation and reporting choices can produce non-replicable, inflated, or incorrectly comparable scores.Examples include unclear ROUGE parameters, unreported pretrained-model hashes, and unspecified handling of multiple references.
- 3.3 Do Benchmarks Help?: Benchmarks are necessary but should remain self-critical and use diverse evaluation approaches because they can amplify data problems and encourage reliance on inadequate metrics.The authors distinguish selecting the best model from characterizing its actual quality and limitations.
4 Challenges of Human Evaluation
Human evaluation is necessary but not a reliable panacea: judgments vary with criteria, framing, instruments, annotators, sample sizes, and text properties. These inconsistencies make evaluations difficult to compare and can obscure meaningful differences between models.
- Measurement: Human evaluations vary across 204 quality dimensions mapped to 71 criteria, with some dimensions defined or applied differently across studies.These disparities complicate comparisons and benchmarking improvements.
- Measurement: Over 50% of 478 evaluation questions did not define the criterion, and 65% did not report the exact evaluator question.Insufficient documentation limits replicability and interpretability.
- Measurement: Question framing, evaluated text, and measurement instruments can change human-evaluation results, while summary length correlates with judgments such as informativeness.Parallel annotation tasks can reduce correlations between dimensions, but substantially increase evaluator cost.
- Statistical significance: Human-rating studies are often underpowered because the median evaluation contains only 100 generated texts, while detecting a 1-point WMT difference may require 10,000 perfect judgments.Small samples make small model differences difficult to detect, and higher-quality models can require more judgments.
- Annotators and subjectivity: Annotation reliability is difficult to establish because most surveyed studies fail reliability tests and rating variability reflects both subjectivity and rater characteristics.Crowdsourcing introduces additional concerns, including variable results, inadequate instructions, technical problems, and ethical issues around compensation and working conditions.
- Annotators and subjectivity: Disagreement does not always indicate evaluator error because text quality depends on values, dialect, context, style, and the representativeness of the annotator population.Increasing annotations can reduce bias only under assumptions about annotator representativeness.
5 Challenges with Datasets
NLG datasets encode consequential choices about representation, communicative goals, collection, and test-set construction, yet these choices and their limitations are often underdocumented. Dataset artifacts, noise, overlap, and task mismatches can inflate results or misattribute failures to models.
- Representation: Among 20 summarization papers from 2021, CNN/DM and XSum were used five and four times respectively, illustrating concentration on a small set of benchmarks.Such concentration reinforces existing design decisions and limits diversity in evaluation.
- Representation: Dataset design choices affect who and which languages are represented, while popular NLG datasets remain concentrated in English and often exclude dialects.The lack of popular corpora for measuring dialect disparities makes these gaps especially difficult to assess.
- Design choices: Dataset creators should specify assumptions and report curation decisions and limitations in structured formats, because there is no one-size-fits-all curation solution.Structured documentation supports interpretation, contextualization, and later analysis of performance results.
- Shortcuts and task fit: Positional bias in news summarization lets models exploit salient information appearing early, inflating results when test sets share the same bias.Selecting the first three sentences was a strong baseline, and only one system significantly outperformed it on DUC-2001.
- Shortcuts and task fit: Reusing datasets for incompatible tasks can make benchmark results misleading, as CNN/DM homepage bullet points were designed for reading comprehension rather than summarization.This reuse occurs alongside concentration on very few datasets.
- Noise and faithfulness: Over 70% of XSum references contain external hallucinations, while WikiBio references realize less than half of attributes and often score low in faithfulness.These problems weaken the correspondence between references and intended communicative goals.
- Noise and faithfulness: Cleaning the E2E NLG dataset reduced slot-error rates by up to 97%, showing that apparent model failures can instead originate in noisy data.Few datasets use multiple post-editing steps to ensure a low noise ratio.
6 Suggestions for NLG Researchers
The paper calls for evaluation practices that document data, metrics, human judgments, errors, and model behavior across contexts. It emphasizes reusable evaluation reports, multiple metrics, extrinsic testing, and audits focused on limitations rather than single performance scores.
- Documentation, Releases and Maintenance: Evaluation progress requires improved documentation of data collection, dataset limitations, social impact, and human evaluation procedures.The paper recommends data cards, human evaluation datasheets, and clearer rater qualifications.
- Documentation, Releases and Maintenance: Datasets should be treated as dynamic resources, with updated versions, broader coverage, evaluation suites, and released model outputs for re-analysis.The authors specifically argue that dataset updates and output releases support metric development, validation, and meta-evaluation.
- Metrics and Extrinsic Evaluation: Researchers should use complementary metrics and avoid relying on a single score, while expanding extrinsic evaluations tied to users, tasks, and situational context.Extrinsic evaluation shifts attention from the appearance of text toward its content, purpose, and usefulness.
- Community and Peer Review: Adoption requires peer-review changes that reward resource papers, rigorous evaluation documentation, phenomenon-specific analysis, and reimplementation when existing benchmarks are inappropriate.The paper also argues that authors should show brittleness and a path toward improvement rather than treating empirical gains as the only contribution.
- Model Audits and Evaluation Reports: Evaluation reports should use model audits, fine-grained error analyses, and stimulus-based comparisons to characterize limitations and eventually support performance guarantees.The proposed reports are intended to document what breaks models across input types and deployment scenarios.
7 Every Cloud has a Silver Lining
An analysis of 66 papers finds partial adoption of the recommendations but no consistent evaluation standard. Papers commonly use multiple datasets and human or multiple-metric evaluations, while dataset documentation, issue remediation, and rigorous human-evaluation design remain weak.
- Overall Analysis: 36.7% of 2046 judgments were positive, while paper scores ranged from 6.5% to 58.1% with an average of 27.3%.The median score was 25.8%, indicating broad variation in recommendation coverage.
- Basic Evaluation Setup: 84% of papers used multiple datasets and 73% reported human evaluation, but only 38% motivated dataset choice and 30% motivated metric choice.The setup is often present, while documentation of evaluation decisions and claims is less consistent.
- Metrics and Analysis: 57% of papers reported metrics from different categories rather than relying only on lexical overlap.Some papers also developed metrics targeting the specific property being claimed.
- Datasets: 29% of papers identified dataset issues, but only one contributed to documentation and only 3/13 worked toward solutions while releasing dataset updates.The authors highlight documentation and maintenance as a major opportunity for future evaluation research.
- Human Evaluation: Among papers reporting human evaluations, 82% stated what was measured and 58% documented who evaluated it, but none estimated the required number of annotations.The authors note that their criteria were permissive, so these results may look more positive than stricter assessments.
- Releases and Reproducibility: Almost none of the papers released model outputs or human-evaluation data, and 37% did not use the same metrics environment for comparison.These omissions can slow evaluation research and hinder comparable re-analysis.
8 Discussion
The paper argues that perfect evaluation is unattainable because datasets and evaluations cover only subjective subsets of possible inputs and effects. It therefore favors practical infrastructure, accountability, interpretability, attribution, and external evaluation while recognizing important scope limits.
- The perfect evaluation is a white whale: Evaluation cannot fully eliminate subjectivity because datasets and evaluations reflect only a small subset of possible inputs.The paper frames perfect evaluation as an unattainable ideal rather than a final endpoint.
- Adoption and Accountability: Evaluation reports and checklists are presented as ways to make improved practices practical and hold model developers accountable through peer review.The authors argue that adoption depends on reviewers requiring these practices.
- Model Interpretability: Interpretability tools can reveal model shortcuts and systematic errors that belong in evaluation reports, although interpretability is treated as orthogonal to evaluation.The paper omits a broader interpretability-literature discussion while acknowledging its usefulness for evaluation.
- NLG is not ML and also not NLU: Because NLG lacks a universal equivalent of accuracy or F1-Score, its evaluation problems extend beyond one-size-fits-all machine-learning analyses.The survey emphasizes the distinctive complexity of evaluating natural-language outputs.
- NLG is not ML and also not NLU: Explicit attribution of generated information to sources is identified as important for future evaluation processes, alongside unresolved planning-stage evaluation.Current practices remain focused primarily on output forms.
- The use of models and external evaluation: Intrinsic evaluation cannot capture all external effects, since model behavior may be acceptable in one context and undesirable in another.Cultural background and norms can affect how generated language is perceived.
- Better metrics will lead to better models: Metrics are not reliable proxies for task performance, limiting approaches that directly optimize them through reinforcement learning.The paper cites evidence that reinforcement-learning improvements can be unrelated to the intended training signals.
9 Conclusion
The conclusion finds that current NLG evaluation is unsustainable because surface-level measures and problematic datasets do not adequately characterize improved models. It proposes actionable evaluation reports and related practices, while noting that adoption requires peer-review accountability and lacks a consistent standard.
- Conclusion: Current NLG evaluation is unsustainable because models increasingly require careful annotation to characterize quality and distinguish between systems.Problems in popular datasets further conflate evaluation results.
- Conclusion: The paper proposes evaluation reports that frame model limitations causally and aim toward performance guarantees across potential deployment scenarios.These reports are part of a broader set of actionable improvements for model developers.
- Conclusion: The 66-paper analysis finds partial adoption of the recommendations but no consistent standard for which evaluation aspects researchers must provide.The authors connect this gap to the need for peer-review changes that hold developers accountable.
A Surveying recent ACL, INLG, and EMNLP papers
The analysis examined how 66 ACL, EMNLP, and INLG papers from 2021 addressed evaluation criteria, using annotation instructions designed to produce upper-bound results without judging evaluation quality.
- The study annotated the presence of evaluation practices in 66 papers from ACL, EMNLP, and INLG in 2021.The criteria were designed so the results represent an upper bound rather than judgments of quality.
A.1 Paper selection
The paper selected generation-focused papers from the proceedings of ACL, EMNLP, and INLG, while limiting translation papers to avoid overemphasizing machine translation.
- Papers were selected when their titles referenced work on a generation problem.Translation papers were included only when their titles concerned generation broadly, yielding about 10–15 translation papers.
Make informed evaluation choices and document them
The evaluation criteria emphasize informed choices and transparent reporting across datasets, languages, and metrics.
- Papers should evaluate on multiple datasets unless the addressed task explicitly has only one available dataset.This criterion distinguishes broader evaluation from task settings with limited dataset availability.
- Papers should explain why each particular dataset was chosen rather than citing only its use in previous work.Introducing a dataset is treated as not applicable for this criterion.
- Papers should motivate the choice of each metric instead of relying only on precedent.The criterion asks for explicit reasons for selecting the metrics used.
- Papers should include evaluation on non-English language whenever at least one evaluated dataset contains non-English language.The criterion is satisfied by any evaluated dataset with non-English language.
Measure specific generation effects
The criteria broaden evaluation beyond aggregate automatic scores by requiring diverse metrics, human-study rigor, targeted analyses, failure reporting, and reproducible evaluation resources. The setup itself is recall-oriented, so it can overlook nuanced shortcomings and uncovered cases.
- Measure specific generation effects: Automatic evaluation should combine metrics from at least two different families, such as QA-based and lexical metrics, rather than ROUGE and BLEU alone.ROUGE and BLEURT count as different families under this criterion.
- Measure specific generation effects: Papers should avoid generic claims about overall “quality” and instead report improvements in specific generation aspects.Abstract-level claims such as “we outperform baselines” fail this criterion when they lack aspect-specific wording.
- Measure specific generation effects: Evaluation should address data issues, document datasets, release updated versions when appropriate, and create or release targeted evaluation suites.The criteria include data cards or updated documentation, updated data releases or loaders, fine-grained splits, and released data or code.
- Measure specific generation effects: The criteria also encourage retrained baselines, consistently recomputed metrics, and releases of validation, test, non-English, and human-evaluation outputs.For a newly introduced dataset evaluated alone, baseline-retraining, metric-recomputation, and some output-release criteria may be not applicable.
- Measure specific generation effects: Human evaluations should report their setup and participants, estimate effect size or power, test significance, assess annotation validity, and discuss rater qualifications.The criteria cover questions and response procedures, platforms and participant counts, statistical testing, agreement or gold-rating comparisons, and required background.
- Measure specific generation effects: Model evaluation should include disaggregated results, non-i.i.d. test sets, causal analyses of modeling choices, and error or failure analysis.Feature-focused ablations count toward causal analysis, while architecture-only ablations do not.
- Measure specific generation effects: The recall-oriented prompts can ignore nuanced errors and omit possibilities not covered by the instructions.For example, documenting a study setup may receive a positive mark even without exact definitions of measurement categories.