Source-linked AI summary

We're Afraid Language Models Aren't Modeling Ambiguity

Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, Yejin Choi

arXiv:2304.14399v2cs.CL

TL;DR

Ambiguity is fundamental to language, but benchmarks and pretrained language models have largely neglected the need to recognize and separate multiple readings. The paper introduces AMBIENT, a linguist-annotated entailment benchmark and evaluation suite, finding that ambiguity remains extremely challenging for LMs while ambiguity-sensitive NLI tools show promise for detecting misleading political claims.

  • Problem

    Pretrained LMs’ ability to recognize and disentangle ambiguity remains unstudied partly because ambiguous instances are systematically excluded from benchmarks, despite ambiguity’s importance for language understanding and communication.

  • Method

    The paper characterizes ambiguity through its effects on NLI entailment relations, builds the 1,645-example AMBIENT benchmark with disambiguating rewrites, and evaluates pretrained LMs and multilabel NLI models.

  • Results

    Ambiguity remains extremely challenging for LMs, including GPT-4, while a multilabel NLI model can recover fact-checker-flagged ambiguous political claims and identify previously unidentified ones.

  • Takeaways & Limitations

    Ambiguity-sensitive evaluation and tools could support clearer communication and help identify misleading language in real-world political claims.

  • Takeaways & Limitations

    AMBIENT’s size and diversity are limited by its data sources and expert-annotation effort, and the study covers only English.

Abstract

from arXiv · show

Ambiguity is an intrinsic feature of natural language. Managing ambiguity is a key part of human language understanding, allowing us to anticipate misunderstanding as communicators and revise our interpretations as listeners. As language models (LMs) are increasingly employed as dialogue interfaces and writing aids, handling ambiguous language is critical to their success. We characterize ambiguity in a sentence by its effect on entailment relations with another sentence, and collect AmbiEnt, a linguist-annotated benchmark of 1,645 examples with diverse kinds of ambiguity. We design a suite of tests based on AmbiEnt, presenting the first evaluation of pretrained LMs to recognize ambiguity and disentangle possible meanings. We find that the task remains extremely challenging, including for GPT-4, whose generated disambiguations are considered correct only 32% of the time in human evaluation, compared to 90% for disambiguations in our dataset. Finally, to illustrate the value of ambiguity-sensitive tools, we show that a multilabel NLI model can flag political claims in the wild that are misleading due to ambiguity. We encourage the field to rediscover the importance of ambiguity for NLP.

1 Introduction

Ambiguity is intrinsic to language and central to effective communication, yet pretrained language models’ ability to recognize and disentangle multiple meanings remains largely unstudied and highly limited. The paper introduces AMBIENT-based evaluations and shows that ambiguity remains challenging even for GPT-4, while ambiguity-sensitive models can help identify misleading political claims.

  • Motivation: Ambiguity supports efficient communication but requires listeners to recognize multiple interpretations, revise readings, and anticipate misunderstanding.It can also be used deliberately to send covert messages or mislead listeners.
  • Motivation: LMs need to handle ambiguous language for more effective dialogue, writing assistance, contextual adaptation, clearer communication, and detection of misleading language.
  • Research gap: Pretrained LMs’ ability to recognize ambiguity and disentangle possible meanings remains unstudied partly because benchmarks systematically exclude ambiguous instances.
  • Contributions: AMBIENT contains 1,645 English examples spanning lexical, syntactic, and pragmatic ambiguity, represented through entailment effects and disambiguating rewrites.Examples include plausible readings that receive different NLI labels.
  • Findings: The AMBIENT test suite evaluates direct disambiguation, interpretation recognition, and continuation distributions, and finds all three tasks extremely challenging, including for GPT-4.
  • Findings: A multilabel NLI model predicted the exact label set in only 43.6% of instances, while a case study recovered fact-checker-flagged ambiguous political claims and identified previously unrecognized ones.The case study indicates promise for ambiguity-sensitive tools in real-world communication.

2 AMBIENT

AMBIENT is an NLI benchmark designed to represent ambiguity through multiple possible label relations and corresponding disambiguations. It combines manually curated and automatically generated examples with expert annotation and validation, producing a 1,645-example dataset for detecting and resolving ambiguity.

  • Benchmark design: AMBIENT represents ambiguity as multiple possible NLI labels for a premise–hypothesis pair, with disambiguating rewrites for each plausible reading.Unambiguous examples retain a single label, enabling separate evaluation of ambiguity detection and resolution.
  • Data collection: The dataset combines manual curation targeting specific ambiguity types with automatic generation, heuristic filtering, and expert annotation to broaden coverage.The generated collection forms the bulk of AMBIENT.
  • Generated examples: Generated examples are produced from reasoning-pattern groups, sampled from InstructGPT, and filtered using a multilabel RoBERTa-large model that retains examples with probability ≥0.05 for multiple NLI labels.
  • Annotation and validation: Two experts annotate each generated example with label sets and rewrites, after which authors validate, consolidate, and optionally add interpretations.Annotators may discard offensive or low-quality examples.
  • Annotation quality: Validation agreement ranged from κ=0.44 for neutral to κ=0.65 for entailment and κ=0.62 for contradiction, indicating moderate to substantial agreement.
  • Dataset composition: 1,645 final examples comprise a 100-example development set and the remaining test set, with 35.2% of examples containing more than one label.The final dataset combines curated and generated-then-annotated examples.

3 Does Ambiguity Explain Disagreement?

The study finds that ambiguity substantially contributes to disagreement under single-label NLI annotation, while disambiguation enables reliable recognition of multiple readings and their entailment labels.

  • Annotation results: Fleiss κ increased from 0.12 for ambiguous examples to 0.67 for corresponding disambiguated examples.The first score indicates slight agreement, while the latter represents substantial agreement.
  • Annotation results: True disambiguations were marked plausible 96.7% of the time, compared with 46.7% for the distractor.On average, 93.7% of annotators accepted all true interpretations.
  • Annotation results: The majority-vote agreement rate for recognizing the full set of ambiguities and verifying labels was 89.7%.This rate was used as a reference point for later experiments.
  • Interpretation: Under single-label annotation, ambiguous inputs produce disagreement, but disambiguation largely resolves it by separating ambiguity from annotator subjectivity.Individual annotators can recognize multiple readings and their corresponding output labels.

4 Evaluating Pretrained Language Models

The evaluation tests whether pretrained LMs can generate disambiguations, recognize interpretations, and model interpretation-specific continuations. Performance remains difficult across these tests, including for GPT-4.

  • Evaluation setup: The evaluation measures direct disambiguation generation, interpretation recognition, and interpretation-sensitive continuation modeling.The tests use ambiguous AMBIENT instances and evaluate pretrained models including GPT-4.
  • Generating disambiguations: GPT-4 achieves 18.0% EDIT-F1 and 32.0% human-judged correctness for generated disambiguations.Human-judged correctness is far below the 89.7% crowdworker agreement rate on AMBIENT.
  • Generating disambiguations: Models often restate ambiguous sentences while affirming or negating the hypothesis instead of revising the ambiguity directly.Some such shortcuts are technically correct under human evaluation, but they may not identify the ambiguity’s source precisely.
  • Recognizing disambiguations: GPT-4 reaches 63.0% True/False accuracy, but answers all four templates correctly for only 2.5% of disambiguations.The all-template result is below the 6.25% random-guessing baseline.
  • Recognizing disambiguations: For 76% of disambiguation pairs, GPT-4 simultaneously accepts both possible meanings and claims each is the only meaning.This inconsistency appears across questions about the same ambiguous sentence.
  • Modeling interpretation-specific continuations: FLAN-T5 achieves the best KL ranking accuracy at 81.0%, while inconsistent trends show that results depend on how ambiguity competence is operationalized.ChatGPT and GPT-4 are excluded from this likelihood-based evaluation because their APIs do not provide likelihoods.

5 Evaluating Multilabel NLI Models

The paper evaluates whether existing NLI models can be adapted to predict multiple labels for ambiguous and unambiguous AMBIENT examples. Although performance exceeds random guessing, it remains well below crowdworker agreement.

  • Methods: The study fine-tunes existing NLI models for multilabel prediction across ambiguous and unambiguous AMBIENT examples.Models produce probabilities, label distributions, or label sets that are thresholded or classified into NLI label sets.
  • Methods: AMBIENT is too small to provide a training split for this evaluation, so the models use existing NLI data.The authors identify future AMBIENT-style annotation as a way to address this limitation.
  • Results: 72.5% is the highest macro F1, achieved by the WANLI-trained multilabel model.Macro F1 is measured on the original NLI examples.
  • Results: 43.6% is the best exact-match accuracy and 37.8% the best group exact-match accuracy, both achieved by the classifier over label sets.Group exact match requires correct labels for the original example and all its disambiguations.
  • Results: The best exact-match accuracy exceeds the 14.3% random baseline but remains below the 89.7% crowdworker rate.The authors conclude that existing-data fine-tuning leaves substantial room for improvement.

6 Case Study: Detecting Misleading Political Claims

The case study uses multilabel NLI on paraphrased political claims to detect ambiguity that can make fact-checking outcomes misleading. The method recalls many ambiguous claims, while false positives reveal additional unmentioned ambiguities.

  • Detection method: The method treats multiple NLI labels between a claim and its paraphrase as evidence that the claim is ambiguous.Paraphrases either preserve ambiguity or express a particular interpretation, making them useful for this detection strategy.
  • Evaluation: The evaluation paraphrases 200 CLAIMDECOMP claims five times each and applies the WANLI multilabel model.Claims are compared with their PolitiFact fact-checks, which are annotated for ambiguity or factuality issues.
  • Results: The method recalls 88.8% of ambiguous claims but achieves 12.4% precision.Inspection of false positives finds ambiguities that the fact-checks did not mention.
  • Implications: The case study suggests that fact-checking may need to account for claims having both true and false interpretations.The authors present political-claim detection as one use case for ambiguity-sensitive models.

7 Related Work

Prior NLP work has studied ambiguity in symbolic language analysis, label variation, question answering, and multimodal settings. This paper frames input ambiguity as distinct from task ambiguity and annotator subjectivity, and evaluates pretrained LMs beyond task-specific confidence modeling.

  • Ambiguity in NLP: Ambiguity has long been studied in syntactic and semantic parsing and coreference resolution but has received less attention in higher-level understanding and reasoning tasks.The paper positions its work as extending ambiguity analysis into pretrained language-model evaluation.
  • Question answering: Open-domain QA research addresses ambiguous event and entity references through clarification questions and natural-language disambiguation.AMBIENT differs by covering more diverse linguistic ambiguities than open-domain questions.
  • Human label variation: Related work distinguishes task ambiguity, annotator subjectivity, and input ambiguity, with this paper focusing on input ambiguity.The paper argues that uncertainty in the input should be characterized through its underlying reasons rather than only judgment distributions.
  • NLI beyond three-way classification: Prior NLI work models label variation with additional annotations, probability estimates, distributions, or disagreement labels; this paper predicts sets of NLI labels.The set-based formulation represents multiple readings directly.

8 Conclusion

The paper develops a benchmark and evaluations for ambiguity in language models, finding the task remains extremely challenging and motivating further study of ambiguity-sensitive applications.

  • The paper develops the first benchmark for evaluating whether language models recognize different readings of ambiguous text.
  • The authors find that recognizing and disentangling ambiguity remains extremely challenging for language models.
  • The paper encourages future work on contextual and emphasis sensitivity, systematic interpretation biases, and real-world ambiguity-sensitive applications.

Limitations

The dataset and analyses have limited size, diversity, and linguistic coverage, while the evaluation results may not generalize across task settings or extraction methods.

  • The dataset’s size and diversity are limited by its data sources and the effort required for expert annotation.
  • The study covers only English, whose ambiguity patterns may differ substantially from those of other languages.
  • The evaluation results do not guarantee that language models will handle ambiguity poorly in other task settings or with other extraction methods.

Ethics Statement

The paper describes dataset construction and annotation practices, including filtering, expert validation, crowdworker compensation, and the possibility that subtly harmful examples remained.

  • The dataset may contain subtly harmful examples that annotators and validators overlooked.
  • Crowworkers received a median hourly rate of $19.13, and worker IDs were not released.
  • Generated examples were filtered with heuristics and a multilabel model before expert annotation and validation.
  • The authors focused validation on linguistic ambiguity and discarded examples involving temporal ordering or vagueness.

C.3 Recognizing Interpretation-Specific Continuations

This test compares language-model continuations conditioned on disambiguations or distractors, using KL divergence estimated from sampled continuations; stylistic mismatch and distractor closeness limit interpretation.

  • The mean over independent samples estimates KL divergence between continuation distributions.
  • The procedure prepends task-specific stems so models generate topical continuations under the evaluation setup.
  • The test generates 100 single-sentence continuations from the full probability distribution for each disambiguation or distractor.
  • Surface-form style and tone confound the test because formal disambiguations can lower continuation likelihood independently of meaning.
  • Distractor closeness affects difficulty, and noun replacement often produces distractors with substantially different plausible continuations.

D.2 Training Details

The experiments replicate prior training details as closely as possible, using RoBERTa-large-based models and several task-specific training procedures. Political-claim paraphrases are generated with InstructGPT, while multilabel predictions use learned logit thresholds.

  • All models from prior work are based on RoBERTa-large, with training details replicated as closely as possible.
  • UNLI is trained first on heuristically regression-mapped SNLI data, then on human-annotated u-SNLI regression data; AmbiNLI uses staged pretraining and finetuning.UNLI training lasts 1 epoch on SNLI and 3 epochs on u-SNLI; AmbiNLI uses 3 epochs on SNLI + MNLI followed by 2 epochs on AmbiNLI.
  • Distribution Distillation trains for 2 epochs on SNLI + MNLI examples re-annotated with a teacher model’s distributional outputs.The teacher is a traditional three-way classifier trained on SNLI + MNLI.
  • The multilabel model is trained on MNLI and ChaosNLI development data, treating a label as present when 20% of annotators select it.The final model is selected by the lowest held-out loss over 30 epochs.
  • Political claims are paraphrased with InstructGPT using a zero-shot prompt and top-p = 0.9 decoding, while multilabel outputs are mapped using logit thresholds.The decoding setting is intended to encourage correctness and diversity; threshold usage is described for multilabel prediction experiments.
Loading 2304.14399v2…