Source-linked AI summary

Counterfactual reasoning: Testing language models' understanding of hypothetical scenarios

Jiaxuan Li, Lang Yu, Allyson Ettinger

arXiv:2305.16572v1cs.CL

TL;DR

The paper asks whether language-model predictions reflect statistical and lexical heuristics or systematic reasoning grounded in world knowledge. It uses counterfactual conditionals and controlled tests across five pre-trained models, finding that models often override real-world knowledge but mostly rely on lexical cues; only GPT-3 shows sensitivity to counterfactual nuances after controls, while remaining affected by lexical associations.

  • Problem

    It remains unclear whether language-model predictions arise from linguistic correlations or robust reasoning grounded in world knowledge.

  • Method

    The paper uses counterfactual conditionals, psycholinguistic tests, and controlled datasets to evaluate five pre-trained language models while manipulating world knowledge and lexical cues.

  • Results

    Models generally override real-world knowledge in counterfactual scenarios, but most rely on lexical cues; after controls, only GPT-3 shows sensitivity to counterfactual nuances, with continued lexical-associative influence.

  • Takeaways & Limitations

    Counterfactual performance can reflect shallow lexical strategies, so separating lexical cues and world knowledge is necessary when assessing models’ reasoning about hypothetical scenarios.

  • Takeaways & Limitations

    The controlled large-scale datasets introduce unnatural sentences and constrain linguistic complexity, while the study leaves open how much GPT-3’s performance reflects robust counterfactual reasoning.

Abstract

from arXiv · show

Current pre-trained language models have enabled remarkable improvements in downstream tasks, but it remains difficult to distinguish effects of statistical correlation from more systematic logical reasoning grounded on the understanding of real world. We tease these factors apart by leveraging counterfactual conditionals, which force language models to predict unusual consequences based on hypothetical propositions. We introduce a set of tests from psycholinguistic experiments, as well as larger-scale controlled datasets, to probe counterfactual predictions from five pre-trained language models. We find that models are consistently able to override real-world knowledge in counterfactual scenarios, and that this effect is more robust in case of stronger baseline world knowledge -- however, we also find that for most models this effect appears largely to be driven by simple lexical cues. When we mitigate effects of both world knowledge and lexical cues to test knowledge of linguistic nuances of counterfactuals, we find that only GPT-3 shows sensitivity to these nuances, though this sensitivity is also non-trivially impacted by lexical associative factors.

1 Introduction

The paper uses counterfactual conditionals to separate language models’ reliance on linguistic heuristics from reasoning grounded in world knowledge. Across controlled tests, models often override real-world knowledge, but most rely heavily on lexical cues, with GPT-3 showing greater sensitivity to counterfactual nuances.

  • Motivation and approach: Counterfactual conditionals test whether language models distinguish hypothetical scenarios from reality when predicting unusual consequences.A counterfactual premise is false in the real world but treated as true in the hypothetical world, such as cats being vegetarians and loving cabbages.
  • Motivation and approach: The study adapts psycholinguistic tests and controlled datasets to probe counterfactual predictions in five pre-trained language models.The tests examine interactions among counterfactual contexts, existing world knowledge, and shallower associative cues.
  • Main findings: Models generally increase their preference for counterfactual-consistent completions in counterfactual contexts, but most rely strongly on simple lexical cues.This pattern appears when models are asked to override existing world knowledge.
  • Main findings: After controlling for world knowledge and lexical triggers, most models fail to capture real-world implications of counterfactuals, whereas GPT-3 shows greater sophistication but remains influenced by lexical associations.The result concerns sensitivity to linguistic nuances of counterfactual language rather than only completion preferences.

2 Exp1: overriding world knowledge

Experiment 1 tests whether language models choose completions consistent with a hypothetical world even when those completions contradict real-world knowledge. Models generally shift toward counterfactual-consistent completions, but the effect is usually weak and is strongest for GPT-3.

  • Experimental design: The experiment compares Counterfactual-World, Real-World, and Baseline Bias conditions using matched completions such as “cabbages” versus “fish.”CW presents a hypothetical scenario that conflicts with world knowledge; RW presents a real-world-consistent scenario; BB measures baseline completion preferences.
  • Evaluation: Models are evaluated by comparing log-probabilities for CW- and RW-congruent completions and computing the percentage of items preferring the CW completion.Higher CW preference is better in the CW condition, whereas lower values are better in RW and BB conditions.
  • Results: GPT-3 prefers the CW-congruent continuation in greater than 70% of items, while other models’ CW-condition preferences are at best slightly above chance.All models prefer CW-congruent continuations more in counterfactual-world contexts than in real-world contexts, although the BERT small-scale difference is negligible.
  • Results: All models show below-chance CW preference in the RW condition, indicating above-chance preference for the correct RW-congruent continuations.The study therefore focuses on whether counterfactual context changes preferences relative to real-world context, not only on raw accuracy.
  • Interpretation and limitations: The authors associate stronger counterfactual override with stronger relevant world knowledge, while noting that world-knowledge definitions and test scale need further operationalization.A comparable relative pattern appears when examining whether models make correct preferences in both CW and RW conditions.

3 Exp2: impact of cue words in context

Exp2 tests whether models use counterfactual cues or merely lexical triggers when selecting completions. Most models show only minor sensitivity to the distinction between hypothetical and real-world continuations, while GPT-3 shows a substantial effect only on synthetic data.

  • Motivation: Exp2 adds a condition to test whether apparent counterfactual reasoning can instead be explained by simple lexical triggers in context.The setup is illustrated in Figure 2 and Table 3.
  • Experimental design: The CR condition preserves the counterfactual context but makes the continuation refer to reality, conflicting with the lexical trigger.Sentences add “In reality” and use present tense so the correct completion aligns with real-world information rather than the counterfactual cue.
  • Measure: Models relying beyond lexical triggers should sharply reduce CW-congruent preferences in CR, where real-world completions are correct.The experiment measures the percentage of items where models prefer the CW-congruent continuation; higher CW values but lower CR values indicate better predictions.
  • Results: Only GPT-3 shows a truly substantial drop in CW-congruent preference from CW to CR, and only in the large-scale synthetic dataset.Most models show a non-zero but minor reduction, suggesting that their preferences are largely driven by simpler lexical triggers.

4 Exp3: Inferring real world state with counterfactual cues

Exp3 removes relevant world knowledge and controls lexical items to test whether models infer the real-world state from counterfactual structure itself. Only GPT-3 shows substantial sensitivity in the small-scale dataset, and even that effect remains influenced by lexical associations.

  • Design: Exp3 removes subject-specific world knowledge and controls lexical items so models must infer the true state from counterfactual language.For example, “cat” becomes the neutral subject “pet,” preventing prior knowledge from determining whether “cabbages” or “fish” is expected.
  • Conditions: In CWC, counterfactual language reverses the logical completion despite neutral context and lexical triggers that favor the opposite answer.RWCA uses the same lexical triggers without counterfactual language, making the trigger-associated completion logical instead.
  • Measure: The experiment compares CWC-congruent completion rates across CWC, RWCA, and BBC conditions, with high CWC and low RWCA values indicating good performance.BBC establishes models’ default preference for the target factual completion.
  • Results: Only GPT-3 shows a substantial CWC–RWCA difference in the small-scale dataset, indicating finer-grained sensitivity to counterfactual structures.This sensitivity is less pronounced in the large-scale dataset and may reflect cancellation of competing lexical triggers in the small-scale items.
  • Interpretation: Lexical associations remain strong: in the large-scale dataset, “vegetarians” favors “cabbages,” biasing against the CWC-congruent continuation.This suggests GPT-3 has genuine sensitivity to counterfactual indicators, but superficial lexical cues still substantially affect performance.

5 Conclusion

Across the experiments, language models can prefer counterfactual-consistent completions over real-world knowledge, especially when that knowledge is stronger. However, most models rely largely on lexical cues; GPT-3 is more sensitive to fine-grained counterfactual structure but remains affected by lexical associations.

  • Conclusion: PLMs can prefer completions that conflict with world knowledge in counterfactual situations, with sensitivity appearing stronger when the relevant world knowledge is stronger.The authors suggest exposure volume may contribute to both stronger world knowledge and better counterfactual performance.
  • Conclusion: Most models’ counterfactual performance is largely driven by simple lexical cues rather than sophisticated understanding of counterfactuals.GPT-3 is the only model showing more sophisticated sensitivity to fine-grained linguistic cues, though lexical associative influences remain strong.
  • Conclusion: GPT-3 distinguishes counterfactual conditions more successfully but may benefit from lexical-trigger cancellation and remains vulnerable to associative cues.This qualifies the interpretation of GPT-3’s apparent counterfactual sensitivity.

Limitations

The datasets control lexical cues and world knowledge to separate statistical heuristics from causal reasoning, but synthetic scaling reduces naturalness and linguistic complexity. The study also leaves open how world knowledge supports counterfactual reasoning, whether GPT-3’s performance reflects robust reasoning, and how findings generalize beyond English.

  • Synthetic scaling introduces unnatural sentences and constrains linguistic complexity, motivating further study with naturally occurring data.Conflicting lexical cues also produce different model behavior across the small-scale and large-scale datasets.
  • The study leaves open how world knowledge benefits counterfactual reasoning and whether GPT-3’s stronger performance reflects robust logical reasoning.Additional systematic analysis of these questions is deferred to future work.
  • Because the experiments use English-specific counterfactual markers, results may not transfer directly to languages where such conditionals are linguistically ambiguous.The paper gives Chinese conditionals as an example requiring world knowledge for disambiguation.

Ethics Statement

The paper uses existing psycholinguistic or synthetically generated datasets and reports no human-subject experiments or anticipated ethical concerns.

  • The datasets came from psycholinguistic researchers or were synthetically generated without harmful information, and the study included no human-subject experiments.
  • The authors report no anticipated ethical concerns for the paper.

A.1 Example items in small-scale dataset

The small-scale dataset adapts psycholinguistic examples from Experiments 1–3, with logical completions marked and relatively weak semantic associations between context words and targets.

  • Semantic association between target words and contextual lexical items is less salient in the small-scale dataset than in the large synthetic dataset.The paper contrasts “language skills” and “talk” with stronger associations such as “vegetarian” and “carrots.”
  • Tables 7 and 8 provide example items from Experiments 1–3, with the logical completion underlined.Table 7 covers Experiments 1 and 2, while Table 8 covers Experiment 3.

A.2 Generation process of dataset

The large-scale synthetic dataset generates counterfactual items by embedding causal event pairs in conditional templates and varying lexical, syntactic, and informational properties.

  • The generation process creates causal event pairs and places them into counterfactual conditional templates.Examples pair events such as “like” and “feed” within templates describing hypothetical relations.
  • Noun classes are selected to satisfy both verb-selection restrictions and relevant world knowledge, such as carnivorous subjects paired with vegetable-related objects.
  • The process varies subject and object lexical items along with modal and tense markers across generated sentences.Table 9 illustrates how changing lexical items in subject and object positions yields different sentences.
  • Experiments 2 and 3 reuse the template while manipulating syntactic structure or subject informativity.

A.3 Correlation with world knowledge

The paper links stronger world-knowledge representations to stronger counterfactual preferences, while follow-up tests show that GPT-3’s apparent sensitivity is shaped by lexical cues and linguistic markers.

  • A.3 Correlation with world knowledge: All models show a significant correlation between world-knowledge robustness and counterfactual preference, with coefficients of at least 0.69.The reported correlation is measured in the CW condition.
  • A.4 Follow-up analysis on GPT-3’s success: GPT-3’s performance varies by dataset scale, favoring the large-scale dataset in Exp2 but the small-scale dataset in Exp3.The paper attributes this asymmetry to differences in trigger distance and directional lexical bias across datasets.
  • A.4 Follow-up analysis on GPT-3’s success: GPT-3 shows counterfactual sensitivity in a baseline dataset without strong baseline completion bias, unlike most models whose preferences remain similar across conditions.This comparison is used to distinguish sensitivity to counterfactual structure from a fixed preference for one continuation.
  • A.4 Follow-up analysis on GPT-3’s success: Adding a conflicting lexical cue increases CWC-congruent continuations in both conditions, indicating a strong influence of the cue.The manipulation inserts a conflicting lexical item while leaving the logical completion unchanged.
  • A.4 Follow-up analysis on GPT-3’s success: Moving conflicting cues farther from the target reduces their effect, showing that linear distance strongly affects lexical-cue salience.The Distance dataset relocates the cue to the sentence beginning using “instead of”.
  • A.4 Follow-up analysis on GPT-3’s success: Sentence boundaries, discourse connectives, and tense alter CWC-congruent preferences, with tense producing the strongest reported effect.Removing “in fact” unexpectedly slightly strengthens GPT-3’s preference for CWC-congruent completions.
  • A.5 Additional metrics on small-scale dataset: Only GPT-3 shows substantial preference for logical completions in both counterfactual and real scenarios in Exp3.In Exp1, GPT-3, RoBERTa, and MPNet show above-chance preference for logical continuations in both conditions.
Loading 2305.16572v1…