Source-linked AI summary

Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models

Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, Yonatan Belinkov

arXiv:2106.06087v3cs.CL

TL;DR

Neural language models can perform subject-verb agreement in difficult contexts, but the mechanisms underlying this behavior are not well understood, particularly in Transformers. The paper applies causal mediation analysis to pre-trained models and finds structure-dependent mechanisms, size- and architecture-related differences, and shared neuron use across syntactically similar sentences.

  • Problem

    The mechanisms underlying neural language models’ syntactic agreement remain poorly understood, especially for Transformer-based models.

  • Method

    The study applies causal mediation analysis to pre-trained neural language models, using interventions to implicate components such as neurons in agreement behavior.

  • Results

    GPT-2 and Transformer-XL use two structure-dependent agreement mechanisms, XLNet uses one unified mechanism, and larger models do not necessarily develop stronger preferences.

  • Takeaways & Limitations

    Agreement mechanisms are similar across model sizes, more distributed in larger models, and rely on overlapping neuron sets for syntactically similar structures.

  • Takeaways & Limitations

    Causal mediation measures are easier to apply in binary settings such as agreement, but extending them to more nuanced phenomena remains challenging.

Abstract

from arXiv · show

Targeted syntactic evaluations have demonstrated the ability of language models to perform subject-verb agreement given difficult contexts. To elucidate the mechanisms by which the models accomplish this behavior, this study applies causal mediation analysis to pre-trained neural language models. We investigate the magnitude of models' preferences for grammatical inflections, as well as whether neurons process subject-verb agreement similarly across sentences with different syntactic structures. We uncover similarities and differences across architectures and model sizes -- notably, that larger models do not necessarily learn stronger preferences. We also observe two distinct mechanisms for producing subject-verb agreement depending on the syntactic structure of the input sentence. Finally, we find that language models rely on similar sets of neurons when given sentences with similar syntactic structure.

1 Introduction

Targeted evaluations show that neural language models can perform subject-verb agreement in difficult contexts, but the mechanisms behind this behavior remain poorly understood, especially in Transformers. This study uses causal mediation analysis to identify those mechanisms and finds architecture-, size-, and structure-dependent patterns.

  • 1 Introduction: Transformer agreement mechanisms have been less extensively investigated than those of LSTM-based models despite Transformers’ stronger syntactic generalization.Prior work analyzed LSTM agreement mechanisms, while Transformer-based models remained comparatively understudied.
  • 1 Introduction: Causal mediation analysis identifies model components involved in Transformer language models’ syntactic agreement behavior.The method treats components such as neurons as mediators between inputs and outputs.
  • 1 Introduction: GPT-2 and Transformer-XL use two agreement mechanisms, one active only when subjects and verbs are adjacent, whereas XLNet uses one mechanism across structures.The findings distinguish architectures by whether agreement processing changes with syntactic structure.
  • 1 Introduction: Larger models assign higher probability to correct inflections more often, but this does not necessarily produce a larger correct-versus-incorrect probability margin.Agreement mechanisms in larger models resemble those in smaller models but are more distributed across layers.
  • 1 Introduction: The most important agreement neurons overlap across structures to varying degrees, with overlap matching human judgments of syntactic similarity.This connects neuron-level analyses to similarities among sentence structures.

2 Related Work

Prior work has evaluated syntactic behavior, probed internal representations, and applied causal mediation to language models. The paper motivates causal mediation as a way to move beyond correlational probing toward component-level evidence about agreement mechanisms.

  • 2 Related Work: Behavioral evaluations test whether language models prefer grammatical completions in subject-verb agreement and other syntactic dependencies.These evaluations measure predictions on minimally different grammatical continuations, including difficult contexts.
  • 2 Related Work: Probing maps model representations to syntactic phenomena, but added classifiers create interpretive confounds and provide correlational rather than causal evidence.A probe may learn the task itself instead of merely reading information already encoded in the representation.
  • 2 Related Work: Causal mediation analysis studies how mediators explain treatment effects and can identify individual neural components involved in model outputs.For language models, interventions modify input sentences and outcomes are functions of continuation probabilities.
  • 2 Related Work: The paper applies the causal-mediation framework to agreement, where grammatical completions should be strongly preferred over incorrect alternatives.This differs from gender-bias settings that seek balanced preferences in ambiguous contexts.

3 Experimental Setup

The experiments use synthetically generated prompts spanning six syntactic structures and compare multiple Transformer architectures and GPT-2 sizes. Counterfactual subject-number interventions measure models’ agreement preferences under controlled structural conditions.

  • 3.1 Data: The dataset contains 300 randomly sampled prompts for each of six syntactic structures, generated from templates and noun-verb combinations.Synthetic generation controls for potential training-set token-collocation confounds.
  • 3.1 Data: The six structures vary whether the subject and verb are adjacent or separated by adverbs, prepositional phrases, or relative clauses.Relative-clause structures are evaluated with and without the complementizer that.
  • 3.1 Data: Each construction defines correct and incorrect continuations using the third-person singular/plural distinction.The evaluation focuses on whether models select the inflection agreeing with the target subject.
  • 3.2 Models: The study primarily evaluates several GPT-2 sizes and additionally compares Transformer-XL and XLNet.The architectures differ in training objectives, effective context, word-order processing, and attention masking.

4 Total Effect: How strongly do models prefer correct forms?

This section defines total effect as a causal measure of models’ preference for correct subject-verb inflections and evaluates how that preference varies across syntactic structures and GPT-2 sizes. Effects are strongest without subject-verb separation, increase with adverbial distractors, and decrease with attractor phrases, while larger models do not consistently show larger effects.

  • Total Effect: How strongly do models prefer correct forms?: Total effect measures the relative change in the correct-versus-incorrect verb probability ratio after swapping the subject’s grammatical number.The swap-number intervention changes the subject to the opposite number, while the null intervention leaves the prompt unchanged.
  • Total Effect: How strongly do models prefer correct forms?: The analysis averages total effects across prompts and verbs to quantify models’ overall preference for correct inflections.Under the ratio definition, values below 1 indicate preference for the correct inflection before intervention, while the swapped condition reverses the expected direction.
  • Total Effect: How strongly do models prefer correct forms?: The measure differs from accuracy because it quantifies the margin between correct and incorrect continuation probabilities rather than only which probability is higher.The study initially hypothesizes that larger models will have larger margins because they often achieve higher agreement accuracy.
  • 4.1 Results: Total effects are near zero for GPT-2 models with random weights, providing a control for effects associated with learned parameters.Random-weight models are therefore omitted from the plotted results.
  • 4.1 Results: 1,000–5,000 total effects occur in simple-agreement and within-relative-clause structures, exceeding the below-250 effects reported for gender bias.These structures contain no separation between the target subject and verb, and larger GPT-2 models do not show larger total effects.
  • 4.1 Results: Adverbial distractors increase total effects, with DistilGPT-2 and GPT-2 Small reaching the highest effects in those contexts.The authors suggest that adverbs cue an upcoming verb and increase the correct verb’s probability more than the incorrect verb’s probability.
  • 4.1 Results: Figure 3 compares total effects across GPT-2 model sizes and syntactic structures, highlighting opposite effects of adverbial distractors and attractor phrases.The figure reports structure-specific effects rather than a single size-based trend.
  • 4.1 Results: Attractor phrases decrease total effects when PPs or relative clauses separate the subject and verb.GPT-2 is more certain across singular than plural relative-clause attractors, and GPT-2 Medium usually has the highest effects in attractor structures.

5 Grammaticality Margin: Is agreement easier for singular or plural subjects?

Grammaticality measures how strongly models prefer the correct subject-verb inflection, revealing systematic differences between singular and plural subjects across syntactic structures.

  • Grammaticality is defined as the reciprocal of the probability ratio for correct versus incorrect agreement resolution, so larger values indicate stronger correct-inflection preference.The metric distinguishes preferences associated with originally singular and plural subjects.
  • Plural subjects consistently have higher grammaticality than singular subjects across structures, suggesting plural verbs may function as GPT-2 defaults.This pattern holds regardless of attractor number or structure.
  • Attractors separating subjects and verbs reduce grammaticality, whereas adverbial distractors have little effect at comparable token distances.The result indicates that the separating structure matters more than linear distance.
  • When subject number is held constant, grammaticality is higher when the attractor matches the subject’s number.This is the expected attractor-number effect described for the grammaticality measure.
  • Preceding attractors have number-dependent effects: plural relative-clause attractors sharply reduce grammaticality for singular subjects but increase it for plural subjects.The within plural RC condition is the only attractor structure exceeding the simple-agreement grammaticality level.

6 Natural Indirect Effect: Which components mediate syntactic agreement?

Causal mediation analysis identifies neuron-level mechanisms underlying syntactic agreement and reveals that these mechanisms vary with syntactic structure, architecture, and model size.

  • 6 Natural Indirect Effect: Which components mediate syntactic agreement?: The analysis measures a neuron’s natural indirect effect by replacing its value with the value induced by an intervention and measuring the relative response change.This attributes part of the total effect of swapping the subject on inflection preferences to specific neurons.
  • 6.1 Results: Subject-verb separation produces a distinct indirect-effect contour from adjacent agreement, even when only one token intervenes.Separated structures peak at layer 0 and upper-middle layers, while adjacent structures show effects increasing in higher layers.
  • 6.1 Results: Randomized GPT-2 weights yield near-zero higher-layer indirect effects, indicating that trained-model effects largely reflect learning rather than architecture alone.
  • 6.1 Results: Larger GPT-2 models have lower maximum neuron effects but distribute those effects across more layers while preserving structure-specific contours.This suggests stronger concentration in fewer neurons for smaller models and more distributed agreement mechanisms in larger models.
  • 6.1.1 Comparing GPT-2 to Other Architectures: GPT-2 and Transformer-XL show similar local-versus-non-local contours, whereas XLNet exhibits similar contours across structures and approaches zero in its final layer.The authors conjecture that XLNet’s exposure to word-order permutations prevents bifurcating local and non-local mechanisms.
  • 6.1.2 Neuron Overlap Across Structures: Neuron-overlap patterns broadly align with syntactic similarity in GPT-2 and Transformer-XL, while XLNet shows noisier, lower overlap despite its more unified indirect-effect contours.GPT-2 Medium’s layer 21 overlap resembles the hypothesized similarity matrix; XLNet shows greater specialization across structures.

7 Conclusions

The study applies causal mediation analysis to interpret syntactic agreement mechanisms in pretrained neural language models. It identifies influential neurons and suggests extending the analysis to neuron groups, attention heads, other phenomena, verb effects, and incorrect predictions.

  • 7 Conclusions: Causal mediation analysis reveals the location and importance of neurons involved in syntactic agreement across pretrained neural language models.
  • 7 Conclusions: Future work should examine groups of neurons, attention heads, additional syntactic phenomena, verb-specific effects, and cases where models predict incorrectly.

Impact Statement

The study’s impact is bounded by its focus on English subject-verb agreement and by its lack of evidence about modifying training procedures. It also highlights challenges in extending causal mediation beyond binary outcomes.

  • The findings are limited to specific syntactic structures and subject-verb agreement in English language models, so they cannot be extrapolated to other tasks or languages.
  • The study does not examine mitigation mechanisms, leaving the consequences of changing language-model training procedures unknown beyond three examples.
  • Causal mediation analysis is difficult to extend beyond binary cases, which may complicate analyses of fairness and bias involving more nuanced outcomes.
  • Attention-head indirect effects under the swap-number intervention show no consistent cross-structure trend, aside from upper-middle-layer involvement.
  • Under the zero intervention, lower-layer heads are consistently implicated, while upper-layer effects are positive but distributed across heads.
  • Adding adverb distractors increases verb probabilities overall and raises correct-inflection probabilities more than incorrect-inflection probabilities.

C The (Non-)Impact of Complementizers

Removing the complementizer that produces only minor total-effect changes across relative-clause structures and does not substantially alter indirect-effect patterns. The findings therefore suggest similar agreement mechanisms with or without the complementizer.

  • The comparison evaluates total effects and neuron indirect effects for across- and within-relative-clause structures, with and without that.
  • Removing that causes only minor reductions in total effects for across-relative-clause structures across model sizes.
  • For within-relative-clause structures, models vary: most are robust to that, whereas GPT-2 Medium more strongly prefers correct inflections when it is absent.
  • Indirect-effect magnitude and contour do not differ significantly when that is included versus excluded, indicating unchanged agreement mechanisms.
  • The study’s syntactic-similarity analysis uses binary, ternary, and numerical features, with distance differences normalized into similarity scores.

D.2 Neuron Overlap Across Layers

Neuron overlap across syntactic structures generally increases through upper-middle layers, then falls sharply in the highest layer. This pattern recurs across DistilGPT-2, larger GPT-2 models, and Transformer-XL.

  • Neuron overlap rises through upper-middle layers and drops sharply to near-zero in the highest layer of DistilGPT-2.
  • The same layerwise pattern holds across larger GPT-2 models and Transformer-XL, with a slight decline in the second-highest layer before the final drop.
  • Figure 16 measures overlap among the top 5% of neurons by indirect effect across structures and layers.
  • The analysis compares hypothesized syntactic similarity with layerwise neuron-overlap patterns using feature-based similarity scores and ℓ1 norms.

E Total Effects Across Architectures

Total effects are generally similar for XLNet and GPT-2 but much smaller for Transformer-XL, with relative-clause exceptions. Model depth and parameterization do not consistently predict total-effect magnitude.

  • Total-effect magnitudes are generally similar for XLNet and GPT-2 but much smaller for Transformer-XL, except in relative-clause structures.
  • Parametrization and model depth do not correlate well with total effects across architectures.
  • The smaller Transformer-XL effects may reflect its longer effective contexts and broader assignment of probability across tokens, but this explanation remains tentative.
Loading 2106.06087v3…