Source-linked AI summary

Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation

Christos Nikolaos Zacharopoulos, Revekka Kyriakoglou, Chara Tsoukala, Théo Desbordes

arXiv:2609.00416v1cs.CL

TL;DR

The paper asks what happens to syntactic information after it peaks in transformer layers and tests this using controlled Greek scrambling stimuli across three Greek-tuned models. Cross-layer probing finds below-chance late-to-early transfer with a strong canonical bias and a coefficient sign reversal around layer 22, supporting directional recoding rather than simple information loss. The authors frame this as a representational change that can be tested with human EEG and MEG using the same stimuli.

  • Problem

    Prior probing shows syntactic information peaking in middle layers and declining later, but does not establish whether the later code weakens or changes in a specific direction.

  • Method

    The study applies cross-layer generalisation to controlled Modern Greek object-relative minimal pairs across three Greek-tuned transformer models, probing canonical SVO versus non-canonical VSO order.

  • Results

    99.3% of true non-canonical sentences were classified as canonical in late-to-early transfer, with below-chance effects and probe coefficients reversing sign around layer 22.

  • Takeaways & Limitations

    Late transformer layers recode syntactic information toward canonical-form representations, generating a directly testable prediction for human EEG and MEG data.

  • Takeaways & Limitations

    Interpretation is limited to linearly decodable information, a design where SVO and canonical status coincide, and one language and construction type.

Abstract

from arXiv · show

Probing studies have established that syntactic information is decodable in early and middle transformer layers, but what happens to that information in later layers remains poorly understood. We apply a cross-layer generalisation analysis to three Greek-tuned large language models evaluated on tightly controlled minimal pairs: object-relative constructions in Modern Greek, where canonical (Subject-Verb-Object; SVO) and non-canonical (Verb-Subject-Object; VSO) orders differ only in within-clause word order, while preserving propositional meaning. When a probe trained on late layers (20-31) is tested on each early layer individually, it produces below-chance transfer (cluster-corrected, p<0.01), classifying 99.3% of non-canonical sentences as canonical. Probe coefficients reverse sign around layer 22, indicating a directional recoding toward the canonical form rather than simple information loss. These findings characterise a representational format change in late transformer layers that goes beyond the well-established decline in syntactic decodability, and they generate a directly testable prediction for human EEG and MEG decoding studies using the same stimuli. Code and stimuli are publicly available on OSF.

1 Introduction

Modern Greek scrambling provides a controlled way to separate word order from meaning while probing how transformers represent syntax across layers. The study addresses whether late-layer syntactic representations merely weaken or change directionally.

  • Motivation: A competent reader resolves both orders to the same meaning, so the study targets representational changes inside transformer models rather than semantic resolution.
  • Research gap: Prior probing established that syntactic information peaks in middle layers and declines toward the output, but not what happens after the peak.Standard per-layer probing cannot distinguish weakening from a directional representational change.
  • Approach: Cross-layer generalisation tests a probe trained at one layer on representations from another, using transfer failure and its direction to characterise format changes.The approach was borrowed from temporal decoding in cognitive neuroscience and had not previously been applied to word-order representation in transformer LLMs.
  • Motivation: Greek case morphology permits SVO and VSO orders with identical propositional content, enabling tightly controlled minimal-pair experiments.The contrast occurs within an embedded object-relative clause while preserving the same words and grammatical roles.
  • Contribution: The analysis spans all layers of three Greek-tuned transformer models and generates a directly testable prediction for human neural responses to the same stimuli.

2 Materials & Methods

The study uses controlled Modern Greek object-relative minimal pairs and layer-wise probing to classify canonical SVO versus non-canonical VSO order. Logistic-regression probes operate on compact statistical summaries of post-clause hidden states and are evaluated with cross-validation and cluster-based permutation tests.

  • Stimuli: Stimuli are Modern Greek object-relative sentences whose embedded relative clauses vary between SVO and VSO while the matrix sentence remains fixed.The matrix verb is sentence-final in both conditions.
  • Stimuli: Template generation holds lexical material constant within each minimal pair, avoiding lexical differences that naturally occurring SVO/VSO pairs would introduce.
  • Probing: Word order is a binary classification task over four per-layer statistics—mean, variance, skewness, and kurtosis—computed from post-clause hidden states.The statistics summarize the token-by-hidden-dimension activation matrix into a compact sequence-level feature vector.
  • Evaluation: Performance and cross-layer generalisation use stratified 10-fold cross-validation with a two-sided cluster-based permutation test over AUC−0.5 at α = 0.01.A single probe pooled layers 20–31 and was evaluated on test layers 0–19 to test transfer rather than maximize per-layer accuracy.
  • Reproducibility: Stimuli, extracted activations, and analysis code are publicly available in an anonymised OSF repository.

3 Results

Sentence-type information is decodable after the relevant clause appears, peaks in middle layers, and shows a distinct late-to-early transfer failure. The below-chance transfer and coefficient sign reversal indicate directional recoding rather than simple loss of syntactic sensitivity.

  • 3.1 Sentence-type information emerges only after the clause boundary: Post-clause activations support above-chance decoding from early layers onward, peaking in middle layers, whereas pre-clause activations remain at chance.The pre-clause result is expected because the structure-defining constituents have not yet appeared.
  • 3.2 Middle layers support broad cross-layer generalisation: Layers 5–19 form a contiguous region where classifiers generalise above chance across many test layers, indicating a shared representational format.Late-layer cross-layer generalisation declines toward AUC 0.50 by the final layers.
  • 3.3 Late-layer decision boundaries do not transfer to earlier layers: The late-to-early transfer failure reflects a systematic canonical bias rather than merely weak transfer.The late-layer decision boundary treats virtually all early-layer items as canonical regardless of their actual word order.
  • 3.3 Late-layer decision boundaries do not transfer to earlier layers: 99.3% of true non-canonical sentences were classified as canonical when a layers 20–31 probe was applied to significant early layers, yielding 49.3% overall accuracy.The same transfer was below chance across a contiguous early-layer cluster with p < 0.01; 98.0% of true canonical sentences were also classified as canonical.
  • 3.4 Probe coefficients reverse sign in later layers: Around layer 22, probe coefficients reverse sign, showing that the decodable contrast changes direction between earlier and later representations.Features predicting non-canonical membership earlier predict canonical membership later, explaining inverted transfer.

4 Discussion

Late layers appear to recode syntactic information toward canonical forms rather than merely losing sensitivity. This interpretation is supported by directional transfer failure and coefficient sign reversal, but remains subject to alternative explanations and needs testing in human neural data.

  • Late transformer layers recode syntactic information toward canonical-form representations rather than simply losing syntactic sensitivity.
  • Around layer 22, dominant probe coefficients reversed sign, linking the transfer asymmetry to a directional representational change.
  • Simple information loss, feature compression, and probe mismatch are alternative accounts, but the authors judge directional recoding most consistent with the evidence.Causal intervention, such as late-layer ablation, would be needed for a conclusive test.
  • The same cross-layer design could be applied to EEG or MEG responses to test for a corresponding neural asymmetry.Cross-linguistic extension and constructions separating canonicality from surface order are identified as future directions.

5 Conclusion

Controlled Greek scrambling isolates syntactic structure from semantic content and reveals how word-order representations change across transformer layers. The resulting directional recoding signature is directly testable in human EEG and MEG data.

  • Controlled Greek scrambling isolates syntactic structure from semantic content while tracking word-order representations across transformer layers.
  • Late layers recode syntactic information toward canonical word order, marked by 99.3% non-canonical→canonical misclassification and a sign reversal at layer 22.
  • The directional recoding signature is directly testable in human EEG and MEG data.

6 Limitations

The interpretation is constrained by the linear-probe method, confounding between canonicality and SVO order, possible lexical regularities, and limited language and construction coverage.

  • Linear probes detect only linearly decodable information, so nonlinear representations may persist in late layers.
  • Because SVO order and canonical status are co-extensive, the probe may track either dimension.
  • Residual lexical regularities in template-generated stimuli may contribute weakly to classification, although balancing minimises this risk.Jabberwocky variants could isolate structural form more fully.
  • The study examines a single language and construction type, leaving cross-linguistic and cross-constructional generalisation open.

7 Ethical Considerations

The study uses synthetic stimuli and pretrained language models without recruiting human participants or collecting personal data.

  • The study uses synthetic stimuli and pretrained language models, with no human participants recruited and no personal data collected.
  • The authors identify no direct participant-related ethical risks.

A Cross-model GAT matrices

Across both comparison models, the expanded 1024-sentence analysis preserves the main representational pattern: broad middle-layer generalization and late-to-early transfer below chance, accompanied by coefficient sign reversal.

  • Both comparison models preserve a broad middle-layer regime of above-chance generalization on the expanded 1024-sentence set.The overall representational organization remains stable under the larger stimulus regime.
  • Late-to-early transfer falls below chance in both models, although the precise cluster extent varies by model.
  • Coefficients show the same broad sign-reversal profile across models in the late regime driving below-chance transfer.Coefficients are predominantly positive in middle-to-late layers and flip sign in the late regime.

B.1 Tokenization & Forward pass

The models process raw Greek stimuli with native Greek-aware subword tokenization and inference-only forward passes, extracting hidden states at every layer and token position.

  • Inputs were raw Greek strings tokenized with the native subword tokenizer, including Greek-specific units and preserved sentence-final punctuation.No prompt or few-shot context was added.
  • Inference used pretrained weights in forward-only passes with caching and dropout disabled.
  • The complementizer που divided each sentence into before and after regions for consistent token–word alignment.This alignment indexed matrix-subject nouns, relative-clause nouns, and the two verbs.
  • Layer-by-position activations were extracted as the sole inputs to the probing analyses and subsequent figures.
  • Stimuli were generated from a fixed lexicon of Greek human-denoting nouns, verbs, and determiners with controlled morphological forms.

B.3 Expanded stimulus set

The expanded stimulus set increases lexicalized instances while preserving the controlled minimal-pair design, and the comparison analyses reproduce the main qualitative findings.

  • Expanded stimulus set: The expanded set contains 1024 sentences across the same 32 fully counterbalanced conditions as the main experiment.N1 and N2 use distinct lexical stems, and each minimal-pair contrast varies only word order while keeping meaning fixed.
  • Expanded stimulus set: The larger sampling regime preserves the reported qualitative decoding and generalization effects, ruling out an artifact of the smaller 128-item set.
  • Cross-model results: Both comparison models retain above-chance middle-layer generalization and below-chance late-to-early transfer on the expanded set.The precise cluster extent varies by model, but the overall organization is preserved.
  • Cross-model results: Coefficient plots for both comparison models show the same broad sign-reversal pattern accompanying the late below-chance regime.
  • Stimulus construction: The generated materials use Greek noun, verb, and determiner inventories documented in the accompanying lexicon tables.
  • Stimulus construction: Table 3 summarizes the two generated stimulus sets used for the main experiment and expanded sensitivity analysis.
  • Stimulus construction: Templates hold lexical material, morphology, and propositional content constant within each minimal pair while changing relative-clause word order.Representative pairs illustrate matched SVO and VSO stimuli.
Loading 2609.00416v1…