Source-linked AI summary

Lost but not erased: Finding traces of a forgotten language in neural speech models

Peter Plantinga, Charlotte Moore, Peter W. Donhauser, Krista Byers-Heinlein, Denise Klein

arXiv:2608.25976v1cs.CLcs.LG

TL;DR

International adoptees retain phonological traces of a birth language despite losing conscious access to it, but human studies cannot separate learning dynamics from biological maturation. The paper models abrupt language replacement in automatic speech recognition systems, finding that early exposure leaves functional traces mainly in low-level representations and supports a learning-based account of critical-period effects.

  • Problem

    Human adoptee studies show persistent birth-language phonological traces, but cannot determine whether these effects arise from biological maturation or ordinary learning dynamics.

  • Method

    The study trains automatic speech recognition models on a pre-adoption language, abruptly switches them to a post-adoption language, and analyzes behavioral and layer-wise representational changes.

  • Results

    Early exposure produced persistent pre-adoption traces concentrated in the lowest, pre-phonemic layers, and these traces contributed to faster relearning of the lost language.

  • Takeaways & Limitations

    Critical-period effects can emerge from entrenchment of foundational representations under fixed plasticity, so persistence alone does not establish a maturational window.

  • Takeaways & Limitations

    The models do not develop like humans, so the study cannot determine how biological maturation interacts with these learning dynamics.

Abstract

from arXiv · show

International adoptees retain phonological traces of a birth language they can no longer speak or comprehend, a persistence typically attributed to a biologically-timed critical period. We asked whether it could instead reflect the ordinary dynamics of learning, using automatic speech recognition models that simulate the international adoptee experience without maturational confounds. Models were trained on one language and then abruptly switched to a second. We found that traces of the first language persisted throughout second-language training, but mainly in the lowest, pre-phonemic layers. These traces were functional, as models with early exposure re-learned their lost first language 14% faster than naive models; this advantage held even against models adopted early from a related language and disappeared when the earliest layers were substituted from a non-adopted model. We argue that these critical-period effects reflect entrenchment of foundational representations rather than a maturational loss of plasticity, and that experience plays a central role in critical periods in language acquisition.

Introduction

International adoptee research reveals that birth-language phonology can persist despite apparent language loss, raising questions about whether critical-period effects require biological maturation. This study uses artificial speech models to isolate learning dynamics and examines how hierarchical representations preserve such traces.

  • Motivation: International adoptees can rapidly acquire a post-adoption language while retaining phonological traces of a birth language they no longer consciously remember.Evidence includes neural activation patterns resembling continuous birth-language speakers and faster relearning of birth-language phonological contrasts.
  • Motivation: Adoptee advantages are selective: they appear for phonology but not for higher-complexity tasks such as grammaticality judgments.A similar asymmetry has been observed in artificial vision systems, where some early perceptual information persists while other representations are overwritten.
  • Hierarchical account: Hierarchical processing may make low-level representations harder to revise because downstream higher-level computations depend on them.By contrast, revising a high-level category may impose comparatively little cost because fewer downstream computations depend on it.
  • Hierarchical account: Adult learning evidence links unfamiliar-sound difficulties to rigid phonetic scaffolding, motivating a critical-period account based on preserving stable statistical foundations rather than biological maturation.The proposed mechanism is that revising foundational sound representations could destabilize higher-level representations built upon them.
  • Study approach: The study addresses an unresolved causal question by simulating international adoption in speech models whose learning conditions can be controlled without maturational confounds.Models were trained first on a pre-adoption language and then abruptly switched to a post-adoption language, including closely related non-tonal languages.
  • Study approach: The artificial adoptee models reproduced human-like behavioral patterns and retained pre-adoption traces selectively in their earliest layers, with consequences for relearning.They rapidly lost pre-adoption proficiency, reached native-like post-adoption proficiency, and showed advantages in identifying and relearning pre-adoption phonemes.

Analytic Approach and Results

The study simulated international adoption in ASR models, tracking behavioral change and internal representations after an abrupt language switch. Lost-language traces persisted mainly in early layers and supported faster relearning, with the effect localized to foundational representations.

  • Analytic Approach: The four-step pipeline trained language-switch and single-language ASR models, generated layer-wise similarity matrices, compared them with RSA, and tracked trajectories over training.RSA used Spearman rank correlation on representations of held-out pre-adoption samples.
  • Behavioral Links: Models rapidly lost pre-adoption recognition while acquiring post-adoption proficiency, reaching 80% French accuracy after 3000 post-switch updates and converging within less than 0.1% of French-only accuracy.German accuracy declined to 25.9% after 4000 updates, near the shuffled-input baseline of 25.4%.
  • Behavioral Links: Phoneme-probe accuracy was highest in middle layers, but adoptee models outperformed French-only baselines most strongly in pre-phonemic layers, with the advantage shrinking at greater depth.Ma exceeded Mctrl at pre-phonemic and phonemic layers (p = 0.016 for both), but not post-phonemic layers (p = 0.28).
  • Representational Traces of Lost Language: Representational traces persisted after pre-adoption behavior fell to chance, stabilizing above the cross-language baseline and concentrating in layers 1–4 while higher layers remained near zero.Across the final 5k post-adoption steps, RSA(Ma, Mpre) averaged 8.3% (95% CI = 2.9% to 13.7%).
  • Representational Traces of Lost Language: The trace nearly doubled from ~4% to ~8% between 3.4k and 5.0k pretraining updates, peaked around 5.0k–6.7k steps, and then slowly declined rather than accumulating indefinitely.This non-linear pattern occurred across 3 of 4 language-switch pairs.

General Discussion

The models reproduced adoptee-like behavioral loss while retaining early-language traces, especially in the lowest representations, which accelerated phonemic relearning. These findings support critical-period effects arising from entrenched foundational representations and learning dynamics, while leaving biological contributions unresolved.

  • Persistent traces: Early exposure left persistent pre-adoption-language traces, particularly in the model’s lowest representational levels, despite behavioral loss and native-like post-adoption proficiency.The traces accelerated relearning of pre-adoption phonemic contrasts relative to models without early exposure.
  • Implications: The models reproduced human adoptee patterns while demonstrating that fixed-plasticity learning systems can retain latent speech structures after overt proficiency disappears.The authors frame this as evidence that lasting effects need not be attributed entirely to biological maturation.
  • Hierarchical representations: The persistent trace and its causal contribution to relearning were confined to the lowest, pre-phonemic layers, while phoneme decoding extended only slightly higher.Higher-level representations were more readily overwritten under the proposed hierarchical account.
  • Behavior and representations: 93.6% versus 93.7% accuracy accompanied distinguishable internal representations, showing that near-identical behavior did not reveal all retained structure.Internal representations remained distinct from monolingual models of either language after the language change.
  • Limitations: The models cannot determine how biological maturation interacts with these learning dynamics, and they directly tested low-level phonology rather than higher-level language.The authors state that the results constrain, rather than eliminate, a maturational account.
  • Limitations: Because real-world attrition usually includes some birth-language contact, the models’ total and abrupt cessation of pre-adoption input limits direct correspondence with human experience.The authors propose testing punctuated re-exposure and varying modelling scale, objectives, architectures, and languages.
  • Biology and experience: The findings give experiential explanations of critical-period effects a null-hypothesis role: biological accounts must specify what maturation adds beyond learning dynamics.The authors argue that persistence alone is not sufficient evidence for a maturational window.
  • Broader significance: Early exposure can create durable latent structures that support future learning, suggesting that experience order may shape entrenchment without requiring a biological window.The discussion connects this possibility to curriculum learning, early immersion, and heritage-language maintenance.

Methods

The study used speech-recognition models trained on one language before an abrupt switch to another, then probed behavior and internal representations. Representational similarity analysis and layer-wise phoneme probes tested whether pre-adoption structure persisted and where it was localized.

  • Data and models: Models used Common Voice speech from English, French, and German, with approximately 1,000 hours extracted across the three languages.Each language had held-out validation and test data.
  • Data and models: The ASR system used a 12-layer encoder and 6-layer decoder with a Conformer architecture and 109.3 million parameters.The encoder processed 80-dimensional mel-spectral features through a CNN.
  • Training design: Training omitted late learning-rate decay after warmup to prevent apparent traces from arising merely because later updates became ineffective.The warmup lasted 12k updates.
  • Training design: To simulate adoption, models trained on Lpre for 2–5 epochs before an abrupt switch to Lpost data and tokenizer, while retaining encoder and decoder parameters.The switch occurred after 3.4k–8.4k updates.
  • Statistical design: At least three random-seed models were trained per language pair and pretraining duration, with 95% confidence bands computed by non-parametric bootstrapping.Statistical tests generally used random seeds, with one phoneme-probe test using language-switch directions.
  • Representational analysis: Representational similarity analysis compared adoptee, pre-adoption, and post-adoption models using 500 held-out pre-adoption stimuli and Spearman correlations between similarity matrices.Time-averaged layer representations were compared after constructing cosine-similarity matrices.
  • Representational analysis: The RSA comparisons defined same-language models as a topline, non-overlapping domains as a baseline, and adoptee similarity above that baseline as a representational trace.Adoptee–post-adoption similarity assessed adaptation to the second language.
  • Analysis choices: Trace estimates averaged RSA scores over the final 5k training steps, while the pretraining-duration analysis used a headroom-normalized baseline-only formulation.The alternative formulation addressed high variance caused by small baseline–topline gaps in early layers.

Author Information

The paper lists authors affiliated with McGill University, Concordia University, the Ernst Strüngmann Institut, the Centre for Research on the Brain, Language, and Music, and Mila.

  • Contributions: Peter Plantinga and Charlotte Moore contributed equally, while Krista Byers-Heinlein and Denise Klein jointly supervised the work.
  • Affiliations: Peter Plantinga and Charlotte Moore are affiliated with McGill University, Concordia University, and the Centre for Research on the Brain, Language, and Music.
  • Affiliations: Denise Klein is affiliated with McGill University and the Centre for Research on the Brain, Language, and Music.
  • Affiliations: Krista Byers-Heinlein is affiliated with Concordia University and the Centre for Research on the Brain, Language, and Music.
  • Affiliations: Peter Donhauser is affiliated with the Ernst Strüngmann Institut.
  • Affiliations: Peter Plantinga is affiliated with the Mila Quebec Artificial Intelligence Institute.
  • Contributions: All authors contributed to conceptual development, writing, and editing; Plantinga and Moore designed protocols, Plantinga and Donhauser wrote code, and Plantinga ran experiments and generated figures.
Loading 2608.25976v1…