Source-linked AI summary

Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing

L. Choy, A. S. Khan, S. Patrizi, D. Ye, J. Gross, M. Cychosz

arXiv:2608.20396v1cs.CLcs.SD

TL;DR

The paper seeks a scalable way to measure spoken-language development from everyday child-centered recordings without labor-intensive transcription. It compares children’s speech embeddings with a caregiver centroid in HuBERT-BASE space and finds that embedding distance decreases with hearing age and relates to multiple speech-language outcomes. The authors present this as a simple, scalable complement to clinical assessment, while noting that its diagnostic utility is unestablished.

  • Problem

    Manual transcription and language-specific expertise limit scalable measurement of speech-language development, while naturalistic recordings add noise, overlapping speech, and diarization errors.

  • Method

    The study extracts HuBERT-BASE embeddings from diarized child and caregiver speech, computes child distance to a global caregiver centroid, and models it with hearing age, f0, vocalization length, and standardized outcomes.

  • Results

    Embedding distance decreased over development and remained associated with speech-language outcomes across models and measures after controlling for child age, f0, and vocalization length.

  • Takeaways & Limitations

    Embedding distance offers a simple, scalable way to measure children’s language development from everyday speech and remained robust to simulated diarization errors.

  • Takeaways & Limitations

    The measure is intended to complement rather than replace clinical analysis, does not yet capture specific language aspects, and has no established diagnostic utility.

Abstract

from arXiv · show

Language development is characterized by a gradual convergence of children's speech toward adult patterns. Measuring this process has traditionally required detailed transcription and language-specific expertise, limiting scalability across languages and populations. Here, we use speech embeddings to capture this convergence directly from the acoustic signal in longform, child-centered recordings, taken as children go about their daily lives. Using HuBERT-BASE, we extracted embeddings from speech vocalizations of children who are deaf/hard-of-hearing and their female adult caregivers ($>$925 hrs. observation). Embedding distance between children and caregivers decreased with hearing age, controlling for pitch and vocalization length, indicating, as expected, that children's speech patterns converge to caregivers over development. This single distance metric likewise related to multiple standardized measures of speech and language from infancy through preschoolhood. These results suggest a path toward scalable, language-neutral assessment of spoken language development from children's everyday lives.

1 Introduction

The paper addresses scalable measurement of children’s speech-language development from naturalistic audio, where transcription is costly and noisy recordings complicate automatic analysis. It proposes caregiver-relative embedding distance as a lightweight alternative that avoids transcription and discrete linguistic units.

  • Motivation: Naturalistic child-centered recordings contain noise, overlapping speech, and diarization errors that make developmental metrics difficult to derive.These conditions remain an open problem for automatic speech recognition research.
  • Motivation: Embedding-based measures could provide objective language-development signals without costly transcription or language-specific expertise.Prior work used HuBERT-derived discretized-unit entropy, but it failed on noisy naturalistic recordings despite recovering expected convergence on clean synthetic speech.
  • Approach: The proposed framework models development as continuous distance between children’s and caregivers’ speech in embedding space, without transcription or discrete linguistic units.The study applies this framework to children who are deaf or hard-of-hearing, for whom frequent reliable assessments are difficult, especially in infancy and toddlerhood.
  • Contributions: The study introduces a raw-audio embedding-distance metric, tests its relation to evaluated speech-language outcomes, and examines robustness to diarization errors.The contributions target fine-grained developmental change, prediction across children aged 10–65 months, and stability under input perturbations.

2 Methods

The study analyzes longform child-centered recordings from children with hearing loss, combining diarized speech embeddings, caregiver references, standardized assessments, and developmental covariates. Its pipeline uses HuBERT-BASE representations and controls for acoustic and utterance-length factors while modeling child speech relative to caregivers.

  • Participants and recordings: 34 children with bilateral moderate-profound hearing loss contributed 59 recording–timepoint observations, with 14 children providing multiple longitudinal timepoints.Most children used cochlear implants, hearing aids, or both, and 32 of 34 were exposed to English more than 50%.
  • Participants and recordings: Each child wore a LENA recorder for up to 16 hours, yielding 925 hours of child-centered audio across recordings averaging 15.7 hours.The recordings captured surrounding speech and the child’s own vocalizations during everyday activities.
  • Assessments: The study modeled produced and receptive vocabulary plus articulation using MB-CDI, PPVT-4, EVT-2, and GFTA-2 assessments.MB-CDI was modeled as words produced because the recordings measured speech production rather than comprehension in a controlled task.
  • Covariates: Mean f0 and vocalization length were included as covariates because anatomical growth and lengthening utterances could confound linguistic convergence with acoustic maturation.f0 was extracted with PYIN; vocalization length came from diarized .its segments.
  • Processing pipeline: Speech was diarized into child and nearby adult-female segments, then represented with HuBERT-BASE embeddings extracted from layers 7–9 and mean-pooled across time and layers.The pipeline used 16 kHz audio, 20 ms frames, and 768-dimensional vectors.
  • Distance metric and robustness: Child embeddings were compared with a global adult-female caregiver centroid, producing one mean distance observation per child-timepoint.Sensitivity analyses simulated speaker misattribution by replacing increasing proportions of child vocalizations with other speaker classes and recomputing distances and models.

3 Results

The embedding-distance metric was evaluated against hearing age and standardized speech-language measures, with robustness assessed through bootstrap resampling and simulated diarization contamination. Smaller CHD–FEM distance was associated with more mature developmental outcomes, although model-fit gains varied by measure.

  • Analysis framework: The analysis used child-level mixed-effects or ordinary least-squares models, controlling for mean f0, hearing age, and vocalization length.Two hundred CHD vocalization embeddings per recording were resampled for 1,000 iterations to obtain bootstrap confidence intervals.
  • 3.1 Hearing age: Embedding distance significantly predicted hearing age (β = −0.50, p < .001) and explained 11.30% additional variance.Including distance improved model fit (∆AIC= −24.68, p < .001).
  • 3.2 Vocabulary: Embedding distance significantly predicted MB-CDI (β = −0.35, p < .001), PPVT-4 (β = −0.29, p < .001), and EVT-2 (β = −0.24, p < .001) scores.Distance explained 3.51–9.47% additional variance across vocabulary outcomes, but significantly improved model fit only for MB-CDI.
  • 3.3 Consonant articulation: Embedding distance significantly predicted GFTA-2 articulation scores (β = −0.41, p < .001) and explained 14.33% additional variance.Adding distance did not significantly improve overall model fit for GFTA-2, consistent with overlap among developmental and acoustic covariates.
  • 3.4 Robustness to diarization errors: Across outcomes, smaller CHD–FEM distance continued to predict more mature developmental and language outcomes as simulated diarization contamination increased.The standardized coefficient remained directionally stable but progressively attenuated; initially significant effects remained relatively robust under moderate contamination.
  • Figures and models: Figure 2 places reduced CHD–FEM distance to the right and shows fixed-effect relationships with pointwise 95% bootstrap confidence ribbons.Markers encode within-child timepoint order, and lines connect successive observations where repeated data were available.

4 Conclusion

The paper proposes acoustic distance between children’s and caregivers’ speech embeddings as a scalable alternative to transcription-heavy language assessment. This distance decreased over development and was associated with speech-language outcomes after controlling for child age, f0, and vocalization length.

  • Manual transcription limits scalability, with only ~103 of the world’s 7,000+ languages represented in major language-acquisition journals.
  • The proposed approach measures acoustic distance between children and caregivers directly from speech embeddings within the same community.
  • Embedding distance decreased over development and remained a significant predictor of speech-language outcomes after controlling for child age, f0, and vocalization length.

Limitations

The authors position embedding distance as a complementary measure rather than a replacement for clinical language analysis. Its diagnostic utility, coverage of language abilities, and sources of acoustic variation remain unresolved.

  • Embedding distance is intended to complement, rather than replace, clinical analysis by trained professionals.
  • The measure does not yet capture specific language features such as morphological productivity, and there is no evidence that it can be used diagnostically.
  • Because it is based on children’s speech production, embedding distance only indirectly indexes receptive capabilities.
  • Despite controls for mean f0 and vocalization length, embedding distance may capture additional acoustic or non-linguistic variation requiring further investigation.

Ethical considerations

The study used HuBERT-BASE representations pretrained on monolingual English audiobook speech, while participants primarily acquired American English. Evaluation in additional languages and language combinations remains necessary.

  • HuBERT-BASE was pretrained on 960 hours of English LibriSpeech audiobook speech, so its representations reflect monolingual English exposure.
  • The participants primarily acquired American English, motivating evaluation of the metric in languages under-represented in the foundation model’s training data.

A.1 Methods

The analyses compared hearing age with chronological age as predictors of developmental outcomes, using model fit and incremental variance explained beyond baseline covariates.

  • A.1 Methods: Hearing age and chronological age were compared using ∆AIC and incremental variance explained.Models included f0, vocalization length, and CHD–FEM embedding distance as baseline predictors, then added either chronological age or hearing age.

A.2 Results

Hearing age generally accounted for developmental outcomes better than chronological age, especially for later vocabulary and articulation measures, while early vocabulary showed little difference between age metrics.

  • A.2 Results: Hearing age generally provided a better account of developmental outcomes than chronological age.The comparison favored hearing age for PPVT-4, EVT-2, and GFTA-2, but not clearly for MB-CDI.
  • A.2 Results: ∆AIC = 4.92, 3.13, and 2.95 for PPVT-4, EVT-2, and GFTA-2, respectively, favored hearing age over chronological age.All three model comparisons had p < .001.
  • A.2 Results: 19.19% vs. 6.38% additional variance for PPVT-4 favored hearing age over chronological age.The corresponding comparisons were 11.06% vs. 1.34% for EVT-2 and 11.14% vs. 2.30% for GFTA-2.
  • A.2 Results: MB-CDI showed little difference between age measures, with ∆AIC = 1.59 and comparable additional variance of 13.14% vs. 12.59%.The model comparison slightly favored chronological age for MB-CDI.

B.1 Methods

The sensitivity analysis simulated speaker-label contamination in child vocalizations across a 0%–100% sweep while keeping the female-adult reference centroid uncontaminated.

  • B.1 Methods: Speaker-label contamination was simulated by randomly replacing increasing proportions of CHD vocalizations with vocalizations from other speaker classes.The sweep ranged from 0% to 100% in 10% increments, with embedding distances and downstream models recomputed at each level.
  • B.1 Methods: Contamination was restricted to CHD samples, while the FEM representation remained a fixed aggregated centroid.The centroid pooled caregiver vocalizations across recordings and was treated as comparatively stable to individual diarization errors.

B.2 Results

Embedding-distance associations remained directionally stable under simulated speaker-label contamination, although effect sizes, model advantages, and explained variance generally attenuated as contamination increased.

  • B.2 Results: The standardized distance coefficient remained negative across outcomes and did not reverse even at ≥80% contamination.Greater distance from the adult reference continued to predict poorer developmental and language outcomes, while statistical significance diminished with increasing contamination.
  • B.2 Results: Negative ∆AIC values showed that the distance-augmented model consistently outperformed the baseline for hearing age and MB-CDI across most contamination levels.This advantage diminished as contamination increased and was not observed for PPVT-4, EVT-2, and GFTA-2.
  • B.2 Results: Incremental variance explained by embedding distance generally decreased as contamination increased, with confidence intervals increasingly encompassing the null value.Embedding distance nevertheless accounted for a non-trivial proportion of variance across several outcomes, sometimes even under high contamination.
  • B.2 Results: Effects remained relatively stable for outcomes with initially significant embedding-distance associations, suggesting they were unlikely to arise solely from systematic diarization errors.The contamination analysis primarily demonstrated robustness for outcomes whose original distance associations were statistically significant.
  • B.2 Results: Figure 3 displays mean standardized coefficients and 95% bootstrap confidence intervals across contamination levels from 0% to 100%.Solid lines represent bootstrap means and shaded regions represent the corresponding confidence intervals.
Loading 2608.20396v1…