Source-linked AI summary

Letters hide the truth from our eyes: English homophones have meaningfully different phonetic realizations

Yu-Hsiang Tseng, Mirjam T. C. Ernestus, Louis F. M. ten Bosch, R. Harald Baayen

arXiv:2608.26749v1cs.CL

TL;DR

The study asks whether English homophones differ in fine phonetic realization beyond duration, and whether contextual meaning predicts those differences. It analyzes speech and semantic representations to test this question, finding systematic, generalizable alignment between homophone meaning and phonetic realization.

  • Problem

    Research had established meaning- and frequency-related differences in homophone duration, but whether other phonetic details differed and tracked contextual meaning remained under investigation.

  • Method

    The study analyzes time-normalized Mel spectrogram vectors and contextualized embeddings from homophone tokens, using quantitative alignment investigations.

  • Results

    79.3% vs. a 2.8% permutation baseline: speech vectors distinguished 35 homophone pairs, while semantic structure aligned with phonetic realizations at type and token levels.

  • Takeaways & Limitations

    Heterographic homophones are only approximately homophonous, and their normalized phonetic realizations are co-determined by meaning in context.

  • Takeaways & Limitations

    Prosody and other factors outside current contextualized embeddings contribute to speech, and model residuals remain non-negligible.

Abstract

from arXiv · show

The distribution of spoken word duration of English homophones is known to co-vary with frequency of use. This study investigates whether other aspects of the phonetic realization of homophones also differ. A series of quantitative investigations of 14,000 homophone tokens in American television news broadcasts revealed that the tokens of homophone pairs such as \textit{weight} and \textit{wait} have different phonetic realizations, and that these can be predicted from their meanings in utterance context. These systematic differences remain even when taking duration-related variation into account. Time-normalized spectrograms emerged as an excellent tool for probing the fine details of phonetic realization, and obviate the need for phonetic transcriptions, which inevitably hide the phonetic truth from our eyes.

1 Introduction

Earlier work linked homophone duration to frequency and meaning, but this study asks whether homophones also differ in other phonetic details and whether meaning predicts those differences.

  • Earlier research found that higher-frequency homophones tend to have shorter spoken durations than lower-frequency homophone twins.
  • Meaning predicts homophone spoken-word duration, with greater support from a word’s meaning associated with longer duration.
  • The study tests whether heterographic homophones differ in phonetic realization beyond their spoken-word duration differences.
  • The study uses quantitative investigations of 14,000 tokens from 35 homophone pairs in the Redhen television-news resource.

2 Homophones

Homophones share pronunciations while differing in meaning, making them useful for studying how lexical representations organize form and meaning.

  • Homophones are words with different meanings that are said to sound the same, including heterographic and homographic examples.
  • Psychological models commonly distinguish syntactic, morphological, and phonological representation layers connected by activation.
  • Levelt’s model places the word-frequency effect at the lexeme level, predicting equivalent effects for homophones if they share a lexeme.
  • Dell’s account gives homophone twins distinct lemma nodes but a shared lexeme node encoding phonological form.
  • Picture-naming studies found homophone reaction times more similar to word-specific frequency controls and no evidence for frequency inheritance.

3 The phonetic realization of homophones

Research shows that segmentally similar or identical words can differ phonetically, with frequency, morphology, grammatical function, and meaning implicated as explanatory factors.

  • More frequent English homophones tend to have shorter spoken durations than lower-frequency twins after controlling for several lexical and utterance variables.
  • German and Dutch word-final obstruents can differ in vowel duration, burst duration, and number of glottal pulses despite apparent neutralization.
  • English word-final /s/ duration varies with morphological status, while grammatical and discourse functions co-determine the realization of “like.”
  • Morphologically distinct homophones such as laps and lapse have been argued to differ in stem and suffix duration despite sharing the same phones.
  • Prior work reports meaning-related duration differences in English homophones and meaning-related tonal or pitch differences in Mandarin and English.
  • The present study investigates whether fine-grained contextual meanings align with phonetic realization beyond spoken-word duration.

4 Theoretical and methodological framework

The study models form–meaning mappings with the Discriminative Lexicon Model and represents speech using time-normalized Mel spectrogram vectors alongside contextualized semantic embeddings.

  • The Discriminative Lexicon Model maps forms and meanings through continuously updated association weights rather than static representations.
  • The model primarily uses linear transformations, enabling direct analysis of systematic covariation between semantic detail and phonetic realization.
  • Contextualized embeddings represent token meanings differently across contexts, using 768-dimensional vectors derived from preceding-word windows.
  • Mel spectrograms are used as compact speech representations, while deep-learning features are avoided because they integrate information across linguistic time scales.
  • Adaptive step sizes produce fixed-length spectrograms across tokens with different durations, preserving comparable spectrogram shapes.
  • Each token is converted into 50 time steps and 21 Mel-frequency banks, yielding a fixed-size 1,050-dimensional speech vector.
  • A 10 ms analysis window corresponds to 160 samples at a 16 kHz sampling rate, with half-spectrum retention yielding 80 spectral rows before Mel compression.

5 Modeling the phonetic realization of homophones

This section introduces experiments that use speech vectors and time-normalized spectrograms to test whether homophone realizations differ, using 14,000 tokens from television news broadcasts.

  • The study investigates whether homophone members can be distinguished from their phonetic realizations, represented as speech vectors.
  • The speech-vector construction uses an L2-norm least-squares approximation that treats dimensions as equally meaningful in Euclidean space.
  • The analysis sequence tests pair identity, individual homophone-member identity, and spectrogram differences.
  • Tokens were selected from the Redhen 2016 Dataset, which combines United States television news videos, subtitles, and word-level forced-aligner timestamps.
  • The dataset contains 14,000 tokens from 215 unique news programs, with a mean token duration of 330ms.
  • The homophone set includes 35 pairs, mostly monomorphemic, alongside pairs differing in morphological structure and irregular verb forms.

5.2 Experiment 1: Predicting homophone pairs from speech vectors

Experiment 1 tests whether speech vectors distinguish homophone pairs, establishing whether the representations contain systematic information about pronunciation differences.

  • Experiment 1 asks whether homophone pairs can be distinguished from their speech vectors.
  • The analysis uses LDA with 35 homophone pairs as response categories and reduces 1,050-dimensional speech vectors to 50 principal components.
  • 79.30% mean accuracy (SE = 0.42%) exceeded the 2.82% permutation baseline (SE = 0.10%) in 10-fold cross-validation.
  • Most errors involved pairs with highly similar pronunciations, including seen/sea, wait/way, and plane/planes.

5.3 Experiment 2: Predicting homophones from speech vectors

Experiment 2 tests whether speech vectors distinguish the individual orthographic members of homophone pairs, despite their shared conventional pronunciation.

  • The classifier assigns each speech vector to one of 70 orthographic words, making this task harder than identifying the 35 homophone pairs.
  • 51.13% mean accuracy (SE = 0.46%) exceeded the 1.43% permutation baseline (SE = 0.11%) in 10-fold cross-validation.
  • Every homophone pair scored at least two standard errors above the 50% chance level in within-pair classification.
  • Pause-position differences did not explain classification accuracy: preceding pauses yielded r = .26, p = .14, and following pauses r = .10, p = .58.
  • The LDA results show that homophone pronunciations differ but do not specify how the members of each pair differ spectrographically.

5.4 Experiment 3: Difference spectrograms

Experiment 3 uses GAMs and difference spectrograms to identify where homophone members differ phonetically, while controlling for duration, speaking rate, and pauses.

  • GAMs statistically evaluate whether estimated spectrogram surfaces for the two members of a homophone pair differ.
  • Each speech vector contributes observations indexed by normalized timestep and frequency band, with duration, pauses, and speaking rate added as predictors.
  • The model represents one homophone with a smooth spectrogram surface and the other with an added difference surface.
  • The GAM includes tensor-product smooths for the baseline and difference surfaces, plus duration, speaking-rate, and preceding- and following-pause controls.
  • For wait and weight, the mean spectrograms share diphthong formants, while the coda t is not clearly visible in spontaneous speech.
  • The difference model reduced AIC by 5208.73 relative to the simpler model for wait and weight.
  • Across all 35 pairs, models with difference surfaces substantially reduced AIC, averaging 3413.92 (SD = 1881.87).
  • LDA and GAM analyses provide converging but complementary evidence for word-specific phonetic realizations.

5.5 Experiment 4: alignment between speech vectors and semantic vectors

Contextualized embeddings align with homophone speech vectors at both word-type and token levels, showing that contextual meaning helps predict phonetic realization. The token-level alignment is weaker but reveals structured, above-chance form–meaning correspondence.

  • 5.5.1 Contextualized embeddings: Contextualized embeddings form well-separated homophone-word clusters, achieving 95% word classification accuracy.Their relative distances also loosely reflect broader semantic similarities, including inflectional relationships and noun–verb distinctions.
  • 5.5.2 LDL with within-word permutation: Linear mappings were evaluated separately for comprehension and production using 10-fold cross-validation against a full row-permutation baseline.Each homophone pair contributed 400 tokens, represented by 50-dimensional speech vectors and contextualized embeddings.
  • 5.5.2 LDL with within-word permutation: Production showed above-chance token-level form–meaning alignment for almost all homophone pairs, whereas comprehension had 12 pairs outside the permutation interval.Only one comprehension pair had an empirical score below the cross-validation interval.
  • 5.5.2 LDL with within-word permutation: Token-level alignment is weaker than type-level alignment because contextual semantic differences within homophone pairs are subtler than differences between homophone members.The comparison concerns mappings between speech vectors and contextualized embeddings rather than only word identities.
  • 5.5.3 Predicted spectrogram contrasts: Predicted spectrogram contrasts closely matched GAM-estimated difference surfaces across examples including wait–weight, sell–cell, son–sun, and council–counsel.The match indicates that homophone-specific meanings co-determine prototypical spectra and that differences extend beyond timing alone.

6 Accounting for word duration

The analyses test whether homophone-specific spectrotemporal differences merely reflect word duration. After duration-related information is removed, time-normalized spectrograms still distinguish homophone members and retain semantic alignment.

  • 6.1 Predicting duration from spectrograms: 52.90% mean R2 was obtained when spoken duration was predicted from speech vectors, versus 24.15% from GPT2 embeddings.The GPT2 permutation baseline was 12%, showing that speech vectors contain richer duration information than embeddings in this analysis.
  • 6.1 Predicting duration from spectrograms: Partial Least Squares removed linear predictor components most strongly related to duration from the speech vectors.The resulting deflated vectors were intended to lack enough duration information for above-chance duration prediction.
  • 6.2 Classifying homophone members: Removing duration from time-normalized spectrograms decreased homophone classification accuracy only somewhat, while performance remained above duration-only classification at M = 58.1%.Thus, spectrogram shape continues to distinguish homophone members after duration is controlled for.
  • 6.3 Duration and form–meaning alignment: 58.39% mean accuracy was achieved when semantic vectors were predicted from PLS-deflated speech vectors, compared with chance-level prediction from duration alone at M = 51.79%.The deflated spectrograms aligned only slightly worse with semantic vectors than the original speech vectors.
  • 6.3 Duration and form–meaning alignment: The study concludes that form–meaning alignment cannot be traced completely to spoken word duration.The remaining evidence implicates the broader shape of homophones’ spectrotemporal realization.

7 General Discussion

Experiments show that heterographic homophones have systematic, generalizable phonetic differences aligned with meaning in context, including differences beyond duration. These findings support time-normalized spectrograms and challenge models that treat abstract phonological units as shielding articulation from semantics.

  • 79.3% accuracy versus a 2.8% permutation baseline separated 35 homophone pairs from vectorized time-normalized spectrograms.
  • An average AIC decrease of 3114 across 35 GAM models indicated substantial fine-grained spectrogram differences between homophone members.
  • Meaning and phonetic realization aligned at both type and token levels, with semantic centroids mapping onto spectrogram centroids and contextual variation reflected in speech.
  • After duration was removed using partial least squares, representations still differentiated homophone members and aligned with semantic vectors.
  • Contextualized embeddings captured aspects of token-level phonetic realization through a simple linear mapping, while prosody and other factors remained outside their scope.
  • The regularities generalized to unseen data and were not fully attributable to utterance position or word duration, although other speech factors may contribute.

Appendix A GAM-estimated difference surfaces of 35 pairs

Appendix A presents average time-normalized spectrograms and difference surfaces for 35 English homophone pairs, including pairs such as aide–aid, ad–add, and band–banned.

  • Figures A1–A6 show average time-normalized spectrograms and difference surfaces for aide–aid, ad–add, band–banned, blew–blue, and phil–fill.
  • Figures A7–A12 cover hire–higher, hole–whole, here–hear, counsel–council, capitol–capital, and corps–core.

Appendix B Partial Least Squares

This appendix describes Partial Least Squares as an iterative procedure that extracts duration-related directions from spectrogram features and quantifies their contribution to speech representations.

  • PLS iteratively finds directions that maximize covariance between spectrogram features C and durations.
  • The first score t1 is the linear combination of spectrogram features weighted by the direction w1 found in the first iteration.
  • Each extracted rank-1 component t1p1⊤ best approximates the current cue matrix before that component is removed.
  • The procedure is repeated three times, producing weights, scores, and loadings used to construct successively deflated matrices.
  • The extracted scores are orthogonal to later deflated matrices, including the final deflated matrix C4.

Appendix C Frobenius norm decomposition under LDL

This appendix justifies interpreting Frobenius-norm decomposition under LDL as variance explained by predicted and residual speech representations.

  • Under LDL, Frobenius norms of the prediction ˆC and residual E can be linearly decomposed.
  • The Frobenius norm equals the sum of squared matrix elements, obtained by representing C as vertically stacked row vectors and applying the trace.
  • The cue matrix C is represented as speech prediction ˆC plus residual E in the least-squares LDL mapping from semantic embeddings.
  • Orthogonality between ˆC and E makes the cross-product term zero, yielding additive decomposition of their squared Frobenius norms.
  • The resulting ratio interprets the prediction norm relative to the total norm as variance explained, with the residual norm representing what remains.
Loading 2608.26749v1…