Source-linked AI summary

Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores

Eichi Uehara

arXiv:2608.26137v1cs.CLcs.LGcs.SD

TL;DR

L2 learners often lack speaking practice, while automated assessment requires interpretable, fair evaluation against an appropriate human benchmark. The paper evaluates a feature-plus-LLM hybrid on isolated learner speech from a richly rated archive without fitting to human labels. The blend reaches ρ = 0.818, while controlled tests find no useful advantage from how pauses are written for the LLM.

  • Problem

    L2 learners face limited speaking practice, and trustworthy automated scoring requires accuracy, interpretability, fairness, and a human benchmark accounting for rater unreliability.

  • Method

    The paper combines deterministic De-Jong speech-timing features with one text-LLM fluency judgment and evaluates them on learner-isolated ICNALE dialogue without fitting to human labels.

  • Results

    ρ = 0.818 for the blend, exceeding 81% of individual trained raters; pause encodings perform about the same, with inline locations showing no reliable gain over aggregate statistics.

  • Takeaways & Limitations

    The measured speech-timing features carry the fluency signal, while the LLM adds a coarse ranking that the continuous composite refines.

  • Takeaways & Limitations

    Generalization beyond ICNALE’s ten Asian L1 groups, interactive 90-second roleplays, and the tested proficiency range is unverified; per-L1 samples are too small for fairness guarantees.

Abstract

from arXiv · show

Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accurate, interpretable, fair, and benchmarked against the right human bar. We build an interpretable feature-plus-LLM hybrid for spontaneous L2 dialogue. We evaluate it without ever fitting to the human labels, against the ICNALE Global Rating Archive: 140 speeches rated by ~80 trained raters on 10 analytic criteria. We score the 130 L2 speeches with usable audio. A deterministic De-Jong speech-timing composite reaches rho=0.764. Blended with a single text-LLM fluency judgment, it reaches Spearman rho=0.818 against the consensus gold. This agrees with the consensus better than 81% of the 80 individual trained raters: above the median rater (rho=0.73) and near the best, and at ~83% of the reliability-corrected maximum (kappa_max=0.99). The blend improves on the composite alone by +0.054 (paired-bootstrap 95% CI [0.017, 0.108], excludes 0); the LLM adds a coarse fluency ranking that the continuous composite refines. We also report a controlled null on pause encoding, bounded to effects below about +/-0.1 rho at this sample size. Holding the LLM and learner words fixed and varying only how pauses are written into the prompt, inline pause locations do not beat aggregate pause statistics (-0.069, CI [-0.15, +0.08]), and a grounded mid-clause criterion gives no reliable gain. The fluency signal comes from the measured speech-timing features, not from how pauses are written for the LLM. We back every claim with two agreeing learner-isolation methods, paired-bootstrap CIs, a monologue negative control, per-feature reproduction of classical measurements, and a per-L1 fairness audit.

1 Introduction

The paper targets trustworthy automated assessment for spontaneous L2 dialogue, where speaking practice is scarce and human agreement is itself imperfect. An interpretable speech-timing composite blended with one LLM judgment outperforms most individual raters, while pause encoding adds no reliable benefit.

  • Motivation: Scarce speaking practice and speaking anxiety motivate scalable automated L2 speaking practice and assessment.Learners often lack partners, classroom talk may be teacher-dominated, and speaking is especially anxiety-laden.
  • Research question: The study evaluates whether speech-timing features and LLMs can fairly score spontaneous L2 dialogue against a rich human archive.The ICNALE Global Rating Archive contains 140 speeches rated by about 80 trained raters on 10 analytic criteria.
  • Main result: ρ = 0.818 for the blended scorer, exceeding the consensus agreement of 81% of individual trained raters and reaching about 83% of κmax = 0.99.The system combines a deterministic De-Jong composite with one LLM fluency judgment without fitting to human labels.
  • Main result: +0.054 improvement over the composite alone, with paired-bootstrap 95% CI [0.017, 0.108] excluding 0.The LLM contributes a coarse fluency ranking, while the continuous composite resolves its coarse ties.
  • Pause-encoding null: −0.069 for inline pause locations versus aggregate pause statistics, with CI [−0.15, +0.08], indicating no reliable pause-encoding gain.A grounded mid-clause criterion likewise gives no reliable improvement, and the three encodings perform about the same.
  • Verification: The evaluation includes reliability-corrected ceilings, independent learner-isolation methods, paired-bootstrap intervals, a monologue control, feature reproduction, and an L1 fairness audit.These checks are presented as central verification components of the paper.

2 Related Work

Classical automated speaking assessment is grounded in interpretable acoustic fluency features, while newer work uses self-supervised speech representations and speech-LLM graders. Together, these lines establish the technical context for the paper’s hybrid approach.

  • Classical assessment: SpeechRater established operational spoken-response scoring by combining interpretable fluency, pronunciation, prosody, vocabulary, and grammar features.Its fluency measures derive from countable speed, breakdown, and repair measures.
  • Classical assessment: Across studies, speech rate correlates with proficiency at r = .76, mean length of run at r = .72, and pause frequency at r = −.59.These correlations are described as the empirical backbone of deployed delivery scorers.
  • Self-supervised speech: Frozen wav2vec 2.0 reaches 77.9% CEFR accuracy versus 53.5% for a BERT text baseline on ICNALE.This comparison indicates that acoustic representations carry proficiency information beyond the transcript baseline.
  • Speech-LLM assessment: Speech-LLM graders directly assess L2 proficiency from audio, with rubric-guided systems and Qwen2-Audio reaching sentence-level fluency PCC up to 0.85.This represents a newer assessment line coupling speech encoders to large language models.

2.4 Textualizing prosody and pauses as text for LLMs

Prior work renders acoustic structure as text for LLM prompting and develops comparative or reliability-aware evaluation methods. The paper builds on these ideas while testing whether pause placement itself improves fluency scoring.

  • Acoustics as text: TextPA augments transcripts with IPA, CMU phones, and inline pause durations for zero-shot LLM pronunciation and fluency scoring.It reports fluency PCC 0.650 on MultiPA and 0.784 when fused with a supervised system.
  • LLM evaluation: Comparative-judgement methods can substantially outperform pointwise LLM scoring, with QWK 0.633 versus 0.021 on TOEFL11 in one reported comparison.The cited LCES result uses Llama-3.1-8B.
  • Reliability-aware evaluation: Reliability-corrected ceilings estimate the maximum correlation a perfect model can reach given finite human-rater reliability.This reframes evaluation around achievable agreement rather than exact consensus matching.
  • Reliability-aware evaluation: The reliability-corrected human-like bar is √r1α = 0.60, closely matching the directly measured ρ = 0.62 within 0.02.The comparison distinguishes a single-rater benchmark from the more reliable consensus.

2.7 Psycholinguistic grounding of L2 pausing

Psycholinguistic and phonetic evidence motivates encoding pause position and word frequency, especially for mid-clause pauses. The paper positions its controlled pause test alongside this literature and its broader verified contribution.

  • Psycholinguistic grounding: L2 speakers pause disproportionately within clauses, and both L1 and L2 speakers pause more before lower-frequency words.These findings motivate a placement-sensitive criterion combining clause position and word frequency.
  • Psycholinguistic grounding: Phonetic manipulation provides evidence that pause location affects perceived fluency, with mid-clause pauses weighing most heavily.The paper uses a 250 ms pause-detection threshold following de Jong and Bosker (2013).
  • Controlled comparison: The paper’s controlled comparison finds inline pause locations do not beat plain pause statistics for an LLM, with difference −0.069 and 95% CI [−0.150, +0.083].A grounded mid-clause × word-frequency criterion gives no reliable gain.
  • Positioning: The contribution combines a De-Jong composite blended with an LLM, reliability-corrected evaluation, and extensive verification without fitting to human labels.The verification includes two agreeing learner-isolation methods and fairness auditing.

3 Data

The study uses interactive L2 English roleplays from the ICNALE Global Rating Archive, scored against a large trained-rater consensus. Because each recording contains an interviewer, learner speech is isolated before scoring.

  • The corpus contains 140 first-90-second, two-party L2 English roleplays spanning ten Asian L1 backgrounds and CEFR levels A2–B2.
  • The analysis scores 130 L2 speeches after excluding four native-English controls and six recordings without usable audio.
  • Each speech receives 0–10 ratings from about 80 trained raters across 10 analytic criteria plus a holistic score.
  • Learner speech is isolated using official learner-only transcripts aligned to timed ASR words, with speaker diarization as a robustness method.The two isolation methods agree in feature rank-correlation at 0.77–0.89.
  • The GRA labels are used only for evaluation, with no model parameter, feature sign, or threshold fitted to human ratings.

4 Method

The method combines a deterministic, label-independent speech-timing composite with a zero-shot LLM fluency judgment, then evaluates their combination against reliability-aware human benchmarks. A controlled pause-encoding comparison tests whether inline pause placement adds information beyond aggregate statistics.

  • Deterministic scorer: The deterministic scorer re-implements De-Jong utterance fluency using five speech-timing features with signs fixed in advance.The features cover pause ratio, mean length of run, speech rate including pauses, long-pause rate, and ASR word-confidence.
  • Deterministic scorer: Each feature is z-scored from its distribution, multiplied by its fixed sign, and averaged without using human labels.Pure articulation rate is excluded and is reported as null on this data.
  • LLM arm: The LLM receives the learner transcript with pause information and returns one pointwise fluency score zero-shot at temperature 0.
  • Blend: The final scorer combines independently constructed continuous feature-derived and coarse discrete LLM signals, with added information tested by paired bootstrap.
  • Pause encoding: The pause experiment holds the LLM, prompt scaffold, and learner words fixed while varying aggregate statistics, inline locations, or a grounded criterion.The contrasts B−A and C−B test inline locations and the grounded criterion, respectively.
  • Human ceilings: The reliability-corrected maximum is κmax = √α, while the human-like bar is √r1 α, representing agreement achievable by a single trained rater.For Fluency, α = 0.979 gives κmax ≈ 0.99 and a human-like bar ≈ 0.60.
  • Human ceilings: The directly measured single-rater-versus-consensus agreement is ρ = 0.621, so 0.60–0.62—not 0.83—is the human comparison bar.The 0.83 value is an upper-bound interpretation after reliability correction, not a human score.

5 Results

Across 130 speeches, the deterministic speech-timing composite and LLM blend outperform the typical individual human-rater benchmark, while pause encoding adds no reliable advantage. The composite also shows delivery-specific construct validity, but per-L1 estimates are too uncertain for fairness guarantees.

  • Per-feature correlations: +0.749 for mean length of run and −0.715 for pause ratio, broadly reproducing classical De-Jong feature directions and magnitudes.Pure articulation rate is null (+0.08), consistent with its advance exclusion.
  • Pause representation comparison: −0.069 for inline pause locations versus aggregate statistics, with 95% CI [−0.15, +0.08], indicating no reliable encoding gain.The grounded mid-clause criterion likewise provides no reliable improvement; at n = 130, detectable effects are roughly ±0.1–0.15 ρ.
  • Composite, blend, and human ceiling: +0.054 improvement over the deterministic composite alone, with paired-bootstrap 95% CI [0.017, 0.108] excluding zero.The continuous composite resolves the LLM’s coarse, near-tied scores.
  • Composite, blend, and human ceiling: ρ = 0.818 for the composite-plus-LLM blend, exceeding the median individual rater (ρ = 0.73) and 81% of trained raters.The blend reaches approximately 83% of the reliability-corrected maximum κmax = 0.99.
  • Construct validity: +0.44 partial Spearman correlation with Fluency after controlling for the holistic score, versus negative correlations with Sophistication (−0.30), Purposefulness (−0.21), and Logicality (−0.17).This supports a delivery-specific rather than general-proficiency interpretation of the composite.
  • Per-L1 descriptive fairness: Per-L1 correlations range from 0.40 to 0.94, but wide overlapping confidence intervals prevent conclusions about differential validity.Four L1 groups have only n = 4 and no stable estimate, so the audit provides no per-L1 fairness guarantee.

6 Verification and Robustness

The paper verifies its scorer and pause-encoding claims through independent isolation, paired bootstrap testing, negative controls, feature reproduction, and ceiling analysis. These checks support the dialogue signal while qualifying fairness and LLM limitations.

  • Independent verification: Two independent learner-isolation pipelines agree, reducing the risk that interviewer speech contaminates learner timing features.The primary pipeline uses learner-only transcripts; the robustness pipeline uses diarization and a linguistic learner classifier.
  • Pause encoding: −0.069, CI [−0.15, +0.08] for inline-versus-aggregate pause encoding includes zero, supporting no reliable encoding advantage.The paired-bootstrap protocol respects the same speeches being scored two ways.
  • Negative control: ∼0 correlation in the monologue negative control indicates that the dialogue signal depends on speaker-matched information.The monologue speakers were disjoint from the dialogue/GRA participants.
  • Fairness and limitations: ρ = 0.40–0.94 across L1 cells is descriptive only because cells are small and uncertainty is wide.The four smallest L1 groups, each n = 4, are omitted because they returned no stable estimate.
  • Feature reproduction: Mean length of run +0.749 and pause ratio −0.715 reproduce established De-Jong measurement directions and broad magnitudes.Speech rate also matches the prior direction, while predicted-null articulation rate is +0.08.
  • Residual issues: 27–36% of pairs are tied across LLM arms, making coarse pointwise LLM scoring a measured residual issue that the composite helps resolve.The reliability-corrected ceiling analysis is paired with an explicit treatment of this tie phenomenon.

7 Discussion

The discussion narrows the human-ceiling claim and locates fluency information in deterministic timing features rather than pause formatting. It also presents interpretability and auditability as deployment advantages.

  • Human comparison: ρ = 0.818 beats the mean and median single-rater bars without fitting to human labels, but does not match an 80-rater panel.The consensus remains the gold, while 0.83 is an upper-bound interpretation rather than achieved agreement.
  • Where the signal lives: 0.764 for the deterministic composite comes within 0.01 of the best LLM arm, while pause encodings perform about the same.Inline locations do not improve on aggregate statistics, and the grounded criterion gives no reliable improvement.
  • Deployment implications: The system uses five transparent fixed-sign features and one zero-shot text-LLM call, allowing components and learner-facing contributions to be inspected.The discussion contrasts these properties with opaque high-correlation systems.

8 Limitations

The paper’s claims are bounded by corpus, task, ASR, fairness, model, sample-size, and diarization constraints. In particular, small L1 cells do not support per-L1 fairness guarantees.

  • Scope: Generalization beyond ICNALE’s ten Asian L1 groups, 90-second roleplays, and observed proficiency range is unverified.The dialogue task is interaction-embedded delivery and differs from monologue or read-aloud tasks.
  • Measurement: ASR timing and confidence errors on accented A2 speech may weaken features unevenly across L1s.The limitation concerns the lowest-proficiency, most-accented speech.
  • Fairness: ρ = 0.40–0.94 across L1 cells of n = 4–20 is too uncertain to support a per-L1 fairness guarantee.The reported variation is explicitly bounded by small cell sizes.
  • Model dependence: The LLM results use one model, DeepSeek-chat, so another model might shift absolute arm values even though controlled encoding contrasts hold across isolation methods.The paper treats the controlled contrasts, rather than absolute arm values, as the stronger finding.
  • Speaker isolation: Within-clip speaker-embedding collapse makes diarization a robustness check rather than the primary learner-isolation method.The primary pipeline therefore relies on official learner-only transcripts aligned to timed ASR words.

9 Conclusion

The paper concludes that a fairly evaluated feature-plus-LLM hybrid can score spontaneous L2 dialogue at a human-comparable bar, while pause formatting adds no useful fluency signal. Its practical design keeps timing grounding deterministic and uses the LLM for holistic judgment.

  • Conclusion: Aggregate statistics, inline locations, and a grounded placement criterion perform about the same for LLM pause encoding.The fluency signal lies in measured speech-timing features, not their textual rendering.
  • Conclusion: The resulting scorer is cheap, interpretable, and auditable, with deterministic timing analysis combined with LLM holistic judgment.The paper supports both claims with isolation checks, bootstrap intervals, a negative control, feature reproduction, and an L1 audit.

10 Ethics and Honesty Statement

The paper distinguishes its supported findings from stronger claims: pause encodings show no practically useful difference within the study’s power, while the human-ceiling comparison uses a single-rater bar. Reproducibility is constrained by the archive’s redistribution license and the absence of released source code.

  • Claim boundaries: The pause-encoding result supports no practically useful difference, not exact equivalence, because the confidence intervals exclude only effects larger than approximately ±0.10–0.15 ρ.The three encodings perform about the same, but the null is power-bounded.
  • Claim boundaries: The “beats the human ceiling” claim refers to the mean single-rater agreement, while the 80-rater consensus remains the gold standard.The consensus mean is more reliable than the single-rater bar and is not itself the ceiling benchmark.
  • Reproducibility: The paper does not release the scoring pipeline as source code, and the validation corpus and derived data cannot be redistributed under its license.A licensed copy must be obtained from the corpus maintainers; the running system is available for inspection at aflo.one.
Loading 2608.26137v1…