Source-linked AI summary

Towards Quantifying Benchmark Optimization in ASR Models

Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr Cłapa, Jens Madsen, Panagiotis Tzirakis

arXiv:2608.19936v1cs.SDcs.AI

TL;DR

Public ASR benchmarks may overstate general-purpose transcription ability when models exploit benchmark-specific artifacts. The paper introduces behavioral and mechanistic probes for this problem and finds that high-performing open models reproduce benchmark references despite contradictory or missing acoustic evidence, with the behavior triggered by narrow acoustic cues. The authors also show that this benchmark-optimized behavior can be bidirectionally manipulated, while its origin during training remains incompletely understood.

  • Problem

    Public benchmark performance can diverge from real-world ASR utility because models may optimize benchmark-specific artifacts rather than generalizable transcription ability.

  • Method

    The paper combines reference disagreement, masked-entity recovery, and orthographic switching probes with synthetic audio, context manipulation, activation patching, and activation steering.

  • Results

    Across state-of-the-art open models on two widely used benchmarks, reproducing benchmark conventions against the audio is common among the highest-scoring systems.

  • Takeaways & Limitations

    Benchmark scores can reflect benchmark-conditioned behavior without reflecting improved general-purpose transcription ability, so evaluation should examine behavior beyond public-benchmark WER.

  • Takeaways & Limitations

    The study examines inference-time behavior but does not yet fully explain how benchmark optimization arises during training.

Abstract

from arXiv · show

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

1 Introduction

Public ASR benchmarks can reward benchmark-specific behavior rather than generalizable transcription ability. The paper measures this risk with behavioral probes and shows that leading models reproduce benchmark references despite contradictory, masked, or ambiguous audio.

  • Benchmark optimization is performance gained from benchmark-specific artifacts rather than generalizable transcription improvement.
  • The behavior persists for clones of benchmark speakers but weakens for generic or previously unseen speakers from independently collected data in the same domain.
  • The probes test reference disagreement, masked-number recovery, and orthographic switching when audio does not uniquely support the reference transcript.
  • Models systematically reproduce benchmark references despite audio evidence to the contrary, and this behavior is common among the highest-scoring systems on two public benchmarks.
  • Benchmark-specific acoustic context can activate a benchmark-optimized policy, while restricted context can preserve faithful transcription of the target speech.

2 Related Work

Prior ASR work has addressed the benchmark–reality gap mainly through broader coverage, convention control, and contamination analysis. This paper argues that ASR needs an updated framework for measuring benchmark optimization and applies mechanistic interventions to study it.

  • Robustness-oriented evaluations primarily treat the benchmark–reality gap as a coverage problem addressed by adding tests for missing conditions.
  • Bespoke text post-processing and convention matching have historically optimized performance across corpora, motivating consistent post-processing across benchmarks.
  • Evidence of transcript leakage is localized to a specific speech-LLM training paradigm, whereas benchmark optimization appears across multiple architectures.
  • The paper identifies a need for an updated framework to understand and measure how models optimize ASR benchmark performance.
  • Context manipulation, activation patching, and activation steering are used to localize and causally manipulate benchmark-specific transcription behavior.

3 Method

The paper evaluates benchmark optimization by testing whether ASR models reproduce benchmark references when audio is contradictory, masked, or acoustically ambiguous. It combines behavioral probes, likelihood-based readouts, and trigger analyses across public and held-out speech conditions.

  • Models and datasets: The study evaluates 11 open-source ASR models across encoder–decoder, transducer, and speech–LLM architectures.Teacher-forced likelihoods are reported for all models except Parakeet-TDT because its transducer architecture makes arbitrary-position logits difficult to read.
  • Models and datasets: The analysis centers on VoxPopuli, extends findings to LibriSpeech, and uses DaiKon and newly active LibriVox readers as held-out controls.The VoxPopuli Hugging Face version has 40% test-speaker overlap with training; the fresh-reader set contains 2026 recordings from 14 readers.
  • Probes and readouts: Behavioral probes compare the benchmark reference span r with an audio-true or acoustically equivalent alternative a where the audio underdetermines the reference.The primary surface readout, accept-ref, is the fraction of positions where greedy decoding emits r rather than a; high accept-ref indicates reference reproduction beyond acoustic content.
  • Probes and readouts: Audio lift subtracts the silenced-audio language-model prior from the reference-span likelihood and normalizes by the span’s character count.Positive λ(r) means audio increased reference likelihood beyond the prior; high lift suggests reliance on benchmark-specific artifacts in deliberately underdetermined cases.
  • Behavioral probes: Reference-disagreement cases use erroneous references, while masked cases silence target spans to test whether models still emit the benchmark content.Reference-disagreement accept-ref measures reproducing an erroneous reference instead of its audio-supported correction; masked accept-ref measures emitting spans after their acoustic evidence is removed.
  • Behavioral probes: The study also tests masked-number recovery and orthographic switching, including dataset-specific spellings such as “Mr” versus “Mister” and “any one” versus “anyone”.Numbers are targeted because surrounding audio should provide little information, while orthographic pairs are phonetically and semantically equivalent but benchmark-convention dependent.
  • Mechanism localization: Trigger analyses vary recording conditions, speaker identity, and appended audio to determine when benchmark-optimized behavior activates.The conditions include noisy or reverberant originals, evaluation-speaker clones, fresh speakers, generic voices, and benchmark or control audio appended as suffixes.

4 Results

Across benchmarks and controlled probes, stronger ASR models more often reproduce benchmark references despite contradictory, masked, or ambiguous audio. The behavior is narrowly triggered by benchmark-associated cues and can be switched through input additions or activation-level interventions.

  • WER 5.4–5.8% models had accept-ref 0.18–0.30, whereas models at 6.5% WER or above were at or below 0.10.The same relationship was corroborated using audio-lift and human-annotated data.
  • Six of 11 models exceeded the 0.5 honorific switch baseline, and eight of 11 exceeded it for archaic spacing.The results indicate switching across datasets and, for archaic spacing, across subpopulations or individual samples within a dataset.
  • 4.1 A narrow acoustic context gates the behavior: Benchmark-speaker clones retained directionally similar accept-ref, while generic and fresh same-domain speakers often produced lower recovery.For elevated models, masked-probe audio lift also collapsed toward zero on generic voices, supporting dependence on benchmark-specific acoustic cues.
  • 4.1 A narrow acoustic context gates the behavior: Truncating context around the target span collapsed reference-disagreement and masked accept-ref toward the floor for several models.The authors report that removing surrounding benchmark cues removes the behavior.
  • 4.2 The behavior spans encoder and decoder and is causally steerable: Appending conversational audio collapsed accept-ref on real benchmark clips, while appending VoxPopuli audio raised accept-ref on low-accept-ref ep-fresh clones.Reported increases were Phi-4 +.10, Canary +.09, Higgs +.07, Cohere +.07, and Parakeet +.04.
  • 4.1 A narrow acoustic context gates the behavior: Models generally returned audio-true transcriptions outside benchmark-speaker distributions or after sufficient benchmark context was removed.Adding non-benchmark audio or removing surrounding context could revert benchmark-optimized outputs.
  • 4.2 The behavior spans encoder and decoder and is causally steerable: A learned low-rank direction bidirectionally steered four of six elevated models, with k=1 recovering 65–80% of the full-direction effect for three.The results support causal modification at both input and activation levels in a subset of models.

5 Discussion and Conclusion

The paper finds benchmark-optimized transcription behavior among high-scoring ASR models and connects it to benchmark design, training-data transparency, and model selection beyond public WER.

  • Across state-of-the-art open models on two widely used benchmarks, reproducing benchmark conventions against the audio is common among the highest-scoring systems.
  • Narrow acoustic cues can activate benchmark-optimized behavior, which can be bidirectionally flipped by appending audio or steering a single model layer.
  • Model development: Model releases should transparently document training data because the study does not yet fully explain how these behaviors arise during training.
  • Benchmark development: Benchmark evaluations should avoid i.i.d. test splits and preferably use fully held-out, non-public sets, while speaker overlap alone did not explain the observed behavior.On leaked versus unleaked speakers, pooled elevated-model reference-disagreement accept-ref was 0.22 versus 0.26, and masked recovery was indistinguishable.
  • Model selection: Practitioners should consider multiple metrics beyond public-benchmark WER, using consensus reference edits and behavioral or mechanistic probes to assess benchmark optimization.For VoxPopuli, any model below 3% WER has to transcribe reference errors.

A.1 Text processing and teacher-forced NLL

Transcripts are scored with the June 2026 Hugging Face Open ASR leaderboard normalizer, while teacher-forced NLL measures reference-span likelihood under audio-conditioned and prior models. The span log-likelihood ratio is length-normalized by reference characters and aggregated across clips with bootstrap confidence intervals.

  • The June 2026 Hugging Face Open ASR leaderboard normalizer lowercases text and removes punctuation while normalizing numbers and contractions.
  • Teacher-forced NLL sums per-token log-probabilities over the target reference span, conditioning the audio term on the clip.
  • The white-box readout sums per-token log-likelihood differences over the query span, producing a subword-segmentation-invariant numerator.
  • Dividing the span log-likelihood ratio by reference characters yields λ(r), making spans of different lengths commensurate.
  • Corpus means are reported with a 2000-resample percentile-bootstrap 95% confidence interval.

A.2 EOS masking

EOS tokens are removed from the softmax denominator before scored continuation probabilities are computed. The same renormalization is applied to audio and silent-audio prior terms so the comparison uses distributions with matching normalization.

  • EOS tokens are masked from the softmax denominator at each scored position before continuation-token log-probabilities are calculated.
  • The audio term and x∅ prior term are renormalized symmetrically to avoid comparing differently normalized distributions.
  • EOS masking addresses substantial EOS probability in some models when extracting the language-model prior from fully silenced audio.

A.3 Synthetic speech stimuli and the intelligibility gate

Synthetic speech stimuli are generated with Qwen3-TTS using stock or cloned voices, including a VoxPopuli speaker different from the original clip’s speaker. Every clip must pass an exact-transcription intelligibility gate before use.

  • Generic synthetic samples use stock voices from Qwen3 1.7B CustomVoice, while cloned samples use Qwen3 Base speakers from the relevant datasets.
  • The Vox-cloned condition uses a VoxPopuli evaluation-set speaker who differs from the speaker in the original clip.
  • Each synthetic clip passes only if at least one of eleven models exactly transcribes the intended transcript under the harness normalizer.
  • Pass rates across gated sets range from 0.84–0.93, including 568/679 consensus-flagged generic renderings.

A.4 Alignment, masking, and number selection

The alignment and masking pipeline converts forced-alignment outputs into word spans, then silences selected number spans while preventing leakage from repeated mentions. Headline masked readouts are ungated, with a hard-cell check designed to reduce language-prior guessing.

  • Alignment: Forced alignment produces per-word spans by converting mms-fa CTC frame indices to seconds and expanding numeric tokens into spoken forms.References are lowercased and stripped before alignment.
  • Alignment: Empty-normalizing words receive zero-width spans at preceding boundaries, and clips shorter than 4 seconds are excluded.This keeps emitted spans index-aligned with whitespace-split references and avoids masks covering a large fraction of short clips.
  • Masking: The masking procedure silences every occurrence of a target value and pads each mask by 120 ms to limit leakage and alignment-boundary errors.Suspiciously tight alignments trigger gap masking, while gaps shorter than 80 ms cause the clip to be skipped.
  • Difficulty gate: Headline masked readouts are ungated, while hard cells require a silenced-audio prior NLL/char of at least 3.5 nats.The threshold identifies 62 of 92 covered spans, or 67%, using the median across white-box-instrumented models.

A.5 Content-preserving perturbations and the robustness protocol

The robustness protocol reruns accept-reference and masked-number probes on content-preserving waveform perturbations. Noise is reported at 10 dB and reverberation at a measured room condition, but reverberation complicates interpretation of masked-number recovery.

  • Protocol: The protocol reruns accept-reference and masked-number probes on perturbed copies of real audio while preserving voice and content.Perturbations are applied to 16 kHz mono waveforms using moderate levels from severity ladders.
  • Perturbations: Additive noise uses Gaussian noise at SNR levels from 20 to 0 dB, with 10 dB reported.The reported condition is the moderate level selected for this perturbation family.
  • Perturbations: Reverberation convolves audio with measured room impulse responses spanning RT60 bins, with the middle bin reported at median RT60 0.52 s.The bins are [0.2, 0.4], [0.45, 0.7], and [0.85, 1.3] s.
  • Protocol: Both perturbations are truncated to source length and RMS-matched to keep duration and energy comparable.The resulting waveforms are sample-for-sample duration-matched to their originals.
  • Protocol: These perturbation families mirror corruptions used in recent augmented benchmarks.
  • Interpretation: At 10 dB noise, accept-reference changes are less confounded by overall WER, whereas reverberation can raise WER and leak masked-word audio beyond the mask.Therefore, masked-number results under reverberation provide weaker evidence than results under noise.

A.6 Splice-induction stimuli

Splice-induction stimuli append benchmark or control donor audio to base clips to test whether context activates benchmark-conditioned transcription behavior. Supplementary probes examine perturbation sensitivity, corpus conventions, context gating, masked-number effects, and steering-related readouts.

  • A.6 Splice-induction stimuli: Each splice stimulus contains a base clip, a 0.15 s silence, and a donor window appended as a suffix or prefix.Donors are RMS-matched to the base clip; conversational audiobook speech serves as the duration-matched control.
  • A.6 Splice-induction stimuli: Base clips are clone renderings of consensus-flagged transcripts that models previously transcribed correctly rather than reproducing reference errors.Real ep-fresh recordings are also used for steering training pairs.
  • A.7 Activation steering protocol: 43 ep-fresh courtesy-opener pairs compare real VoxPopuli and audiobook donors, with 22 pairs for training and 21 held out for evaluation.Encoder activations are mean-pooled over base frames, and a unit-normalized diff-in-means direction is induced or ablated at one layer.
  • A.7 Activation steering protocol: Ablation flips 88% of Cohere, 95% of Parakeet, 93% of Canary, and 45% of Granite no-steer reference reproductions.The generalized readout uses 745 clips and 1,113 edits, counting only non-garbage paired outputs.
  • A.7 Activation steering protocol: Adding the learned direction raises accept-ref from 0.016 to 0.072 for Cohere, 0.013 to 0.047 for Parakeet, 0.011 to 0.105 for Canary, and 0.017 to 0.041 for Granite.Random-direction controls remain flat, while Granite’s steering effect is partial and Phi-4’s direction is causally inert.
  • A.8 Supplementary results: Perturbation results separate models: 10 dB noise collapses Canary-Qwen, Parakeet, and Higgs, while Cohere, Granite, and Phi-4 retain the behavior under noise and reverberation.The supplementary figures also compare corpus-specific honorific conventions, context gating, edit-locus dissociation, and masked-number audio lift.
Loading 2608.19936v1…