Source-linked AI summary

The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev, Grach Mkrtchian

arXiv:2608.15940v2cs.CLcs.LGcs.SD

TL;DR

Encoder–decoder models can hallucinate fluent text from message-free inputs, so this paper tests whether reserved null-token scores support abstention in ASR and NMT. It finds that null-token and decoder-state signals can guide suppression, but stronger intervention trades fabrication reduction for deleted speech or shortened translations.

  • Problem

    The paper asks whether reserved null-token scores provide usable abstention signals for fluent hallucinations from message-free inputs in ASR and NMT.

  • Method

    The study audits native null-token scores and score shifts across models, while probing Whisper decoder states and supervised row edits against external gates.

  • Results

    On the only post-freeze sample, Whisper-small’s trained row moves LOR 91.7% →3.3% and deleted reference words/min 1.58 →33.05.

  • Takeaways & Limitations

    Null-token scores and decoder states serve as diagnostic instruments for moving short-segment abstention operating points, without establishing a safer risk frontier.

  • Takeaways & Limitations

    The comparison uses an illustrative 10% real-speech empty-output cap, and its cost-gap uncertainty is unquantified.

Abstract

from arXiv · show

Modern encoder-decoder systems can produce fluent text even when their input contains no recoverable message. We study this failure in ASR and NMT through the models' reserved null tokens, asking whether the score for ending generation already carries a usable abstention signal. Across speech recognizers and translation models, we audit native null-token scores and scalar logit shifts. In Whisper, we additionally probe decoder states and compare supervised row edits with conventional external gates. The evaluated models often expose a useful abstention signal, but stock decoding does not reliably act on it. Raising the null-token score can sharply suppress fabrication, but aggressive intervention also deletes valid speech or shortens legitimate translations. These findings turn the null token into a diagnostic lens on hallucination and motivate evaluating abstention methods by both suppression and deletion costs, rather than by hallucination reduction alone.

1 Introduction

The introduction frames fluent fabrication from message-free inputs as a shared failure in attention-based speech recognition and neural translation. It asks whether reserved null-token signals can support abstention while accounting for the tradeoff between suppressing fabrication and deleting valid output.

  • Problem: Attention-based speech recognizers and neural translators can produce fluent text from silence, room tone, music, or other message-free inputs.The introduction notes documented downstream harms for speech recognition and analogous failures for translation.
  • Prior approaches: Operational responses include native nospeech scores, VAD or endpointing, and reject options, alongside retraining, feature steering, and teacher-distillation approaches.The introduction contrasts these methods with the models’ already exposed reserved output coordinate.
  • Research question: The paper asks what separation decoder states and native null-token scores provide, and what fabrication–deletion tradeoff results from editing the reserved coordinate.This frames readout edits as instruments for moving the abstention operating point rather than as standalone semantic detectors.
  • Developmental evidence: On the developmental battery, frozen decoder states separate non-speech from speech after the first decoder block, while native null-token scores distinguish conditions even without controlling the output decision.Scalar biasing changes constant prior odds, whereas the analytic graft and trained row are supervised, state-dependent variants of one output coordinate.
  • Evaluation scope: The evidence spans a 1060-clip stress battery, a source-disjoint locked Whisper-small 300/300 evaluation, and a 15-model zoo covering AED recognition, CTC/RNN-T, and neural translation.The analysis is organized around battery separation and readout, checkpoint auditing, and boundary tests across architectures, modalities, languages, and long-form audio.
  • Measurement boundary: Fabrication is measured as deterministic normalized lexical output on spans with exact empty lexical targets, after removing documented non-lexical tags, punctuation, and digits.The score is normalizer-dependent, is not semantic adjudication, and excludes fabrication interleaved with genuine content.

2 Background and Related Work

Prior work frames ASR and NMT hallucination as a rejection, calibration, and no-speech-scoring problem, with related Whisper and decoding interventions. This paper positions its contribution as an audit of direct null-token edits and their fabrication–deletion costs.

  • Related work: ASR and NMT hallucination is documented, with native no-speech scoring, VAD/endpointing, and selective prediction providing closely related reject mechanisms.Complementary Whisper mitigations edit attention heads, encoder features, teachers, or emitted text.
  • Related work: Label-shift correction, EOS calibration, and frozen-state or logit-lens probes provide established tools for characterizing abstention signals.NMT miscalibration particularly affects EOS, and inference miscalibration persists under generated context.
  • Scope: Coverage, length, and lexical constraints alter translation search, whereas this work edits the null-token scoring function rather than using those methods as matched baselines.The experiments test greedy ASR and each NMT backend’s fixed decoder; beam-width invariance and continuous long-form inference are outside the claim.
  • Contribution: The empirical contribution is a diagnostic and falsification audit of direct null-coordinate edits, including fabrication–deletion costs across five Whisper-small checkpoint conditions.The audit identifies where stronger representation, prior-shift, risk-frontier, and cross-family hypotheses fail.

3 Experimental Setup

The experiments use a diverse stress battery and model zoo to evaluate null-token interventions under controlled data partitions. Developmental comparisons are complemented by a locked, non-overlapping evaluation set and explicitly defined sensitivity criteria.

  • Data and metric: The stress battery contains 630 non-speech and 430 speech clips spanning synthetic, environmental, and clean/degraded speech conditions.The total duration is 3.1 h; MUSAN is split deterministically into disjoint calibration and evaluation halves, while FLEURS test and LibriSpeech test-clean are evaluation-only.
  • Models and interventions: The model zoo covers six Whisper scales, Canary-1B, OWSM v3.1, five CTC recognizers, two RNN-T recognizers, NLLB-200, and MarianMT.Additional stock and readout checks include Distil-Whisper, SeamlessM4T-v2-large, and Parakeet-TDT-1.1B; no intervention claim is made for these checks.
  • Models and interventions: Three interventions are compared with the same decoder: scalar null-token biasing, supervised analytic row grafting, and trained EOT-row updates.The first adds a constant to one logit, the second adds a learned hidden-state direction to one unembedding row, and the third updates one output coordinate on balanced calibration data.
  • Analysis and selection pools: After methods and configurations were frozen, evaluation used 300 non-speech and 300 speech clips with no overlap in hashes, recordings, or speakers with fitting or selection pools.The held-out speech set includes clean, low-SNR, WHAM!-degraded, and accented inputs, while a collision checker rejects contaminated manifests.
  • Analysis and selection pools: On the developmental common-cost set, an operating point is within cap when at most 10% of real-speech clips are emptied.This threshold is illustrative rather than application-derived; the locked-set sensitivity family was pre-specified and frozen before decoding.

4 RQ1: Developmental Separation and Readout

Decoder states and native null-token scores separate speech from non-speech, but native readouts do not reliably trigger abstention. State- and score-based interventions can suppress hallucinations, with tradeoffs in speech deletion, implementation, and sequence-position coverage.

  • Developmental separation: From the first decoder block onward, linear probes reach the maximum observed AUROC on all six scales, while layer 0 remains at chance.The permutation null is 0.50, and a top-50 PCA probe reproduces the result; folds are not source-corpus grouped.
  • Native readout: Native abstain readouts distinguish evaluated speech conditions, but final message-free margins remain below the abstention boundary in every model.The native coordinate orders conditions without selecting abstention, and message-free null-token scores often remain below content-token scores.
  • Readout interventions: The trained EOT-row update is weakly aligned with the no-speech direction, and grafting that direction into EOT does not suppress normalized lexical output.The update cosine is comparable to the upper tail of random-direction cosines on every model.
  • Readout interventions: 4.3% residual LOR remains for Whisper’s SOT no_speech_prob detector under the illustrative 10% speech-side cap, versus 89.8% for native EOT-margin thresholding.This indicates that the observed failure is specific to the EOT operating point rather than an absence of native abstention signal.
  • Readout interventions: On the common-cost set, a frozen-state logistic gate reaches 0.0% LOR and FireRedVAD 1.3% within the cap, while row edits show differing speech costs.The constant EOT bias records 24.3% LOR and 7.4% real-speech empty output; the trained row records 0.2% and 67.7%, and the analytic graft 0.0% and 35.6%.
  • Position-wise readout: Calibrated readouts increase margins most at early positions, but turbo and large-v3 remain negative later, limiting the intervention to short, whole-segment abstention.Removing the stock first-position EOT suppression rule does not change stock LOR, although it can block a calibrated EOT decision.

5 RQ2: Null-Coordinate Diagnostic Audit

The audit finds that scalar null-coordinate corrections require strong calibration assumptions and that their no-fit operating-point prediction fails on translation models. Supervised row edits can suppress message-free output, but their deletion and speech-quality costs are model- and method-dependent.

  • Scalar diagnostic assumptions: A scalar correction is exact under label shift only when the neural margin is calibrated unit-scale posterior log odds; stock decoder margins need not satisfy this.Unit base-rate slope and margin-predicted error-balance bias are applicability checks, not evidence of the training prior or a causal explanation for stock behavior.
  • Scalar diagnostic assumptions: The step-0 margin predicts abstention under null-logit bias β, but the β⋆ operating point is Bayes-optimal only under symmetric-margin and equal-cost assumptions.The prediction agrees with Canary at grid resolution, is adjacent to OWSM, and fails on both translation models; these dose-response results reuse the selection subset.
  • Whisper row recalibration: A random-effects analysis of Whisper EOT-row slopes estimates 1.29 with 95% CI [0.94, 1.65] and substantial heterogeneity (I2 = 84%), rejecting pre-specified unit-slope equivalence.Every per-scale slope is positive, but some scales include unit slope while others are steeper.
  • Whisper row recalibration: Supervised row recalibrators modify hundreds of EOT-row parameters, while the analytic graft adds a state-dependent probe direction scaled by λ within calibration-derived bounds.The analytic and trained directions are weakly aligned, indicating distinct recalibrators of the same output coordinate; the analytic method performs well on five Whisper scales but regresses on medium.
  • Whisper row recalibration: At the common step-50 checkpoint, LOR reaches zero on all six Whisper scales; held-out probe WER is unchanged on four scales but rises on turbo and large-v3.The disjointly fitted analytic graft leaves 0.3–8.4% LOR and includes a medium test-clean change from 2.88 →7.60 WER.
  • Locked endpoint evaluation: On the locked frozen Whisper-small set, trained-row LOR falls from 91.7% to 3.3%, but deleted reference words rise to 33.05 per minute versus stock’s 1.58.Scalar bias b = 5 has the lowest point-estimate cost at κ = 0.5, 1, while stock is lowest for κ ≥2 and the trained row never minimizes the displayed grid.

6 RQ3: Boundary Tests Across Evaluated Systems

Across unrelated AED recognizers and translation models, null-token or EOS scores provide controllable abstention signals, but intervention costs vary substantially. CTC/RNN-T blanks and multilingual settings require distinct calibration, and no method is uniformly safest.

  • AED and translation models: Across unrelated AED recognizers and translators, step-0 null-token margins distinguish message-free from meaningful input, while scalar null-token biases reduce output.The evaluated recognizers are Canary-1B and OWSM v3.1; the translators are NLLB-200 and MarianMT.
  • AED and translation models: Canary reduces LOR 79.3 →0.0% with WER 7.02 →6.94, whereas OWSM reduces LOR 90.7 →1.3% with WER 8.19 →20.9.For OWSM, β = 9 is the empirical error-balance point and β = 10 is the lower-LOR endpoint; the reported points differ in operating conditions.
  • AED and translation models: Translation EOS bias reduces lexical output on empty, punctuation-only, whitespace, and random-character sources while also shortening real translations and lowering quality.At the reported NLLB point, lexical-output rate changes 27.5 →1.0% while COMET changes 83.4 →80.9.
  • CTC and RNN-T: CTC and RNN-T blanks are natively unsuppressed, and anchor fine-tuning or blank bias helps different systems, with no uniformly safest intervention.Their non-speech outputs are usually short fragments rather than fluent text, and the results do not establish the AED mechanism for these architectures.
  • Multilingual calibration: Multilingual abstention requires a separately calibrated multilingual row or one scalar per evaluated language rather than reliable zero-shot transfer.The multilingual row is calibrated on seven language tags and evaluated on eight held-out tags; Korean is the largest residual.
  • Whisper developmental analysis: On Whisper-small’s inspected training matrix, a 768parameter EOT row reaches zero developmental LOR at the first checkpoint without changing probe WER, demonstrating parameter efficiency rather than unique safety.Larger and randomly located updates also reach zero, while AudioSAE and reduced LLT provide no safer common operating point.

7 Limitations and Scope

The evaluation is limited to whole-segment message-free inputs, where fabrication is defined by construction, and does not assess interleaved fabrication or speech-containing cases requiring transcript adjudication. Its operating-point and deployment conclusions remain conditional on false-silence costs, target-domain calibration, model scale, decoding setup, and incomplete evaluation coverage.

  • Evaluation scope: The study measures normalized lexical output only on whole-segment empty targets, so every retained lexical token is fabricated by construction.Interleaved fabrication is excluded because it requires per-token decisions on non-empty targets and lacks a label-free detector.
  • Cost and safety: The common-cost comparison caps emptying real-speech clips at 10%, an illustrative operating constraint rather than a measure of application utility.Lexical-output suppression and false silence therefore require joint evaluation; external state and VAD gates achieve lower developmental LOR within this cap.
  • Cost and safety: Model-dependent safety results restrict checkpoint-editing evidence: tiny and base improve degraded speech through insertion and substitution reductions despite more deletions, while four stronger Whisper scales increase WER.Medium and OWSM retain narrow low-cost regions; in long-form decoding, the edited row reduces gap words but raises WER, whereas correctly implemented VAD segmentation dominates both axes.
  • Deployment scope: Deployment requires target-domain calibration, a non-speech base-rate estimate, and explicit false-silence costs because clean-speech constraints do not certify new acoustics.The locked 300/300 evaluation lacks the intended offline spontaneous stratum and is not a complete deployment evaluation.
  • Evaluation scope: Speech-containing inputs remain excluded because hallucination and mistranscription require human transcript adjudication, which was not performed for the headline endpoint.Empty lexical references are used for evaluated message-free spans after documented non-lexical event-tag removal.
  • Deployment scope: Current normalization leaves 13 of 15 models invariant, while Whisper medium shifts 0.31 points and OWSM shifts 6.67 points; LLT remains a reduced proxy and beam search is untested.False silence peaks in locked low-SNR and accented groups, motivating avoidance of small models on degraded or spontaneous speech, long-form audio, or when false silence costs more than fabrication.

8 Conclusion … Statistical, Scoring, and Protocol Audits

The paper concludes that null-token scores and compact logit or row edits diagnose abstention miscalibration and fabrication–deletion tradeoffs, but do not establish universally safe abstention. Post-freeze audits constrain interpretation through locked costs, family-specific effects, protocol limitations, and unresolved statistical and measurement uncertainty.

  • 8 Conclusion: Native null-token scores and decoder states separate evaluated conditions, while scalar bias, analytic grafts, and trained-row edits shift operating points without proving a privileged mechanism.Clip-stratified probes do not establish a source-independent message-absence representation or safer risk frontier.
  • 8 Conclusion: 91.7% → 3.3% LOR accompanied the trained row’s 1.58 → 33.05 deleted reference words/min on Whisper-small’s only post-freeze sample.Bias b = 5 had the lowest displayed Cκ point estimate for κ = 0.5, 1, and stock for κ ≥2, but rank uncertainty was unquantified.
  • Data Roles and Source Disjointness: The validation used 12 pools across four roles, a developmental battery of 1060 clips from 869 recordings, and a locked set of 300 non-speech plus 300 speech clips.Locked non-speech and speech sources were specified, while the planned spontaneous stratum was unavailable and test-clean was not previously unseen.
  • Matched Developmental Cap Comparison: All matched developmental comparisons used 630 non-speech and 430 speech clips; external detectors had lower developmental LOR within the coarse speech-side cap, not a locked superiority comparison.The first four detector rows were separately thresholded, but their exact achieved speech-side rates were unavailable in the summary.
  • Locked Fabrication–Deletion Sensitivity: The locked endpoint defines Cκ as a conditional-rate sensitivity index balancing retained non-speech words against deleted speech reference words, not deployment utility or total WER.Methods and κ ∈{0.5, 1, 2, 5, 10} were pre-specified before the 300/300 decode; external gates were not run on this set.
  • Held-Out Multilingual Speech Cost: Held-out multilingual evaluation calibrated 40 clips and reported on 60 disjoint clips per language, with smaller models reaching low mean LOR by abstaining heavily.Large-v3 and turbo retained real speech at the calibrated point, while Korean was the worst residual language across all six scales.
  • Conditional Prior-Shift Derivation: Prior-shift correction requires fixed class-conditionals and a unit-scale calibrated margin; the evaluated positive, heterogeneous slopes therefore do not establish the classical unit coefficient.The pooled slope was 1.29 with 95% CI [0.94, 1.65], while equivalence to unity under [0.8, 1.2] failed.
  • Statistical, Scoring, and Protocol Audits: Statistical and scoring audits found high slope heterogeneity, normalization sensitivity up to 6.67 points on OWSM, and incomplete annotation, limiting semantic and protocol conclusions.Multilingual NMT points were within cap only when at least 90% of genuine-source translations remained non-empty; NLLB Chinese and Marian Spanish failed this criterion.

B Implementation and Procedural Checks · C Analytic Row-Graft Procedure

The paper identifies implementation defaults that can make fine-tuning appear ineffective or distort abstention evaluation, then specifies an analytic row-graft that recalibrates one output coordinate from supervised decoder states. Together, these procedures define the training, evaluation, intervention, and calibration checks used in the study.

  • B Implementation and Procedural Checks: Training uses a 50/50 example-balanced anchor mix with specified optimization settings, while the study compares sixteen intervention regimes including eot_row and LayerNorm variants.Runs use 400 steps, evaluation every 50 steps, gradient clipping at 1.0, and an EOT-row learning rate of 10−3.
  • B Implementation and Procedural Checks: Four HuggingFace-pipeline defaults can silently reproduce the impression that the model resists fine-tuning, so each requires explicit diagnosis and correction.The paper frames these defaults as implementation failures independent of the model itself.
  • B Implementation and Procedural Checks: Prefix masking can create decoder context [sot, eot, eot, . . . ], yielding teacher-forced loss near ln V ≈10.7; passing decoder_input_ids explicitly fixes it.The collator converts masked −100 labels into pad/EOT tokens during shift_tokens_right.
  • B Implementation and Procedural Checks: 89% to 100%: enabling begin_suppress_tokens can increase hallucination by forbidding EOT as the first token, even when P(EOT | noise) ≈1.Fine-tuned models should be evaluated with begin_suppress_tokens=None; this was verified as a no-op on stock models.
  • B Implementation and Procedural Checks: ∼2%: token averaging gives the anti-hallucination gradient only this share of the weight because non-speech targets contain one EOT token versus roughly forty speech tokens.Under token averaging, −log P(eot) barely moves from 7.3 →6.4 over 400 steps; example-balanced loss fixes the imbalance.
  • B Implementation and Procedural Checks: With use_reentrant=True, checkpointed parameters whose inputs lack gradients receive none, leaving only two final LayerNorms of 156 tensors trained on turbo.Using use_reentrant=False fixes this gradient omission.
  • B Implementation and Procedural Checks: Checkpoint selection by clean-probe WER can favor destructive late checkpoints, leaving up to 62% of spontaneous out-of-domain speech empty; selection must include out-of-domain speech.Scorers must also strip non-speech tags such as [music] before counting fabricated letters.
  • C Analytic Row-Graft Procedure: The analytic graft fits a logistic direction to frozen decoder states, using non-speech and genuine end-of-utterance states as positives and mid-utterance speech states as negatives.It is a supervised linear recalibrator of one output coordinate, not a scalar prior-odds correction.

D Reporting and Reproducibility Details · E Deployment Checklist · F Stress Set, Statistics, and Spontaneous Speech

The paper documents its reproducibility artifacts, deployment safeguards, and stress-test design while emphasizing that abstention can cause substantial false silence and deletion costs. Its evaluation uses controlled subsets, paired statistical tests, capacity-matched verification, and spontaneous-speech checks.

  • D Reporting and Reproducibility Details: The prepared artifact inventory includes immutable model revisions, lockfiles, split manifests, preprocessing code, seeds, configurations, selection logs, learned parameters, outputs, and reproduction commands.The inventory also resolves short model labels, although the supplied passage truncates before listing them.
  • D Reporting and Reproducibility Details: The principal foreseeable harm is false silence: the locked row suppresses 42.7% of low-SNR speech and 38.7% of accented speech in measured groups.These disparities motivate an explicit non-deployment recommendation for the row, but the supplied passage truncates before stating the full recommendation.
  • E Deployment Checklist: Deployment requires target-domain calibration across message-free, clean, and degraded speech, no evaluation overlap, explicit fabrication-versus-silence costs, and verification that decoding permits the null token.The checklist conditions prior-shift estimation on whether its scalar model is appropriate; otherwise it selects from a labeled cost curve.
  • F Stress Set, Statistics, and Spontaneous Speech: The stress battery contains 1060 clips spanning 3.1 hours at 16 kHz mono, with clips no longer than 30 seconds and RMS-normalized speech sources.Non-speech strata have empty references, speech strata are FLEURS-derived, and potentially verbal UrbanSound8K and vocal MUSAN material is excluded.
  • F Stress Set, Statistics, and Spontaneous Speech: Evaluation subsets differ by computational cost, and every intervention is compared with stock re-decoding on the same subset rather than across rows.The first three subsets are fixed and nested only when their listed counts coincide; NS denotes non-speech.
  • F Stress Set, Statistics, and Spontaneous Speech: All learned-patch LORs are 0.0% with Wilson CI [0.0, 0.6], while the independently decoded analytic patch has c = 0 at every scale and exact p < 10^-130.Table 16 reports paired McNemar tests over all 630 non-speech clips, with Holm correction across six scales and Haldane-corrected odds ratios.
  • F Stress Set, Statistics, and Spontaneous Speech: Matched-capacity random-subnetwork tests show reachability at most 2% across scales, but WER safety is not scale-wide because medium is acutely learning-rate-sensitive.The result supports no privileged parameter locus, not interchangeable operating points.
  • F Stress Set, Statistics, and Spontaneous Speech: Spontaneous-speech evaluation exposes checkpoint-selection risk: late checkpoints leave up to 62% of spontaneous speech empty, while tiny incurs +27–+39 pp WER under any zeroing checkpoint.The supplied passage concerns Earnings-22 WER, deletion rates, and empty outputs for LayerNorm and eot_row regimes, but truncates before its full comparison.

G Cross-Lingual and Cross-Architecture Detail

Cross-language and cross-architecture evaluations show that targeted null-token or EOS interventions can improve abstention-related behavior, but their benefits are tightly coupled to transcription and translation costs. Language-identification repair is an exception: it substantially improves detection while leaving transcription unchanged.

  • Held-out languages: Seven held-out tags remain at 1.5–3.0% hallucination, with Korean as the visible 13.3% exception in multilingual Whisper-small calibration.The evaluation is multilingual calibration followed by language holdout, not English zero-shot transfer.
  • Language identification: Mean langid rises 0.29 →0.94 while WER is bit-identical across fifteen FLEURS languages and seven scripts after a scoped language-token repair.The edit changes only the detection step, not transcription; Japanese and Chinese word-WER invariance is the relevant comparison.
  • Cross-architecture steering: At α = 1, SAE steering barely moves HR, while stronger steering destroys transcription before suppressing fabrication; zero HR requires WER of 100%.One replicated SAE layer across twelve encoder layers contributes 113.3M parameters.
  • Cross-architecture EOS bias: Maximum step-0 ∆P(EOS) is only 0.0006 on Canary and 0.0001 on OWSM, so ablating selected decoder heads has little or no effect.The paired OWSM comparison uses maxlenratio=0.4, making its absolute WER incomparable with other OWSM tables.
  • Cross-architecture EOS bias: At β = 9, OWSM reaches the empirical balance point, whereas the 1.3% HR headline at β = 10 more than doubles WER.The bias provides monotone hallucination control but no free operating point on the fixed 150 non-speech plus 60 speech subset.

H Cross-Architecture Logit-Lens Detail

Cross-architecture decoder-depth traces show that similar final abstention failures can arise from different internal trajectories. Canary, OWSM, and NLLB retain negative mean null-token margins throughout, whereas Marian favors EOS mid-stack before reversing that preference in its final block.

  • Cross-Architecture Logit-Lens Detail: Canary, OWSM, and NLLB keep the mean null-token margin below zero throughout decoder depth.Their message-free inputs finish above clean inputs but still below the EOS argmax boundary.
  • Cross-Architecture Logit-Lens Detail: Marian makes EOS the mean argmax through the middle decoder stack for all three source conditions, then reverses that preference in the last block.The reversal is strongest on clean sources.
  • Cross-Architecture Logit-Lens Detail: Message-free input finishes above clean input in every evaluated model, but remains below the EOS argmax boundary.Figure 4 reports 95% confidence bands over n = 120 clips per condition for Canary/OWSM and n = 200 sources per condition for NLLB/Marian.
Loading 2608.15940v2…