Source-linked AI summary
Are Caption Metrics Broken? Latency, Deaf and Hard of Hearing User Ratings, and Bias across Technologies
Bernard Thompson, James Waller, Luz Fanny Calderon Torres, Lu Ming, Mariana Arroyo Chavez, Dante Conway, Raja Kushalnagar, Christian Vogler
TL;DR
The paper asks whether caption-quality metrics reflect DHH viewers’ live-TV experience across captioning technologies. Using a large U.S. survey, it finds that TV and ASR captions receive similar viewer ratings, while metric relationships differ by technology and latency affects experience. The findings caution against treating current metrics as technology-neutral.
Problem
Caption-quality metrics may not adequately reflect DHH viewers’ lived TV experience or support technology-neutral comparisons across captioning methods.
Method
The paper conducts a large U.S. DHH-centric survey and evaluates viewer ratings against WER, ACE2, and NER across TV and ASR captions.
Results
TV and ASR captions received similar viewer ratings, while WER, ACE2, and NER correlated moderately to highly with TV ratings but were substantially biased across technologies.
Takeaways & Limitations
Caption metrics should be interpreted alongside DHH viewer experience rather than assumed to be technology-neutral measures of caption quality.
Takeaways & Limitations
The study used ASR captions from only one vendor, and its findings may differ for other vendors.
Abstract
from arXiv · showhide
Live captions on TV often contain errors and timing issues, making it hard for deaf and hard-of-hearing (DHH) viewers to follow dialog. It is essential that caption quality metrics reflect the lived DHH TV viewing experience. To this end, we describe a U.S.-based large-scale online survey with 216 validated participants, who provided 302 responses containing a cumulative 4,832 data points. Participants viewed videos drawn from a pool of 70 clips recorded from live TV, and were asked to rate the caption quality and subjective understanding of the content across four conditions: TV captions as originally recorded with up to 7-12 seconds delay, TV captions synchronized with audio, Automatic Speech Recognition (ASR)-generated captions synchronized with audio, and ASR captions with an average two-second delay. All captions were evaluated against the Word Error Rate (WER), Automated Caption Evaluation (ACE2) and Number, Edition and Recognition (NER) metrics. Results show that TV and ASR captions were rated similarly. For TV captions, all three metrics were moderately-to-highly correlated with viewer ratings, but far less so for ASR captions, making them far from technology-neutral. Additionally, caption latencies significantly impact the viewer experience, especially typical 7-12-second TV delays. We discuss the implications for the adoption of caption quality metrics.
1 Introduction
The paper addresses whether caption metrics reflect DHH viewers’ lived TV experience and remain valid across captioning technologies. It introduces a large U.S. survey to examine metric alignment, technology neutrality, latency, and individual differences.
- Motivation: The paper targets a validation gap: few studies compare caption metrics directly with DHH viewer ratings for live TV, and NER had not received comparable validation.Prior DHH-centric ACE2 validation focused on transcripts rather than video captions.
- Contribution: The main contribution is a large-scale, DHH-centric U.S. survey with 216 verified participants, 302 responses, and 4,832 video ratings.Participants could customize caption appearance to reduce one-size-fits-all settings as a confound.
- Research questions: The study asks whether WER, ACE2, and NER reflect DHH viewer experience and whether these metrics remain technology-neutral across TV and ASR captions.It also examines demographic effects and the impact of typical broadcast latency.
- Caption metrics: WER counts deletions, substitutions, and insertions equally, whereas ACE2 weights errors by word importance and semantic distance using contextual modeling.ACE2 scores higher when errors are more severe and requires sentence-level transcript alignment.
- Caption metrics: NER evaluates edition and recognition errors while incorporating punctuation and non-speech information such as speaker labels.Its weighted score ranges from 0 for unusable captions to 100 for perfect captions, with 98 or higher classified as good.
2 Related Work
Related work shows that caption quality depends on accuracy, timing, presentation, and viewer diversity. However, direct DHH-rating comparisons across captioning methods and technologies remain scarce.
- Personalization: Caption preferences vary across individuals, genres, cultures, and communication contexts, motivating customizable presentation and flexible accessibility features.Studies examined formatting, speaker identification, non-speech cues, and caption reduction for sports broadcasts.
- Viewer experience: Timing and accuracy are strongly linked to DHH satisfaction, while delays and caption speed can impair readability, understanding, and synchronization with visuals.Prior studies found that shorter delays may be preferred even when captions are verbatim.
- Research gap: Direct comparisons among captioning methods remain limited, especially for DHH viewer ratings and real-world cross-technology evaluation.Existing ASR studies report challenges from unclear speech, noise, speaking rate, and differences between controlled and live conditions.
- Caption metrics: NER research reported average scores above 98 for human captions, lower ASR scores in some evaluations, and near-convergence between newer ASR and broadcast captions.NER has also been deployed in several countries and remains under active evaluation.
- Caption metrics: Prior DHH studies found WER, WWER, ACE, and ACE2 increasingly correlated with viewer ratings, but accuracy alone did not fully capture viewing experience.Highly accurate captions could still be perceived as problematic.
3 Methods
The study uses a within-subjects survey comparing broadcast TV and vendor-generated ASR captions under synchronized and delayed conditions. Participants rated caption quality and understanding across varied live-TV clips with customizable display settings.
- Study design: The 2x2 factorial design crossed caption source—original broadcast TV or AppTek ASR—with latency: synchronized, typical TV delay, or a 2-second ASR delay.Original TV delays could reach 7–12 seconds, while ASR delay averaged 2 seconds with small random jitter.
- Materials: The stimulus set contained 70 live-TV clips across news, finance, sports, talk shows, and competitions, yielding 280 clip-condition combinations.Clips were screened for controversial content and transmission errors, then counterbalanced across participants.
- Data release: The released dataset includes captions, transcripts, WER, ACE2, NER, and participant-level ratings, while demographic and open-ended responses remain managed-access materials.The restriction protects against re-identification risk in a relatively small DHH community.
- Procedure: Each participant viewed 16 distinct clips, four per condition, and rated caption quality and understanding after each video.Video selection and condition order were pseudo-randomized while avoiding repeated videos for the same participant.
- Video playback: The custom player supported adjustable font size, characters per line, colors, opacity, placement, caption replay, and repeated video playback.These controls were refined through focus-group feedback and piloting.
4 Results
Across 4,832 ratings, caption source, latency, and participant hearing status shaped perceived quality and understanding. Delayed TV captions were especially penalized, while TV and ASR captions were similarly rated without delay.
- 4.1 Participant Characteristics: Participants’ hearing status, audio use, and language use were correlated, and White and more highly educated participants were overrepresented relative to U.S. census proportions.Hearing status correlated with audio use at ρ=0.62, while hearing status and English use correlated at ρ=0.47.
- 4.3.1 Caption Quality: Delay significantly reduced quality ratings, with a larger drop for TV captions than ASR captions.The delay × source interaction was significant: F(1,4552.0)=199.5, p<0.001.
- 4.3.1 Caption Quality: A 1.14-point quality gap favored delayed ASR over delayed TV captions, while no-delay TV and ASR ratings differed by only 0.10 points.The delayed-TV versus delayed-ASR difference was significant, whereas the no-delay source difference was marginally significant.
- 4.3.2 Hearing Status and Quality Perception: Hard-of-hearing participants rated delayed TV captions two points lower than synchronized TV captions, while hearing and hard-of-hearing participants were more bothered by delay than deaf participants.Hearing status interacted significantly with both delay and source.
- 4.3.4 Genre and Video Length: Caption quality varied by genre: Financial and Sports clips received lower ratings, while Financial TV captions scored much worse than ASR across delay conditions.For Financial clips, WER was 9.8% for ASR versus 44.6% for TV captions.
4.4 Participant Ratings and Delay Magnitude
Caption delay was associated with larger quality losses, especially for TV captions, while metric–rating relationships differed substantially by caption source. Viewer ratings tracked all three metrics more strongly for TV than for ASR.
- 4.4.1 Delay Magnitude: Larger TV caption delays produced more severe quality drops, with a negative correlation of r=-0.552.TV broadcast delay averaged 8.46 seconds, whereas ASR delay was fixed at 2 seconds.
- 4.4.2 Caption Quality Metrics: Viewer ratings correlated strongly with WER, ACE2, and NER for TV captions but weakly across all three metrics for ASR captions.WER and ACE2 were negatively associated with quality, while NER was positively associated.
- 4.4.3 TV–ASR Metric Comparisons: TV captions had substantially higher WER than ASR captions, with a strong-to-very-strong paired effect size of 0.86.ACE2 also favored ASR with a moderate-to-strong effect size, while NER showed a weak-to-moderate effect favoring TV.
- 4.4 Participant Ratings and Delay Magnitude: Figure 11 plots transcript delay against quality change, and Figure 12 plots each metric against mean quality separately for TV and ASR.The figures visualize the source-specific delay and metric relationships described in the analyses.
4.6 Participant Caption Settings
Participants frequently customized caption presentation and emphasized synchronization, placement, readability, speaker identification, and non-speech information. Delay was repeatedly described as more disruptive than caption errors, while customization features were valued but uncommon on TV.
- 4.6 Participant Caption Settings: Nearly half of participants adjusted caption placement, and most modified font size or characters per line.The default placement was bottom-centered, while many participants preferred more than 32 characters per line.
- 4.7 Participant Comments: Participants described delay as more disruptive than caption errors because desynchronization interferes with comprehension and attention.Several responses preferred incomplete captions in sync with speech over more accurate captions that lagged behind.
- 4.7 Participant Comments: Customization features were positively received but were described as uncommon on television.Participants valued adjusting caption color, style, placement, and other presentation settings.
- 4.7 Participant Comments: Participants identified accuracy, timing, readability, speaker identification, and important sound effects as critical live-caption features.They also noted that caption positioning can obscure on-screen graphics when television channels fix the placement.
- 4.7 Participant Comments: Sports and news generated especially negative comments about confusion, poor captioning, and delayed timing.The paper presents these open-ended responses as illustrative examples rather than a full thematic analysis.
5 Discussion
The discussion finds that caption metrics track DHH viewer ratings better for broadcast TV than ASR captions, while latency substantially harms viewing and should be prioritized in quality oversight. It also emphasizes that participant differences and regulatory incentives complicate the adoption of a single metric.
- 5 Discussion: The study’s larger sample and more diverse demographics increase statistical power and confidence in reported effects, although some findings mainly confirm prior literature.The authors frame these confirmations as valuable because of the study’s scale and participant diversity.
- 5.1 RQ1: Viewer Characteristics and Experiences are Related to Metrics: NER correlated with TV viewer ratings as strongly as WER and ACE2, but its additional penalties for speaker identification, non-speech information, and punctuation did not improve correlation.For ASR captions, NER was only moderately correlated and no better than WER, although better than ACE2.
- 5.1 RQ1: Viewer Characteristics and Experiences are Related to Metrics: TV and ASR captions received negligibly different viewer quality ratings, despite WER and ACE2 favoring ASR and NER penalizing it more harshly.The results therefore reveal source-dependent bias in the metrics rather than a viewer-perceived ASR advantage.
- 5.3 RQ3: Caption Delays Have a Major Impact: Caption latency produced substantial rating penalties: participants strongly disliked broadcast delays, and even 2-second ASR delays reduced ratings.Hard-of-hearing participants penalized delays especially harshly, while the predicted greater tolerance for ASR limitations held only when captions were synchronized.
- 5.3 RQ3: Caption Delays Have a Major Impact: The findings support prioritizing delay reduction over accuracy metrics and checking the full caption-production, encoding, and delivery workflow.The discussion characterizes delay reduction as a relatively accessible regulatory priority.
- 5.1 RQ1: Viewer Characteristics and Experiences are Related to Metrics: All three metrics correlate moderately to strongly with ratings for TV captions but only weakly to moderately for ASR captions.The metrics distinguish good from bad broadcast captions more successfully than good from bad ASR captions.
- 5.4 Metrics, Value Capture and Individual Needs: The authors warn that regulatory metrics could become incentives in themselves, capturing value at the expense of DHH-centered needs.They discuss representing caption content more completely while shifting customization toward playback devices and software.
- 5.4 Metrics, Value Capture and Individual Needs: Ratings varied substantially across participants, with many videos showing standard deviations above 1.5–2 points on a 7-point scale.This variability suggests that caption quality is highly individualized, complicating universal metric choices.
6 Limitations
The study’s limitations concern recruitment friction and sample representativeness, caption-source and latency scope, and uncertainty about production workflows. The authors emphasize that end-user caption quality depends on what viewers actually see, including errors introduced after caption generation.
- Recruitment and sample: Recruitment friction made participation difficult, increased completion time and cost, and created challenges for obtaining representative U.S. DHH samples.Online and in-person recruitment had different completion patterns, while in-person recruitment was expensive and not sustainable.
- Recruitment and sample: The sample skewed toward women and self-identified deaf participants, underrepresenting hard-of-hearing people, non-signers, and participants with less-than-college education.The authors caution that the findings may not capture the full range of education, literacy, and diversity in lived DHH experience.
- Generalizability: The findings are U.S.-specific and should not be generalized to other countries.The authors identify national scope as a direct boundary on interpretation.
- Caption conditions: ASR results used one vendor and a fixed two-second delay with jitter, so other vendors and real-world latency ranges may produce different results.The authors nevertheless state that the concern about metrics not being fully technology-neutral would remain if vendors differed.
- Caption workflows: Broadcast workflows can introduce errors after caption generation, meaning observed stimulus quality may be lower than comparable studies’ quality.These errors can arise at multiple points in the encoder-to-transmission chain, and the study could not identify whether TV captions were produced by steno, respeaking, or ASR.
- Caption workflows: From a DHH user-experience perspective, production-time quality matters less than the caption quality ultimately displayed to viewers.The authors frame end-user output, rather than upstream production quality, as the relevant experience measure.
7 Future Work
Future work should clarify acceptable caption-delay thresholds and examine end-to-end broadcast workflows. It should also improve metrics and develop models that better capture DHH viewers’ perceptual experience.
- Latency: Future studies should identify where caption latency becomes acceptable or unacceptable for broadcast TV and ASR captions.The authors report that delays are detrimental but say more information is needed about the relevant tipping points.
- Latency: Researchers should examine end-to-end broadcast caption workflows to locate and eliminate major sources of delay.This extends latency research beyond viewer ratings to the operational chain producing and transmitting captions.
- Metric development: Better language models could improve ACE2 by estimating word importance and semantic distance more effectively, potentially making it more predictive of user ratings.The proposed direction targets ACE2’s language-sensitive evaluation components.
- Metric development: AI or machine-learning models that emulate DHH perceptual experience remain promising but currently lack sufficient generalization.The authors identify generalization as the key unresolved issue for this modeling direction.
- Metric development: NER could be improved by studying whether its 2008 NCRA guidelines remain aligned with contemporary DHH caption-viewer experiences.The proposed work questions the continuing fit of the guideline basis underlying the evaluation model.
- Metric development: AI and LLM assistance could reduce NER evaluation costs and help ACE2 and NER align transcripts automatically.These proposals address both the cost of evaluation and the effort required for transcript alignment.
8 Conclusion
The study links caption metrics to DHH viewer ratings while showing that metric behavior differs between TV and ASR captions. It also finds that caption latency and customization materially shape the viewing experience and should inform metric adoption.
- The study provides evidence connecting caption-quality metrics with DHH experience through a large-scale survey of carefully vetted participants.The work is presented as evidence based on a large-scale U.S. survey with high statistical power.
- WER, ACE2 and NER show similar moderate-to-high correlations with TV viewer ratings, but differ in computational cost.The study compares the metrics as measures related to viewer ratings and notes trade-offs in calculation expense.
- None of WER, ACE2 and NER is technology-neutral because their behavior differs substantially between TV and ASR captions.Metric differences between caption technologies appear in the metrics but not in participant ratings.
- Government adoption of caption metrics should avoid value capture and preserve the diversity of lived DHH experience.The conclusion explicitly frames this as a regulatory concern.
- TV caption latency has a major impact on viewer experience, especially when captions are delayed relative to audio.The conclusion identifies latency as an issue that needs to be addressed.
- Caption placement and characters-per-line customization options are important for DHH viewer experience in the U.S.Nearly half of participants adjusted caption placement settings, and the study reports customization as an important aspect of the experience.
AI Statement
The authors disclose that AI was not used to write the paper, while limited AI assistance supported copyediting and statistical-analysis code generation under author review.
- AI was not used in writing the paper, but supported post-review copyediting and parts of the R-code and ggplot-generation process.The authors state that generated code and figure attributes were vetted for correctness.
A Distribution Checks and CLMM Models for Ordinal Responses
The appendix checks response distributions and compares linear mixed models with cumulative link mixed models for ordinal ratings. The conclusions are generally robust across model choices, with one adjusted interaction changing significance.
- The analysis reruns the three main linear mixed models as cumulative link mixed models for ordinal responses.The comparison is intended to assess whether the main results are robust to the response-model specification.
- Quality ratings span the full 1–7 scale, with under 21% of responses at either floor or ceiling and low-to-moderate skew.These distributional patterns support similar results from linear and ordinal approaches.
- Understanding ratings are more compressed at the ceiling, with up to 37% of responses there and negative skew across conditions.The reported skew ranges from -0.42 to -1.18.
- The robustness checks use the same fixed effects and crossed random intercepts for participants and videos across the compared models.Models cover quality ratings, quality ratings with hearing status, and understanding ratings.
- The three-way interaction in models including hearing status changes from adjusted p=.056 to p=.018, while other null-hypothesis conclusions remain unchanged.This is the only exception reported at α=.05 across the robustness comparisons.
- Both linear and ordinal models find significant source, delay, and source-by-delay effects for quality and understanding ratings.The corresponding model tables report the same significant main effects and interactions at α=.05.
B Caption Metrics
This appendix provides caption-metric values for each video stimulus across the reported conditions. The tables define metric directionality and explain how delayed TV-caption values are represented.
- Tables 8 and 9 report caption metrics for all video stimuli.
- For TV-related columns with two entries, the values represent no-delay and delayed-TV captions.The tables use the format no delay/delay for these paired values.
- Lower WER and ACE2 scores are better, whereas higher NER scores are better.The table legend also identifies NER values above 98 as exceeding the acceptable threshold.