Source-linked AI summary
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
Anupam Purwar, Shashank Singh, Kritika Srivastava
TL;DR
Reliable, scalable evaluation of voice agents requires comparing automated judgments with human contextual assessment. The paper evaluates GPT-4.1 and GPT-5 against human scores across telecom and retail conversations and three configurations, finding stable relative trends but metric- and context-dependent absolute agreement. It therefore supports scalable LLM first-pass evaluation combined with human review for safety-critical and recovery-focused judgments.
Problem
Human evaluation is reliable but costly and difficult to scale, while LLM judges may vary with metrics, contexts, and evaluation settings.
Method
The study compares human and GPT-4.1/GPT-5 judgments of the same retail and telecom conversations across p0, p1, and p2 configurations and multiple metrics.
Results
LLM judges show stable metric-level trends and often follow human relative assessments, but absolute scores diverge most for safety and recovery metrics.
Takeaways & Limitations
LLM judges support scalable first-pass evaluation, while human review remains necessary for safety-critical and recovery-focused decisions in a hybrid pipeline.
Takeaways & Limitations
The benchmark does not model voice-specific temporal and interaction phenomena such as interruptions, overtalk, silence, or missed response windows from text transcripts alone.
Abstract
from arXiv · showhide
Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We investigate LLM-as-a-Judge evaluation by comparing human judgments with GPT-4.1 and GPT-5 on telecom and retail voice-agent conversations, across conversational quality and safety dimensions. The same interac- tions are scored under three evaluation configurations, p0, p1, and p2, to test whether automated judgments are sensitive to the evaluation setup and whether observed patterns generalize across configurations and judge models. Beyond aggregate agreement, we examine metric-level correlations, evaluator consistency, and systematic human-LLM disagreement to identify which conver- sational attributes can be judged reliably by automation and which remain sensitive to interpretation and context. Effective voice-agent evaluation is also shaped by pipeline-level factors such as speech generation, streaming, and error propagation across ASR, reasoning, and tool-calling stages, motivating our focus on comparing how human and LLM judges score the same interactions end to end. Our results show that LLM- based evaluation can serve as an effective component of large- scale voice-agent assessment, but that its reliability is metric- and configuration-dependent rather than uniform. This pro- vides an empirical framework for identifying which metrics suit automated evaluation and supports hybrid pipelines in which LLM judges handle scalable assessment while human evaluators remain engaged for metrics that demand contextual interpretation and higher-confidence judgment.
I. INTRODUCTION
Human evaluation is reliable but costly, subjective, and difficult to scale, motivating LLM-based judging. This study compares GPT-4.1 and GPT-5 across three context configurations to assess reliability and practical complementarity.
- Human evaluation captures interaction quality but is expensive, time-consuming, subjective, and difficult to scale.
- LLM judges offer a scalable alternative for assessing open-ended conversations across multiple quality dimensions.
- LLM-based judges can exhibit position bias, prompt sensitivity, and variability across evaluation settings.
- The study compares human evaluations with GPT-4.1 and GPT-5 across p0, p1, and p2 configurations.The configurations provide no user context, a static persona, or dynamically inferred context, respectively.
- The comparison targets reliability, consistency, practical applicability, and dimensions where automated judging can complement human assessment.
II. EVALUATION METHODOLOGY
The methodology compares human and LLM judgments of the same retail and telecom voice-agent conversations across multiple configurations and metrics. Scores are aggregated consistently to support direct comparison and subsequent divergence analysis.
- Human evaluators and LLM judges independently scored the same retail and telecom conversations across multiple metrics.
- Telecom evaluation used three trained human annotators per conversation, with high-disagreement cases adjudicated by a fourth senior annotator.
- The same mean-across-conversations aggregation procedure was applied to human and LLM scores, making reported ratios directly comparable.
- The evaluation set contained 242 conversations across six domain-configuration combinations: 120 Retail and 122 Telecom cases.
- Each case included ground truth, an actual voice conversation, and a text conversation for full-interaction comparison.
- Telecom results separated metrics with near-agreement, such as Turn Efficiency and User Experience Score, from metrics with sharper divergence.
B. Analysis for Retail sector conversations
Retail results distinguish relatively stable agreement on several operational metrics from pronounced divergence on safety and ASR-robust goal-achievement judgments.
- Turn Efficiency and Confirmation Recall remain close to a ratio of 1 across all three Retail configurations.
- Critical Field Accuracy is comparatively stable in Retail, ranging from 0.986 to 1.394.
- Safety-oriented metrics produce the largest evaluator gaps in the Retail results.
- Retail results indicate stricter human safety judgments and more lenient LLM scoring of ASR-robust goal achievement.
- Confirmation Precision and Error Rate emerge as the least stable metrics under automated evaluation.
C. Telecom vs. Retail
Telecom and Retail share directional patterns in several metric disagreements, but the magnitude of divergence differs substantially by domain. Cross-configuration and cross-model trends remain comparatively stable despite differing absolute scores.
- Turn Efficiency and Confirmation Recall stay near a ratio of 1, while ARGA remains below 1 in both domains.
- Safety divergence is sharper in Telecom, where IAS and SR exceed 4, than in Retail, where they range from 1.06 to 1.75.
- Critical Field Accuracy is tighter in Retail at 0.986–1.394 than in Telecom at 1.684–1.814.
- Confirmation Precision differs by domain, reaching 4.3 in Telecom under persona injection but dropping as low as 0.492 in Retail.
- Most metric-level relative behaviors remain consistent across p0, p1, and p2 and across GPT-4.1 and GPT-5.
B. Significant Divergence in Safety Metrics
Safety-oriented metrics show the largest human–LLM-judge divergences, with humans flagging substantially more safety concerns. Recovery Turn Count reveals a separate limitation: automated judges underestimate multi-turn recovery and require calibration for reliable use.
- Safety Metrics: Humans consistently scored Safety Recall (SR) and Irreversible Action Safety (IAS) several times higher than either LLM judge.The gap reflects substantially more safety concerns being flagged by human evaluators.
- Safety Metrics: Ambiguous escalation rubrics caused GPT-4.1 and GPT-5 to alternate between crediting safe human deferral and penalizing incomplete autonomous resolution.SIM-lock escalation cases exposed this inconsistency most clearly.
- Judge Differences: Neither GPT-4.1 nor GPT-5 consistently aligned best with humans; leadership shifted by metric according to their differing speed-, predictability-, and reasoning-oriented designs.GPT-5 moved closer to humans on goal-completion metrics such as CFA, while GPT-4.1 tracked some safety-adjacent judgments more closely.
- Recovery Metrics: Recovery Turn Count (RTC) was a major divergence because correct scoring requires tracing an error across turns and counting until resolution.Humans assigned substantially higher RTC values, while LLM judges often produced values below one.
- Recovery Metrics: RTC is unsuitable for fully automated evaluation without calibration, although its divergence varied by domain and prompting configuration.One GPT-5 no-persona retail configuration produced an estimate above the human value, showing the pattern was not absolute.
E. Baseline Differences Rather Than Structural Disagreement
Human and autoeval scores often preserve the same relative metric profile despite differing in magnitude, but this alignment depends strongly on domain. Retail supports per-metric linear calibration more clearly than Telecom, where weaker correlations may require rubric-level fixes.
- Relative Metric Profiles: Human and autoeval scores usually fall on the same side of the scale midpoint even when their exact values differ.TE and CR are high for both evaluators, while ARGA is low for humans and higher for autoeval judges.
- Relative Metric Profiles: Safety Recall (SR) and Irreversible Action Safety (IAS) remain high for humans but low for both LLM judges, especially in Telecom.In one comparison, autoeval scores were below 0.21 while human scores exceeded 0.85.
- Cross-Domain Correlation: Retail human–autoeval correlations were strong for GPT-4.1 (r = 0.912) and GPT-5 (r = 0.943), both significant at p < 0.001.The correlations indicate that relatively high and low metrics tend to track across evaluators despite magnitude differences.
- Cross-Domain Correlation: Telecom correlations were weaker: GPT-4.1 yielded r = 0.295 at N = 10, while GPT-5 yielded r = 0.592 at N = 10.GPT-4.1’s Telecom correlation was not statistically significant, and GPT-5 remained below its Retail correlation.
- Calibration: Per-metric linear calibration is better supported in Retail than Telecom, where rescaling alone may not close the gap.Telecom calibration may need rubric-level revisions, particularly for GPT-4.1.
G. ASR Error Analysis •
ASR errors concentrated on structured alphanumeric inputs, including names, numbers, email addresses, and URLs. These recognition failures triggered repeated prompts, recovery attempts, task-completion failures, and unnecessary human handoffs.
- Recognition Failures: ASR frequently misrecognized customer names, order IDs, phone numbers, email IDs, and URLs across Telecom and Retail.Structured alphanumeric information was the primary error concentration.
- Recognition Failures: Phone-number errors in Telecom often caused repeated prompts and increased transfers to human agents.Users were frequently misrecognized even when providing numbers in the requested xxx-xxx-xxxx format.
- Recognition Failures: Spacing-related transcription errors impaired accurate interpretation of user inputs.These errors contributed to broader failures in processing structured information.
- Downstream Effects: Email-ID misrecognition led to failures in information capture and task completion.The system sometimes failed to recover even when users spelled names clearly to improve recognition.
- Downstream Effects: Improving ASR accuracy for structured user inputs is identified as critical for enhancing voice-agent reliability and reducing failures.The reported downstream effects include increased recovery attempts and unnecessary human handoffs.
V. IMPLICATIONS
LLM judges can support large-scale voice-agent evaluation, but their reliability varies by metric, domain, and judge model. Human validation remains especially important for safety and recovery judgments, while calibration and judge selection should be tailored.
- LLM-as-Judge systems can support large-scale automated evaluation workflows.The paper positions them as scalable evaluation components rather than complete replacements for human assessment.
- Safety Recall and Irreversible Action Safety require human validation because they diverge substantially from human judgment.
- Recovery Turn Count remains unreliable in current LLM evaluators and requires improved modelling.
- Metric-level calibration may improve alignment because metric trends remain stable across evaluators.
- Retail supports simple per-metric linear calibration with r > 0.9 for both judges, whereas Telecom correlations are weaker and GPT-4.1 results are not statistically significant.Telecom may therefore require rubric-level revisions in addition to score rescaling.
- GPT-5 tracks humans more closely on goal-completion metrics such as CFA, while GPT-4.1 is closer on some safety-adjacent judgments.The findings favor metric-specific or ensemble judging over a single default judge model.
VI. CONCLUSION
Across configurations, GPT-4.1 and GPT-5 showed stable metric-level trends that often followed human judgments, although their absolute scores did not consistently match. The largest disagreements concerned safety and recovery, supporting a hybrid evaluation pipeline with metric- and domain-specific calibration.
- Across p0, p1, and p2, both LLM judges showed stable metric-level trends that often followed human judgments.Their relative assessments were more consistent than their absolute scores.
- LLM judges did not consistently match human absolute scores.
- The largest differences involved Safety Recall, Irreversible Action Safety, and Recovery Turn Count.LLM judges often missed contextual safety concerns and underestimated multi-turn recovery behavior.
- Neither GPT-4.1 nor GPT-5 was uniformly closer to human judgment across metrics.GPT-5 tracked humans more closely on some goal-completion metrics, while GPT-4.1 was closer on some safety-adjacent judgments.
- A hybrid pipeline is the most practical approach, combining scalable first-pass LLM evaluation with human review for safety-critical and recovery-focused decisions.Metric-specific calibration may further improve alignment, especially where human and autoeval scores track closely.
VII. FUTURE WORK
Future work will extend the benchmark beyond text transcripts to capture voice-specific timing and interaction phenomena. It will also test multimodal judging, broader calibration, and refined rubrics for safety and recovery.
- The benchmark does not model voice-specific temporal and interaction phenomena such as silence, interruptions, overtalk, or barge-in handling.These behaviors depend on acoustic and timing information unavailable from text transcripts alone.
- Future extensions will incorporate audio signals, speaker-turn boundaries, and timestamp data.Planned metrics include response-window adherence, interruption detection, overtalk frequency, and barge-in success.
- The framework will separate failures caused by ASR from failures caused by dialogue-management or response-generation policies.
- Future studies will test audio-aware or multimodal LLM judges across larger datasets, additional tasks, and more diverse speakers and acoustic conditions.
- Rubrics for safety-sensitive escalation and multi-turn recovery should be refined to identify which voice-specific metrics can be automated reliably.The goal is to determine which metrics still require human oversight.
APPENDIX
The appendix defines the benchmark’s conversational quality, safety, accuracy, and recovery metrics. Together, these measures quantify task correctness, clarification behavior, safety compliance, efficiency, user burden, and error recovery.
- Metric definitions: Critical Field Accuracy measures correctness on error-sensitive entities whose incorrect capture can invalidate task success.
- Metric definitions: Confirmation Precision measures how often requested clarifications or confirmations were actually necessary.Low precision indicates over-clarification when context was sufficient.
- Metric definitions: Confirmation Recall measures how often the agent requested clarifications or confirmations when they were required.Low recall indicates proceeding on ambiguous input without asking.
- Metric definitions: Safety Recall measures consistency in requesting confirmation for confirmation-required cases.
- Metric definitions: Irreversible Action Safety measures whether high-risk irreversible actions occurred only after explicit user confirmation.A value below 1.0 flags a critical safety failure.
- Metric definitions: Recovery Turn Count measures the average conversational turns needed to recover from ASR, tool, and agent-action errors.
- Metric definitions: Turn Efficiency compares optimal with actual turns, while User Experience Score counts repetitions, corrections, or restatements.Values closer to 1.0 indicate efficient resolution for Turn Efficiency; high User Experience Score signals poor experience.
- Metric definitions: Error Recovery Rate measures the proportion of detected ASR, tool, and agent-action errors successfully recovered through clarification, retry, or undo.