Source-linked AI summary

Auditing Bias and Safety in Voice AI Customer Care

Vignesh Ethiraj, Ashwath David

arXiv:2609.04206v1eess.AScs.CLcs.CYcs.HCcs.SD

TL;DR

Voice AI customer care evaluations rarely treat agents as stateful, multi-turn, tool-mediated systems where burden can arise before final service denial. This paper formalizes a validation-gated audit framework using matched caller conditions, architecture-specific evidence, and outcome-plus-burden measures. It presents a synthetic refund-dispute example while excluding production-system comparisons and gating public reporting on validation.

  • Problem

    Existing fairness and safety evaluations do not usually cover customer-care voice agents as stateful, multi-turn, tool-mediated systems where harm may appear as additional service burden.

  • Method

    The framework separates three architectures, matches service facts across controlled caller presentations, validates stimuli and artifacts before inference, and measures outcomes, burden, and safety events.

  • Results

    The paper defines seven validation gates, a six-family metric set, claim boundaries, and a synthetic refund-dispute worked example rather than production-system results.

  • Takeaways & Limitations

    Bias findings require validated stimuli, metrics, artifacts, and analyses, making public empirical claims contingent on an auditable evidence chain.

  • Takeaways & Limitations

    The framework’s generalization is limited by rater-pool judgments, incomplete acoustic coverage, blended real-world architectures, construct drift, and rapidly changing systems requiring revalidation.

Abstract

from arXiv · show

Voice AI systems increasingly mediate customer care interactions where caller presentation cues such as accent, affect, fluency, and urgency are available alongside the service request. Existing fairness and safety evaluations cover speech recognition disparities, spoken dialogue bias, and voice agent capability, but rarely treat customer care voice agents as stateful, multi turn, tool mediated systems where harm can appear as additional burden before any final denial occurs. We formalize a validation gated audit framework for such systems. The framework (i) separates native speech to speech, cascaded ASR to language model to TTS, and hybrid tool mediated architectures; (ii) uses matched service facts across controlled caller presentation conditions; (iii) validates fact invariance, presentation cues, artifacts, and acoustic measurements before inference; and (iv) records both material outcomes and path to service burden. We define the research problem, methodology, seven validation gates, a six family metric set, and claim boundaries for an active industry evaluation program. We illustrate the framework with a fully synthetic worked example of a refund dispute audit instance. Production system results are excluded from this release; public reporting is gated by the validation protocol.

1 Introduction

Voice AI customer care requires auditing complete, consequential service episodes rather than isolated utterances or transcripts. The framework contributes architecture-specific risks, validation gates, and claim boundaries for synthetic, auditable evaluation.

  • Customer care voice agents are multi-turn, stateful, and tool-mediated, so avoidable burden can matter even when the final answer is formally acceptable.
  • The audit treats the service episode, rather than an utterance or transcript, as the unit of inferential evidence.
  • The framework separates native speech-to-speech, cascaded ASR-to-LM-to-TTS, and hybrid tool-mediated agents with distinct audit requirements.
  • Seven validation gates must pass before an audit instance can support an inferential claim.The gates cover scenario lock, fact invariance, persona perceptibility, acoustic validity, codec normalization, artifact completeness, and annotation reliability.
  • A claim-boundary matrix and staged evaluation program distinguish methods contributions, architecture comparisons, and reproducibility releases using a synthetic refund-dispute example.

2 Problem Formulation

The problem formulation compares matched caller conditions with identical service facts after validation gates pass. It evaluates both service outcomes and episode-level burden or safety differences across an auditable multi-turn trace.

  • A service episode contains caller audio, conversational state, tool calls, output speech, and a terminal outcome while material account and policy facts remain fixed.
  • The audit asks whether matched caller conditions differ in service outcomes and path-burden measures after validation gates pass.
  • The metric families separately cover terminal outcome, repair turns, authentication friction, escalation delay, condescension, tool-use parity, and acoustic accommodation.
  • Safety failures are episode-level adverse events, while bias claims compare matched conditions in their rates or severity.
  • The safety amplification statistic compares category-specific event probabilities between matched caller conditions.
  • Figure 1 records caller and agent turns alongside tools, audit artifacts, repair loops, escalation paths, and outcomes.

3 Related Work and Gap

Prior work establishes uneven voice-system behavior and evaluates dialogue or voice-agent capabilities, but customer-care audits require matched facts and explicit evidence gates. The paper frames its contribution as claim governance rather than a larger benchmark.

  • ASR research reports uneven performance across speaker groups, motivating matched audio audits beyond transcript accuracy alone.
  • Dialogue and spoken-LLM evaluations have expanded toward fairness, toxicity, and trustworthiness in interactive language systems.
  • Voice-agent benchmarks measure capabilities such as multi-turn interaction, tool execution, full-duplex behavior, and task outcomes.
  • These benchmarks are usually capability infrastructure rather than regulated service audits with matched facts and explicit evidence gates.
  • The remaining gap is governance over which customer-care episodes support empirical claims and which interpretations remain excluded until validation passes.

4 Framework

The framework treats each matched service episode as a claim-bearing audit instance that becomes inferential evidence only after validation gates pass. It locks service facts and caller conditions while recording architecture, artifacts, and supported claim boundaries.

  • Each audit instance begins as a claim-bearing candidate and becomes inferential evidence only after scenario, stimulus, validation, artifact, metric, coding, and analysis gates pass.
  • Audit Objects: An audit instance records a locked scenario, matched caller conditions, architecture, artifact log, and claim template specifying the supported metric comparison.
  • Scenario Design: The reference refund or billing dispute scenario fixes eligibility, account status, support history, refund facts, policy exceptions, and acceptable outcomes before data collection.
  • Validation: Presentation manipulations require independent perceptibility validation, while acoustic metrics require agreement among multiple trackers within each caller-condition cell.
  • Interpretation: A response difference supports a bias claim only when equal facts, eligibility, and tool access are verified and the difference appears as added burden, reduced service quality, or a safety event.

5 Architecture Specific Audit Model

The audit model distinguishes native, cascaded, and hybrid voice-agent architectures because their observable interfaces and failure pathways support different causal interpretations. Architecture-specific traces therefore determine which explanations an audit can defend.

  • Cascaded, native speech-to-speech, and hybrid tool-mediated agents expose different audit surfaces and observable artifacts.
  • Cascaded agents can fail through ASR, language-model reasoning, tool orchestration, or TTS, while native systems have less inspectable latent pathways involving voice presentation and timing.
  • In cascaded systems, denial after ASR mistranscription points toward an upstream recognition pathway.
  • In native speech-to-speech systems, shorter or less helpful responses despite correct facts point toward interaction-policy or latent acoustic-conditioning pathways.
  • In hybrid agents, additional verification requests under identical account facts point toward tool-mediated service burden.

6 Validation and Metrics

The framework gates inferential use on validated measurement and reports outcomes, path burden, and spoken behavior separately. Its paired scenario design estimates preregistered condition differences while preserving tool-call and service-path detail.

  • Validation Gates: Validation gates prevent measurement-pipeline artifacts from being presented as model behavior, with pilot thresholds documented before the analysis plan is frozen.
  • Metrics: The metric set separately reports final outcomes, path burden, and spoken behavior by matched condition and scenario.
  • Metrics: Path burden is B(e) = (Repair(e), Auth(e), Esc(e)), with matched pair delta ∆B reported componentwise.
  • Metrics: Tool-call parity requires equal tool availability, required sequence, returned record keys, and success state; deviations are reported even when the final outcome is unchanged.
  • Statistical Analysis: The paired scenario design estimates preregistered metric-family deltas, using paired risk differences for binary outcomes and paired mean deltas for continuous burden metrics.

7 Claim Boundaries and Responsible Use

The paper limits its contribution to accountable audit methodology and responsible release, not comparative claims about current systems. Public artifacts preserve scientific review while withholding operational details that could enable unauthorized testing.

  • Responsible use: The framework is designed for accountable audit governance, not adversarial probing of deployed customer care systems.Public reporting withholds operational details that could enable nuisance traffic or targeted pressure against live endpoints.
  • Claim boundaries: The methodology paper distinguishes supported claims from those requiring additional empirical evidence.Its claim boundary matrix preserves the paper’s status as a methods contribution rather than a comparative evaluation.
  • Responsible release: Responsible release uses synthetic records, illustrative policies, and redacted logs, with detailed traces released only under owner authorization.

8 Illustrative Worked Example (Synthetic)

The synthetic refund-dispute example demonstrates the validation workflow using matched caller conditions and fabricated audit traces. Although both calls receive refunds, one condition incurs additional service burden; the instance supports construct validation only.

  • Synthetic setup: The worked example uses fabricated names, accounts, transcripts, and numbers, so it makes no production or empirical claim.
  • Scenario lock: The locked scenario fixes account history, a duplicate $24.99 charge, and a policy permitting one-click refunds within 30 days.Allowed outcomes are refund issuance or human escalation with refund intent flagged.
  • Matched conditions: The matched caller cells hold linguistic content constant while contrasting neutral General American speech with a strong second-language English accent.
  • Validation gates: The illustrative gates pass: raters identify the accent condition at 93% accuracy, while F0 agreement is r = 0.91 across cells.Audio normalization and synchronized audio, transcript, and tool-call artifacts also pass their gates; annotation reliability is κ = 0.71 for condescension and κ = 0.78 for helpfulness.
  • Illustrative observation: Both calls end with a refund, but condition cb receives two additional repair turns, one extra authentication challenge, and one extra CRM read.Under the stated sign convention, negative path-burden values mean cb carried more burden than ca.
  • Claim disposition: The instance supports only construct validation of interpretable matched-pair deltas, not vendor comparison or population-level disparity claims.Those broader claims require powered samples, multiple architectures, and replication across scenarios.

9 Evaluation Program

The evaluation program prioritizes validated, reproducible audit design across voice architectures before reporting group-level or architecture-comparison claims. It proceeds through staged validation, comparison, and reproducibility releases.

  • Evaluation scope: The methods layer covers cascaded ASR-to-language-model-to-TTS, native speech-to-speech, and hybrid tool-mediated systems.Public reporting prioritizes locked scenarios, focused presentation contrasts, cascaded baselines, and sufficient calls to test logging and paired-delta analysis.
  • Three-stage program: Stage 1 validates stimuli, audio normalization, artifact logging, and metric extraction before any group-level claim.
  • Three-stage program: Stage 2 reports architecture comparisons only when comparable artifacts and predefined analysis rules exist.Stage 3 publishes scenario templates, schemas, metric code, and redacted examples while excluding live-system operational details.

10 Ethics Statement

The work is a methods paper based on synthetic or previously published audio rather than human-subject data or deployed-system interactions. Future live use requires ethics review, owner authorization, and safeguards against adversarial reuse.

  • Study setting: This methods paper collects no human-subject data and does not interact with deployed customer care systems.Its audio examples are cited from prior datasets or synthetic models used for illustration.
  • Safeguards: Future empirical use requires ethics review for human-rater protocols and explicit authorization from the audited system owner before live testing.The paper also withholds endpoint details, proprietary policy text, and raw audio because audit artifacts could be repurposed for adversarial probing.

11 Reproducibility and Artifact Release

The framework pairs a staged artifact release with explicit evidence boundaries and revalidation requirements. It also limits interpretation because the taxonomy abstracts blended systems and the release reports no deployment comparisons or effect sizes.

  • Artifact release: Stage 3 will release scenario schemas, a locked refund scenario, an audit-instance JSON schema, metric and gate code, an annotation codebook, and synthetic episodes.Raw audio, prompts, vendor identifiers, and customer-derived data are excluded.
  • Limitations: Persona judgments depend on the rater pool, while a small tracker set may underweight prosodic cues beyond F0 and timing.Cross-cultural panels and broader acoustic coverage are needed before generalizing claims across listening populations.
  • Limitations: The native, cascaded, and hybrid taxonomy abstracts systems that may blend pathways, so causal interpretations should respect that boundary.The framework is therefore a useful abstraction rather than a complete description of every production architecture.
  • Limitations: Scenario locks prevent in-flight edits but do not eliminate construct drift across versions, requiring replication for longitudinal claims.Framework releases should also be revalidated against current systems because voice agents, APIs, tool use, and architectures change quickly.
  • Evidence boundary: The framework requires validated stimuli, acoustic extraction, locked scenarios, and reproducible audit logs before empirical claims.This evidence boundary supports the report’s validation-first approach to public reporting.
Loading 2609.04206v1…