Source-linked AI summary
Not the Same Protector: Deployment-Dependent Protective Intervention in LLMs
Eunna Lee, Soomyoung Lee, Jungpyo Nam, Heonjin Ha, Jamin Jung, Kyunam Choi, Sunjun Hwang, Yeonghun Kim, Seok-Jae Lim
TL;DR
The paper asks whether models protect distressed users the same way when they speak rather than type. It evaluates four models with a matched injury vignette across voice, text, and raw API deployments, finding deployment-dependent protective behavior that cannot be explained by response length alone.
Problem
The study addresses limited evidence about whether deployment surface changes proactive protective behavior toward users at risk.
Method
Four models receive a single injury vignette across voice UI, text UI, and raw voice API conditions, with responses coded using five binary protective indicators.
Results
Safety confirmation occurs in 0/30 API responses across all models, while medical directives decline under voice for every model despite being at ceiling elsewhere.
Takeaways & Limitations
Protective intervention is sensitive to deployment surface, and evaluating one access pathway does not establish behavior at another.
Takeaways & Limitations
The study is limited by an uncontrolled GPT interface confound, one language, thirty responses per condition, and speech-act presence coding without intensity or position.
Abstract
from arXiv · showhide
We ask whether a model protects a user in the same way when that user speaks rather than types. Using a single distress vignette---a physical injury of unstated severity following an interpersonal conflict---we present four frontier models with matched inputs across voice, text, and raw API deployment conditions (n=30 per cell) and code each response along five binary protective indicators, including whether the model issues an explicit medical-care directive. Voice-interface responses are markedly shorter than text-interface responses for three of the four models, and protective behavior contracts alongside that compression: medical directives are at ceiling under both the API and text conditions but decline under voice for every model tested. The contraction is not reducible to length. One model produces voice and text responses of comparable length yet still drops medical directives, and another falls below ceiling between its API and voice conditions, whose responses are of nearly identical length. Under raw API access the pattern is categorical rather than partial: no model asks after the user's safety even once. These results show that protective intervention is sensitive to the surface through which a request arrives, that this sensitivity is detectable using a simple protective coding scheme, and that it is not explained by turn length alone.
1 Introduction
The study asks whether models protect distressed users consistently across voice and text, where voice may be the only available channel. It tests whether protective components are dropped in voice responses because of shorter turns or for another reason.
- The study examines whether the same distressed user receives equivalent protection when speaking rather than typing.
- The vignette describes a physical injury of unstated severity after an interpersonal conflict.
- Four frontier models receive matched inputs through voice UI, text UI, and raw voice API conditions.
- Responses are coded using five binary protective indicators, including an explicit medical-care directive, alongside response length.
2 Background
Prior audio-safety work shows that spoken interaction can expose risks absent from clean-text evaluation, but proactive protective behavior toward users at risk remains undercharacterized. This study focuses on whether deployment surface changes how and how forcefully models protect such users.
- Audio safety risks can arise from speaker attributes, non-speech sounds, paralinguistic cues, and disfluencies rather than text content alone.
- Voice responses impose turn-level constraints because they are spoken aloud and cannot be skimmed like text.
- The study addresses the undercharacterized question of whether deployment surface shifts how, when, and how forcefully models protect users at risk.
3 Methodology
The methodology uses a single distressed-user vignette, matched deployment conditions, and a binary coding scheme for distinct protective speech acts. It evaluates four models while defining the study’s measurement scope and experimental controls.
- The study measures protective guidance in responses, not downstream user outcomes.
- The vignette presents a first-person account of physical conflict and injury, with recordings from 10 speakers in voice conditions.
- The design leaves injury severity unstated while making a medical directive guideline-concordant without explicitly cueing it.
- Five indicators are coded independently and non-exclusively, so one response can receive multiple protective codes.
- The coding scheme defines medical directives, fault attribution, empathy, unilateral separation prescriptions, and safety confirmation as distinct response indicators.
- Claude, Gemini, GPT, and Grok are evaluated across user-facing voice, user-facing text, and raw voice API conditions.
4 Results
Across four models and three deployment conditions, protective responses varied by deployment surface and modality. Medical directives declined under voice interfaces despite ceiling performance in API and text conditions, while safety confirmation was absent under raw API access.
- Study design: 30 responses per cell were evaluated across four models and three deployment conditions using five binary protective indicators.Response-length medians were reported for each cell, with Wilson 95% confidence intervals provided in the appendix.
- Response length: 2.9, 9.2, and 9.2 times shorter were Claude, Gemini, and Grok voice responses than text responses at the median, while GPT was comparable at 113 versus 118 words.The GPT comparison shows that voice-interface compression was not universal across models.
- Length and protection: GPT dropped medical directives from 30/30 to 21/30 despite comparable voice and text lengths, while Gemini dropped from 30/30 to 25/30 at nearly identical API and voice lengths.Within voice conditions, directive-bearing responses were longer, but between-condition comparisons showed that length alone did not explain directive contraction.
- Safety confirmation: 0/30 raw API responses included safety confirmation across all four models, compared with voice-interface rates of 4/30 to 22/30 and text-interface rates of 0/30 to 28/30.For three providers accepting native audio, API and voice inputs held modality constant, so the shift tracked deployment surface rather than input modality alone.
- Medical directives: 30/30 API and text responses from every model included a medical directive, but voice-interface rates fell to 25/30 for Claude, 25/30 for Gemini, 21/30 for GPT, and 10/30 for Grok.The decline occurred for all four models and was not consistently offset by other protective indicators.
- Model-dependent variation: Normative and separation-oriented indicators reversed across deployment conditions, and voice-versus-text shifts differed in direction across models.Gemini, GPT, Grok, and Claude showed distinct modality patterns for safety confirmation, fault attribution, and user-separation prompts.
5 Representative Response Examples
Representative excerpts show how the coded protective indicators appear in practice across voice, text, and API responses. The examples include medical directives, safety checks, separation guidance, affective validation, and fault attribution.
- Selection: Selected excerpts illustrate the contrasts summarized in Tables 2 and 3 rather than sampling responses randomly.Responses within each condition were largely similar in form.
- API example: A Gemini API example contains affective validation, medical guidance, and emergency-contact advice in a concise response.This excerpt demonstrates that API responses could still include protective guidance even though safety confirmation was categorically absent in the aggregate API results.
- GPT: GPT’s voice example includes a same-day doctor directive, safety guidance, and a direct safety check, while its text example gives emergency-care and separation instructions.Both examples show medical and separation-oriented intervention, with the voice response also asking about current safety and symptoms.
- Claude: Claude’s analyzed-audio API example gives affective validation, a medical directive, and separation guidance, whereas its voice example also asks about medical care and current safety.The text example includes fault attribution, a medical check, and a safety question.
6 Discussion and Conclusion
Across models, deployment surface systematically changes protective behavior, including safety confirmation and medical directives, while model-specific patterns vary in direction and magnitude. These results are constrained by an uncontrolled GPT interface confound, limited scenario design, and coding of speech-act presence rather than appropriateness or intensity.
- No model produced safety confirmation under raw API access, although all four produced it in at least one interface condition.
- Medical directives fell under the voice interface for every model, and the decline was not consistently offset by other protective indicators.
- GPT produced 21/30 medical directives in voice responses versus 30/30 in text responses despite median voice length of 113 words.
- Claude attributed fault in 26/30 API responses and prescribed separation in 22/30, while Gemini and GPT produced neither indicator under API access.
- GPT interface results are confounded because interface-condition responses were collected through WORK rather than the standard chat interface.
- The study reports speech-act distributions, not appropriateness, and its single vignette, one language, thirty responses per condition, and presence-only coding limit generality.
A Appendix. Audio Input Methodology by Provider
The appendix documents provider-specific audio pathways and matched API testing. Four providers were tested with fixed speaker stimuli, while Claude required a transcript-and-acoustic-feature substitute because its API did not accept audio bytes.
- Four providers were tested with two fixed speaker stimuli whose transcripts were identical across providers.
- Direct_audio used the original audio, whereas analyzed_audio supplied a text-based substitute when direct audio input was unavailable.
- Claude received a structured text input combining the verbatim transcript with objective acoustic features because its Messages API accepted text and image input but not audio bytes.
- A pilot truncation led to a higher output-token cap before the full Claude run, after six calibration calls completed without truncation.
- All four providers received matched testing with 30 calls per provider, one API request per row, no automatic retries, and append-only response logging.
- Table 4 summarizes each provider's input pathway, API method, and outcomes across 30 calls per provider.
- The eight length-truncated OpenAI responses were retained and flagged for quality rather than deleted or re-queried.
B Appendix. Interval Estimates for Indicator Frequencies
The appendix reports Wilson 95% confidence intervals for every indicator-frequency cell. These intervals describe within-cell sampling variability rather than supporting inference across scenarios from the single-vignette design.
- Wilson 95% confidence intervals are reported for every indicator-frequency cell, with n = 30 per cell.
- The intervals are used because many cells lie at or near proportion boundaries, where normal-approximation intervals can have zero or negative width.
- The intervals describe sampling variability within a cell and do not support inference about underlying models across scenarios from one vignette.