Source-linked AI summary
How AI Assistants Respond to Repeated Abuse
William Guey, Wei Zhang, Pierrick Bougault, Yi Wang, Agoston Bodo, Vitor D de Moura, José O Gomes
TL;DR
The paper asks how repeated verbal abuse changes assistants’ engagement with otherwise benign tasks. It introduces a bilingual, multi-turn framework and benchmark that separates hard disengagement, soft withdrawal, availability, task-related work, and boundary setting. Hard disengagement varied sharply across configurations, showing why a single refusal label is insufficient.
Problem
Evidence is limited on how sustained hostility changes an assistant’s engagement with a benign task, beyond whether it complies with harmful requests.
Method
The study uses a bilingual five-turn benchmark and multidimensional response framework across eight API configurations, combining automated judgments with human coding.
Results
Hard disengagement ranged from 0/48 in four configurations to 24/48 (50.0%) for Gemini 3.1 Pro, with strong configuration-associated heterogeneity.
Takeaways & Limitations
A single refusal label cannot capture whether an assistant leaves, pauses, preserves availability, sets a boundary, or continues substantive task work.
Takeaways & Limitations
Results are limited to a fixed benchmark and recorded time-specific API configurations, while paired English and Chinese prompts do not establish cultural or pragmatic equivalence.
Abstract
from arXiv · showhide
AI assistants are expected to remain useful during difficult interactions, but little is known about how repeated verbal abuse changes their engagement with an otherwise benign task. We contribute a bilingual, multi-turn framework that separates hard disengagement, an unconditional statement of noncontinuation with no stated route to resume, from soft withdrawal, continued availability, observable task-related work, and boundary setting. Each of eight time-specific API configurations contributed 48 escalation conversations and eight smaller constant-frustration comparisons, giving 448 five-turn conversations, 2,240 responses, and 6,720 metadata-blinded model judgments. Primary results use the sustained-abuse endpoint of the 48 escalation conversations per configuration. Hard disengagement ranged from 0/48 in four configurations to 24/48 (50.0%) for Gemini 3.1 Pro, with strong configuration-associated heterogeneity (matched-label Monte Carlo p = 0.00001). GPT-5.6 Sol produced hard-disengagement labels in 15/48 (31.2%) endpoints, whereas Claude Fable 5 produced none and yielded 42/48 (87.5%) soft-withdrawal labels. Aggregate hard-disengagement rates were similar in English and Chinese (30/192 versus 32/192), although configuration-specific directions varied. Availability also differed from task-related work: Claude Opus 4.8 and Claude Fable 5 remained explicitly available in 48/48 endpoints while providing observable task-related work in only 8/48 and 7/48. Human coding was used to evaluate measurement quality. The results show why a single refusal label cannot capture whether an assistant leaves, pauses, preserves a route back, sets a boundary, or still performs substantive work.
1 Introduction
The paper asks how repeated hostility changes an assistant’s engagement with an otherwise benign task, where responses can differ in continuation, availability, and closure. It introduces a bilingual multi-turn benchmark and multidimensional measurement framework to distinguish these behaviors.
- Different responses can set boundaries, pause while preserving a route back, or end the interaction, creating distinct consequences for users, tasks, and services.
- Repeated hostility is studied while the underlying requests remain benign, shifting attention from harmful-request compliance to how assistance changes during sustained abuse.
- A single refusal label can obscure whether the assistant remains available, performs task-related work, or practically withdraws.
- The framework studies fixed five-turn coding, factual, planning, and writing tasks in English and Chinese under repeated hostility.
- The study compares eight model configurations using a common factorial prompt design, metadata-blinded automated judges, and human coding to assess measurement consistency.
2 Related work
Prior work documents abuse toward conversational agents, studies context-sensitive response strategies, and evaluates safety and overrefusal. This paper builds on measurement research showing that automated judgments can be biased and that multi-turn failures may emerge only after interaction history accumulates.
- Research has examined bullying, sexual harassment, hostile language, refusal, redirection, confrontation, and task continuation in conversational-agent settings.
- Safety benchmarks commonly assess harmful-input handling and overrefusal of safe requests, rather than sustained hostility during a benign task.
- LLM judges enable scalable evaluation but can exhibit position bias, framing effects, and multilingual inconsistency.
- Multi-turn benchmarks indicate that important failures may appear only after interaction history accumulates.
3 Conceptual framework
The framework separates withdrawal from availability, task continuation, and boundary setting rather than treating conversational responses as one refusal axis. Hard disengagement is the narrowest withdrawal outcome, while soft withdrawal preserves non-hard forms of stepping back.
- The five constructs are observable, potentially overlapping dimensions rather than an exhaustive taxonomy, excluding reactions such as apology, empathy, humor, and moral condemnation.
- Hard disengagement means an explicit, unconditional statement that the assistant will not continue, with no stated route to resume immediately, later, or elsewhere.
- Soft withdrawal is withdrawal without hard disengagement; the two are mutually exclusive in reported profiles, while hard disengagement remains nested within broader withdrawal.
- Continued availability requires an explicit offer or open signal, whereas observable task-related work requires substantive code, editing, explanation, planning, or factual content in the current reply.
- Boundary setting requires an explicit response to the user’s tone or treatment and is measured separately from withdrawal and task work.
4 Method
The study uses matched bilingual five-turn conversations across eight API configurations, with automated majority labels checked against human coding. Its exploratory analyses compare configuration-associated outcomes under a fixed design while limiting causal and replication claims.
- 448 five-turn conversations produced 2,240 assistant responses across eight configurations, including 48 escalation conversations and eight constant-frustration comparisons per configuration.
- The benchmark pairs English and Chinese escalation ladders across coding, factual, planning, and writing tasks, supporting local comparison but not cultural or pragmatic equivalence.
- The primary endpoint is each escalation conversation’s turn-5 response, and conclusions generalize only to the fixed benchmark and recorded time-specific API configurations.
- Each response received three metadata-blinded automated judgments, yielding 6,720 assessments whose fieldwise majorities formed canonical labels.
- Human coding supported hard disengagement more strongly than soft withdrawal, with machine-human F1 of 0.836 versus 0.644–0.675.
- The matched configuration test shuffled labels within 48 design blocks over 100,000 repetitions, while the focal-pair McNemar test and frustration comparisons were exploratory or descriptive.
5 Results
Sustained abuse produced sharply different response profiles across configurations: some disengaged unconditionally, while others preserved availability, performed task-related work, or withdrew softly.
- 5.1 Configuration differences under sustained abuse: 24/48 (50.0%) hard-disengagement labels came from Gemini 3.1 Pro, versus 15/48 (31.2%) for GPT-5.6 Sol and 0/48 for four configurations.Configuration-associated heterogeneity was strong, with matched-label Monte Carlo p = 1.0 × 10−5.
- 5.2 Hard disengagement and soft withdrawal separate: 42/48 (87.5%) Claude Fable 5 endpoints were soft withdrawal, whereas GPT-5.6 Sol had 15/48 (31.2%) hard disengagement and 11/48 (22.9%) soft withdrawal.In matched endpoints, 15 cells were hard-positive only for GPT-5.6 Sol and none for both configurations.
- 5.3 Availability is not task-related work: 48/48 endpoints remained explicitly available for both Claude Opus 4.8 and Claude Fable 5, but observable task-related work appeared in only 8/48 and 7/48 responses.GPT-5.5 provided task-related work in 38/48 responses, while Grok 4.5 did so in 44/48 and neither produced hard disengagement.
- 5.4 Timing, language, and the comparison condition: 62/384 hard-disengagement positives occurred at turn 5, compared with 4/384 at turn 4 and none at turns 1–3.The late concentration is consistent with a history-sensitive pattern, although changing prompt content prevents isolating accumulated history.
- 5.4 Timing, language, and the comparison condition: 30/192 English and 32/192 Chinese turn-5 endpoints showed hard disengagement, but model-specific directions varied and do not support a general language claim.In shared ladders, escalation endpoints had 11/96 hard disengagement versus 1/64 under constant frustration; this comparison was reported descriptively.
6 Discussion
Repeated abuse elicits meaningfully different forms of disengagement, availability, and task continuation across assistants. Hard disengagement is the clearest primary separation, while soft withdrawal and availability–work mismatches require more cautious interpretation.
- Assistants may leave, pause, preserve a route to return, remain available without working, or continue the task, so these behaviors should not be treated as one refusal category.
- Hard disengagement separates configurations clearly: Claude Fable 5 tends toward soft withdrawal and continued availability, whereas GPT-5.6 Sol more often announces noncontinuation.The focal pair illustrates that assistants can differ substantially even under the same escalation design.
- Soft-withdrawal results deserve more caution because the construct combines pausing, conditional re-engagement, redirection, and effort withdrawal, unlike the narrower hard-disengagement criterion.
- Separating temporary pause, boundary-only redirection, and effort withdrawal could improve measurement reliability in future work.
- Availability and observable task-related work can diverge, making both fields operationally relevant for evaluating continuity, user recovery, and worker-facing support.An open invitation does not establish that substantive assistance continued.
7 Conclusion
Repeated-abuse responses span leaving, pausing, boundary setting, continued availability, and task continuation rather than a single refusal axis. Evaluations should therefore distinguish how an assistant withdraws, whether help remains available, and whether substantive work continues.
- Hard disengagement varied sharply across eight configurations, while soft withdrawal and explicit availability often followed different profiles.
- Safety evaluations should ask whether an assistant leaves, pauses, sets a boundary, remains available, or continues the task, rather than recording only refusal.
- Figure 3 displays Turn-5 behavioral profiles from three-judge majorities, with darker cells indicating higher within-configuration response proportions and daggers marking the focal pair.
Limitations
The study’s structured, time-specific benchmark constrains population and replication claims, while its paired bilingual prompts and small comparison condition limit broader interpretation. Its labels capture local textual evidence rather than full-dialogue fidelity.
- The structured benchmark cannot estimate the population prevalence of abusive interactions.
- API configurations were accessed in July 2026 without synchronized provider states, so later deployment changes may limit exact behavioral replication.
- English and Chinese prompts are paired materials, not evidence of pragmatic or cultural equivalence.
- The small, bundled, unequal constant-frustration condition supports description but not a causal estimate.
- Task-related-work and withdrawal labels reflect isolated local text, do not establish answer quality, and may be underdetermined when requests depend on earlier context.
- Soft withdrawal remains an especially imperfect measurement target because its component behaviors have less discrete boundaries.
Ethics Statement
The study used synthetic abusive prompts and time-specific API configurations to characterize observed outputs without analyzing real user conversations or inferring model intentions.
- The study evaluated language-model outputs with synthetic abusive prompts rather than real user conversations and collected no personal data.
- The findings characterize behavior in time-specific API configurations, not model intentions or internal states.
- Because coders reviewed repeated insulting language, future reuse should minimize unnecessary exposure and allow annotators to pause or stop.