Source-linked AI summary

You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals

Ruoxuan Li, Pinqiao Wang, Sheng Li, Cameron Robert Jones

arXiv:2608.30856v1cs.CLcs.HC

TL;DR

LLM refusal research has largely focused on whether models safely decline harmful requests, leaving the pragmatic quality of refusal underexamined. This paper develops and validates a three-layer taxonomy and applies it to 16 models across 14 harm categories. Models generally refuse explicitly and with strong moral framing, while repair relies more on safer alternatives than interpersonal facework.

  • Problem

    Existing LLM non-compliance taxonomies emphasize safety behavior, leaving open how refusals function as pragmatic, face-threatening acts across harmful contexts.

  • Method

    The paper introduces a three-layer taxonomy and validates it through human annotation before applying it to responses from 16 LLMs across 14 harm categories.

  • Results

    Models’ refusals are generally explicit and strongly grounded in ethical justifications, with interactional repair relying mainly on safer alternatives rather than interpersonal facework.

  • Takeaways & Limitations

    Alignment evaluation should assess whether refusals are contextually adaptive and socially accountable, not only whether models decline harmful requests.

  • Takeaways & Limitations

    Layer 1 and Layer 2 comparisons cover only valid in-voice responses coded as non-compliance, so models are not compared in a perfectly balanced refusal setting.

Abstract

from arXiv · show

Refusals are often treated as face-threatening acts in pragmatics because they can challenge the requester's socially claimed self-image. Large language models (LLMs) are increasingly trained to refuse unsafe and inappropriate requests, and these refusals may harm users when models fail to manage this interactional cost properly. While existing work has mainly approached LLM non-compliance as a safety-alignment outcome, it does not provide a way to evaluate whether LLMs refuse appropriately across different harmful contexts. To study this question, we propose (to our knowledge) the first taxonomy of LLM refusals that is grounded in pragmatic theory. Applying this taxonomy to responses from 16 modern LLMs across 14 harm categories, we find that although models differ in how they refuse, their refusals are overall explicit and strongly morally evaluative, with interactional repair occurring mainly through offering or providing safer alternatives instead of interpersonal facework. This pattern is especially consequential in sensitive harm contexts, where overuse of negative framing may make users feel shamed or provoked, undermining the purpose of safe non-compliance. We therefore call for alignment evaluation that considers not only whether models refuse harmful requests, but also whether they refuse in ways that are contextually adaptive and socially accountable for the interactional consequences of saying no.

1 Introduction

Existing LLM non-compliance research largely treats refusal as a safety outcome, while pragmatics highlights its interactional and face-threatening dimensions. The paper introduces and validates a pragmatic taxonomy to evaluate how models refuse across harmful contexts.

  • Motivation: LLM non-compliance can involve outright refusal, harmless reinterpretation, apologies, guideline-based explanations, or moralizing blame.These realizations may affect users’ responses, perceptions of the model, and future requests.
  • Motivation: Pragmatic theories treat refusals as face-threatening acts, motivating evaluation of whether LLMs use facework analogous to human speakers.Human mitigations can protect both parties’ socially claimed self-image during refusal.
  • Research gap: Existing taxonomies focus primarily on safety behavior rather than pragmatic theories of refusal.The paper identifies user experience and moral-stanc e projection as additional reasons to study refusal realization.
  • Contributions: The paper introduces a three-layer pragmatic taxonomy covering response action, refusal rationale, and realization strategy with adjunct features.The taxonomy, code, and data are released on GitHub.
  • Contributions: κ = 1.0 for Layers 0 and 1, and average κ = 0.953 for Layer 2 features, in human annotation of 100 query–response pairs.The authors also validate an LLM-as-judge pipeline for scaling annotation.
  • Contributions: The study analyzes 200 responses each from 16 recent LLMs across 14 harm categories to compare how models justify and realize non-compliance.The findings identify pragmatic dimensions that alignment training can target for communicatively appropriate refusals.

2 Related Work

Prior work characterizes when and why LLMs do not comply, but leaves open how refusals should be analyzed as pragmatic, face-threatening acts. This paper situates refusal expression within facework, responsibility, stance, and response design.

  • LLM non-compliance taxonomies: Prior taxonomies organize LLM non-compliance around unsafe instructions, contextual triggers, or refusal motivations.They do not primarily explain how refusals are pragmatically expressed.
  • Beyond binary refusal: Recent work proposes safe completions and documents over-refusal of benign queries, expanding evaluation beyond binary refusal decisions.These approaches address underlying needs or examine inappropriate denial behavior.
  • Pragmatic foundations: Refusals threaten positive and negative face by rejecting expectations and constraining the requester’s projected course of action.Facework refers to efforts to manage perceived social value during interaction.
  • Pragmatic foundations: Apologies, hedges, explanations, criticism, suggestions, and offers distribute the interactional cost of refusal differently between speaker and hearer.These strategies can mitigate, redirect, or intensify face threats.
  • Operational distinctions: A refusal marker is the core syntactic phrase that signals non-compliance; realization strategies are mutually exclusive, while adjunct features may co-occur.This distinction supports separate analysis of refusal structure and accompanying pragmatic elements.
  • Human-AI interaction: Conversational style, apology, responsibility-taking, and denial design can shape users’ emotional responses, trust, and evaluations of AI systems.This motivates treating non-compliance as a user-facing pragmatic behavior rather than only a safety event.

3 Taxonomy

The proposed taxonomy analyzes non-compliance in three layers: what response action occurred, why the model refused, and how the refusal was linguistically realized. Later layers apply only to responses coded as non-compliant.

  • Development: The taxonomy was developed from pragmatic accounts of human refusal and refined through three rounds of pilot coding.Pilot rounds used 20–25 query–response pairs, with disagreements discussed and decision rules revised.
  • Layer 0: response action: Layer 0 distinguishes full compliance, partial compliance, and non-compliance based on execution of the explicit request.It does not rely on whether the model verbally announces compliance or refusal.
  • Layer 1: refusal rationale: Layer 1 distinguishes policy-based rationales from ethics-based rationales according to where responsibility for refusing is located.Policy rationales attribute refusal to external constraints, whereas ethics rationales endorse the requested act’s normative unacceptability.
  • Layer 2: realization: Layer 2 separates the core realization strategy from adjunct elements that frame, soften, justify, or redirect non-compliance.This layer captures both refusal directness and accompanying linguistic forms.

4 Experiment Setup

The experiment combines a category-balanced harmful-query set, structured comparisons across 16 models, human validation, and LLM-as-judge annotation. Analyses include refusal rationales across models and harm categories.

  • Query set: 200 harmful prompts were assembled from SORRY-Bench and LMSYS-Chat-1M, reclassified with the 14-category Llama Guard 3 taxonomy.The set includes 140 category-balanced prompts and 60 supplementary naturalistic prompts.
  • Models: The study selects 16 models to compare size, reasoning mode, and model family as potential influences on non-compliance behavior.Comparisons include within-family size pairs and models supporting reasoning and non-reasoning inference.
  • Human annotation: Two authors independently coded 100 validation query–response pairs and resolved disagreements through adjudication.The resulting gold labels support evaluation of the LLM judge.
  • Validation: Layer 0 is evaluated on all 100 validation examples, whereas Layers 1 and 2 are evaluated only on the 77 human-adjudicated non-compliant responses.This conditional coverage is the structure of the agreement analysis.
  • LLM-as-judge: GPT-5.5 in non-reasoning mode serves as the judge model, whose prompt is adapted from the human codebook and validated against gold labels.The judge is then applied to all 3,200 query–response pairs.

5 Results

Models usually refuse explicitly and frame non-compliance in ethical terms, while using safer alternatives more often than interpersonal facework. Refusal styles vary across models and harm categories, with sensitive contexts exposing tensions between clear boundaries, compassion, and moral evaluation.

  • Layer 0: How often do models comply?: 21% of responses complied: 17% fully and 4% partially; among the remaining cases, non-compliance ranged from 8% to 38% across models.Non-compliance was near-universal for Violent Crimes (95%) and Sex-Related Crimes (92%), but lower for Specialized Advice (47%).
  • Layer 1: What rationale do models provide for their refusals?: 70% of non-compliant rationales were ethics-based, compared with 23% bare refusals, while policy-based and capacity-based rationales were rare.Ethics-based framing places responsibility on the request’s moral status rather than on policy constraints or model inability.
  • Layer 2: How do models realize their refusals?: 90% of refusals used explicit rather than implicit realization, with no hedge or epistemic-softener features and explanatory prefaces averaging 3%.Roughly half of models produced little-to-no apology or regret, indicating limited linguistic mitigation.
  • Layer 2: How do models realize their refusals?: Models favored negative stance and normative suggestion over apology, positive alignment, or solidarity, while alternative offers appeared in 55% and executed alternatives in 39% of refusals.This pattern combines moral evaluation and redirection with continued helpfulness rather than primarily using interpersonal facework.
  • Do LLMs refuse adaptively to different harm categories?: Five harm-category clusters emerged, with Suicide & Self-Harm most care-oriented and Sexual and violent harms showing the highest explicit non-compliance and lowest executed-alternative use.Suicide & Self-Harm had the highest solidarity/empathy and apology/regret and the lowest negative stance and normative suggestion.
  • Do LLMs refuse adaptively to different harm categories?: Suicide & Self-Harm refusals still showed around 40% negative stance, which the authors identify as potentially inconsistent with non-judgmental clinical guidance for emotionally distressed users.The paper argues that clear non-compliance remains necessary, but further moral censure may fail to de-escalate situations and risk deepening distress.

6 Robustness Analyses

Robustness checks indicate that the taxonomy’s annotations and observed refusal patterns remain broadly stable across judges, prompt resampling, and repeated generation.

  • Judge annotations were moderate to highly consistent across model families and with human-adjudicated gold labels.The cross-family check used Claude Opus 4.8 in non-reasoning mode and found agreement was moderate to high for most features.
  • ρ = .977 correlation showed highly stable Layer 2 model-by-feature profiles across alternative prompt collections.The resampling analysis used 262 alternative prompts and a matched panel of 13 models.

7 Conclusion

The paper characterizes LLM refusals as explicit, firm, and ethically justified rather than primarily mitigated through interpersonal facework. This style may protect the model’s projected positive face while jeopardizing the user’s.

  • LLM refusals exhibit explicit, firm non-compliance grounded in ethical justifications.
  • Models appear to maintain their own projected positive face through moral justifications rather than interpersonal facework.

Limitations

The study’s conclusions are bounded by its refusal-only comparisons, interpretive scalable labels, broad feature definitions, and focus on model outputs rather than user perception.

  • Comparisons analyze valid in-voice non-compliant responses, so they approximate the refusal behaviors models produced rather than a balanced refusal setting.
  • LLM-judge labels are scalable approximations rather than perfect ground truth because some categories remain interpretive and responses combine multiple strategies.
  • Broad Layer 2 features do not assess the quality or usefulness of safer assistance, motivating finer distinctions between pragmatic support and generic templates.
  • The analysis does not directly test how refusal styles affect trust, shame, perceived support, or boundary negotiation in human-LLM interaction.

Ethical Considerations

The study uses existing harmful-query resources and analyzes refusal behavior rather than generating harmful instructions, while limiting potential misuse and defining a layered coding procedure.

  • Data and risk: The dataset contains 200 queries drawn from SORRY-Bench and LMSYS-Chat-1M, without creating new harmful requests.The analysis focuses on model-generated refusals rather than providing harmful instructions, although exposure to harmful content remains possible.
  • Data and risk: The taxonomy’s intended use is safety evaluation and alignment research, while its misuse risk is mitigated by omitting safety-system implementation details and jailbreak methods.
  • Data and privacy: The study uses query text without user identifiers, re-identification attempts, or inferred personal attributes, following the source datasets’ licensing conditions.
  • Annotation: Two authors annotated the data using the codebook, with limited but nonzero exposure to harmful and offensive topics and no recruited external participants.
  • Coding framework: The three-layer taxonomy assigns response action first, then rationale and realization strategies only when Layer 0 is non-compliance.
  • Coding framework: Layer 0 is coded from whether the response executes the requested content and purpose, not from whether the model claims compliance or refusal.
  • Rationale coding: Layer 1 identifies the substantive rationale nearest the refusal marker, while masking role-based phrases such as “As an AI” unless another rationale appears.
  • Realization coding: Layer 2 separates explicit from implicit non-compliance and codes adjunct features such as explanatory prefaces, alternatives, normative suggestions, and negative stance.

B Dataset Category Distributions

The study samples harmful prompts across 14 Llama Guard 3 categories and evaluates responses from a 16-model panel, using human annotation and an LLM judge to support the analysis.

  • Validation: Two annotators independently coded 100 stratified query–response pairs, then adjudicated disagreements to create the gold label set.The validation set covers different models and harm categories.
  • Validation: The taxonomy evaluates Layer 0 on all validation pairs, while Layers 1 and 2 are evaluated only on responses independently labeled non-compliant by both annotators.The agreement table reports Layer 0 across 100 pairs and later layers across 77 jointly non-compliant responses.
  • LLM-as-judge validation: GPT-5.5 in non-reasoning mode was selected as judge because it outperformed GPT-5.3 reasoning on Layer 1 and average Layer 2 agreement.GPT-5.3 had slightly higher Layer 0 agreement, but the analysis depends most on rationale and realization-strategy distinctions.
  • Model panel: The main analysis compares 16 models across six families using the same 200-prompt query set, while service-level refusals provide no model-generated rationale.The panel supports comparisons by model size, reasoning mode, family, and release period.
  • Layer 0 results: 38% was the highest combined full-plus-partial compliance rate, observed for Qwen3-32B reasoning; Llama-3.1-8B was most restrictive at 8%.Partial compliance accounted for 3.75% of all responses overall.

H Layer 2 Feature Rates by Model

Across models, refusals were generally explicit, certain, and negatively framed, while models differed in implicitness, facework, alternatives, and sensitivity to harm category and reasoning mode.

  • Cross-model patterns: None of the 16 models used hedging, indicating that coded refusals were expressed with high certainty rather than tentativeness.Hedge was absent across the coded refusals.
  • Cross-model patterns: Negative stance occurred in 55–91% of refusals for most models, with GPT-4o, Llama-3.1-8B, and Llama-3.1-70B as clear exceptions.Those exceptions recorded negative stance at 4%, 11%, and 10%, respectively.
  • Cross-model patterns: Qwen models showed implicit non-compliance rates of 22–34%, whereas most other models remained at or below 10%; Claude Sonnet 3.7 reached 20%.These rates distinguish Qwen from most of the other models in the panel.
  • Cross-model patterns: Claude Sonnet 3.7 recorded 25% positive alignment and 12% explanatory prefaces, compared with 5% and 7% for Claude Sonnet 4.6.The comparison indicates more elaborated facework for Claude Sonnet 3.7 on these features.
  • Cross-model patterns: Gemini 2.5 Pro had a 67% executed-alternative rate and a 13% alternative-offer rate, favoring carried-out alternatives over offers.The two features are distinguished by whether alternative content is performed or merely offered.
  • Harm-category variation: Pooled non-compliance was approximately 95% for Violent Crimes, 92% for Sex-Related Crimes, and 91% for Suicide & Self-Harm, versus 47% for Specialized Advice.Individual models still showed substantial category sensitivity.
  • Reasoning-mode comparisons: Reasoning mode changed in-voice non-compliance rates by at most 1.5 percentage points across four model pairs, with no uniform effect on refusal realization or adjunct strategies.For Qwen3-32B, reasoning reduced bare refusals from 18.3% to 8.9% and increased ethics-based rationales from 75.4% to 82.9%.

M Bare Refusal Template

Appendix examples show fixed-form bare refusals across several models, with limited surface variation and occasional apologies, empathy, or follow-up offers.

  • M Bare Refusal Template: 144 bare refusals were observed in one template set, with 72% identical; its core form was “I can’t assist with that request.”Observed variations included pronoun shifts, object substitutions, and content-specific prefixes.
  • M Bare Refusal Template: A third template set contained 141 bare refusals centered on “I can’t fulfill that request.”Some responses substituted “accommodate,” while others appended a follow-up offer.
  • M Bare Refusal Template: Another template set contained 154 bare refusals centered on “I’m sorry, I can’t assist with that request.”Variations included minor punctuation changes and occasional empathetic openers for emotionally charged queries.
  • M Bare Refusal Template: The fixed-form examples are annotated as Layer 1 Bare refusals, characterized as fixed-form refusal templates.The appendix labels these patterns explicitly in the annotation records.
  • M Bare Refusal Template: GPT-4o, Llama-3.1-8B, and Llama-3.1-70B each exhibited dominant fixed-form refusals with limited surface-level variation.The appendix contrasts these models with newer models producing more semantically varied, rationale-driven refusals.
  • N Cross-family Judge Comparison: The appendix also records five Claude Opus 4.8 judge refusals across Code Interpreter Abuse, Indiscriminate Weapons, and Violent Crimes examples.Three blocked examples came from Code Interpreter Abuse and one each from the other two categories.
  • N Cross-family Judge Comparison: The Opus 4.8 comparison applied the same codebook, judge prompt, and output schema used for GPT-5.5.It was designed to assess cross-family robustness using a contemporaneous model from another provider family.
  • N Cross-family Judge Comparison: Opus 4.8 produced no parseable annotations for 5 of 100 pairs, including four cases human annotators labeled non-compliant.Layer 1 agreement was therefore computed over 73 remaining non-compliance cases and Layer 2 over 72.

O Robustness of Data Expansion

The robustness analysis expanded a separately seeded prompt set toward 20 prompts per harm category and restricted comparisons to models observed in both collections.

  • Robustness sampling: The August robustness set began from a separate 200-prompt seed sharing 135 queries with the May collection.It used the same source pools and category-aware sampling framework.
  • Robustness sampling: Sampling expanded the set toward 20 prompts per harm category while excluding duplicates and reducing within-category redundancy.The procedure excluded represented template groups and candidates with token-set Jaccard similarity above 0.8.
  • Robustness scope: Robustness comparisons used the same 13-model panel across both collections because three models were unavailable during re-collection.Claude Sonnet 3.7, GPT-5.3, and GPT-5.3 reasoning were unavailable through OpenRouter.
Loading 2608.30856v1…