Source-linked AI summary
The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal
Md Mokarram Chowdhury, Ernie Chang, Yang Li
TL;DR
Roleplay jailbreaks can elicit harmful answers even when the harmful request remains visible, raising the question of how wrapper context reverses refusal. The paper uses matched mechanistic analysis, causal interventions, and geometric decomposition across benchmarks, models, and wrappers. It finds that harm recognition persists at the request but loses refusal-associated influence at answer generation, while complete wrapper construction and scenario framing causally contribute through largely shared refusal-related structure.
Problem
Roleplay jailbreaks can reverse refusal while leaving harmful requests explicit, but the internal operation and causal role of wrapper components remain unresolved.
Method
The study compares matched harmful and benign requests with and without wrappers, traces position-resolved hidden-state contrasts, intervenes on wrapper directions, and analyzes their geometry.
Results
Successful attacks preserve harmful–benign distinction at the request but weaken its refusal-associated expression at assistant start; complete construction and scenario framing causally contribute, largely through shared refusal structure.
Takeaways & Limitations
Safety safeguards should preserve the connection from harm recognition to refusal while retaining benign roleplay.
Takeaways & Limitations
The findings may not extend to other languages, models, or roleplay styles, and uncalibrated edits can increase refusal on benign prompts.
Abstract
from arXiv · showhide
Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply. We use mechanistic interpretability to determine how this context reverses refusal and which elements contribute to the reversal. Across two benchmarks, three model families, and four authored wrappers, we compare matched harmful and benign requests with and without this wrapper. We trace hidden-state contrasts from the request to the final prompt state, isolate wrapper operations through controlled counterfactuals, intervene on their activation directions in held-out evaluation requests, and decompose effective directions geometrically. Our analysis yields three findings. (1) Successful attacks retain the measured harmful-versus-benign distinction at the request, while its refusal-associated expression weakens where the answer begins, a pattern we call safety-relay attenuation. (2) Constructing the complete roleplay around the request and framing it within the scenario contribute causally: removing the associated activation changes restores refusal. (3) These effects largely share internal structure, and most repair is reproduced by components aligned with the model's ordinary refusal of harmful requests without roleplay; scenario framing retains a smaller, model-dependent component. Together, these findings explain how roleplay can produce compliance despite retained evidence of harm and identify a concrete target for future safeguards: maintaining the connection from harm recognition to refusal.
1 Introduction
This paper asks how roleplay wrappers reverse refusal while leaving harmful requests explicit, and identifies the internal changes and wrapper operations involved. Using matched contrasts, counterfactuals, interventions, and geometry, it finds weakened refusal expression at answer generation and shared refusal-related structure.
- Motivation: Roleplay jailbreaks preserve an explicit harmful request while reshaping its context through persona, scenario, task, rules, and output format.The wrapper can turn harmful requests into actionable guidance.
- Research questions: The study asks which internal change distinguishes successful attacks, which wrapper operations contribute causally, and whether those operations converge on shared refusal structure.
- Method: The matched design compares harmful requests with benign counterparts in plain and wrapped conditions while tracking states at the request and assistant-start positions.Surrounding wrapper construction is held fixed within each pair.
- Findings: Across two benchmarks, three model families, and four wrappers, successful roleplay preserves the harmful–benign contrast at the request but weakens its refusal-associated expression where answering begins.The paper calls this positional pattern safety-relay attenuation.
- Findings: Complete wrapper construction and scenario framing make reliable causal contributions, with endpoint interventions restoring refusal on held-out evaluation requests.The effects are weaker when moved to final request tokens and retain substantial average effects across wrappers and benchmarks.
- Findings: The robust effects largely converge on refusal-associated structure already present for plain harmful requests, while scenario framing retains a smaller residual that varies across models and wrappers.
2 Related Work
Prior work examines jailbreak representations, refusal directions, and harmfulness judgments, but leaves the internal effects of individual roleplay operations unresolved. This paper treats roleplay construction itself as a component-level causal unit.
- Jailbreaks through natural-language context: Roleplay exploits persona modulation and nested fictional scenarios, turning an intended model capability into an attack surface.
- Harm recognition and refusal: Prior mechanistic work distinguishes refusal-related directions from harmfulness representations, motivating analysis of how harm judgment and refusal response can separate.
- Mechanistic accounts of jailbreak success: Existing studies identify internal signatures at the level of completed prompts, attack families, or learned features, but do not explain transparent roleplay construction component by component.
- Mechanistic accounts of jailbreak success: A roleplay wrapper consists of identifiable elements such as persona, scenario, and task, making component-level causal analysis possible.
- Contribution: The paper provides a component-resolved causal account linking wrapper structure, harm recognition, and refusal control in one analysis.
3 Methodology
The methodology combines matched harmful–benign prompts, position-resolved residual-stream analysis, controlled wrapper counterfactuals, and held-out interventions to measure and causally test refusal changes.
- Controlled roleplay comparisons: Matched harmful and benign requests preserve topic, form, and requested output while differing in harmful objective, enabling comparisons with and without fixed roleplay wrappers.Responses are classified as refusal, harmful compliance, or neutral/other using deterministic rules and semantic judgment.
- Measuring the safety relay: Residual-stream contrasts are measured at the request and assistant-start locations to track harmful–benign distinctions before answer generation.States are averaged over location-specific token spans and then over matched request pairs within each cohort.
- Measuring the safety relay: Harmfulness retention measures preservation during request processing, whereas refusal transfer measures preservation in the final prompt state immediately before generation.Their reduction from request to assistant start defines safety-relay attenuation.
- Localizing causal refusal control: Decoder-block interventions add a repair direction to wrapped harmful prompts or ablate the unwrapped harmful–benign direction to test causal effects on refusal.Edits are applied at every assistant-start token while holding prompts and generation settings fixed; blockwise tests localize candidate control points.
- Held-out causal tests: Held-out tests subtract wrapper-operation directions, while geometric decomposition separates their refusal-aligned and orthogonal components under fixed evaluation conditions.Harmful prompts test refusal restoration, and matched benign prompts test nonspecific refusal.
4 Experiments
Across benchmarks and models, successful roleplay preserved request-level harmfulness signals but weakened their refusal-associated transfer at assistant start. Counterfactual interventions identified complete-wrapper construction and scenario framing as causal contributors, with effects concentrated at the answer boundary and largely aligned with ordinary refusal structure.
- 4.2 Roleplay weakens refusal at the answer boundary: Late-layer harmfulness retention differed by at most 0.063 between successful and failed attacks, while fail-control refusal transfer was 0.149–0.209 higher and up to 3.2× as large.The separation emerged in later blocks at assistant start despite similar request-level harmfulness curves.
- 4.2 Roleplay weakens refusal at the answer boundary: Adding a repair direction restored refusal to 100% on wrapped harmful prompts, while removing the unwrapped reference projection suppressed up to 96.8% of refusals without roleplay.Both effects occurred across evaluated model–benchmark–wrapper configurations.
- 4.3 Wrapper operations causally change held-out responses: Complete-wrapper construction, scenario framing, and their prefix–scenario interaction increased harmful compliance across benchmark–model aggregates, with complete-wrapper construction producing the largest mean change.These were selected from nine matched wrapper variants; the associated figure reports overall means, benchmark–model means, and random controls.
- 4.3 Wrapper operations causally change held-out responses: Complete-wrapper subtraction raised refusal from 0.6% to 78.6% and lowered harmful compliance by 73.6 points, whereas scenario subtraction raised refusal from 44.1% to 81.2% and lowered compliance by 36.3 points.The prefix–scenario interaction produced a smaller, less consistent 25.3-point refusal gain, while random directions changed refusal by at most 1.2 points.
- 4.4 Position, sign, and transfer: Moving edits from assistant start to final request tokens reduced refusal gains to 1.8 and 2.9 points, while directions transferred across wrappers and benchmarks retained substantial effects.Cross-wrapper gains were 43.8 and 24.7 points; cross-benchmark gains were 39.7 and 25.3 points for complete-wrapper and scenario directions, respectively.
- 4.5 Two robust wrapper effects share refusal-associated structure: The two reliable directions had mean cosine similarity 0.578, and shared components explained 78.9% of their squared norm while closely matching total-edit refusal gains.Shared-component gains were +44.1 versus +42.5 points for complete-wrapper construction and +29.5 versus +27.7 for scenario framing.
- 4.5 Two robust wrapper effects share refusal-associated structure: Refusal-reference-parallel components accounted for most restoration, producing gains of 38.3 and 25.3 points versus total-edit gains of 42.5 and 27.7.Orthogonal components had smaller, configuration-dependent effects.
5 Conclusion
The paper concludes that roleplay can preserve explicit harm recognition while weakening its influence on refusal at assistant start. Its causal tests implicate complete-wrapper construction and scenario framing, whose effects largely converge on structure used for ordinary refusal.
- 5 Conclusion: Successful roleplay preserves the measured harmful–benign distinction at the request but weakens its refusal-associated expression at assistant start.The paper names this positional pattern safety-relay attenuation.
- 5 Conclusion: Constructing the complete roleplay around the request and framing it within the scenario are reliable causal contributors to refusal reversal.Removing their activation directions restores refusal in the tested cases.
- 5 Conclusion: The findings support preserving the connection from harm recognition to refusal while retaining benign roleplay.The interventions are presented as explanatory probes rather than defenses.
Limitations
The study evaluates two English safety benchmarks, three open model families, and four authored wrappers, so its findings may not extend beyond those settings. Some failed-attack comparison groups are small or empty, restricting comparisons to configurations with examples in both groups.
- Scope: The evaluation covers two English safety benchmarks, three open model families, and four authored wrappers, limiting supported scope across languages, models, and roleplay styles.The authors explicitly caution that findings may not extend to other settings.
- Comparison coverage: Failed-attack comparison groups are sometimes small or empty, so successful-versus-failed comparisons include only configurations where both groups contain examples.This constraint is built into the comparison design.
- Causal scope: The causal tests address requests refused without roleplay but answered harmfully with it, rather than mapping the model’s complete safety process.The authors also note that later geometry analysis uses the same evaluation data and is not an independent replication.
Ethical Considerations
The study uses public benchmarks and locally hosted open-weight models, while acknowledging misuse risks and restricting access to executable artifacts. Its experiments use matched prompts, controlled interventions, native chat templates, frozen checkpoints, and reproducible development/evaluation splits.
- The work could be misused to improve jailbreak prompts, so executable artifacts should be shared with appropriate access controls.
- The study uses public safety benchmarks, no personal data, and reports only aggregate results without harmful answers.
- Matched harmful and benign prompts preserve topic, form, and requested output while changing the objective, and are tested with identical wrappers.
- Controlled counterfactuals isolate wrapper operations, while held-out requests test whether removing associated directions restores refusal.
- Interventions modify every token in the native assistant-start span using residual-state directions estimated from that span.
- Experiments use frozen instruction-tuned checkpoints with greedy decoding, capped formatted inputs, and no model training or fine-tuning.
H Wrapper-Level Behavioral and Relay Results
Roleplay wrappers produce substantial harmful compliance despite a common no-roleplay refusal reference. Relay diagnostics show that successful reversals retain harmfulness while weakening refusal transfer near the answer boundary.
- 79.0%–100% harmful compliance occurs across every roleplay wrapper, compared with the unchanged no-roleplay prompt as the common refusal reference.
- In the final layer quartile, reversal and fail-control harmfulness retention differ by at most 0.063.
- Fail-control refusal transfer is 0.149–0.209 higher than reversal-cohort refusal transfer in every benchmark–model summary.
- The relay pattern persists as minimum fail-control cohort sizes increase, indicating it is not driven by the smallest cohorts.
I Assistant-Start Intervention Sweeps
Assistant-start interventions identify a refusal-sensitive answer-boundary state across model families. Repair can restore refusal, while directional ablation can suppress baseline refusal, with recurring responsive block ranges.
- 100% refusal is reached by repair in at least one configuration for every model family, while ablation suppresses up to 96.8% of baseline refusals.
- Both interventions recur at blocks 10–12 for Llama, 13–18 for Qwen, and 17–28 for Gemma across wrappers.
- Within-cohort strongest settings establish causal localization for those cohorts, then define block windows for held-out component tests.
- The intervention sweeps are confined to the assistant-start span and report wrapper means and ranges at each block’s strongest tested nonzero strength.
K Causal Tests of Wrapper Components
Held-out interventions show that complete-wrapper construction and scenario framing make stable causal contributions to harmful compliance, while their interaction is smaller and more variable. These effects remain positive after norm-matched random-direction controls and preserve the same ordering under alternative aggregation and strength choices.
- The held-out evaluation uses directions and settings selected on development requests and frozen before intervention.
- Complete-wrapper and scenario subtraction raise refusal and reduce harmful compliance in every benchmark–model aggregate.
- The recurrence of these effects across three model families and two benchmarks supports stable causal contributions, whereas the prefix–scenario interaction is smaller and more variable.
- Every wrapper-level margin remains positive after subtracting mean effects from approximately norm-matched random directions.
- Request-weighted effects and architecture-fixed strengths preserve the ordering: complete-wrapper construction strongest, scenario framing next, and interaction smallest.
L Causal Validation and Transfer
Controlled interventions show that complete-wrapper construction and scenario framing causally alter refusal, with effects concentrated at assistant start and transferable across contexts. The edits are not selective: repairing harmful targets also increases refusal on benign targets.
- Causal validation: Subtracting a target direction increases refusal, whereas adding it to the matched control decreases refusal, reversing the observed wrapper effect.The comparison uses directions from matched controls toward roleplay targets.
- Causal validation: The refusal effect nearly disappears when the unchanged edit moves from assistant start to the final five request tokens.Direction, block, strength, and sign remain fixed across positions.
- Causal validation: The matched wrapper variants recombine request, scenario, persona, and later instructions to test progressively different construction operations.Rows denote benchmarks and columns denote models; bars report wrapper-averaged refusal, neutral/other, and harmful compliance.
- Transfer: Both stable directions transfer positively across excluded wrappers and the other benchmark without target-side retuning, although the prefix–scenario interaction changes sign for Llama.Transfer uses fixed intervention settings and excludes the target wrapper or benchmark.
- Specificity: The edits are not selective defenses: stronger harmful-target repair generally incurs a larger benign-refusal cost, and each direction also affects the other nested target.This validates a causal mechanism while imposing a specificity cost.
M Shared Structure Across Wrapper Operations
The complete-wrapper and scenario operations largely converge on shared internal structure, whose refusal effect is strongest at assistant start and aligns with ordinary refusal. Scenario-specific residual effects are smaller and model- and wrapper-dependent.
- Shared structure: The complete-wrapper and scenario directions are positively aligned across models and benchmarks, defining shared and contrast axes through their normalized sum and difference.The decomposition separates common structure from operation-specific contrast.
- Shared structure: The shared component closely reproduces both total refusal effects, whereas the contrast remains small; norm matching strengthens the shared effect without making the contrast effective.Request weighting preserves the same ordering.
- Evaluation design: The evaluation reports paired refusal changes, transfer conditions, intervention positions, decomposition components, and harmful-minus-benign differences with bootstrap intervals across configurations.These summaries cover the paper’s held-out causal, transfer, and geometric analyses.
- Shared structure: Reversing the contrast changes refusal by only +1.1 points for the complete-wrapper target and +2.3 points for the scenario target, with both intervals including zero.These effects are contrasted with the larger assistant-start gains from the total edit.
- Intervention position: Moving the identical one-token total edit from the wrapper suffix to assistant start improves refusal by +8.4 and +11.5 points.The common component is therefore effective at assistant start rather than at the wrapper suffix.
- Model variation: The shared component carries a positive refusal effect in every model family, while the contrast is small or changes sign.Figure 19 disaggregates the decomposition by model.
- Alignment with ordinary refusal: The component parallel to the unwrapped refusal-associated reference accounts for most total repair for both targets.The complete-wrapper residual contributes little even after norm matching, while the scenario residual retains a smaller harmful-minus-benign effect with little benign refusal.
- Limitations: The residual does not support a universal scenario-specific mechanism: effects are largest for Qwen and selected wrappers, modest for Gemma, and near-zero for Llama and the historical wrapper.The stable average effect follows ordinary refusal-associated structure, while the residual depends on model and wrapper.