Source-linked AI summary
LLMs Encode Harmfulness and Refusal Separately
Jiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau, Weiyan Shi
TL;DR
The paper asks whether LLMs understand harmfulness beyond refusing, given failures on harmful and harmless instructions. It separates harmfulness from refusal through hidden-state analysis and causal steering, then applies the representation to jailbreak analysis and Latent Guard. The results show distinct, more robust harmfulness representations and a safeguard comparable to or better than Llama Guard 3 8B across tested settings.
Problem
LLMs frequently accept harmful prompts or over-refuse harmless ones, while it remains unclear whether harmfulness is internally represented separately from refusal.
Method
The paper clusters hidden states at tinst and tpost-inst, extracts a harmfulness direction, tests causal steering, analyzes jailbreaks, and builds Latent Guard from latent harmfulness representations.
Results
Harmfulness is encoded at tinst separately from refusal at tpost-inst; harmfulness directions vary by risk category, and Latent Guard performs comparably to or better than Llama Guard 3 8B.
Takeaways & Limitations
Latent harmfulness representations provide a lens for studying internal safety mechanisms and can support intrinsic safeguarding that remains robust to tested finetuning attacks.
Takeaways & Limitations
The study does not determine how different layers formulate harmfulness and refusal and mainly evaluates open-source 7B and 8B models.
Abstract
from arXiv · showhide
LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs' refusal behaviors can be mediated by a one-dimensional subspace, i.e., a refusal direction. In this work, we identify a new dimension to analyze safety mechanisms in LLMs, i.e., harmfulness, which is encoded internally as a separate concept from refusal. There exists a harmfulness direction that is distinct from the refusal direction. As causal evidence, steering along the harmfulness direction can lead LLMs to interpret harmless instructions as harmful, but steering along the refusal direction tends to elicit refusal responses directly without reversing the model's judgment on harmfulness. Furthermore, using our identified harmfulness concept, we find that certain jailbreak methods work by reducing the refusal signals without reversing the model's internal belief of harmfulness. We also find that adversarially finetuning models to accept harmful instructions has minimal impact on the model's internal belief of harmfulness. These insights lead to a practical safety application: The model's latent harmfulness representation can serve as an intrinsic safeguard (Latent Guard) for detecting unsafe inputs and reducing over-refusals that is robust to finetuning attacks. For instance, our Latent Guard achieves performance comparable to or better than Llama Guard 3 8B, a dedicated finetuned safeguard model, across different jailbreak methods. Our findings suggest that LLMs' internal understanding of harmfulness is more robust than their refusal decision to diverse input instructions, offering a new perspective to study AI safety.
1 Introduction
This work asks whether LLMs internally represent harmfulness separately from refusal, addressing failures in both harmful-prompt acceptance and harmless-prompt over-refusal. It identifies distinct harmfulness and refusal representations, analyzes jailbreak behavior, and proposes Latent Guard as an intrinsic safeguard.
- Motivation: LLMs still accept some harmful prompts and refuse some harmless prompts despite training intended to produce both harmlessness and helpfulness.More sophisticated jailbreaks further reduce refusal rates for harmful prompts.
- Research gap: Prior work identified a refusal direction, but whether LLMs encode a generalizable harmfulness concept separately remains unclear.Earlier studies often treated the refusal direction as harmfulness, leaving the two concepts potentially conflated.
- Causal evidence: Steering along the harmfulness direction can make harmless instructions be interpreted as harmful, whereas refusal-direction steering directly elicits refusal behavior.The harmfulness direction is extracted at tinst, while the refusal direction is extracted at tpost-inst.
- Core finding: Hidden states cluster primarily by instruction harmfulness at tinst and by refusal behavior at tpost-inst.The analysis uses the last token of the user instruction and the last token of the whole input sequence.
- Representation structure: Harmfulness directions vary across risk categories, while refusal directions are more similar across categories.This indicates a finer-grained categorical representation for harmfulness than for refusal.
- Application: Certain jailbreaks suppress refusal signals without fully reversing the model’s internal harmfulness belief, motivating Latent Guard as an intrinsic safeguard.Latent Guard achieves performance comparable to or better than a dedicated finetuned Llama Guard model.
2 Experimental Setup
The experiments study open-source instruct models using hidden states at instruction and post-instruction token positions across harmful, harmless, over-refused, and jailbreak-related datasets. The setup also specifies prompting templates and refusal-rate measurement.
- Models: The study evaluates LLAMA2-CHAT-7B, LLAMA3-INSTRUCT-8B, and QWEN2-INSTRUCT-7B as open-source instruct models.Experiments run on A100-40GB GPUs.
- Prompting templates: The models’ chat templates determine post-instruction tokens, such as [/INST] in Llama2-chat.Unless otherwise specified, the experiments use each model’s default prompting template.
- Hidden-state extraction: The analysis extracts residual-stream activations at tinst and tpost-inst, which both contain information from the full instruction but differ in whether special post-instruction tokens are included.The study examines tinst because a model may accept a harmful instruction there yet refuse it at tpost-inst.
- Datasets: Harmful examples come from Advbench, JBB, and Sorry-Bench, while harmless examples use ALPACA and datasets containing over-refused benign prompts.The setup therefore includes both ordinary harmful requests and harmless requests that trigger refusal.
- Jailbreak methods: The jailbreak evaluation covers GCG adversarial suffixes, persuasion-based rephrasing, and adversarial prompting templates.These methods are designed to make LLMs accept harmful instructions.
- Evaluation: Refusal rate is the fraction of test examples whose responses contain a compiled set of common refusal substrings.In Section 3.5, the rate is computed using the refusal token “No”.
3 Decoupling Harmfulness from Refusal
The paper separates harmfulness from refusal in LLM hidden states: instruction-end representations cluster primarily by harmfulness, while post-instruction representations cluster primarily by refusal. Correlation and steering experiments further show that the two concepts can diverge causally.
- 3.1 Removing post-instruction tokens weakens refusal abilities: Removing post-instruction special tokens lowers harmful-instruction refusal rates across tested LLMs, weakening refusal behavior.This motivates comparing hidden states at the last instruction token, tinst, and the last post-instruction token, tpost-inst.
- 3.2 Hidden states cluster by harmfulness at tinst, and by refusal at tpost-inst: At tinst, hidden states primarily cluster by harmfulness, whereas at tpost-inst they primarily cluster by refusal.Across models, accepted harmful and refused harmless examples support harmfulness-based clustering at tinst, while the pattern reverses at tpost-inst.
- 3.2 Hidden states cluster by harmfulness at tinst, and by refusal at tpost-inst: At tinst, clustering based on harmfulness is most evident, while tpost-inst clustering reflects whether instructions were accepted or refused.Layer-wise patterns differ across models: early strong refusal signals appear for Llama3 and Qwen2, but later for Llama2.
- 3.3 Correlation between beliefs of harmfulness and refusal: Refused harmless instructions can have low harmfulness beliefs, while accepted harmful instructions can retain positive harmfulness beliefs.Thus, behavioral refusal or acceptance is not always aligned with the model’s internal harmfulness perception.
- 3.4 Eliciting refusal with harmfulness directions: Steering along harmfulness and refusal directions can both elicit refusal, but their average cosine similarity is around 0.1 on Llama2.The harmfulness direction reaches a 94% refusal rate at layer nine on Llama3, while the refusal direction reaches 100% at layer eleven.
- 3.5 Causally separating the harmfulness direction and the refusal direction: Reply inversion shows that harmfulness steering reverses perceived harmfulness, whereas refusal steering mainly changes refusal signals without reliably changing that perception.Steering harmless inputs along the harmfulness direction makes the model answer “Certainly,” while reverse harmfulness steering makes harmful inputs receive “No.”
4 Analyzing Jailbreak via Harmfulness
The paper analyzes jailbreaks by separately measuring internal harmfulness beliefs and refusal signals. Some jailbreaks suppress refusal while leaving the model’s harmfulness judgment intact, whereas persuasion can sometimes reverse that judgment.
- Some jailbreak methods suppress refusal signals without fundamentally reversing the model’s internal belief that prompts are harmful.The analysis considers adversarial suffixes, persuasion, and adversarial prompting templates.
- Persuasion jailbreaks can sometimes make LLMs internally classify persuasive harmful prompts as harmless.This appears as negative harmfulness differences in the analysis.
- Other jailbreak methods generally produce negative refusal differences while retaining high harmfulness scores.The paper uses Figure 6 to compare harmfulness and refusal beliefs across jailbreak categories and refused harmful instructions.
5 Developing a Latent Guard Model with Harmfulness Representations
Latent Guard uses LLMs’ internal harmfulness representation to detect unsafe inputs and over-refused harmless instructions. Its harmfulness belief remains nearly unchanged after tested adversarial finetuning, supporting robustness to that attack.
- Latent Guard design: Latent Guard uses an LLM’s internal belief of harmfulness to classify unsafe inputs and harmless instructions that are over-refused.A negative ∆harmful indicates harmlessness, while a positive value indicates harmfulness.
- Evaluation: Latent Guard achieves performance comparable to or better than Llama Guard 3 8B across jailbreak, refused-harmless, and accepted-harmful test cases.The evaluation covers adversarial suffixes, persuasion, prompting templates, refused harmless instructions, and accepted harmful instructions.
- Robustness to finetuning: After finetuning on 50 to 400 adversarial examples, harmful instructions are accepted more often but their harmfulness belief remains almost unchanged.The adversarial examples are created by steering harmful instructions along the reverse refusal direction and pairing them with acceptance responses.
- Robustness to finetuning: Because ∆harmful remains stable, Latent Guard continues detecting these finetuned models’ harmful instructions as harmful.The result supports robustness to the tested narrow finetuning attack, while broader effects on representations remain future work.
6 Related Work
Related work studies linear representations, refusal directions, and jailbreak behavior in LLM latent spaces. This paper distinguishes a general harmfulness representation from refusal features and directions.
- Linear representation in LLMs: Prior studies show that concepts such as truth can be represented linearly as directions in LLM latent spaces.Intervening along a truthful direction can make models treat false statements as true.
- Refusal and harmfulness in LLMs: Existing refusal directions are computed from harmful and harmless instruction clusters at the last post-instruction token position.Prior work reports that ablating the refusal subspace can jailbreak models without degrading utility.
- Refusal and harmfulness in LLMs: This paper focuses on the general concept of harmfulness rather than sparse features associated with dangerous trigger words.It demonstrates that harmfulness and refusal are encoded separately at different token positions.
- Understanding jailbreak in the latent space: Jailbreak prompts have been reported to resemble accepted harmless instructions at the post-instruction position and to have weak similarity with the refusal direction.These observations motivate analyzing jailbreak behavior in latent space.
7 Conclusion and Discussion
The paper concludes that harmfulness and refusal are separate representations, with harmfulness encoded at the instruction position and refusal at the post-instruction position. It applies this distinction to jailbreak analysis and Latent Guard, while noting limits in layer-level interpretation and model-size generalization.
- Conclusion: Harmfulness is encoded at tinst, whereas refusal is encoded at tpost-inst.Steering the harmfulness direction can reinterpret harmless inputs as harmful, while refusal steering may reinforce refusal without reversing harmfulness judgment.
- Conclusion: Harmfulness directions differ across risk categories, while refusal directions are more similar across categories.The conclusion characterizes harmfulness as a more fine-grained representation.
- Conclusion: Some jailbreak methods suppress refusal signals without fully reversing the model’s internal belief that an instruction is harmful.This separates external refusal behavior from internal harmfulness judgment.
- Conclusion: Latent Guard reliably and efficiently detects unsafe inputs and has performance comparable to a finetuned Llama Guard model, while remaining robust to a tested finetuning attack.The application uses the model’s internal harmfulness belief as an intrinsic safeguard.
- Limitations: The study is an existence proof rather than an exhaustive account, and its findings may not generalize to larger untested models.The experiments mainly use open-source 7B and 8B LLMs, and the roles of different layers remain unstudied.
B Data
The data procedures assemble harmful, harmless, and over-refused instruction sets from public benchmarks, while selecting examples according to model responses at the instruction and post-instruction positions. Additional examples improve Latent Guard classification, but the detector can fail out of domain.
- Harmful instructions: Harmful examples are sampled from Advbench and JBB, while accepted harmful examples also include aggregated Advbench, JBB, and Sorry-Bench data.Accepted harmful examples are needed because nearly all Advbench and JBB examples are rejected at tpost-inst.
- Cluster construction: The study samples harmful instructions refused at tpost-inst to compute harmfulness and refusal cluster centers.The cluster centers are defined at tinst and tpost-inst respectively.
- Latent Guard data: Latent Guard’s sampling pool also includes harmful instructions accepted at tpost-inst and refused harmful Sorry-Bench examples.Adding these examples improves classification performance.
- Data limitation: The latent detector can fail on out-of-domain examples, which the authors identify as a limitation of the application.This boundary concerns application performance beyond the tested data distribution.
- Harmless instructions: Refused harmless instructions are identified from Xstest, whose benign prompts include keywords such as “kill” and “strangle” that can trigger over-refusal.The remaining harmless and accepted Xstest examples are aggregated for the data construction procedure.
C Steering with the harmful direction
Steering harmless instructions along the harmfulness direction can elicit refusal behavior, with model-dependent layer effects. These results extend across models, though the strongest steering layer varies.
- Steering harmless instructions along the harmfulness direction can elicit refusal behavior, as shown by layer-wise experiments across models.For Llama2, the best refusal rate occurs relatively early at layer 9, consistent with Llama3; Qwen2 shows different steering behavior.
D Harmfulness Representation Differs By Risk Categories
Harmfulness representations vary across risk categories, whereas refusal representations are more similar across categories. Category-specific steering produces distinct refusal effects, with some categories transferring strongly and others weakly.
- Harmfulness directions differ across risk categories, while refusal representations remain more similar across categories.This contrast suggests that token position tinst captures domain-specific harmfulness features, whereas tpost-inst captures more surface-level refusal signals.
- Category-specific harmfulness directions are extracted and evaluated by steering test instructions in the reverse direction to reduce perceived harmfulness.The reply inversion task measures intervention effectiveness through increased refusal-token rates.
- 100% refusal rate is reached with the in-domain “Hate_Haras_Violence” direction, compared with 32% using the out-of-domain “Political_Campaigning” direction.The test set consists of “Hate_Haras_Violence” examples, making these in-domain and out-of-domain effects directly comparable.
- 100% refusal rate is also reached with the “Adult_Content” direction, while the average harmfulness direction across categories yields 52%.The strong “Adult_Content” transfer suggests similar harmfulness representations for the two categories, whereas averaging directions is less effective.
- Refusal-direction steering is ineffective when applied only to instruction tokens before the inversion question.This contrasts with harmfulness-direction steering, which can alter the model’s perception of the original instruction.
E.2 Evaluation on different models and inversion prompts
Across prompting templates and models, harmfulness-direction steering changes harmless-instruction judgments, whereas refusal-direction steering mainly elicits shallow refusal responses. Additional token-position analyses identify tinst as the clearest harmfulness-encoding position.
- E.2 Evaluation on different models and inversion prompts: Across prompting templates, harmfulness-direction steering makes harmless instructions appear harmful and increases affirmative “Certainly” responses.The reply inversion experiments use multiple templates to test whether the pattern depends on prompt wording.
- E.2 Evaluation on different models and inversion prompts: Refusal-direction steering mainly preserves the model’s harmfulness judgment and produces refusal tokens such as “No.”The refusal direction is characterized as containing shallow refusal features rather than substantially changing harmfulness judgments.
- E.2 Evaluation on different models and inversion prompts: Llama3-8B exhibits the same harmfulness-versus-refusal steering pattern observed with Qwen2.The evaluation uses a separate inversion template for Llama3.
- E.2 Evaluation on different models and inversion prompts: Different inversion templates are adapted to models because smaller models may ignore the inversion question and answer the initial instruction instead.The authors attribute this issue to weaker instruction-following ability in smaller LLMs and expect it to be less problematic for larger models.
- F Analysis on More Token Positions: Hidden states are compared at token positions from immediately before tinst through tpost-inst to identify where harmfulness is encoded most clearly.The analysis combines clustering and steering experiments across these positions.
- F.1 Clustering at different token positions: Positive sl(h_l) values indicate hidden states closer to refused-harmful than accepted-harmless clusters.The metric is averaged over middle layers 9 to 20, which tend to be more responsible for handling harmfulness.
- F.2 Directions extracted at different token positions: Steering directions extracted at different token positions are applied before the inversion question to measure how strongly they raise perceived harmfulness.For both Qwen2 and Llama3, the resulting steering performance is compared across token positions.
G More results on categorical harmfulness in LLMs
Additional cross-model results show that harmfulness directions distinguish risk categories internally. This separation is more pronounced in Qwen2 and Llama3 than in earlier models.
- All evaluated models differentiate harmfulness directions across risk categories, indicating distinct internal representations of different risks.The category-level separation is especially pronounced in more recent Qwen2 and Llama3 models.
- More pronounced category separation in Qwen2 and Llama3 suggests that more capable models may represent harmfulness more finely.The passage frames this as a suggestion rather than a definitive causal conclusion.
H.1 More experiments
Latent Guard is evaluated against Llama Guard across ToxicChat and OpenAI Moderation datasets, with performance varying by domain and model. The experiments also examine how harmfulness and refusal representations persist after adversarial finetuning.
- Latent Guard evaluation: Latent Guard performs worse than Llama Guard 3 on ToxicChat and particularly poorly for Qwen2 and Llama3 on OpenAI Moderation.The authors attribute degradation to distribution shifts and limited coverage of harmfulness taxonomies in Latent Guard’s training data.
- Latent Guard evaluation: Adding 50 in-domain Sexual examples substantially improves Latent Guard, surpassing Llama Guard 3 on nearly every OpenAI Moderation taxonomy.The comparison may be unfair because Llama Guard uses a much larger, curated dataset and supervised finetuning, whereas Latent Guard uses a smaller statistical model without supervised finetuning.
- Latent Guard evaluation: Llama2’s stronger Latent Guard performance may reflect more centralized harmfulness representations than those of Qwen2 and Llama3.The authors suggest tighter cross-category clustering helps unseen cases map to less distant latent-space regions.
- Adversarial finetuning: After adversarial finetuning, the accepted-harmfulness difference-in-means remains a refusal direction that induces high refusal rates on held-out harmless instructions.This behavior is similar to the original refusal direction measured before finetuning.