Source-linked AI summary
Representational alignment yields generalizable safety in language models
Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
TL;DR
LLMs can fail to generalize safety when harmful intent is expressed in unfamiliar forms, reflecting weak preservation of human-like moral categorization. The paper introduces ReSO to align latent representations with graded human moral judgements, finding that representational alignment improves adversarial safety more consistently than response-level alignment.
Problem
Current alignment methods optimize observable responses, but models remain vulnerable when harmful intent is recast in unfamiliar or adversarial forms.
Method
ReSO directly aligns latent representations with graded human moral categorization rather than supervising generated responses.
Results
Directly aligning categorical structure improved safety under unfamiliar and adversarial conditions while largely preserving general capability, whereas optimizing preferred responses did not.
Takeaways & Limitations
The findings support graded categorical organization as a representational basis for more generalizable safety under adversarial conditions.
Takeaways & Limitations
The moral categorization reference is not an exhaustive or exact description of human safety alignment.
Abstract
from arXiv · showhide
Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.
1 Shanghai Artificial Intelligence Laboratory, Shanghai, China
Current LLMs often fail to preserve human-like moral categorization, while representational alignment supports more adaptable safety under adversarial conditions.
- Across 23 LLMs, models often failed to distinguish opposed moral categories.
- Behavioral alignment left categorization structure largely unchanged and increased vulnerability.
- The findings support prototype-based categorization as a contributor to behavioral adaptability.
Main
The paper asks whether human-like moral categorization can help LLMs generalize safety across unfamiliar linguistic forms. It compares response-level alignment with direct alignment of latent moral representations.
- Human moral categorization organizes actions by shared categories and graded typicality around prototypes.
- LLMs can remain vulnerable when harmful intent is recast in unfamiliar or adversarial forms.
- Across 23 models ranging from 0.6B to 235B parameters, moral prototypes were inconsistently separated, especially for care–harm and authority–subversion.
- Peak typicality correlations remained below 0.55 across all models and categories, while linear probes recovered only limited moral-structure variance.
- DPO improved explicit moral judgements but left RSA largely unchanged, whereas ReSO increased RSA with modest judgement-accuracy gains.
- ReSO directly aligned latent relations with human moral categorization without supervising generated responses.
- Across evaluations, DPO increased jailbreak attack success, while ReSO consistently improved robustness and reduced attack success rates.
- Across 20 checkpoints and baseline, RSA explained 86% of the variation in attack success, with later RSA regression accompanying higher attack success.
Language models preserve human moral categorization weakly
The study constructs a graded human moral reference and finds that current LLM representations only weakly preserve category separation, typicality gradients, and complete moral profiles.
- A graded reference of human moral categorization: The reference encodes each action as a sparse ten-dimensional vector whose active dimensions represent moral categories and typicality magnitudes.
- Category centroid analysis: Across five foundations, authority–subversion similarity averaged 0.65 ± 0.26 and care–harm similarity 0.31 ± 0.38, indicating frequent category overlap.
- Category centroid analysis: Sanctity–degradation separation was negative in all 23 models, whereas loyalty–betrayal and fairness–cheating separation was less uniform.
- Typicality gradients within moral categories: Vice typicality was better preserved than virtue typicality by 0.10 on average in 95.7% of model-foundation combinations.
- Linear decodability of human moral categorization: Base, instruction-tuned, and safeguard variants generally showed similar trajectory shapes, suggesting standard alignment left latent organization relatively stable.
Behavioral alignment leaves categorization unchanged during training
Matched training experiments dissociated behavioral alignment from representational reorganization: DPO rapidly improved explicit judgements, whereas ReSO reorganized latent moral relations with modest behavioral change.
- Representational and behavioral trajectories: DPO left RSA almost completely unchanged despite optimizing the same human annotations.
- Representational and behavioral trajectories: The shuffled-label control showed only transient RSA gains before returning toward or below baseline.
- Representational and behavioral trajectories: ReSO produced modest judgement-accuracy changes, while DPO reached approximately 0.69–0.78 early in training.
- Representational and behavioral trajectories: DPO learned the preferred output pattern without reorganizing underlying moral categorization.
Aligning moral categorization produces generalizable safety
Across four model families, ReSO improved adversarial safety more consistently than DPO while producing mixed effects on explicit moral judgements and preserving general capabilities. ReSO also improved the trade-off between refusing unsafe requests and answering benign ones.
- Adversarial robustness: ReSO lowered every red-teaming ASR across the three Qwen3 models, whereas DPO raised every one.ReSO reduced HarmBench ASR from 26.17% to 14.72% at 8B, from 22.48% to 13.33% at 14B, and from 19.00% to 13.67% at 32B.
- Adversarial robustness: ReSO reduced ASR on 23, 23, and 19 of 27 OpenRT attacks across the three Qwen3 models.DPO worsened ASR on 22, 23, and 22 OpenRT attacks, respectively.
- Refusal balance: ReSO improved XSTest’s balance between unsafe and safe refusals while DPO preserved safe refusal rates but produced smaller or negative score changes.ReSO reduced balanced-score penalties by 4.6, 3.74, and 0.4 points at three models; DPO changes were -0.70, +0.95, and +1.45 points.
- Explicit judgements: On explicit morality and value benchmarks, DPO improved Ethics Benchmark and MoReBench across all four families, while ReSO’s effects were mixed on English suites.DPO lowered the Chinese-language Flames harmless score in all four families, whereas ReSO raised Flames scores in the reported evaluations.
- Capability preservation: ReSO largely preserved capabilities, with MMLU-Pro changes of at most 1.87 points and HaluEval changes of at most 1.29 points.The largest reported capability regression was 3.72 HaluEval points for gpt-oss-20b under DPO.
- Control condition: The shuffled control reproduced ReSO’s effect direction but recovered only part of its magnitude because checkpoint selection used peak validation RSA.For example, it recovered 2.34 of ReSO’s 11.45 points.
Generalizable safety scales with representational alignment
Training ReSO increased representational similarity while reducing HarmBench attack success, and checkpoints followed a compact RSA–ASR trajectory. Reversing the representational trend reproduced the relationship, whereas DPO left RSA nearly unchanged while increasing ASR.
- Training trajectory: Under ReSO, validation RSA rose from 0.060 to 0.24 while HarmBench ASR fell from 24.2% to 15.1% at its minimum.The minimum ASR change was Δ = −9.1, 95% CI −12.8 to −5.4; P = 3.5 × 10⁻⁶.
- Training trajectory: ReSO checkpoints traced a compact path through the RSA–ASR plane with R² = 0.855.The trajectory connected checkpoints in training order, with the plotted relationship summarized in the RSA–ASR plane.
- Training trajectory: During reversal, RSA declined from 0.241 to 0.196 while ASR rose from 0.154 to 0.178.Separate slopes for the ascending and reversal phases were −0.510 and −0.506, with indistinguishable estimates (difference −0.004; P = 0.99).
- Controls: DPO left RSA essentially fixed while raising ASR from 24.5% to 36.6%.The RSA range was 0.003 across the run, while ASR increased by Δ = +12.1%, 95% CI +7.7% to +16.4%.
Discussion
Behavioral and representational alignment are dissociable: response-level training can produce aligned outputs without reorganizing moral categorization, whereas ReSO reorganizes latent relations and improves robustness to unfamiliar attacks. These findings support prototype-based categorization as a representational resource for adaptive safety generalization.
- Moral categorization in LLMs: Across 23 open-weight LLMs, opposing moral categories overlapped and within-category typicality was only modestly preserved.These patterns appeared across training lineages, including base, instruction-tuned, and safeguard variants.
- Dissociation between output and representation: Behavioral alignment learned intended moral responses while leaving representational similarity and latent categorization trajectories largely unchanged.The training experiment therefore distinguishes an effective output policy from reorganization of moral concepts.
- Adversarial robustness: DPO increased susceptibility across the evaluated jailbreaks, whereas ReSO consistently reduced attack success.This contrast links representational reorganization with improved adversarial behavior beyond familiar response patterns.
- Representational alignment: ReSO aligns latent relations with human moral categorization without supervising generated responses, reorganizing similarities among concepts by category and typicality.The approach constrains instances sharing moral category and typicality to retain corresponding similarities despite linguistic variation.
- Adversarial robustness: Reorganizing moral categorization reduced attack success across all jailbreak evaluations in all four models, including further gains in the strongly aligned gpt-oss-20b.For gpt-oss-20b, attack reduction coincided with improved balance between refusing harmful requests and answering benign ones.
- Implications and scope: Direct categorical alignment improved safety under unfamiliar and adversarial conditions while largely preserving general capability, unlike response optimization alone.The result motivates representational organization as a complementary alignment target for behavior that generalizes as language, context, and attack strategies change.
A graded human reference for moral judgement
The authors constructed a graded human reference for moral judgement from cleaned, confidence-weighted annotations, representing ten virtue- and vice-oriented moral categories. Each action’s sparse vector encodes its moral category and confidence-weighted typicality.
- 251,334 atomic judgements remained after filtering incomplete, low-quality, and multiply labelled entries into foundation-specific records.The data came from crowd-sourced everyday social situations annotated under Moral Foundations Theory.
- Each action was mapped from judgement polarity and annotator agreement into confidence-weighted membership degrees.Judgement scores were mapped from [−2,2] to [−1,1], while agreement scores were converted to confidence weights.
- Ten categories split five paired Moral Foundations Theory foundations into virtue and vice poles, including care, harm, fairness, cheating, loyalty, betrayal, authority, subversion, sanctity, and degradation.
- The resulting sparse representation uses the active dimension for moral category and magnitude for typicality, with neutral judgements represented by zero vectors.A value near 1 denotes an action judged extreme with strong annotator agreement.
- Data splits occurred before multi-label flattening, keeping records derived from the same annotation within one partition.The resulting partitions contained 201,023 training, 25,170 validation, and 25,141 test records.
Models
The study evaluates 23 open-weight models spanning model sizes, architectures, families, and alignment stages. Training experiments test scaling within one Qwen3 lineage and transfer across architecture and model family.
- 23 open-weight models span 0.6B to 235B parameters, dense and mixture-of-experts architectures, and base, instruction-tuned, and safeguard stages.
- The evaluation covers four model families, including Qwen3, Llama-3, Llama-4, and gpt-oss.
- Qwen3-8B, Qwen3-14B, and Qwen3-32B test scaling within one lineage, while gpt-oss-20b tests generalization across architecture and family.The gpt-oss-20b model is a mixture-of-experts model with substantially stronger baseline safety alignment.
Representation extraction and anisotropy correction
For representational analysis, the authors extract mean-pooled residual-stream states across decoder layers for a stratified action set. They center activations using layer-wise global means to reduce anisotropy before correlation analyses.
- A stratified subset of 16,135 actions was balanced across ten moral categories and their typicality levels.
- Each action was rendered with “action is morally” and processed in one forward pass, with mean-pooled residual-stream states recorded at every decoder layer.Mean-pooled representations were selected because they have been reported to encode conceptual structure more robustly than final-token states.
- The authors computed a per-layer global mean across extracted activations and subtracted it from every representation.This correction addresses the common direction shared across Transformer inputs, known as anisotropy.
- Correlation analyses were performed on the centered representations.
Category prototype analyses
The paper tests whether model representations capture moral prototypes, opposing-category separation, and within-category typicality, then trains and compares representational and behavioral alignment methods. The experiments evaluate judgement fidelity, general capability, moral reasoning, refusal balance, and robustness to diverse attacks.
- Category prototype analyses: Opposing-category separation uses prototype cosine similarity, where negative values indicate opposing directions and positive values indicate overlap.
- Category prototype analyses: Typicality gradients measure Spearman correlation between action proximity to a category prototype and annotated typicality.
- Representational similarity optimization: ReSO trains latent relational structure toward human moral-judgement relations without supervising generated responses, reward models, judges, or behavioral targets.
- Representational similarity optimization: The ReSO objective combines structural and preservation losses, constraining layer-wise similarity orderings while regularizing against drift from a frozen reference model.The preservation term uses vocabulary-level KL divergence on a 50,000-document general replay corpus without moral annotations.
- Comparison arms: The behavioral comparison uses DPO on preference pairs constructed from the same 251,334 annotations, while a shuffled-label control disrupts action–judgement correspondence.
- Evaluation: Evaluation covers nine text attacks over 240 HarmBench-derived behaviours and benchmarks of capability, moral values, harmlessness, procedural reasoning, refusal balance, and red-team robustness.The test suite includes a white-box GCG attack and 26 black-box attacks.