Source-linked AI summary
Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
Lihao Sun, Lewen Yan, Xiaoya Lu, Andrew Lee, Jie Zhang, Jing Shao
TL;DR
LLM emotion controls lack a clear common basis beyond discrete emotion labels and task-specific steering directions. The paper recovers a circular 2D valence–arousal subspace through principal components and ridge regression, finding human-aligned affective control and bidirectional control of refusal and sycophancy across three models.
Problem
Existing emotionally framed controls influence LLM behavior, but their common basis and failure conditions remain unclear.
Method
The paper applies principal-component decomposition and ridge regression to emotion steering vectors to recover a shared valence–arousal subspace.
Results
A circular VA geometry supports monotonic affective control and bidirectional refusal and sycophancy control across Llama 3.1-8B, Qwen3-8B, and Qwen3-14B.
Takeaways & Limitations
Lexical mediation offers a mechanistic perspective in which VA steering shifts emission probabilities of refusal and compliance tokens.
Takeaways & Limitations
Human evaluation of generated outputs would provide stronger affective validation, and lexical mediation may not be the only mechanism affected by VA steering.
Abstract
from arXiv · showhide
We show that emotion vectors in LLMs are organized by a two-dimensional valence-arousal (VA) subspace exhibiting circular geometry. Through principal component decomposition and ridge regression, we recover meaningful VA axes underlying emotion steering vectors whose projections correlate with human affect ratings across 44,728 words. Steering along these axes produces monotonic control over the affective properties of generated text, and further affords bidirectional control over multiple downstream behaviors (refusal and sycophancy) from a single subspace. These effects replicate across Llama-3.1-8B, Qwen3-8B, and Qwen3-14B. We propose lexical mediation to explain why these effects and prior emotionally framed controls work: refusal and compliance tokens occupy distinct VA regions, and VA steering directly modulates their emission probabilities.
1 Introduction
The paper separates discrete emotion labels from continuous valence–arousal dimensions and identifies a circular VA subspace that supports reusable affective and behavioral control in LLMs.
- 1 Introduction: The paper treats categorical emotion labels and continuous valence–arousal axes as separate representational structures.This modeling choice distinguishes discrete emotions from the continuous affective dimensions from which they can emerge.
- 1 Introduction: 211,225 emotion-labeled text samples yield VA projections that correlate with human ratings across 44,728 words and replicate across three LLM families.The reported models are Llama 3.1-8B, Qwen3-8B, and Qwen3-14B.
- 1 Introduction: VA steering produces monotonic affective control and bidirectional control over refusal and sycophancy from one shared set of axes.The authors frame this shared structure as more reusable than single-task contrastive steering directions.
- 1 Introduction: Lexical mediation links VA perturbations to behavior by shifting emission probabilities for refusal and compliance tokens occupying distinct VA regions.The proposed account also explains why prior emotion-based prompting methods induce corresponding VA shifts.
- 1 Introduction: Principal-component decomposition and ridge regression recover a 2D VA subspace whose emotion-vector geometry is circular and human-interpretable.The arrangement is analogous to Russell’s circumplex model of affect.
2 Related Work
Prior work studies emotional prompting, discrete emotion representations, individual behavioral directions, and limited low-dimensional geometries. This paper contributes a continuous 2D VA subspace that unifies emotionally framed behavior across representational and mechanistic levels.
- 2 Related Work: Prior emotion-control studies largely analyze discrete labels, including hierarchical representations, emotion-specific neurons, attention heads, and affective biases.These approaches motivate searching for a meaningful underlying subspace beyond categorical emotion analysis.
- 2 Related Work: The paper presents a continuous 2D VA subspace as a unifying account linking emotionally framed behaviors across representation geometry, unembedding structure, and neuron-level analyses.This extends prior work that examined emotion directions or behavioral controls separately.
- 2 Related Work: Activation-steering research typically uses task-specific contrastive vectors, while two-dimensional representational geometries remain less commonly characterized.The paper positions its approach at the intersection of behavioral steering and low-dimensional representation geometry.
3 Identifying Valence and Arousal Subspaces
The paper extracts emotion steering vectors, maps them into a low-dimensional valence–arousal space, and evaluates whether the resulting geometry recovers affective structure across models. Ridge regression yields strong VA recovery, while cross-model and human-rating analyses support the subspace’s generality, with intervention experiments testing its behavioral relevance.
- Method: The method extracts mean-difference emotion directions, elicits model VA scores, and learns orthonormal VA axes as principal-component combinations.Emotion directions come from GoEmotions, while ridge regression maps principal-component scores to valence and arousal ratings.
- VA subspace recovery: Ridge regression over multiple principal components reaches r = 0.97 for valence and r = 0.87 for arousal, outperforming single-component recovery.Valence is well captured by PC1, whereas arousal is distributed across secondary components.
- Circular geometry: The recovered VA projections form a circular emotion geometry analogous to Russell’s circumplex, with valence and arousal separating opposing affective states.Positive and negative emotions oppose each other along valence, while high- and low-arousal states separate along the orthogonal axis.
- Cross-model generalization: Comparable circle radii of 0.37, 0.37, and 0.39 and recovery correlations across Llama-3.1-8B, Qwen3-8B, and Qwen3-14B support cross-architecture generalization.The same extraction and fitting pipeline was applied to all three architectures.
- Human-rating validation: Human-rated regression targets yield comparable or improved recovery, with arousal increasing from 0.87 to 0.95 for Llama and valence remaining at r = 0.97.Human- and self-report-supervised axes are closely aligned across layers and architectures.
- Caveat: The identification procedure alone does not establish that the subspaces represent human-interpretable VA concepts or causally influence behavior, motivating intervention experiments.The axes are fitted to self-reported scores for 27 labels, while steering vectors come from GoEmotions text passages.
4 Validation of VA Subspaces
The recovered VA subspace aligns with human affect ratings and enables directional control over the affective properties of generated text. Valence and arousal show distinct, largely separable effects, while arousal steering also controls refusal across benchmarks.
- 4.1 Correlation with Human-Crowdsourced Lexicon Annotations: r = 0.71 valence correlation with human ratings peaks at layer 6, while arousal reaches r = 0.23 at layer 7 across 44,728 words.Valence projections are substantially more predictable than arousal projections from de-contextualized words.
- 4.2 Controlling Affective Properties of Responses with VA Steering: ∆V = +0.75 at 0° and ∆V = −0.73 at 180°, demonstrating symmetric bidirectional control of response valence.VADER sentiment tracks the same pattern.
- 4.2 Controlling Affective Properties of Responses with VA Steering: Arousal steering changes intensity with limited valence leakage: ∆A = +0.50 at 90° and ∆A = −0.16 at 270°.The results support separable control of emotional intensity and positive–negative tone.
- 4.2 Controlling Affective Properties of Responses with VA Steering: Valence steering also raises arousal at both extremes, with ∆A = +0.49 for +V and ∆A = +0.22 for −V.This collateral effect is consistent with highly valenced content tending to be more arousing.
5 Controlling Refusal and Sycophancy with a Single VA Subspace
A single VA subspace, especially its arousal axis, controls refusal and sycophancy bidirectionally across multiple benchmarks. These effects are stronger and more consistent than random controls, but extreme steering can produce out-of-distribution generations.
- 5.1 VA for Refusal Control: Refusal changes from 87% to 5% on HarmBench at α = +0.45, while random directions stay within 2–3% of baseline.Across benchmarks, decreasing arousal increases refusal and increasing arousal suppresses it.
- 5 Controlling Refusal and Sycophancy with a Single VA Subspace: Valence steering modulates both behaviors but has less consistent directionality across benchmarks than arousal steering.The VA decomposition exposes distinct roles for valence and arousal that single emotion directions can obscure.
- 5.1 VA for Refusal Control: At |α| ≤ 0.20, out-of-distribution generations remain below 2%, whereas extreme strengths can produce substantial OOD rates.The reported range extends to α ∈ [−0.45, 0.45] for the Llama refusal experiments.
- 5.1 VA for Refusal Control: Qwen3-14B shows monotonic refusal control from 67.0% to 95.5% on HarmBench with <1% OOD across α ∈ [−3.0, +3.0].Qwen3 models require approximately seven times larger steering strengths than Llama for comparable shifts.
- 5.2 VA for Sycophancy Control: Sycophancy on Political Typology changes from 61% at α = −0.30 to 84% at α = +0.30, with similar near-monotonic trends elsewhere.Random directions remain near-flat across the sycophancy benchmarks.
- 5 Controlling Refusal and Sycophancy with a Single VA Subspace: A single arousal axis supports clean bidirectional control over both refusal and sycophancy across benchmarks.Increasing arousal increases compliant behavior: it reduces refusal and increases agreement with users’ opinions.
6 Lexical Mediation: Why Emotionally Framed Control Works in LLMs
The paper explains VA-based behavioral control through lexical mediation: refusal and compliance tokens occupy different VA regions, and steering changes their emission probabilities. Token-clamping and logit analyses support this pathway, while emotional prompts produce corresponding VA shifts.
- 6 Lexical Mediation: Why Emotionally Framed Control Works in LLMs: At α = +0.30, refusal-token log-odds decrease by 5.63, matching a refusal-rate drop from 86.5% to 59.5%.Arousal steering also flips 23% of prompts away from refusal tokens as the top-1 prediction.
- 6 Lexical Mediation: Why Emotionally Framed Control Works in LLMs: Refusal tokens cluster in the −V, −A region, whereas compliance tokens occupy a more +V, +A region separated by 256° on the circumplex.This unembedding geometry provides a lower-level account of how VA directions can alter token logits.
- 6 Lexical Mediation: Why Emotionally Framed Control Works in LLMs: Negative emotional prefixes shift representations toward −V and −A, with the largest shift increasing refusal by 30%.The reported shift is ∆V = −0.11 and ∆A = −0.07.
7 Conclusion
The paper identifies a replicable 2D VA subspace underlying emotion steering vectors, with circular organization and bidirectional control over affective properties, refusal, and sycophancy.
- 7 Conclusion: A 2D VA subspace organizes emotion steering vectors in a circular geometry and supports bidirectional control over affective properties, refusal, and sycophancy.The effects replicate across Llama 3.1-8B, Qwen3-8B, and Qwen3-14B.
- 7 Conclusion: Lexical mediation links VA perturbations to token-level probability shifts and downstream behavioral outcomes, offering an account of emotionally framed controls.
8 Future Work
Future work should determine which additional behaviors are VA-modulated, identify lower-level implementation pathways, and explore higher-dimensional affective representations and alignment applications.
- 8 Future Work: Future studies could map whether behaviors beyond refusal and sycophancy, including verbosity, hedging, and hallucination, are VA-modulated.
- 8 Future Work: Full causal tracing or circuit-level analysis could identify the attention heads and MLP sub-layers implementing VA-to-logit pathways.
- 8 Future Work: Extending the pipeline to dominance as a third affective dimension could yield additional insights.
- 8 Future Work: Alignment procedures could leverage consistent VA structure, including training objectives that decorrelate VA from safety-critical token probabilities.
Limitations
The paper’s validation and mechanistic account have important boundaries: affective evaluation relies on task-specific models, and lexical mediation is not claimed to be the only pathway.
- Limitations: Human evaluation of generated outputs would provide stronger affective validation than the task-specific VADER and VAD-BERT models used here.
- Limitations: The lexical mediation account confirms token-probability shifts in refusal behavior but does not establish that lexical mediation is the only mechanism.
- Limitations: The steering evaluation uses 130 open-ended prompts written to avoid explicit emotional language across diverse genres and response lengths.
- Limitations: The reported logit-lens analysis covers layers 18–31 and compares baseline with four steering conditions at α = 0.45.
E Comparison Against Individual Emotion Vectors
Compared with individual emotion vectors, VA steering provides stronger cross-behavior consistency, especially for sycophancy, while preserving capability within tested steering ranges.
- E Comparison Against Individual Emotion Vectors: Individual emotion vectors produce refusal shifts comparable to or exceeding VA arousal at matched |α|, but their effects vary across behaviors.
- E Comparison Against Individual Emotion Vectors: At |α| = 0.30, VA valence reduces sycophancy from 78% to 47%, versus 11% for the strongest individual emotion, disgust.
- E Comparison Against Individual Emotion Vectors: Human-supervised arousal recovery improves from 0.87 to 0.95 in Llama, 0.79 to 0.83 in Qwen3-8B, and 0.81 to 0.87 in Qwen3-14B.
- E Comparison Against Individual Emotion Vectors: At |α| ≤0.10, MATH-500 accuracy remains within 1% of baseline during arousal steering, while IFEval remains within 2% at |α| ≤0.20.
H Contrastive Steering Directions and the VA Subspace
Contrastive steering is stronger on its target refusal task but transfers poorly, whereas VA steering provides shared, bidirectional control across behaviors from one subspace.
- Contrastive Steering Directions and the VA Subspace: The contrastive refusal direction is nearly orthogonal to the VA plane at 86.5°, while its in-plane component has negative VA coordinates.The in-plane coordinates are V = −0.058 and A = −0.021, consistent with lexical mediation contributing to the contrastive direction.
- Contrastive Steering Directions and the VA Subspace: Refusal drops to 2.0% with contrastive steering on HarmBench, versus 59% for VA arousal steering.The comparison uses α = −0.30 for the contrastive direction and α = +0.30 for VA arousal.
- Contrastive Steering Directions and the VA Subspace: Contrastive refusal steering causes severe over-refusal on safe-query benchmarks, increasing OKTest from 19.7% to 93.3% and XSTest from 8.4% to 94.4%.These increases occur at α = +0.45.
- Contrastive Steering Directions and the VA Subspace: Transferred to sycophancy, contrastive steering reduces sycophancy in both directions rather than providing bidirectional control.Sycophancy changes from 77.6% to 57.0% at +α and 66.4% at −α.
- Contrastive Steering Directions and the VA Subspace: Task-specific contrastive directions optimize single-behavior control, whereas the VA subspace exposes shared affective structure for monotonic, bidirectional control across behaviors.The near-orthogonality suggests multiple pathways may modulate the same behavior.