Source-linked AI summary
Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores
Haoran Jisun
TL;DR
The paper asks whether decodable empathy directions reliably control automated empathy scores and human-perceived change, rather than merely revealing detectable representations. It tests EPITOME Recognition and Resonance across three instruction-tuned LLMs using multiple judges, a classifier, interventions, and positive controls. Resonance yields a model-dependent, sub-maximal automated affective shift, while Recognition steering is unmeasurable under an insensitive cognitive instrument; Gemma Recognition ablation nevertheless lowers the classifier’s cognitive score.
Problem
Decodability, automated-metric control, and human-perceived empathy change are often conflated, despite uncertainty about whether decodable directions provide reliable causal control.
Method
The study probes and intervenes on EPITOME Recognition and Resonance directions in three instruction-tuned LLMs, evaluating outputs with two LLM judges, an EPITOME classifier, human checks, and positive controls.
Results
+0.29: Qwen Resonance at α=0→+8 raises the classifier-measured affective score, while direct contrasts show facet specificity in Qwen and Llama and no Recognition steering contrast survives.
Takeaways & Limitations
Decodable empathy directions do not imply reliable global control; Resonance provides only partial automated affective control, and cognitive claims require measurement-sensitivity checks.
Takeaways & Limitations
The study is limited by modest power, inconsistent cognitive measurement, single-turn interventions, and a context confound that residualization does not eliminate.
Abstract
from arXiv · showhide
A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do not, however, establish a matching human-perceived change. Additive cognitive steering produces no measurable change, but a within-domain control shows the cognitive instrument is too coarse to resolve the differences such steering would produce -- unmeasurable, not a clean null. By contrast, Gemma Recognition ablation lowers the classifier's cognitive score even after adjusting for response length. Detection does not imply reliable control under global interventions, and cognitive-empathy claims warrant an explicit measurement-sensitivity check.
1 Introduction
The paper tests whether decodable empathy directions provide reliable causal control, separating cognitive Recognition from affective Resonance in text-expressed empathy. Across automated measures, affective control is partial and cognitive nulls require sensitivity checks.
- Motivation: Decodability is insufficient for causal claims because probe accuracy need not indicate genuine encoding, model use, or reliable steering.Prior work also reports brittle steering vectors and interpretability illusions.
- Motivation: Empathy is tested as a high-stakes, multidimensional construct with cognitive Recognition and affective Resonance facets.Recognition denotes communicating inferred understanding of a seeker’s experience, while Resonance concerns affective feeling-related responses.
- Findings: Recognition and Resonance remain decodable after residualizing against a sentence-embedding-derived surface score.This establishes detection beyond the tested surface-text control, not reliable intervention control.
- Findings: Affective steering is a graded, sub-maximal automated-metric shift: Qwen Resonance at α=0→+8 rises +0.29, approximately 26% of the natural gap.The shift is facet-differentiated in Qwen and Llama, but it is measured by classifiers rather than established in human perception.
- Findings: Additive Recognition steering is unmeasurable because a within-domain control shows insufficient cognitive-instrument sensitivity, whereas Gemma Recognition ablation lowers the classifier score.The resulting cognitive additive null is not established as a clean absence of effect.
2 Method
The method probes EPITOME-derived empathy facets, removes surface-text confounds, and tests directional interventions with multiple automated instruments and controls. It combines additive steering, production-side directions, ablation, and output-change diagnostics to distinguish intervention potency from metric movement.
- Probes: Probe directions are differences between high- and low-level EPITOME class-mean residual-stream activations, evaluated layerwise with Cohen’s d.For each facet, the procedure samples 100 Level 2 and 100 Level 0 items in Client/Counselor exchanges and extracts last-token, pre-generation activations.
- Surface-text residualization: Surface-text residualization regresses probe projections on a cross-validated sentence-embedding surface score and recomputes residualized effect sizes.Each facet’s peak layer is selected by the maximum residualized d, making residualization a detection control.
- Causal steering: At each facet’s peak layer, additive steering scales the probe direction by its projection standard deviation across α ∈ [−8, 8] and generates deterministic responses.The intervention uses 50 prompts per condition, including emotional and neutral prompts, while a random-direction floor addresses differences in raw perturbation magnitude.
- Evaluation: A blind GPT-4o judge rates cognitive and affective empathy separately, while concept-specific effects must exceed the other dimension and a paired random-direction floor.The analysis uses bootstrap 95% confidence intervals, Bonferroni-corrected lexical analyses, and degeneration checks.
- Production-side direction: The study also estimates a response-token Recognition direction, compares its cosine with the reading-time probe direction, and steers at its own peak layer.This separates directions derived from model reading-time activations from those derived from generated response-token activations.
- Directional ablation: Directional ablation removes the peak-layer unit direction from every layer and token position during generation, with matched random-direction effects tested using paired sign-flip permutation and Wilcoxon tests.The residual stream is treated as one shared space across layers, and the random-direction floor aggregates five directions per prompt.
- Output-change diagnostic: Output-change diagnostics compare α=0 and α=+8 generations using MPNet cosine, random- and cross-prompt floors, and a unique-token-ratio degeneration threshold.Lower cosine indicates larger text change, while collapse is flagged below a unique-token ratio of 0.5.
- Instruments and controls: Every intervention is evaluated with two additional automated instruments and an emotional-vs-neutral positive control, expanded to 40 neutral prompts for cognitive sensitivity.The cognitive control’s original 10 neutral prompts implied a minimum detectable d of approximately 1.0, motivating 30 added Dolly factual prompts.
3 Results
Both empathy facets are decodable, but global steering reliably controls neither facet at full natural range: Resonance produces a partial affective shift, whereas Recognition remains unmeasurable under additive steering. A coarse-scale Gemma Recognition ablation lowers classifier cognitive scores, while human-perceived change is not established.
- Detection: Both Recognition and Resonance remain decodable after surface-text residualization, with model-dependent geometric relations between directions.Residualized effect sizes remain substantial in every model; Recognition peaks later than Resonance, and their geometry is descriptive rather than causal.
- Cognitive steering: Recognition steering substantially changes text while leaving cognitive scores flat, but within-domain controls show the cognitive instrument cannot resolve the relevant differences.The additive cognitive result is therefore unmeasurable rather than a clean null; Qwen lacks cognitive classifier range, while Llama and Gemma show only coarse range.
- Affective steering: +0.29: Qwen’s classifier-measured affective score rises from α=0 to +8, approximately 26% of its natural emotional-vs-neutral gap.The corresponding gains are +0.11 for Llama and +0.05 for Gemma; the full-span gains are +0.43/+0.25/+0.06.
- Affective steering: Resonance produces an ordered, facet-specific affective shift in Qwen and Llama, but not Gemma, with no Recognition contrast surviving.The direct between-direction contrast is the primary specificity test; Gemma’s contrast is −0.01 with p=0.83.
- Interpretation: Automated affective shifts cannot be identified as human-perceived empathy changes because the EPITOME-ER classifier is built from EPITOME-ER labels.The paper states that separating added emotional register from added empathy requires human evaluation.
- Ablation: Gemma Recognition ablation lowers the classifier cognitive score from 0.31 to 0.11, and the drop survives adjustment for response length.The adjusted classifier effect is −0.19 with a 95% CI of [−0.35, −0.03], while affective scores remain at ceiling.
4 Related Work and Positioning
The paper places its findings within work showing that empathy-related directions can be decoded or steered, while emphasizing that global steering may provide only partial control. Its positive-control analysis also reveals substantially stronger affective than cognitive measurement reliability.
- Prior work links affective concepts to decodable directions and activation steering, but probe accuracy alone does not establish genuine encoding or model use.
- The paper distinguishes EPITOME communicative empathy under global interventions from related work on behavioral or localized empathy steering.Cadile studies an “empathy-in-action” direction, while Chebrolu et al. and Tak et al. report localized interventions.
- Affective classifier scores cleanly separate emotional from neutral prompts in all three models, with Cohen’s d=3.1–4.8.The affective positive control uses 40 emotional and 40 neutral responses.
- Cognitive classifier distributions overlap heavily: Llama and Gemma are moderately separable, whereas Qwen is exactly coincident.Cognitive d values are 0.62 for Llama, 0.67 for Gemma, and 0.00 for Qwen.
- The two LLM judges likewise agree more reliably on affective than cognitive scores, supporting measurement checks before interpreting cognitive null results.Inter-judge Krippendorff α is 0.70–0.84 for affective scores versus 0.24/0.44/−0.02 for cognitive scores across Llama, Gemma, and Qwen.
5 Conclusion
The conclusion separates decodability, automated-metric control, and human-perceived change. Recognition steering is unmeasurable with the available cognitive instrument, while Resonance produces only model-dependent partial automated affective shifts.
- Recognition and Resonance directions remain decodable after residualizing against a surface score, but decodability does not yield reliable global control.
- Resonance steering produces a model-dependent, sub-maximal shift in an automated affective metric, whereas additive Recognition steering is unmeasurable.
- Only Gemma shows a length-adjusted ablation drop in the classifier’s cognitive score.
- The paper does not establish a matching human-perceived empathy change, and a six-rater panel could not adjudicate the Qwen α=0 versus +8 contrast.
Limitations
The main limitations are modest statistical power, binding measurement constraints, bounded human and within-domain checks, a context confound, and single-turn scope. These boundaries limit interpretation of cognitive nulls and generalization to deployed dialogue.
- Each facet uses modest EPITOME item counts, and small cross-model ablation samples leave smaller necessity effects unresolved.
- Cognitive range is inconsistent across instruments, and the classifier cannot resolve the within-domain differences probed by additive steering.Accordingly, additive cognitive steering is interpreted as unmeasurable rather than null.
- Human and within-domain checks are bounded by hand-constructed items and one constructor’s construal of “understanding.”
- The affective effect is established only on automated instruments that do not fully agree, so it is framed as control of a steerable automated metric rather than human-perceived empathy.
- EPITOME labels the counselor turn while analysis uses Client+Counselor context, leaving a residualized but unresolved context confound.
- All interventions and measurements concern one generated response to one prompt, leaving multi-turn and long-context steerability unestablished.
Ethics Statement
The study documents its data, classifier implementation, evaluation instruments, and intended-use constraints, emphasizing that automated empathy shifts should not be treated as demonstrated human-perceived change.
- Six adult volunteers participated in a 10–15-minute pairwise-comparison text-rating task after reviewing consent information.
- The study uses emotional-support prompts from ESConv, factual prompts from databricks-dolly-15k, and EPITOME empathy annotations.
- The paper cautions against deploying activation-level empathy steering in emotional-support settings without evidence of reliable human-perceived change.The authors specifically warn that automated metrics may move without demonstrated human-perceived change.
- The retrained EPITOME classifier is a Sharma et al. bi-encoder with an empathy-identification head, trained on a stratified 80/20 held-out split.The test set contains n=617 items per head, using the same scoring configuration.
- The cognitive IP head predicts level 1 for 0/617 test items, making it effectively binary and limiting resolution for within-domain cognitive contrasts.Its macro-F1 is 0.558, while accuracy is 0.822 on the dominant 0/2 levels.
- The positive control shows strong affective separation across instruments, but cognitive separation is inconsistent across judges and the classifier.Affective Cohen’s d is d≥2.6; cognitive results vary by instrument and model, with marginal classifier results for Llama and Gemma.
E Sample sizes and pilot attenuation
Sample sizes determine which effects the experiments can detect, and larger-sample retests showed attenuation of pilot effects.
- Cross-model ablation uses n=15, which can detect only effects ≳1 point at the observed standard deviations.Gemma Recognition is verified at n=50, with minimum detectable dz≈0.40, corresponding to ≈0.16 points at the within-prompt difference SD ≈0.39.
F Human panel: a sensitivity-gated check
The human panel failed its sensitivity gate, so its null-looking steered-output comparisons cannot establish zero human-perceived change.
- Six volunteer raters first had to reliably order hand-built calibration pairs before steered-output nulls could be interpreted.
- 21/47 committed understanding trials selected the higher-understanding response, also near chance with binomial p≈0.56.No rater cleared the understanding gate individually.
- The steered pairs showed no consistent direction, with 16 baseline, 19 “same,” and 13 steered responses.The sign test was p=0.71, but the failed sensitivity gate makes this result uninterpretable as a null.
G Specificity controls: cross-facet matrix and raw-norm floor
Cross-facet and perturbation-magnitude controls indicate that Resonance selectively affects the affective metric, while Recognition changes text without a corresponding generic score inflation.
- Cross-facet steering: The cross-facet matrix tests both steering directions against both classifier heads using paired α=0→+8 changes.This design distinguishes a matched-facet effect from a generic increase across EPITOME-related scores.
- Cross-facet steering: Adding Recognition never raises ER, whereas adding Resonance raises ER and lowers IP in Qwen and Llama, indicating facet-differentiated rather than generic inflation.Gemma shows no reliable IP change, and Qwen Recognition→IP is uninterpretable because Qwen has no cognitive range.
- Raw-norm floor: At matched raw L2 norm, Recognition text change is comparable to random-direction means, with MPNet cosine 0.86/0.78/0.78 versus 0.84/0.82/0.79.The affective score remains direction-specific at matched norm: only Resonance clears the random floor.
H Response-token direction and all-layer ablation
Response-token directions and all-layer ablation largely preserve the detection–control dissociation, with no large effects across models except Gemma Recognition ablation.
- H Response-token direction and all-layer ablation: Response-token directions are near-orthogonal to read-time directions and show no large steering or ablation effects across the three models.The sole departure is Gemma Recognition ablation.
- H Response-token direction and all-layer ablation: Gemma Recognition is the sole reported departure from the generally weak response-token and all-layer intervention effects.Table 6 summarizes this exception alongside the broader dissociation.
I Length control for the Gemma ablation
Length adjustment separates the Gemma Recognition ablation’s judge and classifier effects: the judge-side cognitive drop largely disappears, while the classifier drop remains.
- I Length control for the Gemma ablation: −0.40 to −0.13: the GPT-4o judge’s cognitive drop becomes statistically nonsignificant after controlling for length and prompt.The judge rewards length by +0.0067 points/word, while the affective drop changes from −0.24 to −0.003.
- I Length control for the Gemma ablation: −0.199 to −0.193: the classifier’s cognitive drop remains after length adjustment, with a clustered-SE 95% CI of [−0.35, −0.03].The classifier’s length slope is approximately −0.003 E[level]/word and nonsignificant under prompt-clustered standard errors.
- I Length control for the Gemma ablation: The two instruments dissociate sharply after adjusting for completion length and prompt fixed effects.Length is treated as downstream of the intervention and therefore as a potential mediator.