Source-linked AI summary

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

Haoran Jisun

arXiv:2608.24901v1cs.CLcs.HCcs.LG

TL;DR

The paper asks whether decodable empathy directions reliably control automated empathy scores and human-perceived change, rather than merely revealing detectable representations. It tests EPITOME Recognition and Resonance across three instruction-tuned LLMs using multiple judges, a classifier, interventions, and positive controls. Resonance yields a model-dependent, sub-maximal automated affective shift, while Recognition steering is unmeasurable under an insensitive cognitive instrument; Gemma Recognition ablation nevertheless lowers the classifier’s cognitive score.

  • Problem

    Decodability, automated-metric control, and human-perceived empathy change are often conflated, despite uncertainty about whether decodable directions provide reliable causal control.

  • Method

    The study probes and intervenes on EPITOME Recognition and Resonance directions in three instruction-tuned LLMs, evaluating outputs with two LLM judges, an EPITOME classifier, human checks, and positive controls.

  • Results

    +0.29: Qwen Resonance at α=0→+8 raises the classifier-measured affective score, while direct contrasts show facet specificity in Qwen and Llama and no Recognition steering contrast survives.

  • Takeaways & Limitations

    Decodable empathy directions do not imply reliable global control; Resonance provides only partial automated affective control, and cognitive claims require measurement-sensitivity checks.

  • Takeaways & Limitations

    The study is limited by modest power, inconsistent cognitive measurement, single-turn interventions, and a context confound that residualization does not eliminate.

Abstract

from arXiv · show

A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do not, however, establish a matching human-perceived change. Additive cognitive steering produces no measurable change, but a within-domain control shows the cognitive instrument is too coarse to resolve the differences such steering would produce -- unmeasurable, not a clean null. By contrast, Gemma Recognition ablation lowers the classifier's cognitive score even after adjusting for response length. Detection does not imply reliable control under global interventions, and cognitive-empathy claims warrant an explicit measurement-sensitivity check.

1 Introduction

The paper tests whether decodable empathy directions provide reliable causal control, separating cognitive Recognition from affective Resonance in text-expressed empathy. Across automated measures, affective control is partial and cognitive nulls require sensitivity checks.

  • Motivation: Decodability is insufficient for causal claims because probe accuracy need not indicate genuine encoding, model use, or reliable steering.Prior work also reports brittle steering vectors and interpretability illusions.
  • Motivation: Empathy is tested as a high-stakes, multidimensional construct with cognitive Recognition and affective Resonance facets.Recognition denotes communicating inferred understanding of a seeker’s experience, while Resonance concerns affective feeling-related responses.
  • Findings: Recognition and Resonance remain decodable after residualizing against a sentence-embedding-derived surface score.This establishes detection beyond the tested surface-text control, not reliable intervention control.
  • Findings: Affective steering is a graded, sub-maximal automated-metric shift: Qwen Resonance at α=0→+8 rises +0.29, approximately 26% of the natural gap.The shift is facet-differentiated in Qwen and Llama, but it is measured by classifiers rather than established in human perception.
  • Findings: Additive Recognition steering is unmeasurable because a within-domain control shows insufficient cognitive-instrument sensitivity, whereas Gemma Recognition ablation lowers the classifier score.The resulting cognitive additive null is not established as a clean absence of effect.

2 Method

The method probes EPITOME-derived empathy facets, removes surface-text confounds, and tests directional interventions with multiple automated instruments and controls. It combines additive steering, production-side directions, ablation, and output-change diagnostics to distinguish intervention potency from metric movement.

  • Probes: Probe directions are differences between high- and low-level EPITOME class-mean residual-stream activations, evaluated layerwise with Cohen’s d.For each facet, the procedure samples 100 Level 2 and 100 Level 0 items in Client/Counselor exchanges and extracts last-token, pre-generation activations.
  • Surface-text residualization: Surface-text residualization regresses probe projections on a cross-validated sentence-embedding surface score and recomputes residualized effect sizes.Each facet’s peak layer is selected by the maximum residualized d, making residualization a detection control.
  • Causal steering: At each facet’s peak layer, additive steering scales the probe direction by its projection standard deviation across α ∈ [−8, 8] and generates deterministic responses.The intervention uses 50 prompts per condition, including emotional and neutral prompts, while a random-direction floor addresses differences in raw perturbation magnitude.
  • Evaluation: A blind GPT-4o judge rates cognitive and affective empathy separately, while concept-specific effects must exceed the other dimension and a paired random-direction floor.The analysis uses bootstrap 95% confidence intervals, Bonferroni-corrected lexical analyses, and degeneration checks.
  • Production-side direction: The study also estimates a response-token Recognition direction, compares its cosine with the reading-time probe direction, and steers at its own peak layer.This separates directions derived from model reading-time activations from those derived from generated response-token activations.
  • Directional ablation: Directional ablation removes the peak-layer unit direction from every layer and token position during generation, with matched random-direction effects tested using paired sign-flip permutation and Wilcoxon tests.The residual stream is treated as one shared space across layers, and the random-direction floor aggregates five directions per prompt.
  • Output-change diagnostic: Output-change diagnostics compare α=0 and α=+8 generations using MPNet cosine, random- and cross-prompt floors, and a unique-token-ratio degeneration threshold.Lower cosine indicates larger text change, while collapse is flagged below a unique-token ratio of 0.5.
  • Instruments and controls: Every intervention is evaluated with two additional automated instruments and an emotional-vs-neutral positive control, expanded to 40 neutral prompts for cognitive sensitivity.The cognitive control’s original 10 neutral prompts implied a minimum detectable d of approximately 1.0, motivating 30 added Dolly factual prompts.

3 Results

Both empathy facets are decodable, but global steering reliably controls neither facet at full natural range: Resonance produces a partial affective shift, whereas Recognition remains unmeasurable under additive steering. A coarse-scale Gemma Recognition ablation lowers classifier cognitive scores, while human-perceived change is not established.

  • Detection: Both Recognition and Resonance remain decodable after surface-text residualization, with model-dependent geometric relations between directions.Residualized effect sizes remain substantial in every model; Recognition peaks later than Resonance, and their geometry is descriptive rather than causal.
  • Cognitive steering: Recognition steering substantially changes text while leaving cognitive scores flat, but within-domain controls show the cognitive instrument cannot resolve the relevant differences.The additive cognitive result is therefore unmeasurable rather than a clean null; Qwen lacks cognitive classifier range, while Llama and Gemma show only coarse range.
  • Affective steering: +0.29: Qwen’s classifier-measured affective score rises from α=0 to +8, approximately 26% of its natural emotional-vs-neutral gap.The corresponding gains are +0.11 for Llama and +0.05 for Gemma; the full-span gains are +0.43/+0.25/+0.06.
  • Affective steering: Resonance produces an ordered, facet-specific affective shift in Qwen and Llama, but not Gemma, with no Recognition contrast surviving.The direct between-direction contrast is the primary specificity test; Gemma’s contrast is −0.01 with p=0.83.
  • Interpretation: Automated affective shifts cannot be identified as human-perceived empathy changes because the EPITOME-ER classifier is built from EPITOME-ER labels.The paper states that separating added emotional register from added empathy requires human evaluation.
  • Ablation: Gemma Recognition ablation lowers the classifier cognitive score from 0.31 to 0.11, and the drop survives adjustment for response length.The adjusted classifier effect is −0.19 with a 95% CI of [−0.35, −0.03], while affective scores remain at ceiling.

4 Related Work and Positioning

The paper places its findings within work showing that empathy-related directions can be decoded or steered, while emphasizing that global steering may provide only partial control. Its positive-control analysis also reveals substantially stronger affective than cognitive measurement reliability.

  • Prior work links affective concepts to decodable directions and activation steering, but probe accuracy alone does not establish genuine encoding or model use.
  • The paper distinguishes EPITOME communicative empathy under global interventions from related work on behavioral or localized empathy steering.Cadile studies an “empathy-in-action” direction, while Chebrolu et al. and Tak et al. report localized interventions.
  • Affective classifier scores cleanly separate emotional from neutral prompts in all three models, with Cohen’s d=3.1–4.8.The affective positive control uses 40 emotional and 40 neutral responses.
  • Cognitive classifier distributions overlap heavily: Llama and Gemma are moderately separable, whereas Qwen is exactly coincident.Cognitive d values are 0.62 for Llama, 0.67 for Gemma, and 0.00 for Qwen.
  • The two LLM judges likewise agree more reliably on affective than cognitive scores, supporting measurement checks before interpreting cognitive null results.Inter-judge Krippendorff α is 0.70–0.84 for affective scores versus 0.24/0.44/−0.02 for cognitive scores across Llama, Gemma, and Qwen.

5 Conclusion

The conclusion separates decodability, automated-metric control, and human-perceived change. Recognition steering is unmeasurable with the available cognitive instrument, while Resonance produces only model-dependent partial automated affective shifts.

  • Recognition and Resonance directions remain decodable after residualizing against a surface score, but decodability does not yield reliable global control.
  • Resonance steering produces a model-dependent, sub-maximal shift in an automated affective metric, whereas additive Recognition steering is unmeasurable.
  • Only Gemma shows a length-adjusted ablation drop in the classifier’s cognitive score.
  • The paper does not establish a matching human-perceived empathy change, and a six-rater panel could not adjudicate the Qwen α=0 versus +8 contrast.

Limitations

The main limitations are modest statistical power, binding measurement constraints, bounded human and within-domain checks, a context confound, and single-turn scope. These boundaries limit interpretation of cognitive nulls and generalization to deployed dialogue.

  • Each facet uses modest EPITOME item counts, and small cross-model ablation samples leave smaller necessity effects unresolved.
  • Cognitive range is inconsistent across instruments, and the classifier cannot resolve the within-domain differences probed by additive steering.Accordingly, additive cognitive steering is interpreted as unmeasurable rather than null.
  • Human and within-domain checks are bounded by hand-constructed items and one constructor’s construal of “understanding.”
  • The affective effect is established only on automated instruments that do not fully agree, so it is framed as control of a steerable automated metric rather than human-perceived empathy.
  • EPITOME labels the counselor turn while analysis uses Client+Counselor context, leaving a residualized but unresolved context confound.
  • All interventions and measurements concern one generated response to one prompt, leaving multi-turn and long-context steerability unestablished.

Ethics Statement

The study documents its data, classifier implementation, evaluation instruments, and intended-use constraints, emphasizing that automated empathy shifts should not be treated as demonstrated human-perceived change.

  • Six adult volunteers participated in a 10–15-minute pairwise-comparison text-rating task after reviewing consent information.
  • The study uses emotional-support prompts from ESConv, factual prompts from databricks-dolly-15k, and EPITOME empathy annotations.
  • The paper cautions against deploying activation-level empathy steering in emotional-support settings without evidence of reliable human-perceived change.The authors specifically warn that automated metrics may move without demonstrated human-perceived change.
  • The retrained EPITOME classifier is a Sharma et al. bi-encoder with an empathy-identification head, trained on a stratified 80/20 held-out split.The test set contains n=617 items per head, using the same scoring configuration.
  • The cognitive IP head predicts level 1 for 0/617 test items, making it effectively binary and limiting resolution for within-domain cognitive contrasts.Its macro-F1 is 0.558, while accuracy is 0.822 on the dominant 0/2 levels.
  • The positive control shows strong affective separation across instruments, but cognitive separation is inconsistent across judges and the classifier.Affective Cohen’s d is d≥2.6; cognitive results vary by instrument and model, with marginal classifier results for Llama and Gemma.

E Sample sizes and pilot attenuation

Sample sizes determine which effects the experiments can detect, and larger-sample retests showed attenuation of pilot effects.

  • Cross-model ablation uses n=15, which can detect only effects ≳1 point at the observed standard deviations.Gemma Recognition is verified at n=50, with minimum detectable dz≈0.40, corresponding to ≈0.16 points at the within-prompt difference SD ≈0.39.

F Human panel: a sensitivity-gated check

The human panel failed its sensitivity gate, so its null-looking steered-output comparisons cannot establish zero human-perceived change.

  • Six volunteer raters first had to reliably order hand-built calibration pairs before steered-output nulls could be interpreted.
  • 21/47 committed understanding trials selected the higher-understanding response, also near chance with binomial p≈0.56.No rater cleared the understanding gate individually.
  • The steered pairs showed no consistent direction, with 16 baseline, 19 “same,” and 13 steered responses.The sign test was p=0.71, but the failed sensitivity gate makes this result uninterpretable as a null.

G Specificity controls: cross-facet matrix and raw-norm floor

Cross-facet and perturbation-magnitude controls indicate that Resonance selectively affects the affective metric, while Recognition changes text without a corresponding generic score inflation.

  • Cross-facet steering: The cross-facet matrix tests both steering directions against both classifier heads using paired α=0→+8 changes.This design distinguishes a matched-facet effect from a generic increase across EPITOME-related scores.
  • Cross-facet steering: Adding Recognition never raises ER, whereas adding Resonance raises ER and lowers IP in Qwen and Llama, indicating facet-differentiated rather than generic inflation.Gemma shows no reliable IP change, and Qwen Recognition→IP is uninterpretable because Qwen has no cognitive range.
  • Raw-norm floor: At matched raw L2 norm, Recognition text change is comparable to random-direction means, with MPNet cosine 0.86/0.78/0.78 versus 0.84/0.82/0.79.The affective score remains direction-specific at matched norm: only Resonance clears the random floor.

H Response-token direction and all-layer ablation

Response-token directions and all-layer ablation largely preserve the detection–control dissociation, with no large effects across models except Gemma Recognition ablation.

  • H Response-token direction and all-layer ablation: Response-token directions are near-orthogonal to read-time directions and show no large steering or ablation effects across the three models.The sole departure is Gemma Recognition ablation.
  • H Response-token direction and all-layer ablation: Gemma Recognition is the sole reported departure from the generally weak response-token and all-layer intervention effects.Table 6 summarizes this exception alongside the broader dissociation.

I Length control for the Gemma ablation

Length adjustment separates the Gemma Recognition ablation’s judge and classifier effects: the judge-side cognitive drop largely disappears, while the classifier drop remains.

  • I Length control for the Gemma ablation: −0.40 to −0.13: the GPT-4o judge’s cognitive drop becomes statistically nonsignificant after controlling for length and prompt.The judge rewards length by +0.0067 points/word, while the affective drop changes from −0.24 to −0.003.
  • I Length control for the Gemma ablation: −0.199 to −0.193: the classifier’s cognitive drop remains after length adjustment, with a clustered-SE 95% CI of [−0.35, −0.03].The classifier’s length slope is approximately −0.003 E[level]/word and nonsignificant under prompt-clustered standard errors.
  • I Length control for the Gemma ablation: The two instruments dissociate sharply after adjusting for completion length and prompt fixed effects.Length is treated as downstream of the intervention and therefore as a potential mediator.
Loading 2608.24901v1…