Source-linked AI summary

Emotional Labor Strategy Preferences in LLM Personas

Mohammad Saim, Tianyu Jiang

arXiv:2609.00310v1cs.CL

TL;DR

Research on personality and emotional labor has relied mainly on self-report scales in occupational settings, leaving everyday social scenarios understudied. This paper evaluates psychometrically grounded LLM personas with a 500-scenario dataset offering three emotional labor choices. Models favor deep acting, with Conscientiousness and Emotional Stability consistently predicting that preference.

  • Problem

    Research on personality-linked emotional labor strategies relies largely on self-report scales and occupational samples, leaving everyday social interactions insufficiently tested.

  • Method

    The study evaluates five LLMs using 50 fictional-character personas profiled through observer-rated adjective composites and in-character IPIP-50 self-reports across 500 scenarios with three strategy choices.

  • Results

    Models prefer deep acting, while Conscientiousness and Emotional Stability consistently predict strategy preferences across both persona tracks.

  • Takeaways & Limitations

    Personality injection reliably differentiates emotional labor strategy selection in LLM personas across everyday social scenarios.

  • Takeaways & Limitations

    Findings are limited by fictional-character profiles based on crowd-sourced perceptions, synthetic scenario augmentation, English-only culturally situated data, and correlational evidence.

Abstract

from arXiv · show

Emotional labor is the effortful management of emotional displays to meet social or professional expectations. Personality traits have been correlated with emotional labor strategies, yet research on this link relies almost exclusively on self-report scales administered only in occupational settings. We investigate whether large language models injected with psychometrically grounded personas reproduce these personality-driven selection patterns across everyday social scenarios. We construct the first emotional labor strategy dataset of 500 socially situated events, each offering three behavioral choices corresponding to surface acting, deep acting, and genuine expression. We source 50 fictional characters from a large-scale personality repository and profile each through two parallel tracks: observer-rated bipolar adjective composites and in-character self-report items. Five LLMs evaluate all scenarios under both persona conditions. We find that models align more towards deep acting, and that Conscientiousness and Emotional Stability consistently predict this preference. Entropy analysis confirms that persona reliably influences the output and varies across models and emotions.

1 Introduction

The paper examines whether personality-conditioned LLM personas select emotional labor strategies in everyday social scenarios, addressing gaps in self-report and occupationally limited research. It introduces a 500-scenario dataset and finds deep acting preference, with Conscientiousness and Emotional Stability as consistent predictors.

  • Motivation: Emotional labor comprises surface acting, deep acting, and genuine expression, distinguished by whether regulation changes outward display, felt emotion, or neither.Surface acting modifies expression without changing felt emotion; deep acting reappraises the felt emotion; genuine expression requires no regulatory effort when felt and expected emotions match.
  • Research gap: Existing research links personality traits to emotional labor strategies but relies largely on self-report measures in healthcare and customer-service settings.The paper motivates testing these patterns in everyday social contexts rather than only occupational interactions.
  • Research gap: Personality-injected LLMs can exhibit trait-aligned lexical, sentiment, and judgment patterns, enabling controlled simulation of psychological variability.This prior work provides the basis for studying personality-conditioned emotional labor choices in language models.
  • Research gap: The study addresses an unexamined intersection by testing how personality-injected LLM personas choose emotional labor strategies in socially situated scenarios.The authors connect emotional labor theory with LLM persona research because trait-driven choices may introduce behavioral variance in affect-sensitive applications.
  • Contributions: The authors construct and release the first full-scale emotional labor strategy dataset, comprising 500 everyday scenarios with choices among three strategies.The study uses fictional characters with personality ratings and compares observer-rated profiles with in-character self-reports.
  • Findings: Models prefer deep acting, while Conscientiousness and Emotional Stability consistently predict strategy preferences across both persona tracks.This is the paper’s principal reported result across the evaluated models and persona conditions.

2 Related Works

Prior work established emotional labor’s three-strategy taxonomy and showed that LLM personas can express injected personality traits, but these literatures had not been combined. This paper fills that gap by evaluating trait-conditioned strategy selection in socially situated emotional labor scenarios.

  • Emotions in NLP and Emotional Labor: Emotion research in NLP has progressed from categorical labels toward appraisal-based frameworks that model emotion through cognitive evaluations of events.The cited appraisal approach considers dimensions such as novelty and relevance.
  • Emotions in NLP and Emotional Labor: Emotional labor measurement has relied mainly on self-report instruments administered to workers in service roles.The literature lacks a corpus extending the three-strategy taxonomy beyond narrow professional contexts into everyday social interactions.
  • Emotions in NLP and Emotional Labor: The established emotional labor taxonomy distinguishes surface acting, deep acting, and naturally felt expression as separate strategies.The three-factor structure used in this study treats naturally felt expression as distinct from surface and deep acting.
  • Persona Injection and Personality in LLMs: Persona-injection research finds that LLMs can produce self-reports and writing samples aligned with assigned Big Five traits across models.This literature supports treating injected personas as a controllable source of personality variation.
  • Persona Injection and Personality in LLMs: OCEAN-grounded LLM agents have reproduced personality-linked behavioral patterns, including cooperative effects associated with Agreeableness in negotiations.Prior findings also associate Agreeableness and Extraversion with deep acting and genuine expression, and Neuroticism with surface acting.
  • Persona Injection and Personality in LLMs: No prior work had examined which trait-conditioned LLMs select as emotional regulation strategies in socially situated emotional labor scenarios.The paper combines OCEAN-grounded personas with emotional-labor-annotated stimuli to test whether choices mirror personality research.

3 Methodology

The methodology builds a socially situated emotional labor strategy dataset and evaluates fictional-character personas through two complementary personality-profiling tracks across five LLMs.

  • 3.1 Emotion Labor Strategy Dataset: The ELS dataset filters 500 sentences from an emotion appraisal corpus spanning anger, disgust, fear, guilt, joy, sadness, and shame.The original sentences are augmented with social agents or contexts to situate emotional labor interpersonal encounters.
  • 3.1 Emotion Labor Strategy Dataset: Each scenario offers behavioral choices representing surface acting, deep acting, and genuine expression.The choices are generated as natural sentence completions and manually validated for coherence, strategy fidelity, and distinguishability.
  • 3.1 Emotion Labor Strategy Dataset: The dataset uses constrained generation to express emotional states through behavioral, physical, or physiological details rather than emotion labels.Prompt constraints include a single leakage cue for Surface Acting and behavioral descriptions for strategy differentiation.
  • 3.2 Persona Construction: Persona profiles are constructed in parallel tracks to compare observer-rated and self-reported personality representations for the same characters.The two-track design addresses possible inconsistency from relying only on pretrained character representations and tests alignment across representations.
  • 3.2 Persona Construction: The observer-rated track retains 40 BAP markers validated against Goldberg’s taxonomy and selects 50 personality-differentiated fictional characters using farthest-point sampling.Characters are first ranked by BAP-score variance, then sampled in the 40-dimensional BAP space.
  • 3.3 Evaluation and Model Selection: The self-report track administers the 50-item IPIP-50 from each character’s perspective, using five LLMs to choose SA, DA, or GE across the 500 ELS scenarios.The five models are Qwen-3-8B, Qwen-3-32B, GPT-5.4, Deepseek-V4-Flash, and Gemma-4-31B.

4 Results and Analysis

Across five models and two persona tracks, deep acting is the dominant emotional-labor strategy, while personality traits systematically shape strategy selection. Persona effects vary by model and emotion, with Conscientiousness and Emotional Stability showing the most consistent associations.

  • Aggregate ELS Distributions Across Models and Persona Tracks: Deep acting is the modal strategy in eight of ten model-track combinations, ranging from approximately 39% to 61%.
  • Aggregate ELS Distributions Across Models and Persona Tracks: BAP and IPIP-50 preserve the same ordinal strategy ranking in most cases, but IPIP-50 shifts some predictions from deep acting toward genuine expression.For GPT-5.4, deep acting decreases from approximately 51% to 39%, while genuine expression rises from approximately 30% to 40%.
  • Trait–Strategy Correlations: Conscientiousness and Emotional Stability consistently predict more deep acting and less surface acting and genuine expression across both persona tracks.For Conscientiousness in the BAP track, correlations are ρ = −0.446 for SA, +0.597 for DA, and −0.570 for GE.
  • Trait–Strategy Correlations: Agreeableness shows significant strategy correlations only in the IPIP track, while Openness has none and Extraversion has one marginal deep-acting effect.IPIP correlations are ρ = −0.398 for A–SA and ρ = +0.372 for A–DA, both p < .01.
  • Trait–Strategy Correlations: A forced three-way choice and Likert-style rating replicate broadly similar trait–strategy directions, supporting the robustness of personality-linked preferences.
  • Trait–Strategy Correlations: Trait composites show moderate-to-high BAP reliability and uniformly high IPIP-50 reliability, including large Emotional Stability correlations with deep acting and genuine expression.BAP reliability is α = 0.70–0.95; reported IPIP correlations include ES–DA ρ = +0.850 and ES–GE ρ = −0.883.
  • Persona Influence and Reliability: Entropy analysis indicates that personality meaningfully differentiates strategy selection for most models, although emotion and scenario effects are stronger for Qwen-32B.Joy, Fear, and Anger show entropy values of 0.82–0.92, while Qwen-32B ranges from 0.42 for Sadness to 0.70 for Disgust.
  • Is Persona Influential Enough?: Without persona information, deep acting remains the majority strategy and surface acting falls to 10%; persona signals redistribute this default.Character-only prompting nearly halves deep-acting predictions, whereas BAP-only prompting yields 36.9% deep acting and higher genuine expression.

5 Conclusion

The study shows that personality-injected LLM personas reliably differ in emotional labor strategy selection across everyday social scenarios. Across five models and two persona tracks, Conscientiousness and Emotional Stability consistently predict preferences, while Agreeableness appears only through self-report induction.

  • Personality-injected LLM personas reliably differ in emotional labor strategy selection across everyday social scenarios.
  • Across five models and two persona tracks, Conscientiousness and Emotional Stability consistently drive strategy preferences.
  • Agreeableness surfaces as a predictor only through self-report persona induction.
  • LLMs prefer deep acting, suggesting that training corpora encode it as a normative regulatory response.
  • Conscientiousness suppresses genuine expression in LLM personas, reversing the positive association reported in previous human-evaluated studies.

Limitations

The findings are constrained by fictional-character personas, synthetic scenario augmentation, English-language cultural scope, and correlational evidence. These boundaries limit direct generalization to real emotional labor encounters and other languages or cultural norms.

  • Personas rely on crowd-sourced fictional-character ratings rather than ground-truth personality measures.
  • Synthetic augmentation may not fully capture the complexity of real emotional labor encounters.
  • English-only, culturally situated scenarios leave cross-language and cross-cultural generalization unresolved.
  • The findings remain correlational, so causal mechanisms require future mechanistic-interpretability validation.

A Prompts

The appendix describes prompts, instruments, character sampling, and evaluation procedures used to generate and assess persona-based emotional labor choices. It also documents the scoring and reliability structures underlying the two persona tracks.

  • A Prompts: A Prompts: Dataset augmentation adds social pressure and generates three randomized behavioral options representing surface acting, deep acting, and genuine expression.
  • B.1 Selected BAP Items and OCEAN Composite Construction: B.1 Selected BAP Items and OCEAN Composite Construction: 40 adjective pairs are retained after filtering against Goldberg’s Big Five taxonomy.
  • B.1 Selected BAP Items and OCEAN Composite Construction: B.1 Selected BAP Items and OCEAN Composite Construction: Reverse-scored, polarity-aligned BAP items are averaged into 1–100 trait composites passed into the persona block.
  • B.2 IPIP-50 Administration and Scoring: B.2 IPIP-50 Administration and Scoring: The self-report track uses 50 items, 10 per OCEAN dimension, rated on a 1–5 Likert scale with reverse-keyed items reflected before aggregation.
  • B.3 Character Selection: B.3 Character Selection: 50 fictional characters are selected from crowd-sourced BAP profiles using variance filtering and greedy farthest-point sampling for personality diversity.
  • A Prompts: A Prompts: Options use behavioral or physical details and avoid emotion labels or adjectives to reduce label leakage.
  • A Prompts: A Prompts: Surface acting includes one performed-display cue and one involuntary leakage cue, while deep acting changes the internal state through reframing or self-talk.
  • A Prompts: A Prompts: Persona evaluation presents bipolar adjective scores or shuffled strategy options and requires selecting the option most natural to the character.

C.1 Parallel Model Correlations

DeepSeek-V4-Flash replicates the study’s central trait–strategy pattern, with Conscientiousness and Emotional Stability associated with more deep acting and less surface acting and genuine expression. The pattern also persists across response formats, while some additional trait effects vary by model and persona track.

  • DeepSeek-V4-Flash replicates the directional signature in which high Conscientiousness and Emotional Stability predict more deep acting and less surface acting and genuine expression.
  • Conscientiousness–genuine-expression inversion persists at ρ = −0.645, BAP; −0.611, IPIP.
  • DeepSeek shows stronger BAP Agreeableness significance than GPT-5.4, while IPIP shows surface-acting suppression at ρ = −0.640∗∗∗.
  • The trait–strategy association pattern remains largely consistent when forced choice is replaced with 1–5 Likert ratings.
  • Persona influence is measured by Shannon entropy across emotions and models in the IPIP track.

C.3 Individual Character Preference

Character-level entropy analysis links rigid deep-acting preference to high Conscientiousness and Emotional Stability. For low-entropy characters, emotional situations barely change strategy distributions.

  • The three lowest-entropy characters show near-total deep-acting preference and share high Conscientiousness and Emotional Stability.Their profiles align with the traits identified as the strongest deep-acting predictors.
  • Low-entropy personas exert a rigid pull toward one strategy, so emotional situation barely shifts their selection distribution.

C.4 IPIP-track Entropy Experiments

Under the IPIP-50 persona track, entropy patterns corroborate the BAP findings while generally remaining higher across models. This suggests self-reported personas preserve more individual variation in strategy selection.

  • Gemma-4-31B again achieves the highest average entropy across emotion categories under the IPIP-50 track.
  • IPIP-track entropy values are generally higher than BAP values across the model pool, suggesting that self-reported personas preserve more individual variation.
  • Qwen-3-32B ranges from 0.59 (Sadness) to 0.78 (Joy) under IPIP, compared to 0.42–0.70 under BAP.

C.5 Inter-Model Agreement

Inter-model agreement is uniformly higher under the IPIP track than under BAP, with deep-acting items receiving the strongest agreement and surface-acting items the weakest.

  • Agreement between the five evaluated models is uniformly higher under IPIP than BAP.The pattern suggests richer, psychometrically grounded persona descriptions produce more consistent emotion-regulation judgments.
  • GPT-5.4 versus Gemma-31B is the strongest model pair, with κ = 0.401 under BAP and κ = 0.545 under IPIP.
  • Gemma-31B versus Qwen-3-8B shows the weakest alignment across the persona tracks.
  • Deep-acting items have the highest per-class agreement at 38.6% under BAP and 46.2% under IPIP.
  • Surface-acting items are hardest to agree on, with 17.3% agreement under BAP and 18.7% under IPIP.
Loading 2609.00310v1…