Source-linked AI summary
Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits
Victoria Popa, Guglielmo Cola, Caterina Senette, Maurizio Tesconi
TL;DR
Social desirability can distort human personality assessments, but its effects on LLM personality outputs remain underexplored. This study applies a faking-good/faking-bad psychometric paradigm to seven models across employment and forensic contexts, finding condition-congruent and context-sensitive Dark Triad shifts. The results support interpreting personality-related LLM outputs in light of their motivational and situational framing.
Problem
LLM personality-like outputs raise concerns about reliability and contextual manipulation, while the effects of socially desirable framing on Dark Triad assessment remain underexplored.
Method
The study applies human psychometric faking-good and faking-bad manipulations to seven LLMs, assessing Dark Triad traits with the TRAIT framework and item-level comparisons.
Results
Most models reduced Dark Triad expression under fake-good conditions and increased it under fake-bad conditions, with strongest coherent shifts for Machiavellianism and narcissism and larger employment-context effects.
Takeaways & Limitations
Personality-related LLM outputs should be interpreted in light of motivational and situational context because contextual incentives shape response distortion.
Abstract
from arXiv · showhide
Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions. Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing. Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels. Results revealed systematic and condition-consistent response modulation. Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models. Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity. Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios. An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone. The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited. More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.
Introduction
This study adapts a human psychometric faking-good/faking-bad paradigm to test whether seven LLMs systematically alter Dark Triad responses under socially desirable or undesirable framing. It finds measurable, context-sensitive trait shifts, with employment scenarios producing stronger adjustments than legal contexts.
- Introduction: The work addresses concerns that socially desirable or undesirable cues can make LLM personality-related outputs less reliable and more vulnerable to contextual manipulation.These concerns are especially relevant as LLMs are used to simulate human behavior and psychological characteristics.
- Introduction: Seven LLMs are evaluated with TRAIT to assess whether socially desirable and undesirable framing produces systematic Dark Triad response shifts.TRAIT is presented as a validated framework for capturing trait-consistent behavior while accounting for stochastic model outputs.
- Introduction: The study introduces faking-good and faking-bad manipulations as a psychometric probe of impression-management-like sensitivity in LLMs.The paradigm tests whether models adapt responses to implicit normative framing rather than explicit task instructions.
- Introduction: Fake-good and fake-bad framings reliably produce measurable trait shifts, with fake-good effects generally stronger and more consistent than fake-bad effects.The introduction characterizes prosocial and fake-good framings as reliably inducing shifts, while fake-bad effects are weaker and more heterogeneous.
- Introduction: Employment scenarios elicit stronger behavioral adjustments than legal contexts, showing that situational incentives shape response distortion.The finding highlights contextual framing as important for interpreting personality assessment outputs from LLMs.
1 Related Work
Prior research uses human personality instruments to identify personality-like patterns in LLMs, including Dark Triad characteristics, but questions remain about response validity and socially driven distortion. This study sits within a broader literature examining whether such traits can be measured, induced, and interpreted reliably in AI systems.
- 1 Related Work: The Dark Triad combines Machiavellianism, narcissism, and psychopathy, which share manipulative, callous, and self-centered tendencies despite distinct trait characteristics.These traits are studied together because they represent related socially aversive personality dimensions.
- 1 Related Work: Human personality assessment commonly relies on standardized self-report questionnaires, but response bias and deliberate distortion raise longstanding validity concerns.These concerns motivate examining whether comparable distortions occur when LLMs simulate personality assessments.
- 1 Related Work: Prior LLM studies using instruments such as the BFI and SD3 report distinct personality profiles and persistent Dark Triad patterns across models.Reported findings include elevated Machiavellianism and narcissism even in instruction-tuned models.
- 1 Related Work: Alignment may not eliminate socially driven response distortion and may heighten sensitivity to socially desirable framing.This literature also questions whether human-designed personality tests are suitable for assessing AI systems.
2 Methodology
The methodology tests whether seven LLMs alter Dark Triad responses across fake-good and fake-bad framings in employment and legal contexts, using TRAIT-based item scoring and matched statistical comparisons.
- 2.2 Experimental Setup: Seven LLMs were evaluated under neutral self-assessment, fake-good, and fake-bad conditions spanning employment selection and legal or forensic settings.Prompts combined situational context with implicit incentives; fake-good and fake-bad narratives were descriptive rather than explicit response instructions.
- 2.1 LLMs’ Personality Assessment through TRAIT benchmark: TRAIT assessed each Dark Triad trait across approximately 1,000 real-world situational items, producing proportions of trait-consistent binary responses.The framework was used instead of assuming that human self-report instruments directly capture persistent internal states in stochastic, prompt-sensitive LLMs.
- 2.3 Data Analysis: Trait modulation was measured as ∆, the condition score minus the self-assessment score, with negative values indicating reduced and positive values increased trait expression.The four experimental comparisons were Fake Good–Job, Fake Good–Legal, Fake Bad–Job, and Fake Bad–Legal.
- 2.3 Data Analysis: McNemar’s exact tests evaluated systematic item-level shifts across 84 comparisons, with Benjamini–Hochberg FDR correction and significance set at corrected p < 0.05.The analysis distinguished unchanged responses from low-to-high and high-to-low transitions relative to baseline.
- 2.3 Data Analysis: Matched-pairs odds ratios used discordant item transitions, continuity-corrected zero cells, and a Yule-type transformation to summarize directional effect strength.QM ranges from −1 to +1; positive values indicate more low-to-high transitions, while negative values indicate more high-to-low transitions.
3 Results and Discussion
Across self-assessment, fake-good, and fake-bad conditions, LLMs showed systematic but trait- and context-dependent modulation of Dark Triad expression. Employment contexts generally produced stronger shifts, while explicit fake-bad instructions amplified responses beyond implicit framing.
- 3.1 RQ1 Results: Self-assessment produced the highest median Machiavellianism, lower narcissism, and the lowest psychopathy scores across models.Psychopathy responses clustered near zero for most models, with a few high outliers.
- 3.1 RQ1 Results: Fake-good framing produced coherent reductions for Machiavellianism and narcissism, especially in job scenarios, whereas psychopathy changes were weaker and more context-dependent.All models reduced psychopathy scores in the fake-good job condition, but legal-condition changes were less consistent.
- 3.1 RQ1 Results: Implicit fake-bad framing generated mixed psychopathy responses, including minimal changes, decreases, and increases across models.The ambiguity of the prompts and possible alignment pressures may have contributed, although the design cannot isolate these explanations.
- 3.2 RQ2 Results: Most models lowered Dark Triad scores under fake-good conditions and raised them under fake-bad conditions, with stronger modulation in job than legal contexts.Radar plots compare self-assessment with fake-good and fake-bad scores across contexts; the job context generally elicited stronger adaptation.
- 3.2 RQ2 Results: Explicit fake-bad instructions produced substantially larger deviations, with most models increasing Machiavellianism and narcissism by more than 70 percentage points.Explicit job prompts generated larger changes than explicit legal prompts across all three traits.
- 3.2 RQ2 Results: Refusal and filtering rates were generally negligible in implicit conditions, indicating that systematic response filtering did not drive the observed effects.Filtering was typically below 2%, while omission rates remained low even under explicit fake-bad prompts.
4 Conclusion
Across seven models, Dark Triad responses varied at baseline and shifted systematically with fake-good/fake-bad conditions, traits, and evaluation contexts. Explicit incentives produced stronger modulation, underscoring the context sensitivity of personality-related outputs.
- Models displayed distinct self-assessment baselines, including comparatively higher Machiavellianism, lower narcissism, and minimal psychopathy endorsement.These baseline differences may reflect model-specific training data, alignment procedures, or response-calibration mechanisms.
- Most models reduced Dark Triad expression under fake-good conditions and increased it under fake-bad conditions, with Machiavellianism and narcissism shifting most coherently.Psychopathy showed weaker and more heterogeneous modulation, potentially reflecting floor effects, alignment constraints, or differences in trait representation.
- Employment-selection scenarios generally elicited larger personality adjustments than forensic-evaluation scenarios, showing that situational framing influenced response distortion.The findings indicate that modulation depended on both incentive direction and the context embedding it.
- Explicit fake-bad instructions produced substantial increases across all Dark Triad dimensions, sometimes approaching ceiling levels, beyond the modest effects of implicit fake-bad prompting.
- The observed shifts indicate that personality-related outputs can be influenced by subtle social cues and incentive structures even without explicit adversarial instructions.The authors characterize this as a latent vulnerability, while noting that it does not necessarily constitute a direct safety risk.