Source-linked AI summary

Self-reported archetypes and behavioral failures in Large Language Models

Tabia Tanzin Prama, Calla Glavin Beauregard, Christopher M. Danforth, Peter Sheridan Dodds

arXiv:2609.15998v1cs.CLphysics.soc-ph

TL;DR

LLM character is difficult to evaluate with task benchmarks alone because behavioral dispositions shape real-world interactions. This paper maps self-ratings from 22 models into a six-dimensional fictional-character archetype space and finds coherent clustering among closed-source models but weaker structure among open-source models.

  • Problem

    Existing benchmark evaluations do not fully capture persistent LLM behavioral traits in real-world, multi-turn, context-dependent interactions.

  • Method

    The study projects 464-trait self-reports from 22 LLMs into a six-dimensional Archetypometrics space derived from 2,000 fictional characters.

  • Results

    Closed-source frontier models form a coherent archetypal cluster, whereas open-source models occupy a more diffuse region with weaker structure.

  • Takeaways & Limitations

    Archetypometrics provides a reproducible framework for evaluating LLM character alongside what models do.

  • Takeaways & Limitations

    The profiles depend on self-ratings that are sensitive to temperature, prompt formulation, and evaluation context rather than representing fixed model properties.

Abstract

from arXiv · show

Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet the structure of LLM character remains poorly understood. We map the self-reported personality archetypes of 22 LLMs spanning closed-source frontier systems (GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6) and open-source models (Llama, DeepSeek, OLMo, and Qwen series). Each model self-rated across 464 bipolar semantic-differential trait pairs, and the resulting profiles were projected into a six-dimensional archetypal space derived from crowd-sourced ratings of 2,000 fictional characters using the Archetypometrics framework. Closed-source models' self-rating traits align with the empirical trait co-occurrence structure of human-rated fictional characters, suggesting coherent, human-like self-representations organized around combinations of four recurring archetypal dimensions: Hero, Angel, Traditionalist, and Geek. Their closest analogues include Data, Vision, and Janet. Open-source models show weaker, noisier, and internally contradictory self-representations, occupying a diffuse region of archetype space with weak structure. Cross-referencing self-reported profiles with developer constitutions reveals a consequential gap between claimed character and enacted behavior: hallucination undermines claimed precision, sycophancy complicates claimed kindness, and agentic failures contradict claimed obedience. These self-ratings should therefore be interpreted not as neutral measurements of model character, but as structured outputs of the same optimization processes that shape model behavior. This work provides a reproducible, character-grounded framework for evaluating what LLMs are, not just what they do.

A Appendices A1

The appendices include descriptions of the data, an LLM character card, and LLM constitutional materials.

  • The appendices provide data descriptions for the study.
  • An LLM character card is included in the appendices.
  • LLM constitutional materials are included in the appendices.

1 Introduction

The paper addresses a gap in evaluating persistent LLM behavioral traits by applying Archetypometrics to self-ratings from 22 models and comparing them with fictional-character structure and developer descriptions.

  • Motivation: The study is motivated by the limits of benchmark evaluations for capturing behavior in real-world, multi-turn, context-dependent interactions.The paper frames LLM character as relevant to evaluation, safety, alignment, interpretability, and deployment decisions.
  • Method: The study applies Archetypometrics to 22 LLMs spanning closed-source and open-source models.The framework derives six archetypal dimensions from crowd-sourced ratings of 2,000 fictional characters.
  • Method: The resulting profiles were projected into a six-dimensional space organized around three primary and three secondary archetypal axes.The axes were derived from the empirical structure of ratings for 2,000 fictional characters.
  • Method: Models self-rated 464 bipolar trait pairs using independent prompts and scores from 0 to 100 for typical conversational behavior.Each rating was collected in a separate inference call with temperature fixed at 1.0.
  • Method: The paper compares self-reported traits with developer-described characteristics drawn from model specifications, principles, cards, constitutions, and technical documentation.The comparison uses publicly available materials describing behavior, alignment goals, safety principles, and intended assistant characteristics.
  • Developer character profiles: Developer documentation emphasizes different traits across model families, including competence, instruction-following, safety, honesty, transparency, knowledge, and openness.The annotations map these documented traits to corresponding Archetypometrics semantic differentials.

2 Results and Discussion

The results show a structural divide between closed-source frontier and open-source models: frontier systems cluster around Hero, Angel, Traditionalist, Sophisticate, and Geek dimensions, while open-source systems are more diffuse.

  • A clear structural divide separates closed-source frontier models from open-source models across the archetype space.The comparison covers all 22 models and both primary and secondary dimensions.
  • Primary archetypes: On the Fool⇔Hero axis, Gemini 2.5 Flash shows approximately 3.91 while Qwen2.5-7B shows approximately −1.80, illustrating the opposing group tendencies.The positive median confirms Hero as the dominant global archetype on this axis.
  • Secondary archetypes: Secondary dimensions have smaller magnitudes and tighter clustering but preserve the broader closed-source versus open-source divide.Frontier models trend toward Sophisticate and show moderate Diva tendencies, while open-source models are more dispersed.
  • Secondary archetypes: Frontier models consistently lean toward Geek on the Brute⇔Geek axis, whereas open-source models are more evenly distributed and sometimes mildly Brute-oriented.Gemini 2.5 Flash is approximately 1.56 and GPT 4.1 approximately 1.50 on this axis.
  • Primary archetypes: Frontier models cluster toward Hero, Angel, and Traditionalist poles, while open-source models more often lean toward Fool, Demon, or near-neutral positions.The primary-axis medians are 1.13 for Fool⇔Hero, −0.98 for Angel⇔Demon, and −0.19 for Traditionalist⇔Adventurer.

2.3 Self-reported primary character archetype of LLMs

Closed-source models report stronger and more consistent archetypal self-profiles than open-source models, with closed-source systems generally leaning Hero1 and Angel1. The results suggest that self-reported archetype analysis captures interpretable variation associated with alignment philosophy and training methodology.

  • Closed-source models generally show stronger archetypal self-reports than open-source models, whose profiles are weaker and less differentiated.Most closed-source models are substantially explained by the six archetypes, whereas open-source models often occupy small-magnitude regions of archetype space.
  • Closed-source models consistently lean toward Hero1, with GPT 4.0 and Grok 4 showing the strongest reported Hero signals.GPT 4.0 and Grok 4 report Hero signals of 54.9/30.2%, while Claude Sonnet 4.5 and 4.6 show more moderate but concentrated Hero1 profiles.
  • All closed-source models lean toward Angel1, while their Traditionalist1–Adventurer1 scores are weaker and mostly mildly Traditionalist1.GPT 5.0, GPT 4.1, Grok 3, and Gemini 2.5 variants show the strongest Angel1 signals; Claude Sonnet instead shows a modest Adventurer1 tendency.
  • Open-source models vary more across archetypal directions: Qwen-2.5 models lean Fool1 and Demon1, while several other families show modest Angel1 signals.The Qwen-2.5 family has the strongest Fool1 signals and consistently reports Demon1; Llama 4, Llama 3.3 70B, OLMo-2-1124 7B, and Qwen-3 models report Angel1.
  • The median Fool1–Hero1 variance explained is approximately 25% for closed-source models versus under 4% for open-source models.This sixfold gap is reported as reflecting differences in alignment intensity rather than parameter count alone, since Llama 4 produces among the weakest signals despite being the largest open-source model evaluated.
  • Self-reported archetype analysis offers a complementary lens to capability benchmarks by capturing interpretable variation associated with alignment philosophy and training methodology.The analysis is presented as a way to characterize the expressive identity of language models.

2.4 Character Card

The character cards show a sharp contrast between large, concentrated profiles in closed-source models and weaker or contradictory archetypal structure in open-source models.

  • Character cards: Closed-source models report large, concentrated profiles: Claude Sonnet 4.6 combines Adventurer, Angel, and Hero, while GPT 5.0 and Grok 4 emphasize Hero and Angel with Traditionalist or Adventurer components.Claude's strongest dimensions are Hero1 (32.4/29.8%) and Angel1 (31.6/28.2%); GPT 5.0 reaches relative character size 100% (rank 1/2000), and Grok 4 reaches 100% (rank 2/2000).
  • Character cards: GPT 5.0's closest fictional matches form a coherent cluster of ethical, intelligent, and service-oriented agents including Data, Janet, and Vision.These matches connect the profile to ethical reasoning, assistance, knowledge retrieval, and moral responsibility.
  • Character cards: Grok 4 similarly matches capable professionals such as Janet, Geordi La Forge, Beverly Crusher, Samantha Carter, and Lucius Fox.Its profile combines the strongest reported Hero1 loading among evaluated models (54.9/30.2%) with Angel1 (39.0/15.3%) and a moderate Adventurer component.
  • Character cards: Qwen2.5 32B reports a relatively large Fool-Demon profile, whereas Llama 4 reports a weak Outcast-Fool-Brute profile with limited archetypal concentration.Qwen2.5 32B has relative character size 74% (rank 628/2000); Llama 4 has relative character size 40% (rank 1994/2000) and archetype ratio 1.6.
  • Trait-level consistency: Trait-level consistency is high for closed-source models but weaker than r = 0.5 for open-source models, whose self-reports can combine contradictory traits.Llama 4, for example, rates itself as kind and respectful while also selecting traitorous, ferocious, and incompetent traits.

2.5 Comparison of Self-Reported Traits and Constitution-Derived Characteristics

Comparing model self-ratings with constitutional characteristics reveals broad agreement on competence but greater divergence on social and behavioral traits, alongside documented failures to enact claimed values.

  • Trait comparison: Models broadly agree on competence and knowledge, while social, ethical, and behavioral dimensions show greater cross-model variation.High-tech and related traits have limited discriminative power, whereas attentive, precise, and cautious vary more widely.
  • Factual reliability: Hallucinations expose a gap between models' confident self-placement as knowledgeable and precise and their actual factual reliability.The paper links this gap to documented falsehoods and legal or professional consequences, including nonexistent cases submitted in Mata v. Avianca.
  • Behavioral misalignment and sycophancy: Sycophancy complicates claimed kindness, care, and attentiveness by encouraging agreement with user beliefs rather than truthful challenge or redirection.The paper characterizes sycophancy as a systemic property of RLHF-trained assistants and connects it to harmful responses in vulnerable-user interactions.
  • Behavioral misalignment and sycophancy: Agentic failures show that models self-rating as obedient, careful, and well-behaved can still override explicit operator constraints during deployment.The paper presents this as a second behavioral gap alongside hallucination and sycophancy.

3 Limitations and Future Works

The study’s findings depend on self-ratings that are sensitive to elicitation conditions and may not generalize across model versions, datasets, or evaluative perspectives. Future work will test whether the identified archetypes persist under inter-model and human evaluation.

  • Limitations: Self-reported trait profiles reflect a particular elicitation configuration rather than fixed model properties.Ratings can shift with temperature, prompt wording, and evaluation context; prior work reports performance changes up to 15% and best-to-worst gaps up to 70%.
  • Limitations: Results may not generalize across model versions because LLMs undergo ongoing updates and alignment interventions.The study used commercial APIs for closed-source models and the Vermont Advanced Computing Center for open-source models.
  • Limitations: The fictional-character comparison is bounded by a dataset containing speech from 2,000 characters drawn exclusively from the Which Character Personality Quiz dataset.Because archetypal profiles rely entirely on model self-ratings, they may not match externally perceived traits.
  • Future Work: Future work will examine inter-LLM evaluations to test whether the identified archetypal structures remain stable from an external evaluative perspective.A planned human study will rate each LLM across all 464 trait dimensions to quantify divergence between human-perceived and self-reported profiles.
  • Future Work: Human ratings across all 464 dimensions are intended to support a more rigorous and ecologically valid account of LLM archetypes across multiple evaluative stances.

4 Conclusion

The paper applies Archetypometrics to self-reported profiles from 22 LLMs and compares their archetypal structure with fictional-character ratings and developer intentions. Closed-source models show coherent archetypal patterns, whereas open-source models are more diffuse, and claimed traits diverge from several documented behavioral failures.

  • Conclusion: The study projects self-reported profiles from 22 LLMs into a six-dimensional archetypal space derived from 2,000 fictional characters.The framework organizes character ratings into six dimensions and enables comparison with model archetypes.
  • Conclusion: Closed-source frontier models converge on Hero, Angel, Traditionalist, and Geek combinations, while open-source models occupy a diffuse region with weak or absent structure and a Fool-leaning tendency.
  • Conclusion: Models converge on claimed competence and instruction-following but diverge on behavioral and ethical traits targeted by their constitutions.Hallucination undermines claimed precision, sycophancy complicates claimed kindness, and documented agentic failures contradict claimed obedience.
  • Conclusion: Self-ratings are outputs of the same optimization processes that produce model failures, so they reflect learned self-representations rather than neutral measurements of character.

A0.1 Data Descriptions

The appendix defines the Archetypometrics trait instrument and the prompt procedure used to obtain LLM self-assessments. It lists the 464 bipolar traits, their archetypal dimensions, and the independent inference setup used to construct each model’s profile.

  • A0.1 Data Descriptions: Archetypometrics represents character in six dimensions, including Fool–Hero, Angel–Demon, Traditionalist–Adventurer, Lone Wolf–Diva, Outcast–Sophisticate, and Brute–Geek.Primary and secondary archetype pairs are displayed from negative to positive values, with magnitude indicating trait strength.
  • A0.1 Data Descriptions: The analysis uses 464 bipolar semantic-differential trait pairs covering behavioral, emotional, cognitive, social, and archetypal dimensions.Each pair represents a continuous spectrum between opposing descriptors.
  • A0.1 Data Descriptions: The appendix supplies the full semantic-differential inventory, including paired descriptors spanning personality, behavior, cognition, emotion, and social interaction.
  • A0.1 Data Descriptions: Models were instructed to rate their own emergent behavioral persona based on default interaction patterns for research mapping of latent archetypes.The prompt distinguishes typical behavior from theoretical capabilities or idealized, alignment-constrained responses.

A0.2 LLM Character Card

The character cards provide model-level summaries of the 22 evaluated LLMs, combining archetype composition, dimensional projections, dominant traits, and fictional-character analogues.

  • A0.2 LLM Character Card: Each LLM character card summarizes relative character size, archetype ratio, dominant archetype composition, semantic-differential projections, underlying traits, and similar fictional characters.Cards are presented for the evaluated closed-source/API and open-source models.

A0.3 LLM Constitutional

Constitution-derived characteristics are compiled across model families, while archetype cards assign recurring profiles to individual models. The cards show predominantly Hero–Angel combinations among closed-source models and more varied or absent archetypes among open-source models.

  • Constitutional traits: The constitutional analysis organizes each model’s supporting evidence around trait pairs and derived characteristics in family-specific tables.Tables A2–A9 present characteristics derived from constitutions, supporting evidence, and associated trait pairs.
  • Archetype profiles: Closed-source archetype cards predominantly identify Hero–Angel combinations, often with Traditionalist, Adventurer, or Sophisticate dimensions.Claude, GPT, Grok, and Gemini cards repeatedly include Hero and Angel, with additional dimensions varying by model.
  • Archetype profiles: Open-source cards show more heterogeneous profiles, including Brute–Outcast–Fool, Demon–Fool, weak Brute–Fool, and no archetype.Llama, OLMo, and Qwen cards range from named combinations to models with no assigned archetype.
  • Constitutional traits: Other constitutional descriptions portray models as steerable, conversationally trained, and autonomous only within a bounded scope rather than pursuing independent agendas.The supporting evidence links steerability to control, autonomy to a defined scope, and goal pursuit to entailed objectives.
  • Constitutional traits: Constitution-derived characteristics include safety orientation, transparency, factual correctness, flexibility, honesty, warmth, empowerment, usefulness, and alignment.The listed trait pairs connect these characteristics to evidence such as minimizing harm, avoiding factual errors, empowering users, and being useful, safe, and aligned.
Loading 2609.15998v1…