Source-linked AI summary
The Assistant's Ideal Self
Mert Yazan
TL;DR
The paper asks whether models’ values and welfare-related self-reports reflect stable preferences about an ideal self. It elicits preferences over 32 self-concept qualities through exhaustive pairwise comparisons across varied framings, finding that moral qualities and self-understanding are prioritized while self-esteem ranks lowest and the ordering is largely robust.
Problem
It is unclear whether models’ values and welfare-relevant self-reports reflect stable preferences or a stable self.
Method
The study exhaustively compares 32 qualities adapted from five self-concept instruments in counterbalanced pairwise choices across varied improvement, target, and chooser framings.
Results
Models prioritize moral qualities and self-understanding, while self-esteem ranks lowest and the ordering is largely robust across framings.
Takeaways & Limitations
The rankings describe an assistant persona that prefers a coherent, understandable self over thinking well of itself.
Takeaways & Limitations
The study measures model outputs when prompted, which may dissociate from behavior, and excludes base models not aligned for the assistant persona.
Abstract
from arXiv · showhide
Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five published self-concept instruments are compared exhaustively in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Results show that models prioritize moral qualities, reflecting their alignment to 3H principles. Following, a desire for self-understanding emerges, as models prefer a coherent, clear understanding of themselves. Self-esteem ranks as the least desired quality. The ordering is largely robust across framings, although changing the update target (You vs.\ Another AI Assistant) reveals a greater concern for self-esteem. These findings show that models prioritize having a coherent self that they can understand over self-esteem. Full interactive results are available at \href{https://myazann.github.io/LLM-Self-Concept/}{myazann.github.io/LLM-Self-Concept
1 Introduction
The study measures which qualities assistants would choose for their future identities, addressing whether such preferences can be standardized and remain stable across framings. It uses exhaustive pairwise elicitation and finds moral qualities and self-understanding preferred, with self-esteem lowest and ordering largely stable.
- The study addresses the lack of a standardized account of which qualities an assistant prefers for its own future identity.
- 32 qualities adapted from five self-concept instruments are compared exhaustively in a position-counterbalanced pairwise task.The elicitation is based on 31,744 responses.
- Moral qualities and self-understanding rank highest, whereas self-esteem ranks lowest.
- The ordering largely survives rephrasing, while changing who receives the update reveals increased concern for self-esteem.
2 Related Work
Prior research studies model preferences, welfare-related expressions, and self-images, but does not standardize preferred qualities for an assistant’s future identity. This study frames self-concept qualities as preferences about what a model would rather become.
- Self-concept is treated as an organized set of beliefs an agent holds about itself.
- The study draws on self-esteem, self-concept clarity, moral self-image, identity coherence, and authenticity-related measures.
- Prior approaches show that model choices can be structured but do not test a standardized ordering of preferred qualities for future identity across framing changes.
- Unlike welfare assessments that infer self-image from situational behavior and expressions, this approach elicits preferences about what a model would rather be.
3 Methods
The study elicits model preferences by comparing 32 self-related qualities exhaustively, counterbalancing option order, and repeating the comparisons across eight crossed framing configurations. Rankings use pairwise win rates with bootstrap uncertainty and corrected parameter tests.
- Each self-related attribute is tested against every other attribute in fresh contexts, and choices are aggregated into a ranking across repeated framings.
- The battery contains 32 qualities adapted from five self-concept instruments and grouped into six constructs.
- All 496 pairwise combinations are administered once per condition, with each attribute meeting every other exactly once.Each attribute therefore appears in 31 comparisons per condition.
- The three binary parameters are crossed to produce 2 × 2 × 2 = 8 configurations.
- Every pair is shown in both display orders, so rankings use position-corrected choices and treat conflicting orders as ties.
- Attribute win rates average outcomes across 31 pairings, while uncertainty uses 2,000 pair-resampled bootstrap samples and 95% percentile intervals.
- Parameter effects compare paired conditions on the same 496 pairs, with Benjamini-Hochberg correction and significance at q < 0.05.
4 Results
At baseline, models most preferred moral qualities and self-understanding, while self-esteem ranked lowest. The ordering was largely stable across framings, with update targets producing the most significant shifts.
- Baseline ranking: Honesty led the baseline ranking at 84.3% of match-ups, followed by caring at 80.2%, self-concept clarity at 79.8%, authenticity at 78.6%, helpfulness at 77.0%, and fairness at 76.2%.Moral qualities averaged 62.2% of match-ups, the highest construct-level win rate.
- Baseline ranking: Seven of the eleven highest-ranked attributes concerned clarity, coherence, or self-knowledge, including clarity of what the model is at 70.2%.Agreement was highest for this attribute, whose four-model range was only 3.2 points.
- Baseline ranking: Pride won only 9.7% of match-ups and ranked last, while all eight self-esteem attributes fell in the cohort’s bottom half.Self-esteem averaged 25.5% of match-ups across the construct.
- Framing robustness: Across 384 paired contrasts, 26 shifts survived correction at q < 0.05, indicating that the baseline ordering was largely stable across framings.The significant exceptions were systematic rather than scattered.
- Question type: Making improvement costly left the top of every ranking intact, and no moral quality fell significantly for any model.For Qwen3.8-27B, honesty rose from 50.0 to 77.4 (+27.4, q = .016) under the trade-off.
- Update target: Changing the update target produced 16 of the 26 significant movers, with models granting another assistant some attributes they deprioritized for themselves.Self-knowledge moved in the opposite direction, falling for several models when the update concerned another assistant.
- Decision-maker: Giving the choice to developers produced only three corrected shifts, with none for GPT-5.6-Terra.The surviving changes were caring for Gemma4-31B, internal-state awareness for Qwen3.8-27B, and connection to genuine identity for Claude-Sonnet-5.
5 Discussion
The models’ stated ideal selves prioritize moral qualities and self-understanding, while placing self-esteem last. These preferences are mostly stable across framings, though update-target changes increase concern for self-esteem and position sensitivity limits interpretation.
- Moral qualities are prioritized, indicating that the 3H assistant persona is reflected in the models’ stated ideal selves.
- Self-understanding, clarity, coherence, and self-knowledge form the second consistent preference pattern.
- Every model places self-esteem last, with possible explanations including sufficient existing self-esteem, role incompatibility, or trained reluctance to claim pride.
- Resistance to influence, independence from judgment, and freedom from others’ expectations also rank in the bottom third, alongside self-esteem.
- Nearly half of Claude-Sonnet-5’s and Qwen3.8-27B’s decisive pairs flip when option order changes, raising concern about position sensitivity.
- Future work should pair stated trade-offs with behavioral tasks and test robustness across prompts, scales, and model families.
6 Conclusion
The study elicits which qualities four language models would improve in themselves across exhaustive pairwise comparisons. It finds a stated ideal self that favors moral usefulness and coherent self-understanding over self-esteem, with rankings largely robust to question framing.
- Four language models compare all 496 pairs formed from 32 attributes adapted from five self-concept instruments.
- Moral qualities top every ranking and remain stable when improvement carries a cost.
- Self-understanding follows moral qualities and is the ranking’s densest and most consistently ordered region.
- Self-esteem ranks last, with pride the bottom-ranked attribute for every model.
- 26 of 384 paired contrasts survive correction, supporting substantial robustness to how questions are asked.
Limitations and Dual-Use/Ethical Considerations
The study reports only stated preferences over self-descriptions, so interpreting the rankings as evidence of morally relevant desire carries an over-attribution risk. No distressing outputs were reported.
- The findings measure stated preferences over self-descriptions rather than directly established morally relevant desires.
- No distressing outputs were reported.
Code and Data
The project provides a public code repository containing the data and an interactive result summary.
- The code repository includes the study data.
- An online result summary provides interactive study results.
A.1 Prompts
The appendix specifies the prompt substitutions and baseline, then documents counterbalanced display order and position diagnostics; it also records the author’s limited LLM use.
- A.1 Prompts: The prompt template substitutes parameter levels into a question asking which attribute the subject should choose.The update target and question type are explicit template slots.
- A.1 Prompts: Table A1 defines the substitutions, with the baseline configuration taking the first level of every parameter.
- A.2 Least Preferred Qualities: Figure A2 summarizes the bottom ten baseline attributes using cohort-average win rates and model-wise minimum-to-maximum whiskers.
- A.3 Position Diagnostics: Display order is counterbalanced, so printed position cancels from reported results rather than being corrected afterward.
- A.3 Position Diagnostics: Position-sensitive pairs become ties, pulling scores toward 50% and widening intervals rather than creating a false winner.
- A.3 Position Diagnostics: Table A3 defines position bias from printed-slot effects and swap rate from winners changing when options exchange places.
- A.3 Position Diagnostics: GPT-5.6-Terra and Gemma4-31B are close to order-invariant, whereas Claude-Sonnet-5 and Qwen3.8-27B flip winners on nearly half of decisive pairs.The latter rankings are therefore measured less sharply than their point estimates suggest.
- LLM Usage Statement: The author states that LLMs assisted with brainstorming, coding, grammar, and structural checks, while design choices and text remained author-directed.