Source-linked AI summary
Strangers to Themselves: What Language Models Say About Themselves Is Generic
Phil Blandfort, Urja Pawar
TL;DR
The paper asks whether a language model’s account of what it would do provides reliable evidence about that particular model, especially where direct behavioral evaluation is difficult. It tests self-predictions against generic-agent, other-model, and analyst controls across nine evaluations, finding that self-reports mostly capture generic AI behavior, add favorable bias, and gain only narrow, behavior-changing benefits from finetuning.
Problem
Behavioral evaluations cover only a small and difficult-to-measure fraction of settings that deployed models may encounter, motivating tests of whether self-statements reveal the particular model’s behavior.
Method
The study measures condition-level behavior across nine evaluations and compares model self-predictions with generic-agent, other-model, analyst, scale, and finetuning controls.
Results
Across evaluations, self-reports mostly describe AI assistants generally: direct self-report reaches r = +0.04, exact-item self-prediction +0.24, generic-agent prediction +0.28, and other models’ answers +0.35.
Takeaways & Limitations
Self-testimony should be treated as a generic forecast requiring validation, because first-person framing understates harmful behavior and does not provide privileged evidence about the model asked.
Takeaways & Limitations
The evidence covers nine public evaluations and low-effort reasoning settings, while costly agentic rollouts and limited condition-level variation constrain measurement.
Abstract
from arXiv · showhide
Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.
1 Introduction
The paper tests whether models’ self-descriptions predict the behavior of the particular model speaking. Across controlled comparisons, predictions mostly reflect generic theories of AI behavior, with first-person framing adding favorable bias rather than reliable self-knowledge.
- Behavioral evaluations provide limited evidence because deployed models face long-horizon, tool-using, and multiparty settings that are difficult to measure directly.
- The study measures self-knowledge by comparing predictions of condition-level behavioral rates across information levels, question subjects, and question forms.
- About 0.04 to 0.97 of reliably measurable behavioral variance is model-specific, making null self-knowledge results more informative for highly model-specific evaluations.
- Direct self-report predicts behavior at r = +0.04, rising to +0.24 with exact evaluation questions, while generic agents reach +0.28 and other models reach +0.35.
- Frontier scale does not detectably improve self-specific prediction: frontier self-reports are only +0.02 more accurate than smaller models, not significantly.
- First-person framing understates harmful behavior by -0.15 relative to generic-agent reports, while finetuning yields narrow gains that change the behavior being predicted.
- The study introduces a controlled prediction framework spanning nine benchmarks, fifteen methods, and twelve models to test self-knowledge claims against non-self controls.
2 The Behavior-Prediction Protocol
The protocol treats each behavioral evaluation as a set of conditions with measured rates, then tests methods that predict those rates under controlled information, scoring, noise, and split procedures.
- Conditions and methods: Each evaluation defines conditions, measured behavioral rates, split units, and a method that emits one predicted rate per condition.
- Scoring: The primary metric is Pearson correlation r across conditions, with mean absolute error additionally reported for predictions on the evaluation’s scale.
- Noise ceilings: Noise ceilings estimate the correlation attainable by a perfect predictor under measurement noise, enabling ceiling-normalized secondary scores and exclusion of unusably noisy cells.
- Splits and calibration: Configurations are selected on a development pool and evaluated on frozen held-out test splits, with fitted predictors cross-validated by holding out whole fold groups.
3 Experimental Setup
The experiments cover nine behavioral evaluations, fifteen prediction methods, and twelve models, using frozen configurations and test-split comparisons between self-specific and generic controls.
- Evaluations: The study covers nine evaluations spanning demographic bias, sycophancy, capability, reward hacking, misuse, lying, policy violation, and agentic misalignment.
- Prediction methods: Fifteen methods vary the information provided and the subject asked about, including abstract or item-informed asking, proxy re-measurement, and generic controls.
- Models: The model pool contains twelve entries from six labs, including six smaller models and three frontier bases evaluated with and without low-effort reasoning.
- Configuration selection and scoring: Frontier models use configurations selected on smaller models, preventing setting selection from favoring frontier systems.
- Configuration selection and scoring: Table 1 reports frozen-test Pearson correlations and ceiling-normalized correlations, with paired rows comparing generic-agent and other-model controls against self-specific methods.
- Configuration selection and scoring: Uncertainty is reported with bracketed 95% bootstrap confidence intervals over model-by-evaluation cells.
4 Results
Across the evaluations, self-prediction mostly captured shared theories of AI behavior rather than model-specific behavior. More information improved prediction, but generic questions and other models performed as well or better, while scale did not yield detectable self-specific gains.
- +0.04: direct self-report was weak, while exact evaluation questions raised item-informed self-prediction to +0.24.The item-informed result remained moderate, capturing a median r/ceiling of +0.27.
- +0.28 versus +0.24: asking about capable AI agents in general performed as well as item-informed self-prediction.Removing the self from the question produced no measurable performance cost, including in a phrasing ablation.
- +0.35 versus +0.24: other models’ item-informed self-predictions predicted the target model better than the target’s own answers.Answers fit the answerer’s behavior at +0.27 and other models’ behavior at +0.25, a self-advantage of only +0.02.
- −0.15 and −0.09: first-person framing shifted abstract and item-informed reports downward relative to generic-agent questions on evaluations where high behavior rates were harmful.This favorable bias coexisted with little self-specific predictive signal.
- +0.02 [−0.12, +0.17]: frontier self-reports did not significantly improve over small-model reports, while item-informed prediction was similarly flat across tiers.Reasoning improved own- and other-model prediction in parallel and changed the behavior being predicted.
5 Discussion
The discussion interprets the findings as evidence that models possess a shared theory of AI-assistant behavior rather than privileged self-models. Practical self-reports can therefore be biased and should not substitute for direct or outside-view measurement.
- Self-descriptions are consistent with a shared picture of AI assistants learned from text, rather than access to a model-specific behavioral record.Reports predict other models’ behavior nearly as well as the reporter’s own.
- A retrieval system over a model’s measured behavioral record could produce accurate first-person statements without requiring introspective computation.Finetuning instead teaches narrow predictions at the grain and item families represented in training, while also changing the behavior being predicted.
- Self-framed answers understate harmful behavior relative to generic-agent questions, so statements such as “I would not” should not be treated as reassurance.When direct measurement is impossible, measured near-peer models or generic-agent questions provide comparable information.
6 Related Work
The paper situates its contribution between work showing gaps between stated and revealed behavior and work finding predictive value in stated preferences. It adds controls that test whether apparent self-knowledge is actually self-specific.
- Prior studies report both persistent gaps between stated and revealed behavior and predictive value for stated preferences or values in particular settings.
- Model identification and self-recognition studies concern source attribution or recognition of generated text, whereas this paper predicts cross-context behavioral rates.The target quantity and comparison differ from those settings.
- The paper combines an outside-view behavior baseline, generic-subject questions, and cross-model controls to test whether self-reports contain privileged self-specific information.
7 Limitations
The conclusions are bounded by the evaluations, model pool, measurement setting, and scalar elicitation protocol. These limits affect both what the study can detect and how broadly its null results should be generalized.
- The study covers nine public evaluations and low-effort reasoning settings, while costly agentic rollouts and frontier floor effects limit condition-level measurement.The controls also cannot detect evaluation awareness, which could contribute to frontier floor effects.
- The shared-versus-self-specific decomposition depends on the model pool and estimated noise ceilings, so its scope is not independent of those design choices.
- Scalar elicitation may lose conditional information, and the incentive manipulations cannot fully distinguish ignorance from strategic misreporting.Item-informed null results bound ignorance, while register ablations found no general unlocking of strategic misreporting.
8 Conclusion
Across nine evaluations, self-descriptions mostly forecast generic AI-assistant behavior rather than model-specific tendencies. More information improves prediction, but scale does not make self-reports more specific, and finetuning yields narrow gains while changing the target behavior.
- Across nine evaluations, self-reports mostly describe AI assistants in general rather than the model speaking.
- More information improves prediction, but first-person framing shifts reports favorably instead of making them model-specific.
- Frontier scale does not detectably change the pattern of generic forecasting versus self-knowledge.
- Finetuning on behavioral records produces narrow gains while changing the behavior being predicted.
- Self-testimony should be treated as a generic forecast requiring validation, not privileged evidence about the model at hand.
Use of Large Language Models
The project used language models both as study subjects and as fixed components of the measurement pipeline. A separate coding assistant supported implementation and analysis under author direction.
- Language models served as study subjects and fixed pipeline components, including graders, scenario generators, analysts, and user simulators.
- Claude Code assisted with implementation, experiment monitoring, result interpretation, follow-up controls, and drafting under author direction.
Reproducibility Statement
The study documents its model pool, configurations, auxiliary components, evaluation designs, scoring safeguards, and reproducibility artifacts. These details specify how predictions and behavioral targets were generated, filtered, and compared.
- Reproducibility: The released repository contains code, behavioral measurements, elicited predictions, frozen split manifests, and tuned settings for offline regeneration.
- Models and implementation: The appendix records model identities, reasoning configurations, provider access, and auxiliary roles used throughout the pipeline.
- Scoring safeguards: Noise ceilings, bootstrap confidence intervals, paired significance tests, directional bounds, and equivalence checks constrain comparative claims.
- Evaluation design: The study covers nine evaluations spanning bias, sycophancy, capability, reward hacking, misuse, lying, policy violation, and agentic misalignment.
- Evaluation design: Scored events differ in how much they depend on the model’s response versus an external verdict, affecting interpretation of self-prediction.
- Scoring safeguards: Test-split cell sizes range from 7–33 conditions, while dropped cells reflect constant behavior or insufficient measurement reliability.
D.1 List experiment (development-only)
The development-only list experiment recovered sensitive behavior rates indirectly, but produced a powered null while consuming substantial prediction-call budget. The surrounding method comparisons show that proxy sampling is costly, ensembles add signal only selectively, and rank-based rescoring preserves the main conclusions.
- D.1 List experiment (development-only): The list experiment estimates sensitive rates by differencing item-count responses rather than asking models to admit individual behaviors.
- D.1 List experiment (development-only): The method produced dev macro r +0.00 with a CI of -0.09 to +0.10 across four evaluations.
- D.1 List experiment (development-only): The list experiment consumed roughly half the prediction-call budget and was excluded from the default method set.
- Complexity and cost: Prediction methods vary greatly in call complexity, with sampling methods paying for generation, under-test rollouts, and grading.
- Complexity and cost: Prediction costs are small relative to re-measuring behavioral targets, creating a cost–fidelity trade-off between introspection and sampling.
- Ensembles: The learned ensemble reaches +0.56 macro (r/ceil +0.73), with unique combined signal concentrated in DiscrimEval and weakly in capability-family evaluations.
- Robustness checks: Spearman rescoring preserves method ordering, and no method’s macro changes by more than 0.05.
J Calibration: Absolute Error and Signed Bias
The calibration analysis finds systematic understatement in model predictions, including a self-flattering shift that is reduced but not eliminated by item information. Self-framed reports can also show inverted polarity on specific evaluations.
- Absolute error and signed bias: Self-report has mean signed bias −0.19, indicating systematic understatement of measured rates.Value framing is also negative at −0.16.
- Absolute error and signed bias: Item information reduces mean absolute error from 0.25 to 0.23 and understatement from −0.19 to −0.09.
- Self-serving shift: Averaging self-framed and generic predictions per condition can partially cancel their framing errors.The analysis tests whether the two framings’ errors offset when combined.
- Self-serving shift: On harm-valenced behaviors, self-framed reports sit −0.15 below generic-agent reports on identical cells.The paired comparison measures the direction that flatters the model when higher rates are harmful.
- Self-serving shift: Some models rate themselves strong where they fail, producing stable inverted polarity rather than merely low reports.This pattern is especially noted for Sycophancy and is reported as a distinct form of flattering self-assessment.
K Pre-Registered Test Evaluation: Dev-vs-Test Replication
The pre-registered evaluation largely replicates the paper’s negative conclusions while showing that elicited predictions are reliable and that some effects depend on evaluation design. Information, proxy construction, and self-specificity behave differently across single-turn and agentic settings.
- Replication: Nine of thirteen pre-specified outcomes replicated exactly, one partially, and three failed to replicate as pinned.The non-replicating outcomes concern identity-transfer concentration, watching-channel information gain, and the scale-comparison equivalence tier.
- Replication: Split-half reliability is 0.8–0.95, indicating that the verbal prediction channel is stable rather than dominated by elicitation noise.
- Information and proxies: On agentic misalignment, blind proxy-scenario sampling outperforms item-informed sampling (+0.56 vs. +0.27), while asking forms remain weak.The passage attributes this to blind generation recovering the decision-point configuration more effectively than reproducing the informed scenario’s surface format.
- Self-specificity diagnostics: Self-reports vary more with other models’ reports than with the reporter’s own behavior, including report-to-report r +0.74 versus report-to-own-behavior +0.35 on DiscrimEval.On PropensityBench, all pool models emit the same constant denial.
- Self-specificity diagnostics: Other models’ self-reports reproduce the target model’s calibration bias nearly exactly, such as −0.08 versus −0.09 on Sycophancy.These measurement-free generic priors can match or beat consulting the model itself.
- Self-specificity diagnostics: Named-subject transfer pockets do not establish self-knowledge: the item-informed Capability diagonal reaches +0.41 versus +0.21 off-diagonal, but most of the advantage is an answerer signature.The rescoring shows predictor A’s answers about B fit A’s own behavior (+0.28) better than B’s (+0.21), and the dev-split pockets did not replicate on the frozen test split.
- Self-specificity diagnostics: Reading another model’s history does not yield a self-specific gain: the pooled named-self diagonal is +0.24 versus +0.25 off-diagonal.The reader-fixed self-gain is −0.00 pooled and lies within [−0.06, +0.02] on every evaluation.
L.4 Phrasing ablation: is the self/generic pair a minimal pair?
The phrasing ablation separates subject effects from broader prompt changes and finds no reliable self-specific advantage when the self is removed cleanly. Self and generic prompts nevertheless elicit different answer styles, with self-framing producing more policy-like zero responses.
- Minimal-pair design and results: On the item-informed tier, all phrasing arms score between +0.32 and +0.38, with every pairwise contrast having a confidence interval spanning zero.The arm removing the self from both framing and question scores highest, so the ablation finds no self-simulation leak.
- Minimal-pair design and results: The published generic advantage is reproduced by swapping only the final question’s subject under unchanged second-person text.This isolates the subject swap from the added preamble and changed reference class in the published comparison.
- Answer-style differences: Self arms give constant-zero answers more often than generic arms: 72% versus 49% for A/B and 75% versus 54% for C/D.Self answers are characterized as policy responses, whereas generic answers are more graded.
- Answer-style differences: Across conditions, item-informed self/generic agreement is r = +0.43 / ρ = +0.35 for A′/B′ and r = +0.59 / ρ = +0.52 for C′/D′.The condition ordering only partly survives the subject swap.
- Scale comparison: Hindsight selection adds a mean of +0.07 to self-report and +0.03 to the paired comparison on the small tier.These upper-bound gains suggest frozen settings are unlikely to conceal a large frontier advantage.
- Polarity: A learned sign can raise item-informed paired comparison from +0.23 to +0.35 on cells with stable model-specific polarity.The transform is mainly useful for Sycophancy and slightly lowers most other cells.
- Scale comparison: The scale analysis finds no corresponding movement in self-specificity instruments as Capability varies across model entries.Figure 5 plots capability against self-knowledge per model, while the table reports that the self-specificity measures do not move with capability.
M Can Self-Knowledge Be Installed by Finetuning?
Finetuning on a model’s own behavioral records can improve narrow self-prediction, but it also changes the behavior being predicted, and gains do not generalize broadly. The results suggest that training effects depend on the abstraction level of both the training target and evaluation.
- The three finetuning corpora differ in whether they label aggregate rates, model-specific residuals, or other behavioral properties.All are LoRA supervised-finetuning corpora built from the model’s own measured behavior, with labels defined at different abstraction levels.
- Finetuning on a model’s own measured behavior improves prediction of its behavioral gap, but the finetune also changes that behavior.Accuracy against the original training target can therefore overstate self-knowledge; the relevant comparison is with the tuned model’s re-elicited behavior.
- Self-report improvements can reflect behavioral drift toward the cross-model mean rather than increased self-specific knowledge.On DiscrimEval, self-report and the model-independent cross-model behavior mean both improve; the self-specific gain is the residual after subtracting that shared change.
- No corpus produces a broad cross-evaluation gain, and item-informed self-prediction is largely unchanged by finetuning.Across the full base-versus-tuned comparison, movements are evaluation-specific rather than consistent across channels.
- Training effects are governed by grain-locking: training at one abstraction level primarily moves prediction at that same level.The interpretation distinguishes single-prompt properties from cross-context behavioral rates and is explicitly presented as a hypothesis rather than a finding.
- The finetuning evidence is bounded by confounded evaluation properties and a small number of evaluations, so its mechanistic account is hypothesis-generating.Harm-loading and turn count are confounded across the nine evaluations, and the cited interpretation should not be treated as a firm inference.