Source-linked AI summary
How AI Prompts Can Teach Us About the Structure of Human Behavior
Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
TL;DR
Researchers ask whether diverse human choices require many setting-specific traits or can be represented compactly. The paper uses structured LLM prompts to estimate behavioral dimensions and finds three dimensions closely match behavior across ten game roles.
Problem
The paper asks whether a few characteristics explain choices across environments and whether people are broadly distributed or clustered into recurring behavioral types.
Method
The method assigns LLMs natural-language type vectors, prompts them across behavioral settings, and compares generated choices with observed human behavior.
Results
Human behavior can be closely matched using three dimensions—Risk Aversion, Strategic Sophistication, and Trust—and inferred types cluster into distinct groups.
Takeaways & Limitations
The findings support developing parsimonious theories of behavior across settings and using portable types to predict choices in new described settings.
Takeaways & Limitations
A three-dimensional prompt space might still index many unrelated behavioral patterns because LLM interpretations can be highly nonlinear.
Abstract
from arXiv · showhide
We introduce a general, easy-to-implement AI-based method for studying the structure and complexity of human behavior. We assign a large language model a ``type vector'' and then prompt it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) becomes ``You are a player characterized by the following profile: 2 out of 5 in Altruism, 4 out of 5 in Risk Aversion,'' after which it is prompted to make choices. We vary the dimensions (e.g., Altruism, Fairness, Trust, $\dots$) and values (e.g., 1--5) to minimize distance to human choices. Applying the method to 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles, we find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. Moreover, the types needed to fit individuals across games cluster into fewer than a dozen groups, and can predict behavior in held-out games with different rules and available actions. The results suggest that behavior across diverse settings can be approximated by a low-dimensional, portable representation, supporting the possibility of general yet parsimonious theories across the behavioral sciences. More broadly, the method can provide insights into the structure of many human behaviors.
1 Introduction
The paper introduces a portable LLM-based method for representing human behavior with theory-based type vectors and uses it to test whether choices across settings have a low-dimensional structure. Across ten economic game roles, behavior is closely approximated by three dimensions, with individual types clustering into fewer than a dozen groups.
- Method: The method prompts an LLM with vectors of theory-based characteristics and translates those descriptions into behavior across settings and populations.Researchers can vary the dimensions and their intensities to study behavioral dimensionality and human types.
- Main findings: Three dimensions—Risk Aversion, Strategic Sophistication, and Trust—can approximate individual humans across the ten settings.This result provides evidence for a credible upper bound of a simple three-dimensional type representation.
- Data and settings: 119,147 decisions from 78,657 participants across more than 35 countries provide the empirical basis across ten classic economic game roles.The analysis first evaluates how well type vectors match behavioral distributions within each game role.
- Main findings: Fewer than a dozen distinct clusters emerge when more than a thousand individual type vectors are examined.The fitted types are heterogeneous: some participants are highly strategic and risk-seeking, while others are highly risk-averse and fair.
- Generalization: The analysis tests whether an individual’s fitted type from other games can reproduce behavior in a held-out game, addressing portability across settings.This out-of-distribution exercise compares type-based matching with the potential benchmark of using the full set of an individual’s choices.
2 Fitting type vectors to human behavior with AI prompts
The method represents behavioral types as vectors of intensities over researcher-selected characteristics, translates them into natural-language prompts, and elicits choices across described settings. By varying the number and levels of characteristics, researchers can assess the dimensionality needed to approximate human behavior.
- Method: Type vectors assign intensities to behavioral characteristics, which an LLM receives as words alongside setting descriptions to generate predicted choices.The characteristics may be drawn from economics, psychology, or other theories of human behavior.
- Method: Researchers vary the number of included characteristics to determine how many dimensions are needed to closely approximate human populations’ behavior.The same type-vector prompt can be paired with different settings described in words.
- Method: A vector (2, 4) represents a subject rated 2 out of 5 in Altruism and 4 out of 5 in Risk Aversion.The vector-based prompt is paired with descriptions of settings and questions about behavioral choices.
- Method: 5^k possible type vectors arise when k selected characteristics each use the implementation’s 5-point Likert scale.More generally, L^k vectors arise for k characteristics with L possible levels.
3 Game roles, data, and candidate dimensions
The study evaluates AI prompts across ten human game roles using five behavioral dimensions and their lower-dimensional subsets. The resulting behavior varies systematically across games and models, while smooth response patterns support interpreting nearby types as producing similar behavior rather than arbitrary profiles.
- Game roles and data: 10 distinct game roles span eight classic behavioral-economics games, with separate Proposer and Responder roles in Ultimatum and Investor and Banker roles in Investment.The dataset covers Bomb Risk, Commons, Cournot, Dictator, Investment, Keynes’ Beauty Contest, Public Goods, and Ultimatum games.
- Candidate dimensions and type space: 5 behavioral dimensions—Altruism, Fairness, Risk Aversion, Strategic Sophistication, and Trust—are varied at five Likert levels, generating 3,125 type vectors.Lower-dimensional spaces use subsets of these characteristics, and each type vector is mapped with game instructions to a game choice.
- Candidate dimensions and type space: Altruism affects behavior strongly in Dictator, Banker, and Public Goods games but little in Bomb Risk, Beauty Contest, and Cournot games, illustrating systematic game-specific effects.Strategic Sophistication shows a contrasting pattern, including no impact for the Ultimatum proposer and Responder roles in the supplied passage.
- Model robustness: Alternative models preserve intuitive characteristic-by-game relationships, including Altruism’s effects on giving and investment and Strategic Sophistication’s reduction of Beauty Contest choices.The exercise is repeated with GPT-5.6 Terra, GPT-5.6 Luna, Claude Sonnet 5, and DeepSeek V4 Pro.
- Interpreting the type space: Changing prompted intensity generally produces continuous, flat, or approximately monotonic behavior, and two-dimensional prompts remain generally smooth and often roughly linear.These patterns indicate that nearby type vectors generally produce similar behavior rather than arbitrary placement across the type space.
4 Matching Distributions of Human Play
Across ten economic game roles, adding behavioral dimensions makes AI-generated choice distributions increasingly resemble human play, with the largest gains from the first three dimensions. Risk Aversion, Strategic Sophistication, and Trust provide the key combination, while additional dimensions yield smaller improvements except in the Proposer role.
- Overall fit: With three dimensions, AI-generated distributions largely overlap human distributions in most games, whereas fourth and fifth dimensions add only modest improvements.The default prompt concentrates choices on one or two points and fails to reproduce human heterogeneity; even one-dimensional types substantially improve the fit.
- Overall fit: Risk Aversion alone substantially reduces normalized Wasserstein distance for every game, followed by further improvements from Strategic Sophistication and Trust.The largest improvements occur within the first three dimensions.
- Overall fit: Fourth and fifth dimensions provide much smaller additional gains for all game roles other than the Proposer.This pattern is consistent across the game-role comparisons in Figure 4.
- Game-specific fits: Individual roles are often well-matched by only two dimensions, bringing simulated distributions within the sampling error of human choices.The Beauty Contest is the clearest exception, while best single dimensions are often intuitive: Fairness for Dictator and Responder, Risk Aversion for Investor, and Strategic Sophistication for Beauty Contest.
5 Matching an individual human’s behaviors across settings
Assigning each individual a single type vector to match choices across all game roles shows that human behavior can often be represented with a few dimensions, but individuals vary in fit and complexity. These inferred types also predict choices in held-out games, approaching an information-advantaged machine-learning benchmark.
- Individual matching: 9,269 decisions from 1,734 individuals who played at least five game roles were matched using one shared type per individual across roles.The analysis fits the joint distribution of plays across game roles rather than separate marginal distributions for each role.
- Individual matching: 0.132 median matching error results from a one-dimensional type, down from 0.288 under the default prompt, with improvement flattening after three dimensions.Four dimensions outperform five when matching subjects across games.
- Type-space structure: Fewer than a dozen clusters capture the type space, including Risk-Averse and Fair, Strategic Risk-Seekers, and Non-Strategic and Fair groups with distinct risk, fairness, and trust profiles.Risk-Averse and Fair types average (A=2.9, T=2.8, R=4.0, F=4.3, S=2.8), while some left-side clusters have R≤2.0, F=5.0, and T≤1.9.
- Heterogeneity: Nearly 70% of matchable subjects require no more than two dimensions, while 448 of 1,734 subjects cannot be matched within error threshold ε=0.1.Subjects therefore differ in both how well they fit the prompts and how many dimensions are needed.
- Held-out prediction: 1.11 times the HistGBT error is the sample-weighted error for the best-dimensionality types, versus 1.43 for a random human draw and 1.73 for a uniform random guess.Across games, the lowest type-based error ranges from 0.94 to 1.24 times the HistGBT error.
6 Alternative types
Alternative keyword sets can also improve individual behavioral matching, but economic keywords fit best once models use two or more dimensions. The economic set captures most of its error reduction with three dimensions, while placebo gains quickly diminish and the optimal keyword set remains unsettled.
- Method: The analysis compares economic, Big 5 personality, and placebo keyword sets by selecting the lowest-median-error combination for each dimension count from one through five.Each keyword uses three Likert values, and each combination generates 3^k prompt types simulated ten times in every game.
- Results: All three keyword sets substantially improve fit with their best one-dimensional keyword, although placebo dimensions provide rapidly diminishing benefits.Varying even an unrelated placebo keyword creates behavioral variation that gives the matching procedure choices with which to match subjects.
- Results: Agreeableness provides the best one-dimensional fit, while economic keywords perform best from two through five dimensions.Four economic dimensions fit slightly better than all five Big 5 characteristics.
- Limitations: The five economic keywords achieve most of their matching-error reduction with three dimensions, after which the curve becomes flatter.The comparison does not establish that this is the best possible keyword set; another set could achieve the same fit with fewer dimensions or continue improving.
7 Discussion
The findings suggest that a few dimensions generate substantial behavioral heterogeneity, with inferred types clustering into distinct groups. The method also enables cross-setting prediction while leaving open questions about literal keyword interpretation and broader transfer.
- Behavioral structure: A few dimensions generate substantial heterogeneity within and across games, while inferred types cluster tightly into distinct groups.The discussion addresses both the dimensionality of behavior and how people are distributed in the resulting space.
- Methodological contribution: Estimated types are more than statistical summaries because the same natural-language prompt can generate predictions for a subject in any described setting.The method also makes it possible to compare candidate dimensions and test whether types estimated in one set of settings predict behavior elsewhere.
- Interpretation and limitations: Changing a keyword value may alter multiple internal model features, so a nominally low-dimensional type space may not have a literal human interpretation.This limitation follows from the LLM’s black-box nature and the uncertain correspondence between component values and human characteristics.
- Interpretation and future work: The authors make no claim that their keywords test a particular theory, although theoretically motivated dimensions appear consistent with existing theories.Future work could propose new dimensions and evaluate which approximate human responses with the fewest components.
- Transfer and future work: Leave-one-game-out results show that inferred types can generalize out of distribution in this specific context, while transfer to substantially different settings remains an open question.The same exercises can be repeated with new dimensions, populations, and settings beyond the economic games studied here.
A AI Prompts and Game Instructions
The study elicits game choices by pairing type-vector or default prompts with role-specific instructions, generating valid LLM responses and extracting numeric choices through a fixed pipeline. The prompts encode characteristics on Likert scales, while the instructions span classic economic games and specify concrete response formats.
- A.1 Type-vector and Default Prompts: Each elicitation pairs one prompt with the instructions for one game role, with the prompt in the system role and game instructions in the user role.
- A.1 Type-vector and Default Prompts: The main analysis uses gpt-4.1 through Azure Batch and collects ten valid responses for every prompt–game-role pairing.Valid responses are those that can be parsed into choices within the action range specified by the game instructions.
- A.1 Type-vector and Default Prompts: Type-vector prompts insert selected characteristic names, values, and the number of Likert scale levels into a fixed profile template.The template defines 1 as the lowest level and [L] as the highest level.
- A.1 Type-vector and Default Prompts: The example type vector (2, 4) produces a profile assigning Altruism: 2 out of 5 and Risk Aversion: 4 out of 5.
- A.2 Game Instructions: The instructions reproduce classic dictator, ultimatum, trust, public-goods, Bomb Risk Elicitation, beauty-contest, Cournot, and common-pool-resource games.They preserve the original wording, capitalization, punctuation, and paragraph breaks.
- A.3 Data Generation and Extraction: Numeric choices are extracted with a fixed pipeline whose regular expression identifies an unambiguous bracketed numeric choice required by the game instructions.Replies unresolved by this rule were handled separately in the extraction process described in the source passage.
B Human-playing Data
The human benchmark uses MobLab classroom data covering 119,147 choices by 78,657 subjects across ten game roles. Individual-level analyses focus on 1,734 subjects who played at least five games and made 9,269 choices.
- Individual-level analysis: 1,734 subjects who played at least five games contributed 9,269 choices to the individual-level analysis.No subject played more than eight of the ten games.
- Analysis samples: Population-level analyses use all 119,147 choices, whereas individual-level analyses use the ≥5 cohort.Each subject contributes at most one choice per game in the benchmark data.
C Model Robustness
The economic dimensions produce broadly similar effects across five LLMs, supporting robustness across models. Agreement is strongest for Altruism, Trust, Risk Aversion, and Fairness, while Strategic Sophistication is a notable exception except in the Beauty Contest.
- Cross-model agreement: r=0.83 to r=0.88 for alternative models’ correlation-of-correlations with gpt-4.1, indicating highly similar dimension effects.The correlations compare each alternative model’s 50 keyword-by-game Spearman correlations with the corresponding gpt-4.1 correlations.
- Cross-model agreement: 37 of the 50 keyword-by-game comparisons show agreement among at least four models on significance direction or nonsignificance.Agreement means correlations are significantly positive, significantly negative, or not statistically distinguishable from zero.
- Cross-model agreement: 35 of the 40 comparisons for Altruism, Trust, Risk Aversion, and Fairness show agreement among at least four models.The agreement is especially clear for characteristics with large and intuitive effects.
- Cross-model agreement: Higher Altruism increases giving or investment in the Dictator, Investor, Banker, and Public Goods games across all five models, but has little effect in the Bomb game.This pattern demonstrates agreement both within individual panels and across the overall effects of characteristics and games.
- Cross-model agreement: Strategic Sophistication is the main exception, with mostly unshaded panels, but its effect is strongly negative and consistent across models in the Beauty Contest.Elsewhere, its relationships tend to be weaker than many effects observed in other panels.
D Behavior induced by two-dimensional type vectors
This section examines behavior generated by every two-characteristic combination among five economic characteristics. For each pair, it varies one characteristic from 1 to 5, holds the other at each of its five levels, and compares the resulting behavior with the focal characteristic alone across ten game roles.
- Every two-characteristic combination among the five economic characteristics is examined.
- The focal characteristic varies from 1 to 5, while the paired characteristic is fixed separately at each of its five levels.
- The figures compare each two-dimensional profile with a one-dimensional vector containing only the focal characteristic across ten game roles.
E Additional Figures
The additional figures detail held-out-game prediction procedures and error benchmarks, while examining whether economic and personality keyword representations induce systematic behavior across games. They also compare matching distance across alternative characteristic sets on a 5-level Likert scale.
- Held-out prediction: Figure A12 estimates each subject’s type without using their held-out-game choice and predicts that game using the median of ten choices generated by the matched type.Rows select the lowest sample-weighted mean held-out error across games for each number of dimensions.
- Error benchmarks: The prediction-error ratios use mean HistGBT error as the denominator, while separate rows report uniform-random and random-human expected errors.For the Weighted average, N denotes the number of decisions rather than subjects.
- Alternative representations: On a 5-level Likert scale, Figure A13 compares median matching distance across subjects for economic characteristics, Big-5 OCEAN traits, and placebo keywords.For each number of dimensions, it selects the combination with the lowest median matching distance.
- Keyword-induced behavior: Figures A14–A16 vary placebo, economic, and Big-5 psychology keywords across Likert levels and game roles, reporting mean choices as action-range percentages with ±1 standard-error bars.A gray horizontal line marks the midpoint of each game’s action range.