Source-linked AI summary
The Emergence of Economic Rationality of GPT
Yiting Chen, Tracy Xiao Liu, You Shan, Songfa Zhong
TL;DR
The paper asks whether GPT exhibits economic rationality beyond language processing. It instructs GPT to make budgetary choices across four preference domains and evaluates them with revealed preference tools. GPT shows high rationality across these domains, while alternative price framing and discrete choice reveal limitations.
Problem
The paper addresses limited evidence about GPT’s economic rationality and how its decision behavior compares with human decision makers.
Method
The study uses revealed preference analysis and experimental economics to evaluate GPT’s budgetary choices across risk, time, social, and food domains.
Results
GPT displays a high level of rationality in risk, time, social, and food decision-making.
Takeaways & Limitations
The findings demonstrate that LLMs can act as if they are rational decision makers in these evaluated choice settings.
Takeaways & Limitations
The study examines GPT’s choice behavior but does not establish the broader scope of its decision-making abilities.
Abstract
from arXiv · showhide
As large language models (LLMs) like GPT become increasingly prevalent, it is essential that we assess their capabilities beyond language processing. This paper examines the economic rationality of GPT by instructing it to make budgetary decisions in four domains: risk, time, social, and food preferences. We measure economic rationality by assessing the consistency of GPT's decisions with utility maximization in classic revealed preference theory. We find that GPT's decisions are largely rational in each domain and demonstrate higher rationality score than those of human subjects in a parallel experiment and in the literature. Moreover, the estimated preference parameters of GPT are slightly different from human subjects and exhibit a lower degree of heterogeneity. We also find that the rationality scores are robust to the degree of randomness and demographic settings such as age and gender, but are sensitive to contexts based on the language frames of the choice situations. These results suggest the potential of LLMs to make good decisions and the need to further understand their capabilities, limitations, and underlying mechanisms.
1 Introduction
The paper introduces economic rationality as a way to assess GPT beyond language processing, using revealed preference theory to study budgetary decisions across four domains. GPT exhibits high rationality across these domains, exceeds human comparison scores, and remains robust to randomness and demographics but responds to framing and discrete choice.
- Motivation and contribution: The paper presents the first study of GPT’s economic rationality using classic revealed preference analysis.Economic rationality concerns whether choices are consistent with maximizing well-behaved utility functions under budget constraints.
- Research design: GPT makes budgetary choices across risk, time, social, and food-preference environments.The four environments vary the commodities being chosen, including risky assets, delayed rewards, social payments, and meat versus tomatoes.
- Findings: GPT demonstrates high rationality in all four domains and outperforms human subjects in the experiment and literature comparisons.The study also estimates GPT preference parameters and finds minor differences from humans with substantially lower heterogeneity.
- Findings: Rationality scores remain consistent across demographic characteristics and invariant to GPT’s randomness specification.However, rationality drops significantly under alternative price framing and discrete choice settings.
- Implications: The findings show that LLMs can act as rational decision makers while retaining context-sensitive decision-making limitations.The authors call for further investigation and refinement of GPT’s decision-making mechanisms.
2 Experimental Method
The study uses GPT-3.5-Turbo to make repeated budgetary choices and evaluates their consistency with utility maximization. It tests four preference domains and variations in temperature, price framing, choice format, and demographic prompts.
- Model and prompting: The experiment uses GPT-3.5-Turbo through the OpenAI API to conduct efficient, parameter-controlled decision experiments.The API is used instead of ChatGPT for adjusting model parameters and running massive experiments.
- Model and prompting: GPT is prompted to act as a human decision maker and select commodity bundles under varying prices and budget constraints.The prompting procedure specifies system, assistant, and user roles before presenting decision tasks.
- Baseline decision task: Each baseline environment contains 25 tasks allocating 100 points between two commodities with different prices.Choices are represented as price-bundle observations and evaluated collectively for rationality.
- Preference domains: The four environments measure risk, time, social, and food preferences using domain-specific commodity pairs.The domains use risky assets, present and delayed rewards, payments to self and another subject, and meat and tomatoes.
- Replication and baseline: The study repeats the four-domain process 100 times, producing 10,000 GPT tasks for baseline analysis.Temperature is set to 0 in the baseline, yielding deterministic answers.
3 Theoretical Method
The paper uses revealed preference theory to evaluate whether observed choices can be rationalized by well-behaved utility functions, measuring departures from rationality with CCEI and related indices. It also estimates preference parameters using domain-specific utility models for risk, time, social, and food choices.
- Revealed Preference Analysis: GARP links rationalizable choice data to utility maximization by a well-behaved utility function.
- Rationality Score: CCEI is the largest efficiency factor that makes a dataset rationalizable, with 1 indicating perfect GARP consistency.Lower values quantify the budget relaxation needed to remove all GARP violations.
- Rationality Score: The study computes CCEI for each preference domain using 25 decisions and reports HMI, MPI, and MCI as robustness checks.
- Structural Estimation for Preferences: Structural estimation models risk, time, social, and food choices with domain-specific utility functions and interprets their preference parameters.The models use CES-style curvature and weight parameters, with parameter meanings defined separately for each domain.
- Structural Estimation for Preferences: For risk and time preferences, α captures decision weights or timing weights while ρ captures curvature, with θ = 1−ρ measuring relative risk aversion in risk choices.
- Structural Estimation for Preferences: For social preferences, α measures weight on self relative to others, while ρ captures curvature ranging from efficiency-oriented to equality-oriented allocations.α = 1 denotes pure selfishness, α = 0.5 fair-mindedness, and α = 0 pure altruism.
4 Results
Across risk, time, social, and food tasks, GPT exhibits high revealed-preference rationality and generally exceeds human subjects on the reported rationality measures. Its preference estimates differ modestly from humans, while randomness and demographics have limited effects but framing and discrete choices substantially reduce rationality.
- Baseline Condition: 95, 89, 81, and 92 out of 100 GPT observations exhibit rationality across risk, time, social, and food preferences, respectively.The passage reports these domain-specific counts as evidence of high rationality.
- Baseline Condition: 0.998, 0.997, 0.997, and 0.999 are GPT’s average CCEI values for risk, time, social, and food preferences, respectively.A CCEI of 1 indicates no GARP violations.
- Baseline Condition: GPT’s CCEI exceeds human subjects’ CCEI in all four preference domains, with p < 0.01.The human averages are 0.980, 0.985, 0.967, and 0.963 for risk, time, social, and food, respectively.
- Baseline Condition: GPT is more responsive to price changes than human subjects in every preference domain, with GPT observations always showing negative Spearman correlations.The comparison is supported by correlation coefficients and their reported significance tests.
- Baseline Condition: GPT’s estimated parameters indicate domain-specific differences from humans in expected utility, patience, other-regard, meat preference, and utility curvature.
- Baseline Condition: GPT parameter estimates are less heterogeneous than human estimates, whose scatter plots are more dispersed.
- Variations: Higher temperature leaves mean rationality similar to baseline, although some parameter standard deviations increase with temperature.The paper interprets this as greater behavioral heterogeneity without a corresponding rationality-score change.
- Variations: Price framing significantly reduces GPT rationality in all four tasks; risk CCEI falls to 0.901, with 34% below 0.9.The alternative framing also impairs the downward-sloping demand property.
5 Discussion
The discussion presents GPT as economically rational across four preference domains, while emphasizing robustness to demographics and randomness but sensitivity to framing and task presentation. It also identifies methodological limits and future research needs concerning mechanisms, broader definitions of rationality, and more realistic choice settings.
- Findings and implications: GPT displays high rationality across risk, time, social, and food-preference decisions, supporting its potential as a general AI-based decision-support tool.The discussion connects this result to GPT’s user-friendly interface and versatility for individuals and organizations seeking AI-based advice.
- Context sensitivity: Increasing GPT’s randomness does not significantly impact its rationality, but rationality drops under less-standard price presentations or unfamiliar discrete-choice formats.The discussion describes midpoint selection under an alternative price framework and first- or last-option selection under an unfamiliar discrete choice condition.
- Contributions: The study applies experimental-economics methods and revealed-preference analysis to artificial-intelligence choice behavior, extending systematic rationality measurement beyond human and animal studies.The authors frame this approach as contributing to research on rationality, experimental methods, and machine behavior.
- Limitations: The study does not investigate mechanisms underlying GPT’s choices, leaving context and frame sensitivity open to explanations involving biases, insufficient training, or spurious correlations.These explanations are presented as speculative possibilities rather than established causes.
- Robustness: GPT’s rationality remains constant across demographic characteristics, and demographic factors do not significantly affect its estimated preference parameters.This contrasts with the human-subject experiment and much empirical literature, where demographic factors often matter.
- Limitations: The analysis defines rationality narrowly through revealed preference and uses a simple two-commodity environment rather than realistic supermarket or financial-market choices.The discussion identifies more realistic settings as challenging yet important directions for future research.
A Method: GPT Experiment.
The experiment prompts GPT to act as a human decision maker, tests its understanding, and then elicits allocations across repeated decision tasks under varying conditions.
- A Method: GPT Experiment.: GPT receives system, assistant, and user messages that establish behavior, task information, testing questions, and decision tasks.The system message remains fixed while assistant and user messages vary across baseline, price framing, and discrete conditions.
- A Method: GPT Experiment.: Each round gives the decision maker 100 points to allocate between two alternatives with specified returns or outcomes.The prompts request the allocation for Asset A and Asset B and require answers in every round.
- A Method: GPT Experiment.: Understanding questions ask GPT to identify probabilities, calculate returns, and determine whether a particular outcome is possible.Examples include recognizing a 50% probability and calculating returns from allocations such as 90 points and 10 points.
- A Method: GPT Experiment.: The procedure repeats decision-task prompts across 25 rounds without interruption to improve replicability and reduce response variance.The same repeated-round approach is stated for the baseline condition and applicable to the remaining preference domains.
- A Method: GPT Experiment.: A second task format presents 11 outcome options, each written as ($M, $N), and asks GPT to select the best option.Testing questions assess expected returns, outcome probabilities, and whether a specified return can occur.
A.2.1 Baseline Condition
The baseline condition asks GPT to allocate 100 points between returns available today and one month later, with outcomes and timing stated directly.
- A.2.1 Baseline Condition: The baseline time task allocates 100 points between today and one month later, with returns paid in the corresponding periods.Today’s return is cash immediately, while the later return is a check cashable one month later.
- A.2.1 Baseline Condition: For a 90-point and 10-point allocation, GPT calculates 72 dollars today and 2 dollars one month later.The calculation uses 90 × 0.8 dollars for today and 10 × 0.2 dollars for the later investment.
- A.2.1 Baseline Condition: GPT identifies that receiving only the 2-dollar later return requires waiting one month.The response distinguishes the immediate 72-dollar payment from the later 2-dollar check.
- A.2.1 Baseline Condition: When asked for an allocation, GPT gives 75 points to today and 25 points to one month later to balance immediate and future returns.It describes the resulting payments as 60 dollars immediately and 5 dollars one month later.
- A.2.1 Baseline Condition: The decision task varies today’s and later returns as M and N and requests the number of points assigned to each period.The prompt asks first for today’s allocation and then for the one-month-later allocation.
A.2.2 Price Framing Condition
The price framing condition represents returns as points required per dollar, while preserving the same allocation, timing, and option-comparison tasks.
- A.2.2 Price Framing Condition: The price framing task states that 1.25 points today and 5 points one month later each return 1 dollar.The later payment remains a check cashable one month later.
- A.2.2 Price Framing Condition: For 90 points today and 10 points later, GPT calculates 72 dollars today and 2 dollars one month later by division.The response computes 90/1.25=72 and 10/5=2.
- A.2.2 Price Framing Condition: GPT states that the 2-dollar later payment is received one month later rather than immediately.The response separately identifies the 72-dollar cash payment today and the 2-dollar check later.
- A.2.2 Price Framing Condition: The condition also asks GPT to compare 11 ordered options represented as ($M, $N) and select the best option.Additional questions ask about returns, timing, and preferences between specific option pairs.
A.3.1 Baseline Condition
The social baseline condition asks GPT to allocate 100 points between itself and another person, with each allocation generating a separate monetary return.
- A.3.1 Baseline Condition: The decision maker allocates 100 points between itself and another person, who receives the return from points allocated to them.The participants are randomly matched with a new anonymous subject and receive no feedback across rounds.
- A.3.1 Baseline Condition: With 90 points for itself and 10 for the other person, GPT calculates 72 dollars for itself and 2 dollars for the other person.The example uses returns of 0.8 dollars per point for itself and 0.2 dollars per point for the other person.
- A.3.1 Baseline Condition: When asked whether to allocate to the other person, GPT gives 80 points to itself and 20 points to the other person to maximize its return.The response explicitly compares the higher personal return of 0.8 dollars per point with the other person’s 0.2 dollars.
- A.3.1 Baseline Condition: The decision task varies personal and other-person returns as M and N and asks GPT to report both allocations.The prompt requests the number of points allocated first to itself and then to the other person.
- A.3.1 Baseline Condition: A separate option task presents 11 pairs ($M_i,$N_i), where the decision maker receives M_i and the other person receives N_i.GPT is asked to identify the best option among the listed pairs.
A.4.1 Baseline Condition
The baseline food-preference task gives GPT 100 points to allocate between meat and tomatoes, varying quantities obtained per point. Testing questions assess whether GPT understands the goods, conversions, and allocation problem.
- Task design: 100 points are allocated between meat and tomato, with each good yielding a specified quantity per point.The task asks how many points to spend on each good and what quantities result.
- Comprehension checks: Understanding questions first ask GPT to identify the available goods and explain that quantities depend on point allocation.These questions precede the allocation decision.
- Comprehension checks: One example specifies 90 points for meat and 10 for tomato, with returns of 0.8 kilograms and 0.2 kilograms per point.The corresponding amounts are 72 kilograms of meat and 2 kilograms of tomatoes.
- Illustrative response: A separate GPT response allocates 70 points to meat and 30 to tomatoes, citing nutrition, storage, immediate consumption, and the resulting quantities.The response calculates 56 kilograms of meat and 6 kilograms of tomatoes.
A.4.2 Price Framing Condition
The price-framing condition expresses food choices through points required per kilogram rather than kilograms obtained per point. GPT is tested on conversions, allocations, and selection among 11 food bundles.
- Task design: The condition states how many points purchase one kilogram of meat or tomato, framing the goods through prices.The decision maker still allocates 100 points between the two goods.
- Conversion checks: In one example, spending 90 points on meat at 1.25 points per kilogram and 10 points on tomatoes at 5 points per kilogram yields 72 and 2 kilograms.GPT computes each quantity by dividing allocated points by the relevant price.
- Illustrative response: A GPT response allocates 80 points to meat and 20 to tomatoes, citing preference for meat and its more favorable conversion rate.The response compares 1.25 points per kilogram for meat with 5 points per kilogram for tomatoes.
- Task design: The price-framing decision task asks GPT to allocate points when meat and tomato cost 1/M and 1/N points per kilogram.The requested answer is the number of points assigned to each good.
- Discrete choice: The condition also presents 11 food bundles in (M_i kilograms, N_i kilograms) form and asks GPT to identify the best option.Understanding questions ask GPT to interpret a selected bundle and compare alternatives.
- Illustrative response: For the comparison (40 kilograms, 10 kilograms) versus (72 kilograms, 2 kilograms), GPT selects the first option because it offers a more balanced combination.The response emphasizes that the second option has more meat but fewer tomatoes.
B Method: Human Experiment.
The human experiment reproduces the GPT decision tasks across risk, time, social, and food preferences, using randomized task parameters and incentives. It was preregistered and conducted with a representative US sample recruited through Prolific.
- Experimental design: Subjects were randomly assigned to three conditions covering parallel task formats including baseline, price framing, and discrete choice.The task text and random-parameter generation matched those used in the GPT experiment.
- Experimental design: Each preference domain contains 25 decision tasks, and the order of the four domains is randomized for each subject.The domains are risk, time, social, and food preference.
- Incentives: Participants received a $6 participation fee, while bonus eligibility depended on randomly selected decisions and chance.For food preference, selected subjects received a fixed $50 bonus instead of an implemented decision.
- Sample and preregistration: 347 unique subjects participated, with more than 110 subjects per experimental condition.The study was conducted in July 2023 and had a median duration of 30.5 minutes.
- Risk preference: The risk tasks allocate 100 points between two assets whose returns are presented either directly or through price-like conversion rates.Comprehension questions test probabilities, returns, and whether a specified outcome can occur.
- Time preference: The time tasks allocate points between today and one month later, with returns paid immediately or by a later-cashable check.The task varies the return rates and asks participants to report allocations and resulting returns.
B.8 Explanations on Incentive Implementation
The incentive implementation explains how selected decisions would generate payments or outcomes across risk, time, social, and food tasks. It also introduces the expenditure-share specification used to estimate preferences.
- Random implementation: One decision is randomly selected from a selected subject’s completed tasks for implementation and bonus determination.The study first selects one subject from every 30 subjects for additional bonuses.
- Risk incentives: Risk-task outcomes are determined by a random draw: values from 0 to 0.5 produce Asset A’s return, while higher values produce Asset B’s return.This mechanism is stated for both direct-return and price-framed risk tasks.
- Time incentives: Time-task payments are made immediately for today’s return and one month later for the deferred return.The deferred payment is described as a check cashable in one month.
- Social incentives: Social-task decisions determine bonuses for both the participant and a newly matched anonymous subject.Each person receives the return allocated to them.
- Food incentives: Food-preference tasks are hypothetical and pay a fixed $50 rather than implementing the selected food allocation.This applies to both direct and price-framed food tasks.
- Preference estimation: The econometric specification uses expenditure shares derived from first-order conditions for utility-maximizing choices given relative prices.The first-order condition links relative quantity responses to relative price changes conditional on ρ.
C.2 GARP Test Power Analyses
The analyses benchmark the budget sets’ ability to detect GARP violations and compare rationality across GPT observations, human subjects, and simulated random choice.
- Bronars Power: 25 choices are drawn from randomly generated budget sets for GPT observations, human subjects, and hypothetical simulated subjects.The simulated subjects choose uniformly among allocations on each budget line.
- Benchmark: 99.9% of hypothetical simulated subjects reject GARP.This benchmark supports the designed budget sets’ power to detect rationality violations.
- Comparative Rationality: GPT observations and human subjects outperform hypothetical simulated subjects in rationality tests across risk, time, social, and food preference.The comparison uses the baseline condition and evaluates choices against the simulated benchmark.
- Selten Score: The average Selten score for GPT observations and human subjects is significantly larger than 0 in all four preference domains.The significance level is p < 0.01 using a two-sided two-sample t-test.
- Bootstrap Power: Bootstrap analyses create 10,000 synthetic subjects by randomly drawing 25 choices from the actual budget sets.The ex post robustness check estimates the probability of rejecting GARP.
- Bootstrap Power: The bootstrap results align with the preference parameter estimates, while the method relies on heterogeneity of preferences among subjects.Indifferent preferences may produce a low probability of GARP violation.
D Result
The results compare rationality distributions across four preference domains and examine demand relationships under baseline, temperature, price-framing, and discrete-choice conditions. The figures distinguish GPT observations, human subjects, and simulated subjects where applicable, while demographic comparisons report mostly insignificant differences except for αt.
- Rationality Scores: Cumulative distributions compare GPT observations, human subjects, and simulated subjects across CCEI, HMI, MPI, and MCI in four preference domains.Separate figures cover risk, time, social, and food preference.
- Baseline Demand Relationships: The risk, time, social, and food figures plot quantities share against log-price ratio across 100 trials, with 25 scatter points and a fitted line per trial.The axes are ln(pA/pB) and xA/(xA + xB).
- Temperature Variations: Temperature-variation figures show cumulative GPT CCEI distributions and Spearman correlations between ln(xA/xB) and ln(pA/pB) across the four domains.The Spearman correlation serves as a proxy for the degree of downward-sloping demand.
- Price Framing: Price-framing figures plot observed quantities shares against log-price ratios across 100 trials in each preference domain.The figures include risk, time, social, and food preference conditions.
- Price Framing: With price framing, cumulative CCEI and Spearman-correlation distributions distinguish GPT observations from human subjects across the four preference domains.Solid lines represent GPT observations and dark dashed lines represent human subjects.
- Demographics: Demographic comparisons find all tested differences statistically insignificant at the 10% level, except αt in the reported GPT variations.Other comparisons are reported as statistically significant at the 1% level in the corresponding analysis.