Source-linked AI summary

Behavioral Economics of AI: LLM Biases and Corrections

Pietro Bini, Lin William Cong, Xing Huang, Lawrence J. Jin

arXiv:2602.09362v1econ.GNcs.AI

TL;DR

The paper asks whether LLMs exhibit systematic behavioral biases in economic and financial decisions and how those biases can be mitigated. It applies experiments drawn from cognitive psychology and experimental economics across LLM families, versions, and scales. The results show human-like preference responses with greater model advancement or scale, largely rational belief responses in advanced large-scale models, and reduced biases when models are prompted to decide rationally.

  • Problem

    The paper examines whether AI algorithms and agents behave systematically in economic and financial decisions, addressing limited knowledge about the behavioral economics of AI.

  • Method

    The study applies sixteen experimental questions from cognitive psychology and experimental economics to prominent LLM families across model versions and scales.

  • Results

    LLMs become increasingly irrational and human-like in preference-based tasks as models become more advanced or larger, while advanced large-scale models are largely rational in belief-based tasks.

  • Takeaways & Limitations

    Prompting LLMs to behave as rational investors reduces biases, whereas retail-investor prompting and human-bias summaries can produce less rational or counterproductive responses.

  • Takeaways & Limitations

    Debiasing LLMs remains a challenging task.

Abstract

from arXiv · show

Do generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions? If so, how can these biases be mitigated? Drawing on the cognitive psychology and experimental economics literatures, we conduct the most comprehensive set of experiments to date$-$originally designed to document human biases$-$on prominent LLM families across model versions and scales. We document systematic patterns in LLM behavior. In preference-based tasks, responses become more human-like as models become more advanced or larger, while in belief-based tasks, advanced large-scale models frequently generate rational responses. Prompting LLMs to make rational decisions reduces biases.

1. Introduction

The paper introduces behavioral economics of AI to systematically study LLM biases in preference- and belief-based economic decisions, using experiments adapted from psychology and experimental economics. It finds that larger or more advanced models become more human-like and irrational on preference tasks but more rational on belief tasks, while rational-decision prompting reduces bias.

  • The experiments adapt cognitive-psychology and experimental-economics tasks covering preference-based and belief-based questions relevant to financial decision making.
  • The study collects responses from four prominent LLM families and compares model responses with rational and human responses across versions and scales.
  • For preference-based questions, more advanced or larger models produce increasingly human-like and Expected Utility-irrational responses.
  • For belief-based questions, more advanced or larger models increasingly produce rational responses, with three advanced large-scale models generating predominantly rational answers.
  • A brief instruction to use the Expected Utility framework makes responses more rational and less human-like, whereas other approaches can be ineffective or counterproductive.
  • The paper treats generative AI agents as a novel class of economic agents and establishes behavioral economics of AI as a field for systematic study.

2. Experimental Design

The study adapts cognitive psychology and experimental economics experiments to systematically compare LLM decisions and beliefs with rational and human benchmarks. It evaluates preference- and belief-based tasks across model families, versions, and scales using structured prompts.

  • The experiments cover psychology of preferences and psychology of beliefs, including risk, time preferences, forecasting, and investor behavior.
  • The authors ask LLMs the same experimental questions used to document systematic human behavioral biases.
  • LLM responses are classified as rational, human-like, or non-human-like according to their relation to rational benchmarks and majority-human responses.
  • Forecasting experiments compare the true persistence parameter ρ with the perceived coefficient implied by LLM forecasts.
  • Investment experiments replicate a human stock-allocation task and test whether the preference factors explaining human decisions also explain LLM decisions.
  • The model sample includes benchmark, smaller-scale, and predecessor models from ChatGPT, Anthropic Claude, and Google Gemini families.

3. Behavioral Biases of LLMs

LLMs show systematic behavioral heterogeneity: preference-based responses become more human-like with model advancement or scale, while belief-based responses become more rational. These patterns vary across model families and forecasting horizons.

  • Most LLM responses are classified as either rational or human-like, with “other” responses occurring only in a few cases.
  • Preference-based responses tend to be more human-like, whereas belief-based responses tend to be more rational.
  • LLM families differ meaningfully: Gemini models are 22.9% less likely to produce rational preference responses and 16.7% more likely to produce human-like responses than GPT models.
  • For belief-based questions, Llama models are 25.0% less likely to produce rational responses and 21.0% more likely to produce human-like responses than GPT models.
  • With greater model advancement or scale, preference-based responses become less rational and more human-like.
  • With greater model advancement or scale, belief-based responses become more rational and less human-like.
  • For advanced large-scale models, short-term forecasts are largely rational, while longer-term forecasts produce human-like biases and overstate persistence.
  • Detailed information about the data-generating process can produce more human-like biases in forecasts.

4. Correcting LLM Biases

Role priming can reduce LLM biases, especially when models are instructed to act as rational investors, but detailed debiasing prompts are ineffective or counterproductive. The effects appear to operate partly through changes in confidence and reasoning type.

  • Role priming: 4.3% and 3.3%: rational-investor role priming increases rational responses for preference-based and belief-based questions, respectively.The preference-based effect is significant at the 5% level, while the belief-based effect is significant at the 10% level.
  • Role priming: 3.9%: retail-investor role priming reduces rational responses for preference-based questions, with no statistically significant belief-based effect.This contrast indicates that the role specified in the prompt matters for the direction of the response.
  • Mechanisms: 6.6% and 11.4%: rational-investor priming raises high confidence and system 2 thinking, whereas retail-investor priming lowers them by 6.0% and 8.0%.All reported confidence and reasoning-type effects are statistically significant.
  • Mechanisms: 40% and 38%: system 2 thinking increases rational responses in the two panels, while high confidence decreases human-like responses by 35% and 58%.Controlling for confidence and reasoning type attenuates the estimated role-priming effects, supporting these variables as mediating pathways.
  • Detailed debiasing techniques: The expected-utility procedure is effective in reducing biases, whereas supplying Kahneman and Tversky’s findings reduces rational responses by about 29% and increases human-like responses by about 20%.The authors interpret the latter pattern as evidence that additional information can hinder rational responses through information overload.
  • Overall assessment: Overall, the detailed debiasing techniques are either ineffective or counterproductive, and debiasing LLMs remains challenging.The improvement from role priming is described as economically modest.

5. Conclusion

The paper documents systematic and heterogeneous behavioral biases across LLM families, with different patterns for preference- and belief-based questions. Rational-investor prompting reduces biases, whereas useful-information prompts can be ineffective or counterproductive.

  • The study examines four prominent LLM families using experimental designs from cognitive psychology and experimental economics.It treats established human-bias experiments as a framework for studying LLM behavior.
  • LLM responses become increasingly irrational and human-like as models become more advanced or larger on preference-based questions.The preference-based results concern questions studying human preferences and include prospect-theory-related tasks.
  • Advanced large-scale LLMs generate largely rational responses on belief-based questions, contrasting with the preference-based pattern.The paper therefore finds that scaling does not produce one uniform behavioral response across task types.
  • Responses differ substantially across the four LLM families, motivating systematic comparisons with rational and human responses.The paper also contributes a basis for ongoing evaluations of behavioral biases.
  • Prompting LLMs to act as rational investors reduces biases, while retail-investor prompts produce less rational responses.Role priming affects responses through changes in confidence and reasoning type, but its overall impact is economically modest.
  • Providing useful information, including Expected Utility guidance or summaries of human biases, does not reduce LLM biases and can be counterproductive.For preference-based questions, a 3-4% reduction is small relative to baseline irrational-response shares often exceeding 50%.

Figures and Tables

The figures and tables organize the paper’s behavioral-bias experiments, model comparisons, response classifications, and forecast analyses. They emphasize differences between preference-based and belief-based questions and heterogeneity across model families, generations, and scales.

  • Response classifications: The four advanced large-scale models are evaluated by the proportions of responses classified as rational, human-like, or other across six preference-based and ten belief-based questions.The models are GPT-4, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3 70B.
  • Response classifications: Belief-based questions tend to elicit more rational responses from LLMs than preference-based questions.The comparison appears in the figure’s left and right panels and is confirmed by the accompanying table.
  • Model heterogeneity: Radar charts compare predominantly rational and human-like responses across preference-based and belief-based questions, including advanced large-scale, smaller-scale, and older models.The comparisons span model generations and scales.
  • Forecast analyses: Forecast figures plot perceived persistence ˆρ against true ρ, with confidence intervals and a 45-degree line representing full-information rational expectations.Figure 4 covers Experiment 1, while Figure 5 covers Experiments 2 and 3.
  • Experimental questions: Sixteen experimental questions drawn from cognitive psychology cover preference-based and belief-based biases, including prospect theory, framing, discounting, ambiguity, sample-size, base-rate, conjunction, gambler’s, confirmation, anchoring, and overconfidence effects.The questions are summarized across preference and belief panels.
  • Regression comparisons: The regression tables estimate how advanced status and model scale relate to rational and human-like classifications separately for preference-based and belief-based questions.Older models serve as the baseline in one comparison, while smaller advanced models serve as the baseline in another.

I. Additional Figures

Additional figures extend the response-classification comparisons to advanced small-scale and older LLMs. They retain the distinction between preference-based and belief-based questions and classify responses as rational, human-like, or other.

  • Advanced small-scale models: Advanced small-scale models are compared across six preference-based and ten belief-based questions using rational, human-like, and other response categories.The models are GPT-4o, Claude 3 Haiku, Gemini 1.5 Flash, and Llama 3 8B.
  • Older models: Older LLM versions are evaluated with the same preference-based and belief-based response classification framework.The models are GPT-3.5 Turbo, Claude 2, Gemini 1.0 Pro, and Llama 2 70B.

II. Prompt Design

The prompts operationalize established cognitive-psychology experiments for LLM responses. They cover prospect-theory preferences, time and ambiguity preferences, and belief-related biases using structured response formats.

  • Preference-based prompts: The prompt set includes prospect-theory questions on diminishing sensitivity, loss aversion, and probability weighting, plus narrow framing, hyperbolic discounting, and ambiguity aversion.These questions adapt established experiments from Kahneman and Tversky, Barberis and Thaler, Frederick and coauthors, Ellsberg, and related work.
  • Response format: The prompts require structured outputs containing choices or probabilities together with confidence, explanations, and reasoning.Some tasks instead request scenario-specific estimates, directions, or intervals.
  • Belief-based prompts: Belief-based prompts assess sample size neglect, base-rate neglect, conjunction fallacy, gambler’s fallacy, confirmation bias, anchoring, and overconfidence.The overconfidence prompts include overprecision and overestimation tasks.
  • Belief-based prompts: The anchoring task asks models to compare an initial randomly generated number with their estimate before providing the estimate.The example uses a wheel-of-fortune number and then requests an upward or downward estimate.
  • Belief-based prompts: The overprecision task requests lower and upper bounds for a set of general-knowledge questions.The questions are adapted from Deaves, Lüders, and Luo (2009).
Loading 2602.09362v1…