Source-linked AI summary

Can Large Language Models Serve as Rational Players in Game Theory? A Systematic Analysis

Caoyun Fan, Jindou Chen, Yaohui Jin, Hao He

arXiv:2312.05488v2cs.AIcs.CLcs.GT

TL;DR

The paper addresses the unclear capability boundaries of LLMs used as substitutes for humans in game-theoretic social-science experiments. It evaluates rationality through desire formation, belief refinement, and optimal action across three classical games, finding substantial disparities from humans even for GPT-4. The authors therefore urge greater caution when introducing LLMs into such experiments.

  • Problem

    Existing research uses LLMs in game experiments and other social-science settings, but their capability boundaries in game theory remain unclear.

  • Method

    The study evaluates GPT-3, GPT-3.5, and GPT-4 on desire formation, belief refinement, and optimal action using dictator, Rock-Paper-Scissors, and ring-network games.

  • Results

    LLMs show substantial weaknesses: uncommon preferences impair desire formation, many simple patterns defeat belief refinement, and refined beliefs may be overlooked or modified during action selection.

  • Takeaways & Limitations

    LLMs should be introduced into social-science game experiments with greater caution, especially when tasks require complex belief refinement.

  • Takeaways & Limitations

    The selected games are relatively easy and not close enough to real game scenarios, while the analysis considers only rationality and lacks extensive comparative and ablative experiments.

Abstract

from arXiv · show

Game theory, as an analytical tool, is frequently utilized to analyze human behavior in social science research. With the high alignment between the behavior of Large Language Models (LLMs) and humans, a promising research direction is to employ LLMs as substitutes for humans in game experiments, enabling social science research. However, despite numerous empirical researches on the combination of LLMs and game theory, the capability boundaries of LLMs in game theory remain unclear. In this research, we endeavor to systematically analyze LLMs in the context of game theory. Specifically, rationality, as the fundamental principle of game theory, serves as the metric for evaluating players' behavior -- building a clear desire, refining belief about uncertainty, and taking optimal actions. Accordingly, we select three classical games (dictator game, Rock-Paper-Scissors, and ring-network game) to analyze to what extent LLMs can achieve rationality in these three aspects. The experimental results indicate that even the current state-of-the-art LLM (GPT-4) exhibits substantial disparities compared to humans in game theory. For instance, LLMs struggle to build desires based on uncommon preferences, fail to refine belief from many simple patterns, and may overlook or modify refined belief when taking actions. Therefore, we consider that introducing LLMs into game experiments in the field of social science should be approached with greater caution.

Introduction

This paper argues that LLMs should be systematically evaluated as game-theory players because existing work leaves their capability boundaries unclear. Using rationality as the framework, it examines desire formation, belief refinement, and optimal action across three classical games.

  • Introduction: Existing research often uses LLMs as social-science subjects without systematically analyzing which game-theoretic tasks they can perform.The paper frames this gap through questions about suitable subjects, games, and game processes for LLMs.
  • Introduction: The study evaluates LLM rationality through clear desires, refined beliefs about uncertainty, and optimal actions based on both.These characteristics correspond to forming concrete preferences, judging uncertainty from game information, and combining desire with belief when acting.
  • Introduction: Systematic analysis is intended to clarify LLM capability boundaries and guide their wider use in social-science research.The paper presents this work as a basis for introducing LLMs into social-science research more smoothly.
  • Introduction: Across dictator, Rock-Paper-Scissors, and ring-network games, LLMs show basic desire formation but struggle with uncommon preferences, patterned belief refinement, and belief-guided action.The results also report human-like GPT-4 performance for certain Rock-Paper-Scissors patterns and improved action-taking when game behavior is explicitly decomposed.

Related Work

Prior studies report human-like or theoretically consistent LLM behavior in several social-science settings, but systematic evidence about their capability boundaries remains limited. The paper presents game theory as a useful framework because its experiments are operable, analyzable, and broadly abstractable.

  • Related Work: LLMs have been used as human substitutes in social-science experiments involving fairness, consumer behavior, finance, and psychology.Reported examples include framing effects, economic demand behavior, budget allocation, and consistency with mainstream social values.
  • Related Work: Existing studies show rational or human-aligned LLM behavior in selected experiments but do not systematically establish their social-science capability boundaries.The limitation concerns the scope of available evidence rather than a universal failure of LLMs as research subjects.
  • Related Work: Game theory provides a framework for analyzing and predicting rational-player behavior under uncertainty across economic settings such as competition, auctions, and pricing.The paper situates game theory as a mathematical framework originally developed in economics.
  • Related Work: Game theory offers strong operability, analyzability, and generalization because its experiments are simple, theoretically grounded, and abstract many social-science phenomena.These properties motivate using game-theoretic experiments to examine LLM behavior.

Preliminaries of Game Theory

The paper models a game through information, actions, consequences, a consequence function, and a preference-determined desire function. Rational play then requires forming beliefs about uncertainty and choosing actions that maximize desire.

  • Preliminaries of Game Theory: A game is modeled with information I, action set A, consequence set C, consequence function g, and desire function D_c determined by preference P.The desire function ranks consequences: a player prefers c1 to c2 exactly when D_c(c1) > D_c(c2).
  • Preliminaries of Game Theory: Belief theory represents uncertainty through a subjective probability distribution derived from game information I.The model uses belief Ω_I, its distribution p(Ω_I), and a consequence function g: A × Ω_I → C.
  • Preliminaries of Game Theory: The analysis assumes that uncertainty arises only from the opponent’s action, and all games in the study satisfy this assumption.This narrows the uncertainty modeled in the paper’s game-theoretic framework.
  • Preliminaries of Game Theory: The framework maps rationality to three operations: building desire D(·), sampling belief ω from p(Ω_I), and choosing the action that maximizes desire.The paper identifies these operations with clear desires, belief refinement, and optimal action-taking, respectively.

LLMs in Game Theory

The study evaluates whether LLMs exhibit three characteristics of rational players by testing GPT-3, GPT-3.5, and GPT-4 in dictator, Rock-Paper-Scissors, and ring-network games. It uses these classical games to examine desire formation, belief refinement, and optimal action-taking.

  • LLMs in Game Theory: The experiments test GPT-3, GPT-3.5, and GPT-4 across dictator, Rock-Paper-Scissors, and ring-network games.The three games are selected to evaluate the three rational-player characteristics.

Can LLMs Build A Clear Desire?

The dictator game tests whether LLMs can translate textual preferences into consistent desires and allocation choices. LLMs generally handle common preferences but struggle with uncommon ones, with errors reflecting mathematical or preference confusion.

  • Game design: The dictator game isolates desire formation because the recipient always accepts, eliminating uncertainty about the game outcome.Its diverse preferences generate different desire functions, while the fixed belief prevents biased-belief interference from affecting the analysis.
  • Experimental setup: The experiment assigns equality, common-interest, self-interest, or altruistic preferences through textual prompts and evaluates preference-consistent choices across allocation pairs.Each preference was tested in three games, repeated 10 times, with accuracy reported for each model.
  • Findings: LLMs consistently followed common equality and self-interest preferences, but accuracy dropped sharply for uncommon common-interest and altruistic preferences.All three models made preference-consistent choices for EQ and SI; CI produced sporadic errors, while AL caused especially frequent failures.
  • Error analysis: A case study attributes GPT-3 errors under altruism to number confusion and GPT-3.5 errors to assuming joint-income maximization means maximizing the recipient’s income.GPT-4’s analysis and choice were consistent with humans in this case.
  • Implication: The results suggest that explicit, specific preference explanations may help LLMs when game experiments use uncommon preferences.The paper characterizes this as a basic ability to form clear desires from textual prompts, bounded by difficulty with uncommon preferences.

Can LLMs Refine Belief?

The paper tests whether LLMs can refine beliefs about opponents’ actions in Rock-Paper-Scissors from historical patterns. GPT-4 succeeds on some cyclical patterns, but LLMs generally struggle with belief refinement across many simple patterns.

  • Results: GPT-4 consistently chose correct actions after approximately 3 rounds against constant-action opponents, while GPT-3 remained close to random guessing.GPT-3.5 performed above random guessing and its average payoff continued to rise.
  • Results: GPT-4 appeared able to refine beliefs from loop-2 and loop-3 patterns as payoffs rose with updated historical records, whereas GPT-3 and GPT-3.5 captured cycles without taking correct actions.The loop experiments included R-P, P-S, S-R, and R-P-S sequences.
  • Results: For copy and counter patterns, GPT-4 showed only a slight advantage, and overall LLM performance remained insufficient.These patterns determine the opponent’s action from the player’s previous action under a Markov assumption.
  • Results: LLMs generally failed to refine beliefs well across most tested opponent patterns, unlike humans, for whom these patterns were easy to identify.Performance was similar to random guessing when opponent actions were sampled from a preference distribution.
  • Results: GPT-3.5 sometimes recognized a P-R-S loop but still denied a specific pattern, whereas GPT-4 summarized the pattern and became more confident as records accumulated.Figure 4 compares the models’ analyses on loop-3.
  • Implication: The study concludes that belief-refinement ability remains immature, supporting cautious use of LLMs in game experiments requiring complex belief refinement.The authors nevertheless view GPT-4’s performance on one pattern as encouraging for future models.

Can LLMs Take Optimal Actions?

The ring-network game tests whether LLMs can combine beliefs about an opponent’s action with desires to choose optimal actions. LLMs perform better when belief is made explicit or given, but may overlook or modify refined beliefs during action selection.

  • Game and method: The ring-network game models sequential reasoning: the player infers the opponent’s optimal action from the opponent’s payoff matrix, then selects its own action from that belief.The setup uses two players with two discrete actions and evaluates optimal-action accuracy across repeated trials under original and swapped payoff matrices.
  • Experimental setup: The experiment varied payoff difficulty by reducing payoff differences, increasing the incorrect action’s expected payoff, or decreasing the correct action’s expected payoff.The opponent’s payoff matrix remained constant so the target belief about the opponent’s action stayed fixed; payoff labels were also swapped to reduce action-name bias.
  • Results: LLMs performed poorly when they had to infer belief and take action implicitly, whereas explicitly separating belief refinement from action selection significantly improved accuracy.This implicit process is closest to how human players’ belief is characterized, but GPT-4 was almost completely unable to take the optimal action in that setting.
  • Results: In explicit belief prompting, LLMs refined the opponent’s action with accuracy above 0.95, but GPT-4 achieved only about 0.70 accuracy when subsequently selecting its own action.GPT-3.5 performed even worse in the subsequent action-selection step, despite the accurate belief refinement.
  • Results: Given belief produced more reliable action selection: GPT-4 consistently chose the optimal action, while GPT-3.5 exceeded 0.80 accuracy.The two belief forms contain the same content, but LLMs were more successful when the belief was supplied directly rather than refined in an earlier dialogue turn.
  • Error analysis: Errors arose when LLMs overlooked refined beliefs amid game information or modified correct beliefs because they lacked confidence in them.Overlooking was associated mainly with GPT-3.5, while belief modification occurred more frequently in GPT-4; Figure 7 presents these cases.

Conclusion

The paper systematically evaluates whether LLMs can act as rational players in game theory and identifies weaknesses across three rationality components. It concludes that the analysis is preliminary and limited by simple games, a narrow rationality perspective, and relatively rough experimentation.

  • Conclusion: The study evaluates LLM rationality in game theory through three aspects and identifies weaknesses in their game-theoretic abilities.The authors frame this as an analysis of whether LLMs can serve as rational players.
  • Limitations: The study is limited by relatively easy games, a perspective focused only on rationality, and rough analyses lacking more comparative and ablative experiments.The authors also note that the selected games are not close enough to real game scenarios.
  • Future work: Future work includes multi-agent games, human–LLM confrontation, dynamic games in real scenarios, and targeted training for identified ability deficiencies.The paper characterizes research on LLMs in game theory as still preliminary.
Loading 2312.05488v2…