Source-linked AI summary

Assessing mentalization in humans and large language models

Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang

arXiv:2608.26291v1cs.AIq-bio.NC

TL;DR

The paper asks whether LLMs can use mentalization to guide adaptive strategic behaviour, rather than merely perform well on story-based theory-of-mind tasks. It tests LLMs and humans in two economic games using cognitive computational models of recursive strategies. LLMs showed provider- and size-dependent mentalization, strategic prompting generally improved performance, and GPT-5 adapted its reasoning depth to increasingly sophisticated opponents.

  • Problem

    Whether LLMs use mentalization to guide adaptive behaviour remains unclear because story-based theory-of-mind tests can obscure the mechanisms underlying their choices.

  • Method

    The study compares LLMs and humans in the inspection game and rock-paper-scissors, fitting computational models of reinforcement learning, first- and second-order recursive strategies, and mixed equilibrium.

  • Results

    Across both games, LLMs showed distinct mentalization strategies by provider and size; Social Chain-of-Thought generally improved performance, while GPT-5 flexibly adapted recursive reasoning to opponent sophistication.

  • Takeaways & Limitations

    Cognitive computational modeling provides a formal framework for detecting differences in LLM mentalization and evaluating prompting effects beyond surface-level choice patterns.

  • Takeaways & Limitations

    The study cannot provide a mechanistic explanation of mentalization because the evaluated models are closed-source, and its findings may depend on the selected providers, parameters, and prompts.

Abstract

from arXiv · show

Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.

Main

The paper examines whether LLMs use mentalization to guide strategic behaviour, addressing uncertainty left by story-based theory-of-mind assessments. It uses economic games and cognitive computational modeling to distinguish behavioural and latent reasoning strategies.

  • Motivation: Story-based tests show LLMs can interpret mental states, but failures on modified or complex theory-of-mind tasks leave the underlying mechanism uncertain.These results raise the possibility that apparent mentalization reflects shallow heuristics rather than flexible reasoning.
  • Method: Economic games place participants in direct strategic interactions and allow recursive mentalization to be tested by varying opponents’ reasoning depth.Unlike story-based tasks, players must act against an opponent rather than infer a fictional character’s beliefs as an observer.
  • Method: Cognitive computational modeling quantifies latent cognitive processes from choice behaviour and can reveal trial-by-trial adaptation in social decision-making.Prior human work used these models to identify adaptive mentalization, in which recursive depth changes with the inferred opponent strategy.
  • Study design: The study tests LLMs in the inspection game and rock-paper-scissors, compares them with human participants, and evaluates Social Chain-of-Thought prompting.The design includes 2,099 LLM agents and 251 human participants across both tasks.

Results

Across the inspection game and RPS, model-based analyses revealed distinct mentalization strategies across humans and LLMs. Prompting shifted several LLMs toward more sophisticated opponent-tracking, while GPT-5 adapted effectively to increasingly sophisticated opponents and outperformed humans.

  • Assessing mentalization using the inspection game: SCoT prompting increased inspection-game payoffs across LLMs, with benefit size varying by model.The reported effect was smallest for DeepSeek, intermediate for Gemini, and largest for GPT-4.1.
  • Assessing mentalization using the inspection game: GPT-4.1-SCoT predicted opponent choices above chance, Gemini-SCoT below chance, and DeepSeek-SCoT at chance.Mean switch frequency positively correlated with mean payoff across groups (r = 0.735, R2 = 0.540).
  • Assessing mentalization using the inspection game: Human inspection-game choices were best fit by second-order mentalization and influence play, whereas SCoT shifted DeepSeek toward fictitious play and GPT-4.1 toward influence play.GPT-5 was best fit by influence play without an explicit Chain-of-Thought cue.
  • Assessing mentalization using the inspection game: Human influence-play parameters mapped significantly onto switching, task performance, and adaptation to the employer’s last move.Second-order weight κ was associated with switching and performance, while first-order learning rate η tracked the employer’s latest action.
  • Assessing adaptive mentalization using rock-paper-scissors: GPT-5 earned significantly higher RPS payoffs than humans and both DeepSeek variants, outperforming every group at each bot level.GPT-5’s advantage over humans was large (d = 1.58).
  • Assessing adaptive mentalization using rock-paper-scissors: With increasing opponent sophistication, humans and DeepSeek declined in performance, whereas GPT-5 improved and DeepSeek-SCoT showed a less severe decline.GPT-5 was the only group with a positive performance slope (β = 0.057).
  • Assessing adaptive mentalization using rock-paper-scissors: Opponent sophistication significantly interacted with model group, indicating that adaptation patterns differed across agents.The model × opponent-level interaction was significant (F(4,740) = 64.55, p = 3.51×10-43).
  • Assessing adaptive mentalization using rock-paper-scissors: Humans adapted strongly to zero-order opponents, partially to first-order opponents, and not to second-order opponents.The cited comparisons show reliable adaptation at k=0, partial adaptation at k=1, and no significant adaptation at k=2.

Discussion

Across strategic games, LLMs displayed model-dependent recursive mentalization, while computational modeling revealed both adaptive reasoning and important limits on mechanistic interpretation. GPT-5 adapted across opponent sophistication, but prompting benefits and model capabilities varied by task and provider.

  • LLMs showed varying mentalization strategies across two strategic games, with significant differences across model providers.
  • SCoT prompting improved performance and reflected more sophisticated reasoning, but its mentalization benefits were context-dependent and differed across tasks.
  • Interpretations of latent mentalization remain limited because generative accuracy differed across groups and closed-source models prevent mechanistic explanation.
  • Cognitive computational modeling quantified latent social-cognitive processes and detected differences in mentalizing behaviour across model providers and prompting interventions.
  • GPT-5 accurately calibrated recursive reasoning across all opponent levels, whereas DeepSeek struggled with complex reasoners and was not significantly improved by SCoT.
  • Model configuration, prompt settings, fixed ten-trial histories, and the exclusion of open-source models constrain generalization beyond the tested conditions.

Methods

The study used repeated strategic games with algorithmic opponents to assess recursive mentalization in humans and LLMs. Computational models represented increasingly sophisticated strategies, including adaptive recursive reasoning.

  • Experimental procedure: Participants played repeated inspection-game and rock-paper-scissors tasks against algorithmic opponents implementing different reasoning levels.In the inspection game, participants were employees choosing to work or shirk; in rock-paper-scissors, opponents varied across k = 0, 1, 2.
  • Computational modeling: Computational models ranged from reinforcement learning to fictitious play and influence play, representing zero-, first-, and second-order recursive strategies.The models were fitted to participants’ choices; influence play integrated first- and second-order beliefs, while the mixed-equilibrium model provided a non-updating reference.
  • Computational modeling: The CHASE model represented strategic agents through iterative best responses and adaptive agents through beliefs about the opponent’s reasoning level.Higher recursive levels simulated the opponent’s response process, while adaptive agents integrated evidence about different levels over time.

Competing interests

The authors declare no competing interests and report using agentic workflows for text screening and AI-assisted programming, with programming outputs verified by the lead author.

  • The authors declare no competing interests.
  • Agentic workflows using Claude Code were used to screen supplied text and assist programming, with all programming outputs verified by the lead author.
Loading 2608.26291v1…