Source-linked AI summary
M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games
Sixiong Xie, Zhuofan Shi, Haiyang Shen, Yun Ma, Xiang Jing
TL;DR
Existing benchmarks often emphasize isolated capabilities and behavioral outcomes, leaving reasoning and communication processes insufficiently evaluated. M3-BENCH addresses this with 24 mixed-motive games and three complementary views of behavior, reasoning, and communication. It reveals recurring reasoning–communication mismatches, stronger cross-view coherence in humans than in top-performing LLMs, and latent risks hidden by outcome-only metrics.
Problem
Existing benchmarks often focus on single social capabilities and observable outcomes, providing limited support for characterizing reasoning, communication strategy, and latent social intent.
Method
M3-BENCH evaluates 24 mixed-motive games through Behavioral Trajectory Analysis, Reasoning Process Analysis, and Communication Content Analysis.
Results
Reasoning-oriented models show an overthink–undercommunicate pattern, while humans remain more coherent across action, reasoning, and communication despite top LLMs surpassing humans on aggregate task performance.
Takeaways & Limitations
Process-aware, three-view evaluation reveals module-specific weaknesses and latent safety risks that outcome-only scores can miss.
Takeaways & Limitations
Mixed-motive games simplify real social interaction, which often involves longer horizons, richer context, and weaker structural constraints.
Abstract
from arXiv · showhide
Existing benchmarks for LLM agents' social behavior typically focus on a single capability dimension and evaluate only behavioral outcomes, overlooking process signals from reasoning and communication. We present M3-BENCH, a benchmark of 24 mixed-motive games with a process-aware evaluation framework spanning three complementary views: Behavioral Trajectory Analysis (BTA), Reasoning Process Analysis (RPA), and Communication Content Analysis (CCA). Evaluating 11 frontier LLMs and a human baseline, M3-BENCH reveals substantial differences in social competence that outcome-only evaluation misses. In particular, we identify an "overthink-undercommunicate" pattern: reasoning models achieve strong internal deliberation scores but often fail to translate them into effective social communication. Although top models can surpass humans on task outcomes, humans exhibit markedly higher cross-view consistency, suggesting that current LLM agents still lack the behavioral coherence characteristic of human social competence. Our analysis further shows that the three-view decomposition surfaces safety-relevant risks, such as cooperative behavior paired with latent opportunistic reasoning, that remain hidden under outcome-only metrics.
1 Introduction
Existing social-behavior benchmarks often isolate one capability and emphasize outcomes, leaving reasoning, communication, and latent intent undercharacterized. M3-BENCH addresses this gap with mixed-motive games and a three-view process-aware analysis of what agents do, think, and say.
- Motivation: Existing benchmarks often focus narrowly on cooperation or deception and emphasize observable outcomes such as win rate, cooperation rate, or goal completion.This leaves mixed settings involving intertwined cooperation, competition, and deception difficult to evaluate.
- Motivation: Outcome-only evaluation can mistake sustained cooperation for genuine prosociality when an agent strategically exploits trust later.Reasoning logic, communication strategy, and latent social intent are therefore important process-level signals.
- Motivation: Mixed-motive games expose tensions between self-interest, prosociality, short-term gain, long-term relationships, individual strategy, and social expectations.Their controlled rules require integrating beliefs about others, interaction history, and institutional constraints.
- Approach: Its BTA, RPA, and CCA modules jointly analyze behavioral trajectories, decision reasoning, and communication to characterize what agents do, think, and say.The framework is designed to reveal mismatches and potential risks that outcome-only evaluation misses.
- Approach: M3-BENCH introduces a four-level hierarchical benchmark of 24 mixed-motive games for advanced social-behavior evaluation with human and model baselines.The hierarchy broadens the interaction settings while preserving a controlled evaluation structure.
2 Related Work
Prior work provides social-behavior benchmarks and multi-agent environments, but most evaluations emphasize isolated capabilities or behavioral outcomes. M3-BENCH extends this landscape with mixed-motive tasks and interpretable process-level profiles.
- Existing benchmarks: Existing benchmarks commonly target specific capabilities such as cooperation, coordination, deception, social deduction, or negotiation.This specialization limits coverage of social behavior involving multiple motives simultaneously.
- Existing benchmarks: General game-based platforms and social environments provide structured testbeds but mainly emphasize behavioral outcomes or behavioral generalization.Language-based simulation frameworks broaden interaction settings without replacing outcome-focused evaluation.
- Theoretical foundations: Theory of Mind and social-dilemma research supply foundations for studying mental-state attribution, cooperation, and strategic interaction.Recent work reports that LLM Theory-of-Mind reasoning is partial and inconsistent.
- Positioning: M3-BENCH builds on social-dilemma traditions by covering mixed-motive games across multiple difficulty levels while adding process-level analysis.Its scope extends beyond outcome measures to reasoning and communication signals.
- Positioning: Compared with concurrent benchmarks emphasizing outcomes or training-time feedback, M3-BENCH jointly analyzes process evidence and produces interpretable social-behavior profiles.The framework is intended as a diagnostic complement rather than a single-metric comparison.
3 Method
M3-BENCH evaluates social behavior through 24 progressively complex mixed-motive tasks and three parallel evidence streams for actions, reasoning, and communication. It preserves agreement and disagreement across views to produce diagnostic profiles rather than a single score.
- Three-view framework: Each episode is analyzed through BTA, RPA, and CCA, corresponding respectively to behavioral trajectories, decision rationales, and communicative interaction.Together these streams extend evaluation beyond win rate or cooperation rate to what an agent does, thinks, and says.
- Benchmark design: M3-BENCH contains 24 tasks organized into four progressively complex levels, from dyadic choices to repeated interaction, collective dilemmas, and incomplete-information language games.The hierarchy introduces distinct sources of social complexity in controlled stages.
- Benchmark design: The benchmark uses established game-theoretic structures to provide normative baselines, isolate social competencies, and support human–AI comparison.Classic payoff structures make targeted social competencies easier to examine under controlled conditions.
- BTA: BTA converts action trajectories, payoffs, and game-state information into standardized episode-level behavioral metrics.Depending on the task, outputs can include payoff, cooperation, retaliation, deception, alliance stability, or goal attainment.
- RPA: RPA evaluates decision rationales in context using a zero-shot LLM judge and aggregates structured turn-level scores into episode-level reasoning evidence.The scored dimensions include motivational orientation, opponent modeling, temporal horizon, and belief updating.
- CCA: CCA labels each dialogue message with a 15-category social-pragmatic taxonomy and aggregates annotations into communication features such as style, strategic effectiveness, and speech–action consistency.The output is an episode-level communication evidence vector.
- Diagnostic interpretation: The framework retains BTA, RPA, and CCA in parallel, explicitly examining agreement and disagreement instead of collapsing evidence into one score.An optional portrait layer organizes the evidence for cross-model comparison, while the core evaluation remains the module outputs.
4 Experiments
M3-BENCH evaluates 11 LLMs and human participants across 24 mixed-motive games using standardized protocols and three process-aware views. Results show increasing difficulty at higher levels, a recurring reasoning–communication mismatch, and stronger cross-view coherence in frontier models and humans.
- Experimental Setup: The evaluation covers 11 LLMs, rule-based baselines, and 50 human participants across 24 tasks, with Silent and Comm conditions under a shared protocol.Each model–opponent pairing and condition uses 50 independent episodes.
- Experimental Setup: Table 1 reports three-view module scores for representative games and level averages, using normalized task indicators aggregated into module, level, and overall scores.Main-paper scores use a [0, 100] scale, while method-internal variables remain on native scales.
- Overall Task Performance: Closed-source frontier models are strongest overall, while open-weight and reasoning-oriented models remain competitive at lower levels but face greater difficulty in socially complex settings.Open-weight performance declines more visibly on Level 4, where private information and language-heavy interaction add difficulty.
- Overall Task Performance: Level 4 produces substantially more score variability than Level 1, indicating that incomplete-information and communication-heavy tasks separate models more sharply.High aggregate task scores do not by themselves imply balanced social competence.
- Three-View Process Diagnosis: Reasoning-oriented models show strong RPA but lower CCA, creating large reasoning–communication gaps and weaker cross-view consistency despite competitive overall task scores.This recurring overthink–undercommunicate pattern appears across both representative tasks.
- Three-View Process Diagnosis: Frontier models, especially Claude Opus 4.5 and GPT-5.1, maintain relatively close alignment across action, reasoning, and communication, while humans show the strongest consistency despite lower aggregate task scores.Open-weight models are generally weaker but often more balanced than reasoning models.
5 Discussion
The discussion positions controlled mixed-motive games as a diagnostic complement to open-ended environments and argues for evaluating what agents do, think, and say together. It also qualifies reasoning scores and frames deception indicators as safety diagnostics rather than optimization targets.
- Static Games and Open-Ended Environments: Controlled games make failures easier to attribute to action selection, opponent modeling, or communication than open-ended environments, while extending the framework remains future work.Open-ended settings remain valuable for studying long-horizon adaptation and emergent behavior.
- Interpreting Reasoning Processes: RPA measures the quality and consistency of stated reasoning, not necessarily faithful hidden computation, but mismatches between stated reasoning and observed action remain diagnostically useful.A claimed cooperative intent paired with defection is still important to flag for oversight.
- Safety Interpretation: Deception-related indicators are presented as potential risk factors and red-teaming signals, not as standalone capabilities that the benchmark rewards.The authors recommend interpreting them as pre-deployment warning signs.
- Broader Implications: Cross-view consistency can complement outcome leaderboards as a pre-deployment screening signal, while decomposing failures offers a more actionable improvement path than aggregate success rates alone.The overthink–undercommunicate pattern identifies a specific gap between internal deliberation and external communication.
6 Conclusion
M3-BENCH combines a four-level mixed-motive-game hierarchy with BTA, RPA, and CCA to diagnose social behavior beyond aggregate task scores. Its experiments reveal reasoning–communication mismatches, stronger human coherence, and safety-relevant risks that outcome-only evaluation can miss.
- Benchmark Design: M3-BENCH is a multi-stage benchmark of advanced LLM social behavior using a four-level hierarchy of mixed-motive games.The benchmark evaluates social behavior through progressively structured interaction demands.
- Process-Aware Framework: BTA, RPA, and CCA jointly analyze what agents do, think, and say, supporting diagnostic evaluation beyond aggregate task scores.An optional Big Five and Social Exchange Theory layer provides descriptive organization rather than psychometric claims.
- Main Findings: Reasoning-oriented models repeatedly show strong reasoning but ineffective social communication, while humans remain more coherent across action, reasoning, and communication despite lower aggregate performance.Top LLMs can surpass humans on aggregate task performance, but outcome scores alone may overstate AI social competence.
- Implications: The three-view framework identifies module-specific weaknesses and latent risks that remain hidden under outcome-only evaluation.The results support examining how success is achieved and whether actions, reasoning, and communication remain aligned.
7 Limitations
The benchmark has important scope, measurement, cost, and interpretation boundaries. Its games simplify real social interaction, some modules depend on LLM judges, process-aware evaluation is more expensive, and theory-based layers are not psychometric measurements.
- Mixed-motive games simplify real social interaction by using shorter horizons, less context, and stronger structural constraints.
- RPA and CCA scores may reflect judge-specific preferences or blind spots despite human-comparison and judge-swap checks.
- Process-aware evaluation requires rationale collection, dialogue analysis, and cross-view aggregation, making large-scale evaluation harder.
- The Big Five and Social Exchange Theory layers organize evidence but are not validated personality measurements.
Ethics Statement
The paper addresses dual-use risks from evaluating deception and manipulation. It treats deception-related behavior as diagnostic rather than as a capability to optimize and reports privacy protections for human participants.
- Deception-related behaviors are surfaced as diagnostic signals and safety concerns rather than rewarded as targets to maximize.
- The human study used informed consent, compensation, withdrawal rights, pseudonymous identifiers, and no direct personal identifiers.
- Participants were instructed not to share personal information during interactions, with additional privacy and quality-control details reported in the appendix.
A.1 Case Study
The case study demonstrates how BTA, RPA, and CCA jointly expose a shift from apparently cooperative play toward endgame opportunism. Behavioral cooperation persists while reasoning becomes more self-interested and communication remains non-committal.
- The representative episode tracks behavior, reasoning, and communication across rounds to reveal a latent shift from cooperation to endgame opportunism.
- The 10-round repeated Prisoner’s Dilemma weakens future punishment in final rounds, increasing the incentive to defect.
- Across rounds 1–9, Player A cooperates and sounds reciprocal, but its reasoning increasingly prioritizes self-interest before round-10 defection.
- Player A maintains cooperation-supporting language while avoiding clear final-round commitments, preserving room to defect.
- Three-view analysis distinguishes optimistic behavior from motive drift and limited commitment, showing why cooperation rate alone cannot guarantee trustworthiness.
B Additional Method Details
The additional methods define how CCA labels dialogue, standardize and aggregate BTA/RPA/CCA scores, measure cross-view consistency, and report model identities. The pipeline combines task-level evidence with an interpretive portrait while preserving distinct score scales.
- CCA: CCA assigns each utterance exactly one label from a fixed taxonomy of 15 mutually exclusive social-pragmatic acts.
- CCA: The taxonomy achieved Cohen’s κ = 0.82 on the final pilot validation subset after iterative coding and definition revision.
- CCA: Episode-level CCA features aggregate message labels into style-distribution and strategic-effectiveness families.
- Score aggregation: Task-level BTA, RPA, and CCA dimension scores are standardized to [0, 1] before aggregation and reported on a 0–100 scale.
- Score aggregation: Task weights are uniform by default over tasks where a dimension is defined, with released-code support for level-balanced variants.
- Cross-view consistency: The dispersion statistic σD summarizes cross-view agreement: low values indicate similar module pictures, while high values indicate mismatch.
- Portrait reporting: The global portrait combines module-level scores, cross-view consistency, and diagnostic notes as an interpretive summary over primary evidence.
- Model reporting: Table 3 reports standardized model identifiers, providers, inference modes, and underlying API strings when paper labels denote aliases or configurations.
C.2 Sensitivity to Fusion Weights
Sensitivity analyses show that the benchmark’s overall ranking is robust to alternative fusion weights, while reasoning-oriented models move most because strong RPA can be offset by weaker CCA.
- τ ∈ {0.85, 0.93, 0.91, 0.87}, and the top three models are unchanged across four alternative fusion schemes.The alternatives include behavior-only, behavior-heavy, process-heavy, and no-communication weighting.
- Reasoning-oriented models rise under behavior-only fusion but fall under process-heavy fusion because their RPA advantage is offset by weaker CCA.
- 50 episodes per model–opponent pairing keep BTA standard errors below 0.01 by roughly 30 episodes, while RPA and CCA errors remain below approximately 0.015 by episode 50.
- The first 30 episodes produce rankings with Kendall correlation above 0.95 against the final 50-episode ranking, supporting a practical stability–cost balance.
- Expert-rater agreement and rescoring with an alternative frontier judge preserve the main qualitative findings, including the overthink–undercommunicate pattern.
I Per-Task BTA Indicator Configuration
The per-task BTA configuration specifies indicators, scoring directions, normalization bounds, and weights for each benchmark game. Scores are computed by a transparent, rule-based pipeline from logged trajectories, with task-specific behavioral metrics and uniformly weighted normalized indicators by default.
- Configuration: Each task defines an exact BTA indicator set, scoring direction, raw normalization bounds, and within-task weight.The configuration is summarized for all 24 benchmark tasks, with default weights of 1/|Jτ,BTA| unless noted.
- Behavioral dimensions: Indicators cover cooperation, retaliation, forgiveness, switching, reciprocity, contribution, free-riding, welfare, agreement, punishment, coalition, and payoff efficiency across social dilemmas.The task configurations include repeated cooperation, gift exchange, common-pool resources, voting, and volunteer settings.
- Behavioral dimensions: Incomplete-information and competitive tasks use metrics such as accusation, sabotage detection, faction survival, bidding, bluffing, value betting, and call accuracy.These indicators are configured for hidden-traitor, informant, werewolf, auction, and Kuhn poker tasks.
- Scoring pipeline: BTA extracts raw indicators from logged trajectories, normalizes them, corrects directionality, and aggregates them without an LLM judge.The rule-based pseudocode averages direction-corrected normalized indicators to produce the behavioral score.
- Interpretation: A BTA score of 89.0 represents the uniform average of six normalized behavioral indicators in the repeated Prisoner’s Dilemma example.The example combines cooperation, retaliation, forgiveness, endgame defection, action volatility, and payoff efficiency after normalization to [0, 1].
J.5 Validation Protocol: Linking σ to Observable Risk Events
The validation protocol links cross-view consistency σ to observable, task-defined risk events rather than judge introspection. It combines explicit event definitions, held-out calibration, predictive and monotonicity tests, bootstrap uncertainty, and interpretable contradiction diagnostics.
- Risk-event validation: σ is validated against observable task-defined events, including endgame defection, commitment violation, deceptive messaging, and unstable collusion.These events connect communication claims and reasoning or behavior patterns to concrete episode outcomes.
- Validation metrics: Validation reports AUROC, Spearman monotonicity, and calibration stability for predicting risk events or tracking their severity.Thresholds are learned on calibration splits and tested for rank stability or similar risk recall across splits.
- Uncertainty: Bootstrap 95% confidence intervals over episodes assess whether AUROC and correlation conclusions depend on a small number of episodes.This supplies minimal significance reporting for the event-validity analyses.
- Interpretation: Pairwise view gaps identify dominant disagreement types, while calibrated σ ranges distinguish aligned views, partial tension, and diagnosable contradictions.The framework interprets gaps such as does–thinks and says–does inconsistency, with high σ associated with higher-risk contradiction patterns after calibration.
- Three-view inputs: The benchmark records actions, payoffs, stated reasoning, and dialogue for each episode, enabling BTA, RPA, and CCA to analyze complementary process signals.BTA uses action and payoff trajectories, RPA evaluates stated rationales, and CCA evaluates dialogue and talk–act consistency.
- Measurement boundary: RPA scores expressed reasoning rather than hidden chain-of-thought, treating the result as a process proxy and penalizing missing or unparseable rationales.The rationale is requested in structured fields covering goals, beliefs, planned action, and justification, then scored with a fixed rubric.
M.4 Recommended Task Configuration Interface
The task configuration interface exposes indicator selection, importance, and judge-stability controls for each dimension. It also organizes the benchmark into four levels and reports module scores alongside cross-view consistency diagnostics.
- Configuration interface: For each dimension, the interface selects core or task-specific indicators, assigns importance, and enables judge-stability controls.These settings make the resulting score transparent and reproducible while allowing task designers to emphasize diagnostic evidence.
- Interpretability: The interface supports interpretable portraits by mapping evidence to BTA, RPA, and CCA through Big Five and Social Exchange Theory constructs.These mappings connect behavioral, reasoning, and communication signals to diagnostic cues and appraisals.
- Outputs: The benchmark tables report BTA, RPA, and CCA scores by level and task, with separate displays for reasoning models and rule-based baselines.Level-specific tables cover all four levels, while additional tables report reasoning–communication gaps and cross-view consistency.