Source-linked AI summary

A Comprehensive Survey on Multi-Agent Cooperative Decision-Making: Scenarios, Approaches, Challenges and Perspectives

Weiqiang Jin, Hongyang Du, Shixiang Tang, Biao Zhao, Guang Yang

arXiv:2503.13415v2cs.MAcs.AI

TL;DR

Multi-agent cooperative decision-making requires multiple agents to coordinate in complex simulated and real-world environments. This survey systematically reviews simulation environments and five methodological families, focusing on MARL and LLM-based approaches. It synthesizes their applications, capabilities, challenges, and future directions, while highlighting scalability and coordination issues.

  • Problem

    Multi-agent cooperative decision-making must address coordination across complex dynamic environments and diverse applications.

  • Method

    The survey analyzes simulation environments and classifies decision-making into rule-based, game theory-based, evolutionary algorithms-based, MARL-based, and LLM-based approaches.

  • Results

    The survey identifies MARL and LLM-based methods as central advanced paradigms and reviews their applications, methodologies, and challenges.

  • Takeaways & Limitations

    Simulation environments connect theoretical advances with real-world implementation, while intelligent collaboration and adaptability support complex cooperative tasks.

Abstract

from arXiv · show

With the rapid development of artificial intelligence, intelligent decision-making techniques have gradually surpassed human levels in various human-machine competitions, especially in complex multi-agent cooperative task scenarios. Multi-agent cooperative decision-making involves multiple agents working together to complete established tasks and achieve specific objectives. These techniques are widely applicable in real-world scenarios such as autonomous driving, drone navigation, disaster rescue, and simulated military confrontations. This paper begins with a comprehensive survey of the leading simulation environments and platforms used for multi-agent cooperative decision-making. Specifically, we provide an in-depth analysis for these simulation environments from various perspectives, including task formats, reward allocation, and the underlying technologies employed. Subsequently, we provide a comprehensive overview of the mainstream intelligent decision-making approaches, algorithms and models for multi-agent systems (MAS). Theseapproaches can be broadly categorized into five types: rule-based (primarily fuzzy logic), game theory-based, evolutionary algorithms-based, deep multi-agent reinforcement learning (MARL)-based, and large language models(LLMs)reasoning-based. Given the significant advantages of MARL andLLMs-baseddecision-making methods over the traditional rule, game theory, and evolutionary algorithms, this paper focuses on these multi-agent methods utilizing MARL and LLMs-based techniques. We provide an in-depth discussion of these approaches, highlighting their methodology taxonomies, advantages, and drawbacks. Further, several prominent research directions in the future and potential challenges of multi-agent cooperative decision-making are also detailed.

A Comprehensive Survey on Multi-Agent Cooperative Decision-Making Applications: Scenarios, Approaches, Challenges and Perspectives

The paper surveys multi-agent cooperative decision-making and its applications across intelligent systems research.

  • The paper presents a comprehensive survey of multi-agent decision-making methods.

1. Introduction

Multi-agent cooperative decision-making has evolved from single-agent research toward complex collaborative systems with broad applications. Existing surveys leave gaps in methodological breadth, simulation environments, and implementation detail, which this survey addresses systematically.

  • Multi-agent cooperative decision-making involves interacting agents jointly completing tasks in dynamic simulated environments and real-world systems.
  • The field supports applications including smart agriculture, collaborative robots, autonomous driving, navigation, and rescue.
  • Existing reviews often emphasize reinforcement learning while overlooking broader methods, simulation environments, and implementation details.
  • This survey expands coverage by treating environments, methods, techniques, and implementation perspectives as connected parts of multi-agent research.
  • The survey focuses especially on deep MARL and LLM-based methods, while also covering applications and simulation environments.

2. Multi-Agent Decision-Making Taxonomies

Multi-agent cooperative decision-making methods are broadly organized into five categories, with MARL and LLM-based approaches receiving particular attention.

  • The survey classifies cooperative decision-making into rule-based, game theory-based, evolutionary algorithms-based, MARL-based, and LLM-based methods.
  • The taxonomy is presented as a framework for comparing mainstream cooperative decision-making paradigms and their associated techniques.

2.1. Mainstream Paradigms of Multi-Agent Cooperative Decision-Making

The survey reviews mainstream multi-agent paradigms, emphasizing their classifications, applications, strengths, and limitations. It gives particular attention to MARL and LLM-based systems, including scalability and transparency challenges.

  • Mainstream multi-agent decision-making includes rule-based, game theory-based, evolutionary algorithms-based, MARL-based, and LLM-based systems.
  • Rule-based: Rule-based systems provide structured, transparent, reliable, and interpretable decision-making, but face adaptability and scalability challenges.
  • Rule-based: Rule-based MAS are classified into fuzzy logic-based, reinforcement learning-integrated, expert system-based, hybrid intelligence, and ensemble learning-based approaches.
  • Game theory-based: Game theory supports strategic decisions in cooperative, competitive, and mixed scenarios through equilibrium-based optimization.
  • LLMs-based: LLM-based multi-agent systems use role assignment and collaboration to support complex decision-making, but large-scale coordination and reasoning transparency remain open challenges.

2.2. MARL-based Multi-Agent Decision-Making Taxonomies

MARL taxonomies organize multi-agent decision-making around information sharing and execution centralization, with CTDE, CTCE, and communication-based methods addressing different coordination constraints. The survey reviews value decomposition, PPO-based algorithms, communication mechanisms, and representative applications and limitations.

  • MARL paradigms: MARL research is commonly organized into CTCE, DTDE, and CTDE according to how information is shared during learning and deployment.CTDE combines centralized training with decentralized execution, while CTCE integrates learning and execution more centrally.
  • Value decomposition: Value decomposition algorithms simplify joint value estimation by decomposing high-dimensional joint state-action values into individual agent values.This decomposition lets agents select actions using more manageable individual Q-values during execution.
  • PPO-based algorithms: PPO-based methods use clipped policy updates to balance optimization stability and efficiency in CTDE-based MARL.MAPPO, HATRPO, and HAPPO extend PPO for multi-agent coordination and scalable cooperative and competitive tasks.
  • PPO-based algorithms: HATRPO and HAPPO sequentially update one agent while holding others fixed, supporting convergence guarantees and tailored policies without parameter sharing.They perform competitively on benchmark tasks and scale to high-dimensional state-action spaces.
  • Limitations and directions: CTCE can struggle with scalability, DTDE can lack global coordination, and CTDE remains limited by communication bottlenecks during execution.These limitations motivate continued work on communication strategies and communication-based MARL algorithms.
  • Communication-based MARL: Communication-based MARL is categorized by broadcasting, targeted, and networked protocols, and by value-function, policy-search, or communication-efficiency objectives.CADP adds advice exchange during centralized training and gradually prunes communication to preserve decentralized execution; related methods such as CommNet learn shared communication protocols.

2.3. LLMs-based Multi-Agent System Taxonomies

LLM-based multi-agent systems are surveyed through their architectures, autonomy levels, application domains, and use in simulating complex social, natural, and engineering systems. The taxonomy distinguishes adaptive autonomy from self-organizing autonomy and spans economic, social, robotic, scientific, and software-development applications.

  • Taxonomy dimensions: LLM-based multi-agent system taxonomies cover architectural design, application domains, evaluation methods, and future research directions.Architectural design focuses on mechanisms enabling agents to interact, adapt, and make decisions in complex and dynamic environments.
  • Autonomy levels: Adaptive autonomy lets agents adjust behavior within predefined frameworks according to task requirements or context.Examples include changing search strategies based on retrieval relevance or communication style based on social context.
  • Autonomy levels: Self-organizing autonomy allows agents to adapt dynamically without predefined structures, including task allocation based on environmental state and individual skills.Emergent behavior is identified as another feature of this higher autonomy level.
  • Social and operational applications: LLM-based agents are applied to economic modeling, social-network simulation, user behavior analysis, and multi-robot coordination.Applications include market simulation, information propagation, recommender-system user simulation, and search-and-rescue robotics.
  • Natural-science applications: In natural sciences, LLM-based agents support macroeconomic simulation and generative agent-based modeling of complex social and environmental dynamics.These approaches combine realistic agent decisions or mechanistic models with generative AI to study system behavior over time.
  • Engineering applications: In engineering, LLM-based agents assist software development through code generation, bug detection, and system optimization.These capabilities are described as improving software-development productivity and quality.

3. Simulation Environments of Multi-Agent Decision-Making

Multi-agent cooperative simulation environments provide the foundation for testing, validating, and understanding coordination in dynamic settings. The survey covers widely used MARL platforms and LLM-powered environments spanning diverse tasks, technologies, and interaction settings.

  • Purpose and scope: Simulation environments enable researchers to test decision-making algorithms and study how agents coordinate and adapt in dynamic settings.They support theoretical evaluation while providing insight into robustness and efficiency for real-world applications.
  • MARL environments: MPE supports cooperative, competitive, and mixed cooperative-competitive tasks in a time-discrete, space-continuous 2D platform compatible with the Gym interface.Its scenarios include adversarial interactions, cooperative crypto, object pushing, and team-based navigation.
  • MARL environments: SMAC evaluates decentralized micromanagement under local partial observations, while SMACv2 adds procedural content generation and stochastic enemy masking through EPO.These changes address insufficient stochasticity and weak partial observability in the original benchmark.
  • MARL environments: GFootball combines physics-based 3D football simulation with full-game and skill-specific Football Academy scenarios of varying difficulty.Football Academy ranges from single-player scoring tasks to team-based formations involving passing, shooting, and goalkeeper defense.
  • MARL environments: Unity ML-Agents and Gym-µRTS provide extensible training platforms, with Gym-µRTS supporting efficient RTS research on limited hardware and producing agents that defeat competition bots.Unity ML-Agents supports reinforcement learning, imitation learning, and neuroevolution through a Python API; Gym-µRTS supports PPO, invalid-action masking, action composition, and diverse opponents.
  • LLM-powered environments: LLM-powered environments span embodied transport, social interaction, and interactive multi-agent simulation, supporting research on collaboration, reasoning, and task execution.Examples include ThreeDWorld Transport, AgentScope, Generative Agents, and SocialAI-oriented settings.

4. Practice Applications of Multi-Agent Decision-Making

Multi-agent decision-making is applied across transportation, aerial systems, disaster response, military simulation, traffic management, autonomous driving, and collaborative robotics. The surveyed work emphasizes MARL for adaptive coordination and LLM frameworks for role-specific collaboration and task execution.

  • Application scope: Applications span agriculture, disaster rescue, military simulations, traffic management, autonomous driving, UAV systems, and multi-robot collaboration.The survey presents these domains as practical settings for advanced multi-agent methods in dynamic and evolving environments.
  • MARL-based applications: MARL applications address energy-aware agricultural exploration, UAV connectivity in disaster rescue, military decision-making, limited-bandwidth communication, and anti-jamming relay selection.The surveyed methods use reinforcement learning to optimize coordination under resource, terrain, communication, or adversarial constraints.
  • MARL-based applications: UAV pursuit-evasion, traffic control, and autonomous driving studies use decentralized or hierarchical MARL methods to coordinate agents in non-stationary and large-scale settings.Examples include heterogeneous UAV pursuit, curiosity-inspired traffic control, scalable A2C, and models of other drivers’ social value orientations and personalities.
  • MARL-based applications: Collaborative robotics research combines federated learning, learning from demonstration, MADDPG, and simulation to support scalable coordination and evaluation in complex environments.SCRIMMAGE specifically targets efficient simulation of large numbers of aerial vehicles for mobile robotics research.
  • MARL-based applications: MAPD, dynamic parameter sharing, mutual-information rewards, and PMIC target policy diversity, robustness to changing team compositions, and improved cooperative performance.These methods address policy comparison and coordination when agent teams vary across tasks.
  • LLM-based applications: LLM-based frameworks support role-specific autonomy, flexible workflows, tool use, and multi-agent collaboration for complex goals.CrewAI exemplifies this approach, while broader frameworks emphasize dynamic coordination, reasoning, and human-in-the-loop workflows.

5. Challenges in MARL-based and LLMs-based approaches

MARL and LLM-based multi-agent decision-making face unresolved challenges involving stochasticity, coordination, scalability, multimodal adaptation, hallucinations, evaluation, and privacy. The survey links these limitations to future directions such as meta-learning, graph neural networks, multimodal frameworks, verification, and improved evaluation.

  • MARL-based approaches: MARL remains constrained by environmental stochasticity, difficult strategy learning, non-stationarity, scalability, and complex reward design.These factors make stable policy learning, convergence, and effective coordination more difficult.
  • MARL-based approaches: Multi-agent cooperation requires communication and coordination mechanisms that keep agents’ actions globally consistent and complementary.Disaster-rescue drones illustrate the need to coordinate coverage and resource use.
  • Future directions: Future MARL research may combine meta-learning and graph neural networks to improve adaptation, robustness, and scalability in dynamic large-scale settings.These directions are presented as promising ways to address the interplay between robustness and scalability.
  • LLM-based approaches: LLM-based multi-agent systems remain early-stage, with bottlenecks in multimodal interaction, scalability, hallucination control, evaluation, collective intelligence, and privacy protection.The survey presents these areas as current limitations and opportunities for future research.
  • LLM-based approaches: Multimodal applications require agents to process and fuse images, audio, video, and sensor inputs while producing coordinated multimodal outputs.The paper calls for unified multimodal frameworks to support this cross-modal collaboration.
  • LLM-based approaches: Hallucinations can propagate through interconnected agents, allowing one misjudgment to trigger a system-wide chain reaction.The proposed response combines better training with information verification and propagation-management mechanisms.

6. Future Research Prospects and Emerging Trends

The survey presents LLMs-enhanced MARL as an emerging direction that combines language-based reasoning with multi-agent learning for more efficient collaboration, multimodal processing, multitask learning, and long-term planning. It also emphasizes that technical complexity, robustness, privacy, security, accountability, and explainability remain important research priorities.

  • Theoretical development: LLMs-enhanced MARL combines language understanding and reasoning with MARL to address low sample efficiency, difficult reward design, and weak generalization.The framework is presented as a new direction for collaboration in complex dynamic environments.
  • Theoretical development: The framework assigns LLMs four roles: information processor, reward designer, decision-maker, and generator.These roles translate task descriptions, design rewards, support decisions, and generate policies or outputs.
  • Technical integration: LLMs and MARL improve handling of multimodal information, multitask learning, and long-term planning.LLMs can unify data processing, transfer knowledge across tasks, decompose complex tasks, and provide virtual samples through simulation.
  • Application expansion: LLMs-enhanced MARL is positioned for autonomous driving, collaborative robotics, smart grids, healthcare, and disaster relief.The cited applications involve sensor-language integration, dynamic task division, domain-informed rewards, and adaptive role allocation.
  • Challenges and prospects: Unfamiliar environments, computational demands, biases, hallucinations, and real-time responsiveness remain technical constraints.The survey identifies robustness and resource efficiency as continuing research needs.
  • Human-society coordination: Ethical development requires stronger privacy protection, adversarial safeguards, accountability, transparency, and explainability.These concerns are especially relevant to sensitive domains such as healthcare and disaster response.
  • Human-society coordination: Addressing technical and ethical challenges is necessary for LLMs-enhanced MARL to become effective and ethically sound in diverse real-world applications.The survey frames this as a condition for realizing the systems’ broader potential.

7. Conclusion

The survey reviews the evolution of multi-agent cooperative decision-making from traditional approaches to MARL and LLMs, while connecting methods, environments, applications, benchmarks, datasets, and future directions. It presents simulation environments as a bridge between theoretical progress and real-world implementation.

  • Conclusion: The survey traces multi-agent decision-making from rule-based and game-theoretic methods toward MARL and LLM-based paradigms.It compares their capabilities, challenges, and applications in dynamic and uncertain environments.
  • Conclusion: Simulation environments connect theoretical advances with real-world implementation by shaping agent interaction, learning, and decision-making.The survey emphasizes their role across the reviewed methods and applications.
  • Conclusion: Applications in autonomous driving, disaster response, and robotics illustrate the practical scope of multi-agent systems.The survey also consolidates methodologies, datasets, benchmarks, and future research directions for researchers and practitioners.

Declaration of Generative AI and AI-assisted Technologies in the Writing Process

The authors used generative AI and AI-assisted technologies for proofreading and improving readability and language clarity in some sections. They reviewed the resulting content for accuracy and completeness while acknowledging risks of incorrect, incomplete, or biased output.

  • Declaration: Generative AI and AI-assisted technologies were used for proofreading and improving readability and language clarity in certain sections.The authors state that these tools were not presented as substitutes for their review.
  • Declaration: The authors reviewed the AI-assisted content for accuracy and completeness while acknowledging that generated output can be incorrect, incomplete, or biased.This declaration identifies the principal reliability caveat associated with the tools’ use.

Appendix A. Technological Comparisons between Single-Agent and Multi-Agent (Under Reinforcement Learning)

Single-agent reinforcement learning is modeled with fully observable MDPs, whereas multi-agent reinforcement learning uses partially observable models for interacting agents. The multi-agent setting requires joint actions, inferred state information, and coordination under reward interdependencies.

  • Multi-Agent Reinforcement Learning: MARL involves multiple agents interacting within a shared environment rather than a single agent making decisions independently.Agents operate simultaneously within a common dynamic environment.
  • Partial Observability: POMDPs extend MDPs to multi-agent settings where each agent receives only a partial observation instead of the full state.Agents maintain beliefs or infer missing state information from observations generated by the environment.
  • Single-Agent Reinforcement Learning: MDPs model single-agent decision-making in fully observable environments using states, actions, transitions, and rewards.The agent selects an action from the current state, the environment transitions probabilistically, and the policy seeks to maximize cumulative reward.
  • Multi-Agent Interaction: In the POMDP formulation, individual actions form a joint action that influences state transitions and produces individual rewards.The observation function maps resulting states and actions to the information available to each agent.
  • Multi-Agent Challenges: Partial observability and interacting agents introduce coordination, competition, and reward interdependencies that make decision-making more complex.These challenges distinguish the multi-agent setting from fully observable single-agent decision-making.
Loading 2503.13415v2…