Source-linked AI summary
Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities
Timothy Merritt, Alejandro Jarabo-Peñas, Juan Bravo-Arrabal, Maria-Theresa Bahodi, Anders Lyhne Christensen
TL;DR
Safety-critical multi-drone systems need agentic behavior that fits professional workflows, where uncertainty, accountability, and scale make autonomy alone insufficient. This position paper synthesizes NAMUR and PERSIST, proposes participatory iterative design with stakeholders, and presents governability as an architectural and evaluation priority.
Problem
Safety-critical multi-drone operations require autonomy that is understandable, reliable, governable, and adoptable, while agentic AI raises unresolved questions about supervision, override, trust, and operational fit.
Method
The paper synthesizes lessons from NAMUR and PERSIST and uses stakeholder-engaged, iterative prototyping to shape governable agentic behavior and interaction requirements.
Results
The paper presents an LLM-MAS architecture that uses constrained tool interfaces, explicit operator preview and approval, specialized agents, persistent task management, and long-horizon memory.
Takeaways & Limitations
Governability should be a core, auditable system property, built through constrained inspectable tools and evaluated using operational measures beyond task success.
Abstract
from arXiv · showhide
Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure monitoring. Yet, real-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites. We argue that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms. We outline a human-centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof-of-concept systems and evaluation strategies for other safety-critical contexts.
1. Introduction
Safety-critical multi-drone autonomy must be understandable, reliable, governable, and adoptable, not merely technically functional. Agentic AI adds interaction challenges around representing intentions, supervising decisions, and guaranteeing correct behavior in field operations.
- Safety-critical work combines uncertainty, high accountability, professional workflows, and risk management, making technical autonomy alone insufficient.
- Scaling from one vehicle to multiple robots increases coordination overhead, weakens shared situation awareness, and can create cascading safety issues.Operators need both group-level mission control and the ability to inspect, redirect, or control individual drones or subsets.
- Agentic AI and LLMs enable natural-language tasking, mixed-initiative planning, and autonomous execution, while raising questions about intention representation, supervision, override, and behavioral guarantees.
- The paper introduces NAMUR and PERSIST as motivating contexts and develops shared interaction requirements for multi-drone control.
2. Case Contexts and Project Objectives
NAMUR addresses time-critical emergency response, whereas PERSIST addresses persistent monitoring and security at critical infrastructure sites. Together, they expose distinct operational demands while motivating a socio-technical approach centered on governability, stakeholder engagement, and realistic validation.
- NAMUR: Agentic support for SAR and firefighting: NAMUR explores LLM-supported agentic interaction for SAR and firefighting under time pressure, dynamic hazards, incomplete information, and cross-role coordination.
- NAMUR: Agentic support for SAR and firefighting: NAMUR seeks higher-level tasking, mixed-initiative human authorization, interpretable behavior, and realistic exercises that reveal failure modes and governance needs.
- PERSIST: Persistent multi-drone operations for critical infrastructure: PERSIST targets persistent monitoring, inspection, and security workflows spanning days and shifts, including routine missions, anomaly response, and organizational integration.
- PERSIST: Persistent multi-drone operations for critical infrastructure: PERSIST aims to replace ad hoc flights with repeatable delegated operations, lower professional-use barriers, scale supervision, and validate transferable proofs of concept.
- Cross-cutting themes and contrasts: NAMUR emphasizes rapid decision cycles, while PERSIST emphasizes persistent activity, repeatability, and integration across shifts.
- Cross-cutting themes and contrasts: Both contexts require clear authority and auditability, but emergency response prioritizes confirmation and escalation while infrastructure operations require routines, handovers, and post hoc accountability.
- Cross-cutting themes and contrasts: The research treats agentic AI as socio-technical design material, iteratively studying coordination, decision-making, verification, and governable prototypes with stakeholders.
3. Agentic AI for Multi-Drone Operations in Safety-Critical Work: Lessons, Tensions, and an Agenda
The paper frames agentic AI for safety-critical multi-drone work as a governability problem involving attention, explanation, uncertainty, trust, and operational fit. It proposes participatory, iterative prototyping to shape constrained agent behavior with stakeholders.
- Lessons and opportunities: Agentic AI can help manage multiple vehicles, data streams, and stakeholders through selective attention, summarization, mixed-initiative planning, and long-horizon support.
- Safety-critical tensions: Safety-critical systems must operate without constant attention while remaining explainable, constrainable, overridable, and resistant to attention tunneling from strong alerts.
- Safety-critical tensions: Trust requires calibration under imperfect perception and dynamic conditions, supported by conservative defaults, uncertainty cues, verification workflows, and auditable authorization trails.
- Research approach: participatory shaping of agentic autonomy: The approach iteratively shapes agentic AI as socio-technical design material through stakeholder engagement and functional prototypes rather than targeting one final autonomy concept.
- Research approach: participatory shaping of agentic autonomy: Needs discovery and successive prototypes map work practices, decision points, and cognitive bottlenecks while negotiating authorization boundaries and required evidence or uncertainty cues.
- Research approach: participatory shaping of agentic autonomy: The resulting requirements include tunable constraints, explicit intervention points, and observability features for diagnosing behavior and learning from incidents.
- Research approach: participatory shaping of agentic autonomy: The proposed architecture decomposes intent into reviewable subtasks, constrains actions through deterministic tools, supports scheduling and memory, and requires operator preview and approval.
4. An LLM-MAS Architecture for Multi-drone Control
The functional prototype uses a unified LLM-based multi-agent architecture for NAMUR and PERSIST, combining specialized agents with deterministic tools and persistent task management. Its design aims to make natural-language mission execution transparent, interpretable, and safe under operator supervision.
- The functional prototype is an LLM-MAS with five specialized agents: Coordinator, Event, Spatial, Swarm, and Summarizer.
- Spatial and Swarm agents connect to deterministic tools for spatial computation, multi-drone motion execution, memory storage, and task scheduling.
- Available tools depend on deployment, supporting compliant coverage planning for SAR and flight-path or image-acquisition planning for biomass estimation.
- The Coordinator receives operator commands, infers intent, decomposes them into concrete sub-queries, and delegates execution requests to subordinate agents.
- The Events Agent manages periodic, delayed, and future directives, including recurring estimates and scheduled returns to base.
- Separating agent responsibilities and constraining behavior through prompts and tool interfaces supports interpretable mission translation, persistent interaction history, multi-turn dialogue, and long-horizon autonomy under supervision.
5. Conclusion
The paper frames agentic AI as promising for scaling multi-drone operations while requiring governable autonomy, constrained tool use, and explicit operator oversight. It argues that future evaluation must include operational and accountability-related factors beyond task success.
- Agentic AI can scale multi-drone operations in search and rescue and critical-infrastructure monitoring, but raises oversight, trust-calibration, operational-fit, and evaluation challenges.
- Governability should be a core, explicit, and auditable system property, including authorization points, constraints, and intervention mechanisms.
- Constrained, inspectable tool use can support predictable failure modes, reproducible debugging, and clearer responsibility boundaries.
- Safety-critical evaluation should measure situation awareness, trust calibration under uncertainty, coordination overhead, and accountable after-action review in addition to task success.
Declaration on Generative AI
The authors used Grammarly and GPT 5.2 for grammar and spelling checks, then reviewed and edited the content and retained responsibility for its publication.
- Grammarly and GPT 5.2 were used for grammar and spelling checks, followed by author review and editing.