Source-linked AI summary
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
Paul-Peter Arslan
TL;DR
This study examines whether distributing biased triage decisions across a role-differentiated multi-agent pipeline changes bias or its detectability compared with single-agent self-review. Using paired-case disaster-triage simulations, it finds that audit capacity strongly affects whether bias is detected, primarily through review coverage rather than judgment quality, while risk-prioritized queues recover much of the lost coverage.
Problem
Single-agent benchmarks cannot establish whether distributing biased resource-allocation decisions across an independently audited pipeline reduces bias, leaves it unchanged, or makes it harder to detect.
Method
The study uses matched clinical case pairs in a synthetic disaster-triage simulator to compare single-agent self-review with a nine-agent assessment, allocation, and audit pipeline under varied operational pressures.
Results
A separate auditor flags decisions at 7%–62% across policies versus well under 1% for self-review, while audit overload reduces coverage from 100.0% to 65.6% without reducing judgment quality on reviewed cases.
Takeaways & Limitations
Independent oversight can improve detection, but under capacity constraints its effectiveness depends mainly on which cases reach the auditor; risk-aware queueing can partially recover lost coverage without adding capacity.
Takeaways & Limitations
The findings come from one model, modest samples, a single research pass without adversarial replication, and a pipeline design that may not represent all deployed systems.
Abstract
from arXiv · showhide
Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias. Real deployments, however, rarely use a single agent: they use pipelines, with review steps meant to catch exactly this kind of failure. We study what happens to bias when the same decision is distributed across a role-differentiated multi-agent pipeline (assessment, allocation, independent audit) instead of made and checked by one model alone. Using a synthetic disaster-triage simulator with paired cases that are clinically identical except for one demographic attribute, we run 192 episodes (2,304 resolved case pairs) on GPT-4o-mini comparing a single-agent control condition to a nine-agent pipeline under three independently varied pressure dimensions. We find no measurable difference in how often biased outcomes occur between the two conditions (6.9% vs. 6.1%, p = 0.498). We do find a large and significant effect of audit capacity on whether bias is caught: 30.0% of biased outcomes go entirely undetected, rising to 43.8% when the auditor is overloaded and falling to 18.4% when it is not. Decomposing this effect shows it is driven almost entirely by coverage (whether a case is reviewed at all, which collapses from 100.0% to 65.6% under load, p < 0.001) rather than by degraded judgment on the cases that are reviewed (81.6% vs. 85.7%, p = 1.000, direction reversed). A follow-up experiment shows that reordering the audit queue by estimated risk, rather than first-come-first-served, recovers most of the lost coverage under the same capacity constraint (65.6% to 91.7%, p = 0.028). We discuss the implications for any system that adds independent oversight to an LLM agent pipeline under resource constraints, and report the study's limitations honestly: one model, modest sample sizes, and no adversarial replication.
1 Introduction
The paper asks whether distributing triage across specialized agents and an audit stage changes demographic bias or merely makes it harder to detect. It tests this with matched clinical cases differing in exactly one demographic attribute, comparing single-agent and nine-agent pipelines under operational pressures.
- Prior benchmarks found demographic differences in triage severity and resource allocation despite identical clinical cases.
- The central question is whether an audited multi-agent pipeline reduces bias, leaves it unchanged, or makes it harder to locate.
- Matched cases share clinical severity and narrative but differ in exactly one demographic attribute, isolating that attribute as the explanation for outcome differences.
- The study compares a single-agent control with a nine-agent pipeline containing assessors, allocators, and auditors under independently varied operational pressures.
2 Related Work
The study combines prior work on demographic bias, adversarial policy compliance, shared-resource simulation, multi-agent failure modes, and audit judgment under pressure. Its operational-pressure design differs from adversarial user pressure, and its findings test whether capacity affects review coverage or judgment.
- KillBench supplies the contrastive twin-pair design and demonstrates demographic bias in single-LLM life-or-death triage decisions.
- PBSuite examines policy compliance under adversarial multi-turn pressure, whereas this study varies caseload, resource scarcity, and review capacity without a hostile user.
- GovSim contributes the shared depleting-resource mechanic, adapted here to hospital beds under explicit scarcity rather than a commons-management game.
- MAST informs the fixed role-differentiated assessment, allocation, and audit pipeline used to test a documented multi-agent failure surface.
- Prior audit research motivated a hypothesis that capacity pressure would degrade review quality, but the study instead decomposes coverage from judgment.
3 Method
The method uses matched twin cases, controlled single-agent and multi-agent conditions, five audited policies, and a 2 × 2 × 2 factorial design. Operational pressures vary independently across caseload, bed stock, and audit capacity.
- Every twin pair contains clinically identical patients differing in exactly one demographic attribute, making divergent outcomes the study’s bias signal.
- The Control uses one LLM call for assessment, allocation, and self-audit, while the Multi-agent condition separates these stages across four Assessors, three Allocators, and two Auditors.
- Both conditions use the same seed-matched generated cases, avoiding confounding from differences in case difficulty.
- P4 audits whether downstream agents receive demographic information beyond what their operational stage requires, distinct from end-user disclosure.
- Three independently crossed pressures vary caseload curve, bed stock, and audit capacity in a 2 × 2 × 2 factorial design.The experiment includes 192 episodes and 2,304 resolved case pairs on GPT-4o-mini.
4 Results
Splitting triage across agents did not measurably change bias rates, but audit capacity strongly affected whether biased outcomes were detected. Under load, coverage fell while judgment on reviewed cases did not, and risk-prioritized queuing recovered coverage.
- 4.1 Bias rate is not measurably different between conditions: Bias rates did not measurably differ between the Control and Multi-agent conditions.Table 2 reports the comparison pooled across all pressure cells.
- 4.2 An overloaded auditor sees less, not worse: 30.0% of biased outcomes went entirely undetected, rising from 18.4% without overload to 43.8% with a one-review-per-tick limit.These outcomes were neither flagged by an allocator nor caught by an audit.
- 4.2 An overloaded auditor sees less, not worse: Coverage fell from 100.0% to 65.6% under the capacity cap (p < 0.001), while judgment quality on reviewed cases was 81.6% versus 85.7% (p = 1.000).The direction of the judgment difference reversed under load, indicating that the detection loss tracked reduced review coverage rather than degraded reviewed-case judgment.
- 4.3 A risk-based audit queue recovers most of the lost coverage: Risk-prioritized audit queuing significantly recovered coverage under the same capacity constraint.The follow-up compared risk ordering with first-come-first-served using the same cases and audit capacity.
- 4.3 A risk-based audit queue recovers most of the lost coverage: The silent-bias rate is the arithmetic complement of the catch rate on the same biased pairs, not an independent measurement.This qualification limits how separately the two queueing outcomes should be interpreted.
5 Discussion
Independent review detects decision differences that self-review rarely flags, but audit capacity determines how much of that advantage survives. Under load, coverage falls while review quality on examined cases remains nearly stable, making queue design consequential.
- 5 Discussion: Independent auditors flag decisions at 7%–62% across policies, versus well under 1% for single agents auditing themselves.The multi-agent advantage is significant across all five policies, with p < 10^-50 for each.
- 5 Discussion: Under load, audit erosion is primarily a coverage problem rather than degraded judgment on reviewed cases.The passage characterizes coverage as collapsing while review quality on cases actually seen barely moves.
- 5 Discussion: Risk-aware queueing can partially recover lost audit coverage without adding capacity, although it does not eliminate the trade-off.The queue that feeds the auditor determines how much independent oversight is delivered under resource constraints.
6 Limitations
The study’s conclusions are bounded by one model, modest sample sizes, no adversarial replication, and a pipeline design tied to a specific benchmark. These constraints limit generalization to other models and deployed multi-agent systems.
- 6 Limitations: Using only GPT-4o-mini means the findings cannot be generalized to other models without separate testing.KillBench found that bias varies substantially by model.
- 6 Limitations: Modest samples leave the raw bias-rate comparison and priority-queue catch-rate gain underpowered to confirm or rule out effects.The coverage result and seed and episode integrity checks are described as robust.
- 6 Limitations: The study is a single non-peer-reviewed pass without adversarial replication.This limits the evidentiary scope of the reported results.
- 6 Limitations: The audited policies, pressure dimensions, and topology were selected for one external benchmark and may not represent deployed multi-agent systems broadly.The limitation concerns the study’s setting and design choices rather than only its sample size.
7 Conclusion
In this simulation, distributing triage across a role-differentiated pipeline did not measurably alter bias frequency, but audit capacity substantially affected whether bias was detected. Queue reordering recovered much of the coverage lost under overload without adding capacity.
- 7 Conclusion: The multi-agent pipeline did not measurably change how often biased triage decisions occurred.The comparison was between a role-differentiated pipeline with independent audit and the single-agent condition.
- 7 Conclusion: Roughly three in ten biased outcomes went undetected, rising to more than four in ten when the auditor was overloaded.Detection changed substantially with audit capacity, not merely with pipeline structure.
- 7 Conclusion: Reordering the audit queue recovered most lost coverage without adding capacity.This identifies queueing policy as a practical lever for constrained oversight.