Source-linked AI summary

DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning

Zijie Meng, Xiwei Dai, Yixuan Tang, Jin Hao, Yang Feng, Fudong Zhu, Xiaoqiang Liu, Shaosheng Cao, Zuozhu Liu

arXiv:2608.18878v1cs.AIcs.MA

TL;DR

Dental AI systems often handle individual modalities or tasks without explicitly linking heterogeneous observations to conclusions. DentAgent coordinates five specialist agents through an Evidence Blackboard, achieving leading performance across four benchmarks and exceeding senior specialists by 17.3 percentage points on multi-label diagnosis.

  • Problem

    Dental AI remains largely modality- or task-specific and lacks a common mechanism to coordinate heterogeneous evidence while making observations supporting conclusions explicit.

  • Method

    DentAgent coordinates five modality-specific specialists through an Orchestrator and Evidence Blackboard that structures, links, and tracks source-attributed evidence.

  • Results

    Across four benchmarks, DentAgent achieves leading scores in various scenarios and exceeds senior specialists by 17.3 percentage points on multi-label diagnosis.

  • Takeaways & Limitations

    The results support DentAgent’s applicability and traceability as a unified framework for multimodal dental reasoning.

Abstract

from arXiv · show

Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D dental data. Most existing dental AI systems remain modality- or task-specific. Although recent vision-language models support flexible dental question answering, directly generated response leaves evidence implicit and untraceable. To address these limitations, we introduce DentAgent, an evidence-centric multi-agent framework, in which the Orchestrator coordinate five specialized agents spanning various modalities. Each specialist utilizes domain tools to convert observations into structured evidence records. The Evidence Blackboard manages these records as a shared evidence state, tracking coverage, gaps, and conflicts before response generation. This standardized evidence representation integrates isolated dental capabilities into a unified agentic workflow. Across four benchmarks, DentAgent demonstrates leading performance, even surpassing the senior specialists by 17.3 percentage points on multi-label diagnosis, which supports its value for broadly applicable and traceable multimodal dental reasoning, and highlights its potential as a technical foundation for population oral health assessment and management.

I. INTRODUCTION

DentAgent addresses the challenge of integrating heterogeneous dental evidence by coordinating modality-specific specialists through an evidence-centric multi-agent workflow. Across four benchmarks, it achieves leading performance, including a 17.3-percentage-point advantage over senior specialists on multi-label diagnosis.

  • Motivation: Oral assessment requires integrating domain knowledge, radiographs, intraoral photographs, and 3D dental data with differing representations and evidential roles.These heterogeneous sources vary in representation, anatomical granularity, and evidential role.
  • Limitations: Most existing dental AI systems remain modality- or task-specific, while generated answers often leave supporting evidence implicit and difficult to trace.Specialized models operate as isolated pipelines, and generative models commonly map inputs directly to answers or reports.
  • Framework: DentAgent coordinates an Orchestrator, five modality-specific Specialists, and an Evidence Blackboard within a unified workflow.The Orchestrator activates relevant specialists to acquire response-critical evidence, while specialists iteratively select tools and examine observations.
  • Evaluation: 17.3 percentage points: DentAgent exceeds senior specialists on multi-label diagnosis while achieving leading scores across four heterogeneous dental benchmarks.The benchmarks cover bilingual dental knowledge, panoramic radiograph QA, 2D dental diagnosis, and 3D IOS reasoning.
  • Framework: Its evidence-centric mechanism decouples evidence acquisition from answer generation and stores source-attributed findings in a shared Evidence Blackboard for iterative sufficiency assessment.This mechanism supports target-driven reasoning across heterogeneous specialist capabilities.

II. RELATED WORKS … C. Agentic Dental AI

Dental AI has progressed from modality-specific perception models to dental foundation, generative, and agentic systems. However, existing approaches provide limited explicit coordination and management of evidence across specialized components.

  • A. Task-Specific Dental Perception Models: Early dental AI focused on well-defined perception tasks within individual imaging modalities.Applications included caries detection, periodontal bone-loss assessment, panoramic tooth detection and numbering, and instance-level tooth segmentation.
  • A. Task-Specific Dental Perception Models: Deep learning addressed caries detection on periapical and bitewing radiographs.
  • A. Task-Specific Dental Perception Models: Dental AI also supported periodontal bone-loss detection and measurement, tooth detection and numbering in panoramic radiographs, and instance-level tooth segmentation.
  • B. Dental Foundation and Generative Models: Recent work expanded dental AI toward domain-specialized foundation and generative models using dental-specific data and modality-aware representations.OralGPT introduced instruction data and evaluation protocols for panoramic X-ray understanding, while DentalBench provided a bilingual dental-knowledge benchmark.
  • B. Dental Foundation and Generative Models: These foundation and generative models remained largely centered on individual-model capabilities.
  • B. Dental Foundation and Generative Models: They provided limited support for explicitly coordinating and managing evidence from multiple specialized components.
  • C. Agentic Dental AI: Recent agentic dental systems began coordinating multiple sources of specialized expertise rather than only improving individual model capabilities.OralAgent combined visual tools with dental knowledge retrieval, while OPGAgent used specialized perception modules, hierarchical evidence collection, and consensus-based reporting for panoramic radiographs.

III. METHODS · A. Framework Overview · B. Global Orchestration

DentAgent is an evidence-centric hierarchical multi-agent framework that interprets dental tasks, delegates modality-specific evidence acquisition, manages shared evidence, and verifies sufficiency before generating responses. Its orchestration adapts to unresolved evidence gaps and conflicts across text, clinical, radiographic, photographic, and 3D dental inputs.

  • A. Framework Overview: The framework separates task interpretation, specialist execution, evidence organization, and response generation through an Intent Detector, Orchestrator, specialists, and Evidence Blackboard.Supported inputs include text questions, clinical documents, panoramic and cephalometric radiographs, intraoral photographs, and IOS meshes.
  • A. Framework Overview: Each specialist performs bounded modality-specific reasoning, tool use, and observation, with activation limited to at most Kmax tool calls.This bound prevents unbounded or non-terminating tool-use loops.
  • A. Framework Overview: Specialist observation packets enter the Evidence Blackboard through Evidence Normalization, Coverage Tracking, Evidence Linking, and Conflict Resolution.The updated blackboard directs later rounds toward unresolved or contradictory evidence, and execution terminates when verification succeeds, no target remains, or the maximum global rounds are reached.
  • A. Framework Overview: DentAgent maps each dental case to a task specification defining the objective, output format, granularity, and relevant input modalities.The specification remains fixed throughout inference and contains neither a diagnostic prediction nor a candidate response.
  • B. Global Orchestration: The Orchestrator runs an Identification–Delegation–Management–Verification loop that derives task-specific evidence targets from the shared blackboard.Targets may represent clinical findings, anatomical locations, quantitative measurements, structural relationships, or knowledge claims.
  • B. Global Orchestration: After each evidence update, the Orchestrator removes resolved targets, retains gaps and conflicts, and routes subsequent execution toward the revised target set.Verification ends acquisition when the evidence state is sufficient; otherwise, the updated blackboard conditions the next identification stage while τ remains fixed.
  • B. Global Orchestration: Specialist assignments depend on evidence type, available case inputs, and capability boundaries, using one specialist for modality-specific targets or multiple specialists when several modalities are required.The Orchestrator assigns current target subsets to an appropriate specialist subset during each delegation stage.

C. Specialist Sub-Agent Execution

DentAgent executes five complementary specialist sub-agents across clinical text, radiographs, intraoral photographs, and IOS dental structures. Each specialist selects relevant tools based on assigned targets and context, then returns structured findings or measurements with provenance and validity information.

  • Specialist coverage: Five specialists cover clinical text and dental knowledge, panoramic radiographs, intraoral photographs, cephalometric radiographs, and IOS-based dental and occlusal structures.They are activated as needed without a fixed execution order.
  • Tool selection: Each specialist Ak has a corresponding tool inventory Tk and selects at most Kmax relevant tools for its assigned targets Rk.Tool selection is conditioned on the assigned targets, available inputs, current blackboard state, and tool relevance.
  • Execution: Specialists execute their selected tools and return observations for the assigned targets.The execution produces the observations used to populate the specialist output.
  • Structured outputs: Each output Ok contains findings or measurements together with anatomical location, source, quality, and conditions of validity.This output structure makes the specialist evidence explicitly contextualized.

D. Evidence Blackboard

The Evidence Blackboard is a shared, traceable state that organizes evidence from specialists and reasoning rounds around common targets. It standardizes records, tracks coverage, links relationships, and resolves conflicts without discarding source evidence.

  • Evidence Blackboard: The Evidence Blackboard organizes evidence from specialists and reasoning rounds around shared targets while recording each item’s source and applicability conditions.It serves as shared state for Specialist Sub-Agents, the Orchestrator, and the Response Generator.
  • Evidence Normalization: NormalizeEvidence converts specialist outputs into standardized records containing the target, finding or measurement, anatomy, provenance, quality, and limitations.Normalization aligns clinical terminology, anatomical labels, and detail levels for direct comparison and later evidence linking.
  • Coverage Tracking: TrackCoverage classifies evidence for each target as covered, partially covered, conflicting, or uncovered according to the requested output’s requirements.Coarse evidence may establish an abnormality but remain insufficient for its exact location, severity, or size.
  • Evidence Linking: LinkEvidence connects records for the same target as supporting, complementary, redundant, or conflicting, while retaining them separately.These relationship types distinguish independent support, different target aspects, repeated observations, and incompatible findings.
  • Conflict Resolution: ResolveConflicts weighs incompatible records by relevance, anatomical correspondence, applicability, quality, provenance, and conditions rather than majority voting.All source records remain retained for traceability, and no preference is imposed when the evidence does not justify reliable resolution.

E. Response Generation

After orchestration terminates, the Response Generator produces the final output from the task specification τ and final blackboard state B. It selects relevant evidence and formats responses according to task type without additional tool calls or blackboard modifications.

  • Response Generation: After orchestration terminates, the Response Generator produces the final output from task specification τ and final blackboard state B.
  • Response Generation: The Response Generator selects evidence relevant to requested targets and presents the result according to τ without calling additional tools or modifying the blackboard.
  • Response Generation: For closed-set tasks, outputs are restricted to predefined labels, while structured and open-ended responses are generated from blackboard evidence records.

IV. EXPERIMENTS · A. Implementation

DentAgent is implemented in LangGraph around a shared evidence-and-status blackboard, five modality-specific specialists, and 33 tools. Its orchestration uses configurable capability cards, parallel specialist execution, bounded tool selection, and a 15-round global budget.

  • A. Implementation: DentAgent is implemented in LangGraph with a shared blackboard tracking accumulated evidence and execution status across orchestration rounds.The blackboard supports evidence management throughout orchestration.
  • A. Implementation: Qwen3.5-9B, served through vLLM, handles orchestration and final response generation unless otherwise specified.Tailored system prompts are used for both roles.
  • A. Implementation: Reasoning mode is enabled during evidence acquisition and disabled during final response generation.The implementation separates reasoning behavior between evidence collection and answer generation.
  • A. Implementation: The framework includes N = 5 modality-specific specialists and 33 tools.The specialists provide modality-specific capabilities within the framework.
  • A. Implementation: Each specialist is registered through a YAML capability card describing its tools, modality, target coverage, schema, anatomical scope, and applicability constraints.These cards define when and how specialists can be used.
  • A. Implementation: Compatible specialists can execute in parallel during each orchestration round, while each invocation selects at most Kmax = 3 tools.Parallel execution and per-invocation tool limits constrain round-level execution.
  • A. Implementation: The global execution budget is Tmax = 15 rounds, with earlier termination when verification succeeds or no feasible specialist remains.The stated stopping conditions can end orchestration before the maximum budget.

B. Benchmarks · C. Main Results

DentAgent is evaluated across four multimodal dental benchmarks covering textual, photographic, radiographic, and native 3D reasoning. It achieves leading results through evidence-grounded agentic reasoning, surpassing strong specialized baselines and senior specialists on several tasks.

  • B. Benchmarks: DentAgent is evaluated on four benchmarks spanning bilingual dental knowledge, clinical photographs, panoramic radiographs, and IOS-based 3D reasoning.The benchmarks cover closed-set classification, structured prediction, and open-ended clinical question answering.
  • B. Benchmarks: DentalBench contains 7,332 English and Chinese questions, reporting accuracy for close-ended questions and BERTScore for open-ended questions.It includes 5,408 close-ended and 1,924 open-ended questions.
  • C. Main Results: 33.31% BERTScore is DentAgent’s best open-ended DentalBench result, outperforming the strongest dental-adapted baseline by 2.74%.On close-ended questions, DentAgent reaches 66.32% accuracy, trailing DeepSeek-R1 by 2.12% while outperforming GPT-4o and all dental-adapted baselines.
  • C. Main Results: 16.13% is DentAgent’s advantage over OralGPT-Omni on MMOralBench, while it also surpasses all evaluated medical agent frameworks.The result is attributed largely to specialized perception tools that extract fine-grained pathological findings.
  • C. Main Results: 17.3 percentage points separate DentAgent from senior specialists on the DentVLM multi-label diagnostic task, with DentAgent outperforming all baseline models.Comparisons include general-purpose MLLMs, medical-specific MLLMs, and human readers with varying clinical experience.
  • C. Main Results: 6.12% and 14.98% are DentAgent’s gains over the strongest 3D-specific IOSVLM and best 2D multiview baseline, respectively, on IOSVQA.DentAgent grounds reasoning in geometric evidence extracted directly from the native 3D mesh, preserving metric and relational cues.

D. Analysis

The analysis attributes DentAgent’s gains primarily to coordinated agentic orchestration, with iterative verification further improving evidence quality. It also finds that the workflow remains compatible with stronger reasoning backbones and benefits from their advances.

  • Agentic Orchestration Drives the Improvement: Agentic orchestration produces a 17.48% gain over direct inference before iterative verification is enabled.The comparison fixes Qwen3.5-9B as the reasoning backbone across variants.
  • Iterative Verification Further Improves Evidence Quality: Iterative verification raises the score from 48.40 to 52.83 with the backbone and toolset unchanged.The complete Identify-Delegate-Observe-Verify loop revisits insufficiently examined targets and addresses evidence gaps and conflicts.
  • Iterative Verification Further Improves Evidence Quality: Reflection further enhances the completeness and consistency of DentAgent’s supporting evidence.This complements the dominant improvement attributed to agentic tool use.
  • Compatibility Across Reasoning Backbones: DentAgent exhibits steady performance gains as the capability of its underlying reasoning backbone increases.The workflow is fixed across various backbones, indicating compatibility with different models and continued benefit from advances in their reasoning.

V. CONCLUSION

DentAgent is an evidence-centric multi-agent framework that separates evidence acquisition from response generation while preserving source-attributed findings throughout inference. It coordinates five specialist agents and demonstrates strong performance across four benchmarks spanning diverse dental modalities and tasks.

  • Framework: DentAgent separates evidence acquisition from response generation and preserves source-attributed findings throughout inference.The framework uses an Orchestrator and Evidence Blackboard to coordinate this process.
  • Framework: Five specialist agents are coordinated through an Orchestrator and an Evidence Blackboard.These components form the framework’s multi-agent coordination mechanism.
  • Evaluation: Across four benchmarks, DentAgent demonstrates strong performance across diverse dental modalities and tasks.The conclusion reports broad performance across the evaluated benchmark and task settings.
Loading 2608.18878v1…