Source-linked AI summary

DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models

Wenxin Duan, Hanwei Wang, Zhongying Peng, Zhonghua Lu, Jiayi An, Fan Song, Yong Liang

arXiv:2608.18103v1cs.CLcs.AI

TL;DR

Explaining herbal formula mechanisms while integrating traditional Chinese medicine theory with modern science remains difficult. DeepTCM1.0 addresses this gap through multi-expert collaboration and achieved the highest scores across five evaluation dimensions.

  • Problem

    Scientific elucidation of herbal compound mechanisms and integration of traditional theories with modern biology remain limited.

  • Method

    DeepTCM1.0 uses a multi-expert agent collaboration framework to map TCM concepts to measurable indicators in modern biology.

  • Results

    DeepTCM1.0 achieved the highest scores across all five evaluation dimensions, with its mean total score significantly exceeding other evaluated frameworks.

  • Takeaways & Limitations

    The framework provides a multi-expert approach for mechanistic analysis of herbal formulas, illustrated through Guizhi Decoction.

  • Takeaways & Limitations

    The proposed mechanism hypotheses still require experimental verification.

Abstract

from arXiv · show

Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are insufficient for achieving deep integration between classical TCM theory and modern scientific research. In addition, direct question-answering using general-purpose artificial intelligence large language models is limited by inadequate adaptation to TCM theoretical frameworks and susceptibility to reasoning hallucinations. Consequently, there is an urgent need to develop intelligent analytical methods aligned with the holistic principles of TCM. Objective: To establish a multi-expert intelligent agent framework integrating classical TCM theory with modern life sciences, thereby enabling systematic and interpretable mechanistic analysis of TCM compound formulas, with Guizhi Decoction serving as a representative validation case. Methods: The DeepTCM1.0 framework was constructed based on the general-purpose large language model DeepSeek V3.2. It adopts a three-tier collaborative architecture and a three-round iterative quality-control workflow, simulating the collaborative analytical process of 11 interdisciplinary intelligent agents. The framework was applied to the mechanistic interpretation of Guizhi Decoction from the dual perspectives of classical traditional Chinese medicine theory and modern scientific research. Framework performance was comprehensively evaluated through double-blind five-dimensional scoring, intraclass correlation coefficient (ICC) reliability testing, Mann-Whitney U tests, and effect size analysis. The evaluation employed four independent large language models as evaluators, each conducting five rounds of repeated scoring on five anonymized reports, resulting in a total of 100 independent scoring assessments.

1.Introduction

Mechanistic elucidation of TCM formulas remains challenging because existing approaches inadequately capture synergistic, dynamic, multi-scale mechanisms and integrate classical theory with modern evidence. DeepTCM1.0 addresses this gap through a multi-expert intelligent-agent framework based on general LLMs and aligned with TCM’s systemic principles.

  • Research motivation: Mechanistic elucidation of herbal compound prescriptions and integration of traditional theories with modern evidence remain central challenges in TCM modernization.
  • Limitations of existing methods: Traditional data mining remains largely descriptive, failing to reveal complex component synergies and molecular regulatory networks.
  • Limitations of existing methods: Network pharmacology is constrained by fragmented, infrequently updated datasets, predominantly static models, and insufficient integration across molecular, cellular, and tissue scales.
  • Limitations of existing methods: Molecular docking is difficult to align with pattern differentiation, holistic regulation, and formula compatibility, while limited functional compatibility, high costs, and low efficiency restrict its applicability.
  • Rationale for agent-based analysis: Multi-agent systems fit TCM’s holistic and systemic characteristics, while consensus-seeking, evidence-based querying, and critical reflection can reduce LLM hallucinations.
  • Study contribution: The study proposes DeepTCM1.0, a multi-expert intelligent-agent framework using general LLMs to explore TCM mechanisms, with DeepSeek V3.2 as its central AI-agent brain.

Medicine Mechanism Research · 2 Methods · Expert Agent Framework

DeepTCM1.0 is a TCM-oriented, multi-expert agent framework that integrates classical theory with modern life sciences through traceable, interdisciplinary reasoning. Its three-tier architecture decomposes research tasks, coordinates expert collaboration, and validates integrated outputs, with Guizhi Decoction used as a representative case.

  • Medicine Mechanism Research: DeepTCM1.0 aligns mechanistic analysis with pattern differentiation-based treatment, holistic regulation, and compound prescription while connecting TCM concepts with measurable modern biological indicators.The framework establishes a traceable research evidence chain and supports efficient decomposition of complex interdisciplinary tasks.
  • Medicine Mechanism Research: The framework simulates collaboration among classical TCM, formula compatibility, pharmacology, network pharmacology, molecular biology, and molecular docking experts.These perspectives are used to elucidate herbal-formula mechanisms comprehensively.
  • 2.2.1 First Tier: Task Decomposition and Team Formation: For Guizhi Decoction, task decomposition spans classical theoretical tracing, formula compatibility analysis, and pharmacodynamic-material-basis and pharmacokinetic evaluation, producing six core sub-tasks.The example illustrates how user questions are converted into specialized research domains.
  • Medicine Mechanism Research: The framework is intended as a reproducible technical basis for classic-formula research and syndrome-mechanism elucidation, supporting TCM’s transition toward intelligence-driven scientific innovation.Guizhi Decoction serves as the representative case for demonstrating mechanistic analysis.
  • 2.2 Three-Tiered Collaboration Framework: The resulting three-tier framework comprises task decomposition and team formation, multi-expert collaborative reasoning and emergent insight, and integrative validation and output.It is implemented as a modular closed-loop operational architecture.
  • 2.2.1 First Tier: Task Decomposition and Team Formation: The first tier receives research questions, decomposes them into analytical pathways, and assembles a dedicated interdisciplinary expert team.The Principal Investigator and Recruiter form the core decision-making and control hub.
  • 2.2.1 First Tier: Task Decomposition and Team Formation: The Recruiter combines predefined fixed experts with dynamically recruited specialists according to task requirements to provide comprehensive disciplinary coverage.This configuration supports systematic integration of traditional TCM with cutting-edge technologies.

Integration

DeepTCM1.0 operationalizes TCM-guided interdisciplinary integration through specialized expert roles, iterative critic–expert deliberation, and a final report-integration stage. The workflow produces traceable, logically coherent analyses with implementable research schemes and consensus conclusions.

  • Expert integration: A hybrid team of fixed and dynamically recruited experts spans core TCM theory, formula systems, modern medicine, pharmacology, immunology, neuroendocrinology, metabolomics, and microbiomics.The framework characterizes this paradigm as TCM-guided, multidisciplinary collaboration with specialized division of labor.
  • Expert integration: Standardized role prompts define each agent’s knowledge background, responsibilities, and interpretation criteria, supporting focused perspectives and compliance.The Classical Literature Expert prompt links classical sources and commentaries with modern molecular mechanisms, biomarkers, and experimental directions.
  • Iterative deliberation: A critic uses reverse-evidence-based reasoning to examine argument integrity, evidence sufficiency, and compatibility between classical theory and modern findings, reducing hallucinations.Bidirectional critic–expert debate also supports emergent reasoning and mechanistic insight generation.
  • Iterative deliberation: Three-round deliberation proceeds from independent interpretation to TCM-first integration and cross-validation, then targeted optimization using supplementary evidence and multidisciplinary analysis.The final round applies AI-assisted disciplinary reasoning to generate novel research hypotheses addressing core issues and existing gaps.
  • Outcome delivery: The final integration stage consolidates multi-expert deliberations into logically complete reports, implementable and testable research schemes, and consensus conclusions.This closes the compound-mechanism analysis loop and supports traceable, progressively structured reasoning.

3 Results

DeepTCM1.0 completed a full-process analysis of Guizhi Decoction’s efficacy mechanism using three-layer collaboration and three-round iteration. The system decomposed the task into six sub-tasks and dynamically assembled an 11-expert interdisciplinary team without manual intervention.

  • The system completed a full-process analysis of Guizhi Decoction’s efficacy mechanism under a three-layer collaboration framework and three-round iteration mechanism.
  • The Principal Investigator decomposed the analysis into six sub-tasks spanning classical theory, formula compatibility, pharmacokinetics, network regulation, metabolomics and microbiomics, and clinical evidence.
  • The system dynamically assembled 11 experts, including seven fixed core experts and four specialized experts selected according to task demands.
  • Specialized roles covered immunology and neuroendocrinology, metabolomics and microbiomics, systems biology and network pharmacology, and pharmacokinetics and pharmaceutical analysis.
  • The process operated without manual intervention, supporting comprehensive disciplinary coverage and deeper interdisciplinary collaboration.
  • In the first round, 11 experts independently developed viewpoints, underwent critical inquiry and quality-control audits, and contributed to an initial multidimensional evidence-based report.

Interpretation

The interpretation integrates classical TCM theory with modern biomedical mechanisms through Guizhi Decoction’s compatibility principles, multi-system regulation, and dynamic compound–metabolite–microbiota basis. It also identifies evidence limitations while proposing measurable physiological and molecular links.

  • Classical–modern integration: Guizhi–Shaoyao compatibility balances dispersion and restraint through complementary TRPV1 activation and NF-κB inhibition.Cinnamaldehyde acts as a TRPV1 agonist, whereas paeoniflorin inhibits TLR4/NF-κB signaling, linking classical compatibility to immune–inflammatory regulation.
  • Systems mechanism: The proposed “three-tier, dual-axis, one-foundation” model connects NEI regulation, vascular–microcirculatory effectors, and gut–metabolic support through thermoregulatory and barrier-immunity axes.Ying–wei disharmony is modeled as acute NEI maladaptation involving inflammatory imbalance, autonomic dysregulation, and abnormal HPA-axis rhythmicity.
  • Evidence and limitations: Clinical evidence is graded B–C, indicating moderate-quality support for treating the exterior deficiency pattern of common cold, with reported reductions in IL-6 and TNF-α.The interpretation also clarifies that “moderate” describes reduced vascular tension rather than a slowed pulse rate, while direct evidence linking the TCM concept remains absent.
  • Biomarker translation: The framework identifies TNF-α/IL-10 ratio, salivary cortisol rhythmicity, HRV indices (HF/LF ratio), and salivary sIgA as core biomarkers of NEI maladaptation.The proposed abnormalities include reduced HF power of HRV and disrupted cortisol circadian rhythm.
  • Pharmacokinetic and microbiota basis: Guizhi Decoction’s in vivo bioactive basis comprises prototype compounds, metabolites, and microbiota-derived products, forming a three-tier link from classical syndromes to molecular targets.Cinnamaldehyde has less than 1% oral bioavailability and predominantly acts through cinnamic acid, while paeoniflorin is converted by gut microbiota into more active metabolites.

Execution Example: Statement by the Classical Literature Expert (Optimized Following · Critic Feedback)

Following critic feedback and three rounds of iterative discussion, DeepTCM 1.0 refined classical citations, standardized hypotheses, and integrated classical and modern evidence into a mechanistic consensus and testable research plans for Guizhi Decoction.

  • Critic Feedback: Critic feedback refined classical citations, corrected hypotheses, recalibrated interpretations, and incorporated commentator support for Guizhi Decoction’s therapeutic principles.The optimized conclusions emphasize harmonizing ying–wei and restoring yin–yang self-regulation.
  • Integration, Verification and Output: After three rounds of iterative collaborative discussion, DeepTCM 1.0 produced structured consensus findings, mechanistic hypotheses, and corresponding research plans.The Principal Investigator, critic, and report integration specialist participated throughout the discussion and output process.
  • Core Positioning at the Classical Theory Level: Guizhi Decoction fundamentally addresses disharmony between ying and wei, with a classical 1:1 Guizhi-to-Shaoyao ratio balancing dispersion and astringency.Ginger, jujube, licorice, and warm-porridge administration are described as integral to therapeutic efficacy.
  • Mechanism Analysis at the Modern Science Level: The modern mechanism is organized as a three-tier network harmonization model spanning NEI regulation, vascular–microcirculatory effects, and gut–metabolic support.Reported mechanisms include TRPV1, NF-κB, HPA-axis, cholinergic nerve–sweat gland, vascular tone, PGE2, SCFAs, and intestinal-barrier regulation.
  • Material Basis and In Vivo Process: Guizhi Decoction’s in-vivo bioactive basis is a dynamic ensemble of prototype compounds, metabolites, and microbiota-derived products.Cinnamaldehyde has oral bioavailability of less than 1%, while component interactions involving intestinal CYP3A4 may increase active-constituent bioavailability.
  • Resolution of Core Controversies and Blind Spot Remediation: Cross-disciplinary deliberation resolved the prototype-compound versus metabolite debate by defining efficacy as arising from a dynamic ensemble with different forms predominating.A three-level Classic Syndrome–Intermediate Physiological Phenotype–Microscopic Molecule framework linked syndrome indicators to measurable physiology and molecular targets.
  • Standardization of Research Scheme: Three rounds of discussion established two validation approaches: a cold-exposure/LPS composite animal model and a deep-phenotyping clinical study.The standardized warm-porridge procedure used 200 mL at 60–70 °C followed by 32–34 °C for 1 h.
  • Innovative Mechanism Hypotheses and Research Plans: Based on the consensus, DeepTCM 1.0 proposed testable mechanistic hypotheses and corresponding validation research plans.One explicitly standardized hypothesis links Guizhi Decoction’s porridge-administration method to gut microbiota-derived metabolites and remains experimentally unvalidated.

Hypothesis 1: Three-Layer Network Harmonization Hypothesis and Its Research Plan

The hypothesis proposes that Guizhi Decoction achieves ying–wei harmonization through coordination among core regulatory, core effector, and foundational support layers. Its research plan traces a closed-loop sequence from peripheral initiation through central integration to foundational support.

  • Three-layer network harmonization hypothesis: Guizhi Decoction is hypothesized to harmonize ying–wei through cross-level coordination among the NEI, vascular–microcirculatory, and gut–metabolic layers.These are defined respectively as the core regulatory, core effector, and foundational support layers.
  • Research plan: Cinnamaldehyde and paeoniflorin initiate peripheral signaling by activating cutaneous TRPV1 receptors.The passage identifies these compounds as key bioactive compounds.
  • Research plan: NF-κB-mediated modulation of immune–inflammatory balance is followed by stabilization of neuroendocrine rhythms through the HPA axis.The proposed sequence links immune–inflammatory regulation with neuroendocrine stabilization.
  • Research plan: Gut microbiota-derived SCFAs optimize metabolic homeostasis, completing a closed-loop system of peripheral initiation, central integration, and foundational support.The system is characterized as “peripheral initiation-central integration-foundational support.”

Hypothesis 2: Compatibility-Metabolism-Microbiota Synergy Hypothesis and Its … 2. Double-Blind Scoring Procedure

The paper proposes complementary mechanisms for Guizhi Decoction, combining herb compatibility, metabolism, microbiota regulation, and a dynamic ensemble of active substances. It evaluates DeepTCM1.0 against four single-model reports using a double-blind, five-dimensional scoring procedure with 100 independent assessments.

  • Hypothesis 2: Compatibility-Metabolism-Microbiota Synergy Hypothesis and Its Research Plan: The proposed compatibility and microbiota mechanisms synergize with warm, diluted porridge intake to promote SCFA production, improve metabolic homeostasis, and support the spleen–stomach source of ying–wei transformation.This hypothesis is presented as a dual immune–metabolic mechanism.
  • Hypothesis 3: Dynamic Effector Substance Collection Hypothesis and Its Research Plan: Guizhi Decoction’s in vivo effects are attributed to a dynamic ensemble of prototype components, metabolites, and microbiota-transformed products rather than a single component.Examples include cinnamic acid and paeoniflorin derivatives among metabolites and glycyrrhetinic acid among microbiota-transformed products.
  • Hypothesis 3: Dynamic Effector Substance Collection Hypothesis and Its Research Plan: This dynamic ensemble is proposed to act synergistically on cholinergic synapses, NF-κB, and the HPA axis, exemplifying TCM’s “Multi-Component-Multi-Target-Multi-Pathway” regulatory paradigm.The hypothesis is illustrated in Figure 6.
  • 2. Double-Blind Scoring Procedure: The reports were scored anonymously by four independent large language model evaluators using a double-blind, five-dimensional Likert 5-point rubric.The five dimensions covered classical TCM content, modern scientific content, interdisciplinary integration, systematic integrity, and academic expression.
  • 1. Evaluation Subjects: The evaluation compared one DeepTCM1.0 multi-agent report with four independently generated reports from DeepSeek V4, Doubao, Kimi k2.6, and Qwen3.5.All five reports addressed modernization-oriented mechanistic interpretation of Guizhi Decoction.
  • 2. Double-Blind Scoring Procedure: 100 independent scoring assessments resulted from four evaluators scoring each of five anonymized reports five times under randomized presentation order.Unified prompts, newly initialized dialogue contexts, cumulative dimension scoring, and post-scoring disclosure were used to reduce bias and ensure traceability.
  • 2. Double-Blind Scoring Procedure: The repeated large-model evaluation framework was further assessed using four complementary validation metrics to establish reliability and validity.These metrics were introduced after the double-blind scoring design.

1. Reliability Testing · 2. Validity and Difference Testing

Reliability was assessed through intra-evaluator ICC and inter-evaluator Krippendorff’s α, while validity and group differences were tested using repeated-measures ANOVA, η², and Mann–Whitney U analyses. These methods showed strong scoring consistency, significant report discrimination, and a large multi-agent versus single-model effect.

  • 1. Reliability Testing: Krippendorff’s α ranged from 0.842 to 0.867 across five dimensions, with an overall total-score coefficient of 0.868.The overall coefficient exceeded the 0.8 threshold for excellent consistency, indicating strong evaluator agreement.
  • 1. Reliability Testing: ICC(2,1) reliability analyzed total scores across five repeated scoring rounds using a two-way random-effects absolute-agreement model.The analysis focused on intra-evaluator reliability.
  • 2. Validity and Difference Testing: Repeated-measures one-way ANOVA treated report identity across five reports as the independent variable and all 100 scores as dependent observations.The analysis evaluated discriminative validity across reports.
  • 2. Validity and Difference Testing: F=66.969,P<0.001 and η²=0.738 demonstrated highly significant intergroup differences and excellent discriminative validity.The η² result indicates that 73.8% of total score variance could be explained by report quality differences.
  • 2. Validity and Difference Testing: Mann–Whitney U testing compared total scores between multi-agent and single-model groups under relatively small-sample and non-normal-distribution assumptions.The test used the Wilcoxon rank-sum framework for intergroup comparison.
  • 1. Reliability Testing: ICC values for Doubao, DeepSeek, Kimi, and Qwen were 0.867, 0.819, 0.789, and 0.798, respectively.All exceeded the acceptable threshold of ICC ≥0.7, indicating stable repeated scoring performance and controllable random error.
  • 2. Validity and Difference Testing: r=0.5403 reached the threshold for a large effect, and the multi-agent framework significantly outperformed single general-purpose large language models.The result was described as having highly robust statistical significance.

Report · Name · 4 Discussion

DeepTCM1.0 enabled multidimensional, cross-disciplinary, and logically coherent mechanism analysis of Chinese medicine formulas while maintaining alignment with classical TCM principles. Its lightweight design lowers implementation barriers, but knowledge coverage, prompt dependence, and experimental validation remain limitations.

  • 4 Discussion: DeepTCM1.0 achieved multidimensional mechanistic analysis of Guizhi Decoction by integrating cross-disciplinary knowledge, constructing coherent hypotheses, and aligning with TCM principles.The framework combined classical TCM theory with modern scientific perspectives in a representative formula case.
  • 4.1.1 Completeness of Cross-Disciplinary Knowledge Integration: Compared with conventional network pharmacology, DeepTCM1.0 integrated classical theory, clinical practice, pharmacology, and systems biology into a “classical theory–microscopic mechanisms–clinical application” chain.The framework used specialized agents and dynamic collaboration to address inefficient cross-domain knowledge integration.
  • 4.1.2 Logical Closed-Loop Nature of Mechanism Analysis: DeepTCM1.0 organized hypotheses through “Classic Traceability - Evidence Support - Logical Deduction,” with Critic-based quality control supporting traceable conclusions and argument reliability.The framework also allowed further validation through pharmacokinetic–pharmacodynamic modeling.
  • 4.2 Implications for TCM Methodology: The study’s primary methodological innovation was a lightweight, low-cost, reproducible pathway that required neither additional database construction nor model fine-tuning.Shareable prompts and process designs could be migrated to analyses of other classic formulas.
  • 4.2 Implications for TCM Methodology: DeepTCM1.0 used role prompt engineering to transform general large models into TCM domain experts and may support a shift toward intelligence-driven scientific TCM research.The proposed approach was intended to facilitate adaptation across different classic formulas.
  • 4.3 Limitations and Future Directions: The framework depends on pretrained general-model knowledge, making coverage incomplete for rare classic texts and region-specific traditions, while agent depth depends on prompt quality.Prompt engineering therefore requires iterative exploration.
  • 4.3 Limitations and Future Directions: The mechanism hypotheses remain primarily knowledge-based inferences requiring experimental verification; proposed next steps include RAG, automatic prompt optimization, collaborative validation, and testing other classic formulas.Suggested resources include TCMSP and SymMap, with Mahuang Decoction and Xiao Chaihu Decoction proposed for generalizability testing.

5 Conclusion · Availability of data and materials · Funding Declaration

DeepTCM1.0 is presented as a multi-expert agent framework that enabled multidimensional mechanistic analysis of Guizhi Decoction by integrating TCM theory, formula compatibility, and modern pharmacology. The authors describe it as a lightweight, reproducible, interpretable paradigm, with future work focused on experimental validation and extension; data are available from the corresponding author, and no specific funding was received.

  • 5 Conclusion: DeepTCM1.0 used multi-expert agent collaboration to achieve multidimensional mechanistic analysis of TCM formulae through the representative case of Guizhi Decoction.
  • 5 Conclusion: Its three-layer collaboration architecture and three-round iterative discussion mechanism integrated TCM classics, formula compatibility, and modern pharmacology.
  • 5 Conclusion: The framework generated innovative mechanism hypotheses with a logical closed loop and demonstrated advantages in cross-disciplinary knowledge integration, TCM theory adaptability, and interpretability.
  • 5 Conclusion: DeepTCM1.0 provides a lightweight and reproducible paradigm for analyzing complex TCM systems.
  • 5 Conclusion: Future work will focus on experimental validation and framework extension to promote application in the modernization of TCM research.
  • Availability of data and materials: Code and data from the current study are available from the corresponding author upon request.
  • Funding Declaration: The authors received no specific funding for this work.

Appendix

The appendix presents Table A1, which reports detailed scoring results across dimensions and total scores from five rounds of repeated scoring.

  • Table A1 reports detailed scores for each evaluation dimension.
  • The table includes a total score alongside dimension-level scores.
  • The reported scores were obtained from 5 rounds of repeated scoring.
Loading 2608.18103v1…