Source-linked AI summary
Agentic AI uncovers conserved cross-tissue protein co-abundance programs inaccessible to single-dataset analysis
Runyu Guan, Dehao Wu, Qiqi Xie, Yang Li, Haohan Wang
TL;DR
Prior cross-tissue studies focused on selected tissue pairs, leaving most possible combinations unexplored. The paper applies an evidence-grounded LLM-agent framework to all 820 pairwise combinations of 41 tissues and fluids. It identifies 1,833 conserved clusters across 406 pairs, distinguishes broad connectivity from deep pairwise conservation, and generates testable mechanistic hypotheses.
Problem
Prior studies selected tissue pairs from biological knowledge, leaving most possible combinations unexplored despite their potential to reveal shared disease mechanisms and therapeutic targets.
Method
The framework constructs tissue-specific co-abundance networks, derives pairwise consensus clusters across all 820 combinations, and integrates evidence from curated biological resources and literature.
Results
The analysis identified 1,833 conserved co-abundance clusters across 406 tissue pairs and found distinct axes of broad connectivity and deep pairwise conservation.
Takeaways & Limitations
The resulting landscape provides a hypothesis-generating resource for investigating disease programs that may be detectable or modifiable outside their primary pathology sites.
Takeaways & Limitations
Co-abundance does not establish physical interaction, causal regulation, or direct inter-tissue communication, and shared clusters may reflect composition, sampling depth, or systemic physiology.
Abstract
from arXiv · showhide
Protein co-abundance clusters preserved across tissues can reveal shared disease mechanisms and candidate therapeutic targets, particularly when proteins implicated in organ-confined diseases converge in peripheral or accessible tissues. However, previous cross-tissue studies have focused on biologically pre-selected tissue pairs, leaving most possible combinations and non-obvious relationships unexplored. We present an LLM-agent framework for large-scale, evidence-grounded comparison of tissue-specific protein co-abundance networks. The framework constructs tissue networks, derives pairwise consensus clusters, and integrates evidence from expression atlases, protein interaction and complex databases, pathway annotations, disease catalogues, and literature. Applied to all 820 pairwise combinations of 41 human tissues and fluids, it identified 1,833 conserved co-abundance clusters across 406 tissue pairs. Colon, synovial fluid, blood, cerebrospinal fluid, and bone marrow were the most broadly connected tissues, while the most cluster-rich pairs were dominated by bone marrow. The analysis also highlighted non-obvious relationships: skin-bone marrow exceeded the anatomically adjacent bone-bone marrow pair, while colon-breast contained cancer-relevant clusters involving extracellular-matrix remodeling, lipid metabolism, and immune modulation. Cluster-level analyses generated further mechanistic hypotheses, including a brain-gut extracellular-vesicle/redox/serotonin-cofactor axis and a liver-bone marrow stress-response axis involving genes linked to white matter disease. These results provide a global, comparable landscape of conserved protein co-abundance and a hypothesis-generating resource for mechanistic and therapeutic exploration. Code and data are available at https://github.com/Gry1005/AgenticAI-conserved-cross-tissue-protein-co-abundance.
Introduction
Prior studies selected biologically motivated tissue pairs, leaving most of the 820 possible combinations unexplored. The paper introduces an evidence-grounded LLM-agent framework to systematically compare all pairwise combinations and identify conserved protein programs relevant to disease mechanisms and therapeutic targets.
- Exhaustive comparison could reveal shared disease mechanisms and candidate therapeutic targets when disease-associated proteins converge in peripheral or accessible tissues.Such relationships may identify intervention points outside the primary pathology site.
- Prior cross-tissue studies focused on biologically pre-selected tissue pairs, leaving most combinations among 41 tissues unexamined.The unexplored space comprises 820 possible tissue pairs.
- Manual analysis was infeasible because each cluster required integrating expression, interaction, complex, pathway, disease, and literature evidence across thousands of candidates.The practical bottleneck was evidence integration at scale rather than the conceptual difficulty of comparing tissues.
- The proposed LLM-agent framework compares all 820 pairwise combinations of 41 human tissues and biofluids using evidence integration and hypothesis organization.The agent is presented as infrastructure for organizing evidence rather than as an autonomous source of biological conclusions.
Cross-tissue landscape of protein co-abundance conservation
The analysis constructed tissue-specific co-abundance networks and found widespread but structured conservation across tissues. Conserved cluster sharing differed from baseline expression similarity, indicating that it captures coordinated regulatory programs rather than simply similar protein abundance.
- 406 of 820 tissue pairs retained at least one shared protein co-abundance cluster after housekeeping-gene removal.This indicates that conserved protein coordination is widespread across the analyzed tissues and fluids.
- The pairwise landscape showed structured groups with high cluster sharing alongside pairs with little or no overlap.Tissues were organized using hierarchical clustering of shared cluster counts.
- Expression similarity was broadly high across tissue pairs, whereas conserved cluster sharing displayed more differentiated patterns.Figure 1 compares cluster sharing in the lower triangle with Pearson correlation of log-transformed mean iBAQ values in the upper triangle.
- The contrast between the two views indicates that conserved cluster sharing captures coordinated regulatory programs rather than baseline protein-abundance overlap.This interpretation follows the distinct patterns observed for cluster sharing and expression similarity.
- Colon shared clusters with 40 tissues, followed by synovial fluid and blood with 39 each, and cerebrospinal fluid and bone marrow with 38 each.Stomach, lymph node, and uterus were connected to 7, 8, and 9 tissues, respectively.
From co-abundance clusters to biological mechanisms
Cluster annotation recovered established biological programs and identified a functionally coherent liver–adipose pattern absent from current interaction and complex databases. Together, these results support the clusters’ biological coherence while showing that they can expose coordination not captured by existing resources.
- 1,833 conserved clusters were annotated with Gene Ontology, KEGG, STRING, the Human Protein Atlas, and hu.MAP 3.0.The annotations assessed both recovery of established programs and previously uncharacterized coordination patterns.
- The bone marrow–prostate pair shared 21 clusters spanning vesicle trafficking and extracellular exosome biology, distributed metabolism, and chromatin/RNA-processing regulation.These programs correspond to established roles in intercellular communication, metabolism, and gene-expression regulation.
- The recovery of established programs from co-abundance alone supports the biological coherence of the recovered clusters.
- The liver–adipose pair contained a 7-protein cluster linking calcium-associated proteins with mitochondrial and NADP-dependent redox metabolism.Enrichment included mitochondrial calcium ion homeostasis (p = 2.61 × 10^-2) and NADP annotation (STRING FDR = 1.52 × 10^-2).
- The liver–adipose cluster had no documented STRING interactions among its 21 protein pairs and no corresponding hu.MAP 3.0 complex.Its functional coherence despite database absence suggests an undercharacterized cross-compartment coordination pattern.
Systematic ranking of cross-tissue protein coordination
The study ranked tissues by broad connectivity and pairs by depth of conserved cluster sharing, revealing distinct forms of cross-tissue centrality. Colon and synovial fluid were broadly connected, while bone marrow dominated the deepest pairwise relationships and non-obvious disease-relevant pairings emerged.
- Connectivity significance used a binomial test with baseline p0 = 0.495, derived from 406 of 820 pairs sharing at least one cluster.Benjamini-Hochberg FDR correction classified tissues into connectivity tiers.
- 12 tissues showed significantly elevated connectivity, 18 showed reduced connectivity, and 10 were indistinguishable from random expectation.Colon, synovial fluid, blood, cerebrospinal fluid, and bone marrow were among the elevated-connectivity tissues.
- Colon and synovial fluid ranked above blood and cerebrospinal fluid in cross-tissue cluster sharing.This was unexpected because blood and cerebrospinal fluid are canonical systemic fluids.
- Nine of the top 10 tissue pairs included bone marrow, despite bone marrow ranking fifth in tissue-level connectivity.This separates broad connectivity from deep pairwise conservation.
- Skin–bone marrow shared 26 clusters, exceeding the anatomically adjacent bone–bone marrow pair’s 22 clusters.The top pair was significantly enriched relative to degree-matched random networks (p < 0.001).
- The 26 skin–bone marrow clusters spanned chromosome condensation, mitochondrial metabolism, vesicle transport, and chromatin remodeling, with none mapping to known hu.MAP 3.0 complexes.
- Colon–breast was the only top-10 pair without bone marrow and contained cancer-relevant clusters involving extracellular-matrix remodeling and apoptosis regulation.Its shared proteins included AKR1B10, MMP9, PYCARD, and SDC1, which are independently implicated in colorectal and breast carcinogenesis.
Novel cross-tissue biological mechanisms
Cluster-level analyses uncovered biologically coherent but previously uncharacterized cross-tissue programs in brain–gut and liver–bone marrow comparisons. These findings connect extracellular-vesicle, metabolic, redox, and disease-associated stress-response signals while remaining hypothesis-generating.
- Brain–gut mechanisms: The brain–gut consensus cluster was enriched for extracellular exosomes and blood-associated tissue signatures rather than brain or gut.Extracellular exosome was the strongest cellular-component signal, while hematopoietic system and blood plasma were the dominant tissue-enrichment signals.
- Brain–gut mechanisms: The cluster combined EV biogenesis and cargo sorting with complement biology relevant to brain synaptic pruning and gut barrier maintenance.The ensemble included CD9, VPS4B, ATP6V1A, TRIM25–UBE2N, and C3.
- Brain–gut mechanisms: The brain–gut cluster linked BH4-dependent neurotransmitter cofactor biosynthesis with thioredoxin-mediated antioxidant defense through SPR and TXNRD1.SPR supports BH4 production for serotonin synthesis, while TXNRD1 is central to the thioredoxin antioxidant system; their co-abundance was not documented in STRING or hu.MAP 3.0.
- Brain–gut mechanisms: Only 2 of 190 possible protein pairs in the brain–gut cluster had documented STRING interactions, and no corresponding hu.MAP 3.0 complex existed.The sparse prior interaction evidence highlights why co-abundance analysis can surface ensembles not captured by current interaction or complex databases.
- Validation: Experimental validation is required to determine whether the brain–gut proteins are co-packaged in circulating extracellular vesicles.Suggested tests include co-enrichment in plasma-derived EVs and EV-mediated transfer under gut-dysbiosis conditions.
- Liver–bone marrow mechanisms: The liver–bone marrow analysis recovered a 23-protein cluster enriched for thioredoxin-mediated redox regulation and methionine salvage/polyamine biosynthesis.Protein-disulfide reductase activity was the most distinctive molecular-function signal, and SMS–MTAP provided the highest-confidence pairwise interaction.
- Liver–bone marrow mechanisms: AIMP1, CAPN1, and EIF2B5 co-occurred in the liver–bone marrow cluster despite independent links to distinct heritable white matter diseases.AIMP1 and EIF2B5 connect to the GCN2–eIF2α integrated stress response, which also regulates bone marrow hematopoietic stem-cell proteostasis.
Discussion
The study reframes cross-tissue proteomics as an exhaustive search across the whole-body combinatorial landscape rather than comparison of biologically preselected tissue pairs. Its results distinguish broad connectivity from deep pairwise conservation and generate peripheral-tissue disease hypotheses that require experimental validation.
- Interpretation: The framework searches the whole-body combinatorial landscape instead of restricting analysis to tissue pairs selected by prior biological intuition.This expands discovery beyond relationships motivated by targeted hypotheses.
- Interpretation: Broad tissue connectivity and deep pairwise conservation are distinct axes of cross-tissue organization.A tissue can share clusters with many partners without dominating the most cluster-rich pairings, and another can form unusually deep pairwise relationships without broad connectivity.
- Implications: Proteins implicated in organ-confined diseases may converge in peripheral or anatomically distant tissues, creating hypotheses about detection or modulation outside the primary pathology site.These candidates remain hypotheses rather than established mechanisms.
- Limitations: The LLM-agent framework is evidence-integration infrastructure, not an autonomous source of biological truth.Mechanistic models remain hypotheses grounded in co-abundance clusters and curated external evidence.
- Limitations: Co-abundance does not establish physical interaction, causal regulation, or direct inter-tissue communication.Shared clusters may instead reflect cell-type composition, sampling depth, or systemic physiological states.
Appendix A Methodology
The methodology uses a four-stage LLM-agent pipeline to construct tissue networks, analyze consensus modules, integrate biological evidence, and generate hypotheses with novelty assessment.
- Appendix A Methodology: The pipeline comprises adaptive network construction, consensus network and module analysis, RAG-enabled knowledge synthesis, and hypothesis generation with novelty assessment.These stages are summarized in Figure A1 and elaborated in the following subsections.
- Appendix A Methodology: Stage 1 constructs tissue-specific co-abundance networks with agent-driven threshold refinement.
- Appendix A Methodology: Stage 2 derives pairwise consensus networks, applies Leiden community detection, and filters modules.
- Appendix A Methodology: Stage 3 annotates modules through RAG-based searches of HPA, STRING, hu.MAP 3.0, GO, and KEGG.
- Appendix A Methodology: Stage 4 infers module functions, profiles genes, and assesses novelty using Open Targets and Europe PMC.
A.1.1 Data Acquisition and Preprocessing
The preprocessing workflow normalizes protein expression profiles into tissue-specific matrices, constructs weighted co-abundance graphs, and uses an agent-driven iterative procedure to refine sparse network topology.
- A.1.1 Data Acquisition and Preprocessing: Normalized protein expression profiles were aggregated within each tissue and truncated to a common minimum length, producing matrix X ∈ R^N×Lmin.Each row x_i represents the expression vector of protein p_i.
- A.1.1 Data Acquisition and Preprocessing: Tissue-specific co-abundance networks represent proteins as vertices and expression-profile similarity as weighted edges.The graphs are modeled as undirected weighted networks.
- A.1.1 Data Acquisition and Preprocessing: Hard thresholding retains an edge only when the absolute similarity score exceeds the predefined threshold τ.
- A.1.1 Data Acquisition and Preprocessing: An autonomous code agent coordinates data retrieval, heuristic filtering, and quality assurance through an iterative reasoning and execution loop.It selects filtering parameters for iBAQ data and dynamically defines Lmin and θ_iBAQ using current literature guidelines.
- A.1.1 Data Acquisition and Preprocessing: The agent adjusts τ and regenerates the network until predefined modularity criteria or step limits are reached.It evaluates whether the topology has a single giant component and sufficiently dense protein-complex structure.
A.2 Consensus Network Analysis and Module Annotation
Pairwise consensus networks retain shared proteins and edges across tissues, after which communities are filtered and annotated through curated databases before agent-generated functional and novelty analyses.
- A.2 Consensus Network Analysis and Module Annotation: Consensus networks are defined on the intersection of proteins present in two tissue-specific networks, retaining edges found in both networks.The consensus edge weight is assigned as the minimum of the two original weights.
- A.2 Consensus Network Analysis and Module Annotation: Leiden community detection partitions consensus graphs into tightly connected communities at resolution 1.0.A subgraph topology filter retains communities meeting specified size and density constraints.
- A.2 Consensus Network Analysis and Module Annotation: The workflow removes overlapping housekeeping proteins, maps gene symbols through UniProtKB, and serializes finalized modules in JSON and markdown formats.
- A.2 Consensus Network Analysis and Module Annotation: RAG retrieves expression overlap, known complexes, pathways, and functional interactions from HPA, hu.MAP 3.0, GO, KEGG, and STRING.
- A.2 Consensus Network Analysis and Module Annotation: The agent integrates module topology with database profiles to infer functions, assess novelty through literature searches, annotate uncharacterized genes, and retrieve disease association scores.Novelty is quantified using publication volume in Europe PMC, while disease associations come from Open Targets.
Appendix B Case Study Supporting Evidence
The case-study evidence combines curated enrichment tables with unedited agent-generated reports, including a brain–gut module with complete HPA coverage and selected pathway, activity, and interaction signals.
- Appendix B Case Study Supporting Evidence: Each case study is supported by an integrated GO/KEGG/STRING/hu.MAP enrichment table and a raw markdown report produced by the LLM-agent pipeline.The reports are shown verbatim as single grey-shaded artifacts without editorial cleanup.
- Appendix B Case Study Supporting Evidence: Brain–Gut Module 0 contains 20 proteins, with HPA coverage for all 20 and no hu.MAP 3.0 match.STRING reports 2 of 190 protein pairs, or 1.05%.
- Appendix B Case Study Supporting Evidence: The enrichment evidence includes extracellular membrane-bounded organelle annotation with p/FDR 2.25 × 10^-8 and alcohol dehydrogenase activity with p/FDR 5.30 × 10^-3.
- Appendix B Case Study Supporting Evidence: Reported STRING protein interactions include SORD ↔ DCXR at 0.945 and UBE2N ↔ TRIM25 at 0.990.
Brain-Gut Module Analysis Report
The brain–gut module comprises proteins associated with extracellular vesicles, metabolic processes, and cytosolic functions, with STRING identifying specific protein interactions. Individual proteins were further linked to characterized cellular roles and disease associations.
- Module overview: 20 proteins in the module are associated with metabolic processes and extracellular regions.The module was interpreted as potentially supporting extracellular vesicle-mediated communication between brain and gut tissues.
- Functional enrichment: GO/KEGG enrichment identified extracellular exosome, extracellular space, and cytosol components alongside uronic acid metabolism and xylulose 5-phosphate biosynthesis.STRING analysis supported associations with cytosol and extracellular exosome components.
- Interaction evidence: Specific protein–protein interactions included SORD–DCXR and UBE2N–TRIM25.These interactions were identified by STRING analysis within the module.
- Protein-level interpretation: CNDP2 was annotated as a nonspecific dipeptidase involved in proteolysis and was concluded to participate in brain–gut communication, particularly extracellular vesicle-mediated processes.Its listed disease associations included duodenal ulcer, refractive error, myopia, and kidney disease.
- Module annotation: 23 genes in the related module were associated with cytosol, cytoplasm, extracellular exosome, extracellular vesicle, and extracellular organelle functions.The module was marked as present in STRING and as a novel finding because no complexes were found in hu.MAP 3.0.