Source-linked AI summary
DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents
Yubin Wang, Xingjian Wei, Jiang Wu, Yinfan Wang, Boyu Zhu, Lin Zhang, Jianing Yu, Huazheng Zeng, Ruiyi Ding, Junyuan Gao, Jiaxing Sun, Lingli Ge, Haote Yang, Jingchao Wang, Aijia Guo, Qian Jiang, Yurui Zhao, Wenjian Zhang, Chen Zhu, Lijun Wu, Xiaolei Yang, Haodong Chen, Junjie Yuan, Zichao Ye, Shaowei Hou, Jing Ye, Jia Yu, Shan Wang, Lijun Wu, Jiantao Qiu, Chao Xu, Yuqiang Li, Guangyu Wang, Bowen Zhou, Dahua Lin, Conghui He
TL;DR
Organic-reaction knowledge is dispersed across patent text, images, and reaction schemes, limiting uniform machine-readable records for AI4Chem and retrieval. DianShi-RxnDB uses a fully automated extraction and normalization pipeline to build provenance-linked, fine-grained reaction records and shared researcher/AI-agent interfaces; it contains approximately 24 million instances, with approximately 14.8 million (61.7%) qualified, and five evaluated fields reached 92.95% micro-averaged accuracy.
Problem
Organic-reaction knowledge is distributed across heterogeneous patent and literature sources, making complete, uniform, machine-readable records difficult to construct for retrieval and AI4Chem research.
Method
DianShi-RxnDB uses a fully automated pipeline integrating patent text, images, and reaction schemes to extract, normalize, structure, qualify, and provenance-link single-step reaction records.
Results
Approximately 24 million Reaction Instances were produced, approximately 14.8 million (61.7%) passed automated qualification, and five evaluated fields achieved 92.95% micro-averaged field-level accuracy.
Takeaways & Limitations
The platform provides a shared structured data foundation with a Web workbench for researchers and an MCP service offering composable retrieval tools for AI agents.
Takeaways & Limitations
The reported manual quality evaluation applies only to five fields in qualified instances and does not establish overall instance correctness, unevaluated-field accuracy, or database recall.
Abstract
from arXiv · showhide
High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platform built via a fully automated extraction and normalization pipeline integrating patent text, images, and reaction schemes. Its corpus covers organic synthesis patents from the USPTO and EPO published between 1976 and 2025, yielding approximately 24 million reaction instances, of which approximately 14.8 million (61.7%) pass automated qualification checks. Each instance represents a specific single-step experiment recording participants, roles, quantities, temperatures, reaction times, yields, experimental procedures, and provenance links to source patents. In a manual evaluation of 1,300 sampled qualified instances, the micro-averaged field-level accuracy was 92.95%. A matched comparison with Pistachio further indicated advantages in deduplicated record counts, representation granularity, and field-level exact agreement. The platform provides a Web research workbench for searching, filtering, comparing, and source-verifying records, and a Model Context Protocol (MCP) service offering AI agents composable structured retrieval tools. DianShi-RxnDB is available at https://dianshi.opendatalab.org.cn/ .
1 Introduction
DianShi-RxnDB addresses fragmented organic-reaction knowledge by organizing patent-derived information into fine-grained, provenance-linked records and providing researcher- and AI-agent-facing access. Its fully automated platform combines broad coverage, structured experimental detail, and interfaces for retrieval, comparison, and source verification.
- Existing reaction resources have limitations in scale, patent coverage, instance-level detail, provenance localization, and data consistency, while professional databases commonly restrict access through subscriptions or licenses.
- DianShi-RxnDB applies a fully automated extraction and normalization pipeline to patent text, images, and reaction schemes, producing approximately 24 million Reaction Instances.The corpus draws primarily from USPTO and EPO organic-synthesis patents published from 1976 through 2025.
- Approximately 14.8 million Reaction Instances, or 61.7%, pass the automated qualification assessment.Qualified instances are those that pass implemented checks including atom-conservation checks for participants with usable structural information.
- Each single-step Reaction Instance records participant roles, quantities, conditions, yields, procedures, and provenance links, with structured relationships connecting substances, reactions, groups, and references.These relationships support comparison of related experiments and verification against source patents.
- The Web research workbench supports interactive search, filtering, comparison, linked exploration, and source verification, while the MCP service provides composable tools for AI-agent multi-step retrieval.Both interfaces operate over the same structured reaction-data foundation.
- The report evaluates five core fields at 92.95% micro-averaged field-level accuracy and reports higher deduplicated counts, finer-grained organization, and higher exact agreement than Pistachio in matched comparisons.The manual evaluation samples 1,300 qualified instances; the Pistachio comparison assesses six evaluated fields against source-grounded references.
2 Data foundation, construction, and quality evaluation
DianShi-RxnDB constructs a provenance-linked reaction database from USPTO and EPO patents through automated extraction, normalization, qualification, and deduplication. The resulting records show high manual field-level accuracy and advantages over Pistachio in matched comparisons.
- Data sources and construction scope: 1,580,939 patent documents entered extraction, with 608,309 linking to at least one retained Reaction Instance.
- From patent documents to structured Reaction Instances: The pipeline parses patent text, images, and reaction schemes, normalizes experimental information, links records to source locations, and applies automated qualification.
- Database scale and counting conventions: Approximately 24 million Reaction Instances were retained, including approximately 14.8 million qualified instances that passed checks based on atom mapping, atom conservation, and other rules.
- Manual field-level quality evaluation: 92.95% micro-averaged field-level accuracy was obtained across 6,500 judgments from 1,300 sampled qualified instances.
- Manual field-level quality evaluation: Catalyst accuracy was 97.31%, while reagent accuracy was 84.62%, reflecting context-dependent boundaries between reagent and other participant roles.
- Quality limitations: Workup and purification information, participant-role assignment, cross-paragraph extraction, and multi-step descriptions remain directions for data-quality improvement.
- Matched comparison with Pistachio: After deduplication in 100 matched US patents, DianShi-RxnDB retained 4,093 records versus Pistachio’s 2,992, a ratio of 1.368.
- Matched comparison with Pistachio: The matched comparison indicated larger deduplicated counts, finer representation granularity, and higher source-grounded exact agreement across six evaluated fields.
3 Web research workbench
The Web research workbench supports object-specific search, linked exploration, comparison of single-step reaction records, and direct source-patent verification.
- Three search entry points cover substances, reactions, and patent References, with query conditions and results organized by object type.
- Reaction search distinguishes specific Reaction Instances from Reaction Groups that aggregate records sharing a reactant–product combination without merging their experimental information.
- Researchers can browse from substances or References to related Reaction Instances, then inspect participants, groups, and source References through linked object relationships.
- Within a Reaction Group, users compare separately recorded participant roles, conditions, yields, procedures, and source information across multiple instances.
- Each Reaction Instance links to its source Reference and relevant patent text region, enabling side-by-side structured-record and source-context verification.
4 MCP service for AI agents
The MCP service exposes DianShi-RxnDB through composable structured retrieval tools for substances, reactions, instances, and patent References, supporting multi-step searches by AI agents.
- The MCP service provides structured retrieval for Substance records, Reaction Groups, Reaction Instances, and source References.
- Substance retrieval supports name, identifier, chemical-representation, similarity, and substructure queries.
- Reaction retrieval returns groups by representation or target product and provides associated instances with participant roles, conditions, yields, and source information.
- MCP responses expose identifiers and chemical representations that can feed subsequent calls, forming multi-step paths from candidate retrieval to source lookup.
- Target-product retrieval may return multiple reactant combinations, so agents must identify the target Reaction Group before obtaining its associated instances.
- MCP preserves links to patent References, while the Web workbench supports source-text positioning and contextual verification.
5 Reaction-precedent retrieval case study with Web and MCP
The aspirin case study demonstrates complementary Web and MCP workflows for retrieving, filtering, comparing, and source-verifying reaction precedents. Both workflows preserve structured experimental information and provenance links.
- The case study uses aspirin to demonstrate direct Web retrieval and AI-agent MCP retrieval over the same target reaction and database objects.
- 5.1 Retrieving reaction instances through the Web workbench: The Web workbench returns 84 aspirin-product Reaction Instances, after which the researcher selects the salicylic acid–acetic anhydride acetylation.
- 5.1 Retrieving reaction instances through the Web workbench: 53 Reaction Instances belong to the selected aspirin Reaction Group, allowing comparison of single-step records from different source patents.
- 5.1 Retrieving reaction instances through the Web workbench: Source verification confirmed Instance A’s catalyst, temperature, and time, while the patent also reported cooling, precipitation, filtration, washing, and vacuum drying.
- 5.2 Retrieving reaction instances through MCP: The MCP path retrieved 11 candidate groups, identified 1 target group, confirmed 53 instances, and returned 20 instances for the condition query.
- 5.2 Retrieving reaction instances through MCP: The agent organized experimental conditions, yields, source patents, and database identifiers while explicitly marking fields not returned by the tools as missing.
- MCP supports structured multi-step retrieval, whereas the Web workbench supports comparing records, viewing source context, and verifying key information.
6 Limitations, availability, and responsible use
The platform’s scope is bounded by its patent corpus and update rules, while extracted records and quality estimates require source verification and professional review for consequential uses.
- Data and use boundaries: The corpus primarily covers USPTO and EPO organic synthesis patents published from 1976 through 2025, excluding journal articles and unpublished laboratory knowledge.
- Data and use boundaries: Manual quality estimates apply only to qualified instances and five evaluated fields, not to failed instances, unevaluated fields, individual-record correctness, or database recall.
- Data and use boundaries: Automated extraction can still produce omissions, duplicate records, role errors, and step-boundary errors, so critical records should be checked against source patents.
- Responsible use: Provenance linkage enables verification but does not replace qualified professionals’ judgment from complete context.
- Responsible use: Experimental safety, clinical or regulatory matters, patent assessment, and production-condition selection require further review by appropriately qualified professionals.
- Availability: Web and MCP access is free for non-commercial use, while commercial use requires a separate license.
7 Conclusion and outlook
DianShi-RxnDB organizes large-scale, fine-grained reaction data through automated extraction and provides provenance-linked interfaces for researchers and AI agents. The authors report strong sampled-field accuracy, favorable matched comparisons with Pistachio, and plans to expand sources, quality, retrieval, and applications.
- Approximately 24 million Reaction Instances were constructed, with approximately 14.8 million (61.7%) passing automated qualification assessment.The platform covers organic synthesis patents from the USPTO and EPO and records experimental information and structured relationships among database objects.
- Manual evaluation of 1,300 qualified records achieved 92.95% micro-averaged accuracy across five fields.
- The matched Pistachio comparison found more deduplicated records, finer-grained information organization, and higher field-level exact agreement across six evaluated fields.
- The Web workbench and MCP service support researcher retrieval and source inspection alongside structured, multi-step retrieval by AI agents.Both interfaces operate on the same provenance-linked reaction-data foundation.
- Future work will broaden data sources, improve quality and retrieval, strengthen source verification, and explore additional organic-chemistry applications.A subsequent technical report will examine retrosynthesis applications and initial findings.
A Data snapshot, object definitions, and counting conventions
The appendix defines the database snapshot, object-counting conventions, core objects, and operational fields used to represent reaction and process information. It distinguishes processing-stage counts and illustrates structured experimental and provenance records.
- Data snapshot and counting conventions: Database-scale statistics use a snapshot dated 2026-06-25, and counts for References, Substances, Reaction Instances, Reaction Groups, and Reaction Templates are different object types.These object counts cannot be added into a single reaction count.
- Patent-record processing stages and counting conventions: 20,256,438 source patent records formed the pre-deduplication acquisition pool.
- Patent-record processing stages and counting conventions: 1,580,939 patent documents entered the reaction-information extraction pipeline after domain filtering, consolidation, and full-text availability checks.
- Patent-record processing stages and counting conventions: 608,309 References were source patent documents linked to at least one retained Reaction Instance in the database snapshot.The three reported values correspond to different processing stages and are not interchangeable.
- Reaction-process details: Reaction-process details encode stepwise operations and action descriptions, with the current reference definition containing 38 operation types.Examples include adding, stirring, filtering, concentrating, and purifying materials.
C.1 Evaluation population and scope
The quality evaluation estimates field-level correctness only for five fields in the qualified-instance population. Its pooled metric is based on manually adjudicated field judgments rather than record-level accuracy or database recall.
- Evaluation population and scope: 14,808,205 qualified Reaction Instances defined the evaluation population, from which 1,300 records were randomly sampled.
- Evaluation population and scope: The evaluation covered Yield, Reactant, Reagent, Catalyst, and Solvent, producing 6,500 field-level judgments.
- Evaluation procedure: Each record was independently assessed by two annotators, with disagreement-triggered third annotation and majority-vote consolidation.The reported statistics used the final labels from this procedure.
- Statistical definitions: 92.95% was the five-field micro-averaged accuracy, pooling correct field judgments over all field-level denominators.For each field, accuracy is the number of correct judgments divided by the evaluated-record denominator.
- Evaluation scope: The reported accuracy is a pooled field-level measure, not record-level accuracy, entity-matching precision, recall, F1, or database recall.The estimates do not apply to non-qualified instances or unevaluated fields.
D External comparison with the Pistachio Reaction Dataset
The external comparison separates record scale, representation richness, and field-level agreement, using aligned patent samples and within-patent deduplication. A representative text–image overlap illustrates why channel-aware normalization is necessary.
- Comparison design: The comparison evaluated record count, representational richness, and field-level agreement separately rather than collapsing them into one score.Record-level Pistachio analyses used the 2025Q2 release available to the authors.
- Comparison design: The sampling frame comprised US patents for which both datasets contained extraction results, with 100 patents randomly selected for scale analysis.
- Text–image channel overlap: Pistachio supplied separate text- and image-channel records that could describe the same underlying patent transformation.Channel-specific records could differ in identifiers, roles, names, or paragraph text.
- Text–image channel overlap: Counting both channels independently could inflate Pistachio totals by conflating reaction-record scale with overlapping representations.
- Representative overlap example: In the representative US07795244 example, text- and image-channel records represented one underlying reaction under the stated deduplication rule.Their reactant and product sections matched after normalization, while the middle agent section differed.
- Deduplication rule: Within-patent duplicates were identified when normalized reactant and product sets matched after atom-map removal, component normalization, and sorting.Component order and the middle agent section were ignored.
D.2.4 Matched-sample results
In the matched-sample comparison, DianShi-RxnDB retained more deduplicated records and exposed finer-grained information in an inspected shared-reaction example. The comparison used source-paragraph-matched pairs, harmonized roles and names, and field-level exact agreement criteria that depend on representation policies.
- Deduplicated record counts: 4,093 of 4,222 DianShi-RxnDB records remained after deduplication, versus 2,992 of 3,630 Pistachio records, a ratio of 1.368.The matched sample comprised 100 US patents; DianShi-RxnDB retained more records in 51 patents, Pistachio in 42, and 7 were tied.
- Representation granularity: The inspected shared-reaction example exposed finer-grained participant roles and independently retrievable process, workup, provenance, and validation dimensions in DianShi-RxnDB.Both records described oxidation of the same alcohol to the same ketone, but the record-level evidence does not generalize to all Pistachio records.
- Comparison design: 660 source-paragraph-matched reaction pairs from 58 patents were used for the field-level comparison, with identical normalized source paragraphs establishing shared passages rather than extraction correctness.The matching procedure removed markup, punctuation, whitespace, and line-break differences after lowercasing.
- Comparison design: The comparison covered Reactant, Reagent, Catalyst, Solvent, Product, and Yield after mapping both datasets to a shared role vocabulary and normalizing names and aliases.Workup-only entities, filtration aids, and atmospheric gases were excluded from reaction-participant references.
- Field-level agreement: Yield exact agreement reflects both extraction behavior and representation policy because DianShi-RxnDB records explicit percentages while Pistachio may infer values from mass and stoichiometry.Under the source-text exactness criterion, an inferred value was not counted as an explicit percentage.
D.5 Limitations
The matched-sample and field-level comparisons are bounded evaluations rather than extrapolations to the complete corpora. Their results are also sensitive to reference construction, role harmonization, ambiguity treatment, and Yield representation policy.
- Scope boundary: The deduplication results cover 100 US patents, while field-level results cover 660 matched reaction pairs from 58 patents, and neither was extrapolated to the complete corpora.The internal evaluation of 1,300 qualified DianShi-RxnDB records should not be pooled with this field-level comparison.
- Evaluation boundary: Field-level comparison results depend on source-grounded reference construction, role harmonization, ambiguity treatment, and Yield representation policy.These methodological choices constrain direct interpretation of the cross-dataset field-level results.