Source-linked AI summary

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang, Pengze Li, Encheng Su, Jun Yao, Jiabei Xiao, Yuqi Shi, Jielan Li, Hongxia Hao, Zhangyang Gao, Fang Wu, Ben Fei, Xiangyu Yue, Pan Tan, Bozitao Zhong, Jinouwen Zhang, Aoran Wang, Yan Lu, Jiaheng Liu, Xinzhu Ma, Liang Hong, Mingyue Zheng, Phil Torr, Bowen Zhou, Wanli Ouyang, Lei Bai

arXiv:2607.07708v1cs.CLcs.AIcs.CEcs.LG

TL;DR

Scientific AI often separates native structural representation from evidence-linked reasoning, limiting mechanistic interpretation across proteins, molecules, and crystals. SciReasoner unifies domain-native structural tokens with language reasoning and achieves state-of-the-art performance on 67 of 86 benchmarks while producing inspectable scientific inference traces.

  • Problem

    Scientific AI systems often compress biological, chemical, and crystallographic organization into text or separate structural representations from evidence-linked reasoning.

  • Method

    SciReasoner integrates domain-native tokens for coordinates, topologies, and crystallographic lattices with language instructions in one multimodal autoregressive model.

  • Results

    State-of-the-art performance was achieved on 67 of 86 benchmarks, including Cellular Component Fmax increasing from 0.42 to 0.55 for low-homology and orphan-like proteins.

  • Takeaways & Limitations

    SciReasoner links accurate prediction with inspectable scientific inference across proteins, small molecules, and inorganic crystals.

  • Takeaways & Limitations

    The authors are still collecting more human judgments.

Abstract

from arXiv · show

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through scientific principles and physical constraints, from stereochemistry and bonding to symmetry, energetics and periodic order. However, applying artificial intelligence to this process presents a joint challenge of representation and reasoning: models must preserve domain-native structural information while showing how specific evidence supports predictions under these constraints. Here we introduce SciReasoner, a multimodal scientific foundation model for native structural reasoning across proteins, small molecules and inorganic crystals. SciReasoner discretizes coordinates, topologies and periodic connectivities into a unified structure-aware vocabulary, treating structural tokens as addressable evidence units during reasoning. In homology-controlled Gene Ontology prediction, SciReasoner improves Cellular Component annotation for low-homology and orphan-like proteins, increasing $F_{\max}$ from 0.42 to 0.55. In chemistry, it raises single-step retrosynthesis accuracy from 0.63 to 0.72 while generating fragment-level disconnection and precursor-verification traces. In materials science, its representations separate elemental and compound phases and resolve high- and low-band-gap regimes. Across 86 benchmarks, SciReasoner achieves state-of-the-art performance on 67 tasks. Double-blind expert evaluation rates its reasoning traces as preferred or at least comparable to those of a frontier large language model in 98% of cases. By making structure an inspectable substrate for reasoning under scientific constraints, SciReasoner connects accurate prediction with interpretable scientific inference.

1 Introduction

SciReasoner introduces native structural reasoning to connect structure–property prediction with inspectable, evidence-linked inference across proteins, small molecules, and periodic crystals. Its unified structure-aware vocabulary addresses the limitations of text-centered scientific AI and supports strong performance in low-homology protein annotation and broad benchmark evaluation.

  • Mechanistic structure–property reasoning requires integrating local motifs, non-local contacts, chemical environments, conformational geometry, and long-range periodic order under scientific constraints.These heterogeneous cues support functions, reactions, and material properties across proteins, chemicals, and crystalline materials.
  • Text-centered scientific AI can compress structural organization into strings or descriptions, causing explanations to depend primarily on linguistic associations rather than directly addressable physical evidence.This motivates native structural reasoning as a foundation-model paradigm for structure–property analysis.
  • SciReasoner represents proteins, small molecules, and periodic crystals with a unified structure-aware vocabulary whose tokens serve as addressable evidence units rather than auxiliary language descriptors.This enables inspectable reasoning chains in which intermediate claims can be traced to explicit structural evidence.
  • SciReasoner improves Cellular Component Gene Ontology prediction for low-homology and orphan-like proteins, raising Fmax from 0.42 to 0.55.Its attention is enriched at contact-defined DNA-binding residues and localizes to protein–DNA interfaces, linking predictions to residues that physically mediate molecular interaction.
  • Across 86 benchmarks, SciReasoner achieves state-of-the-art performance on 67 tasks, supporting the generality of its structure-grounded modeling strategy.The benchmarks span proteins, DNA, RNA, small molecules, inorganic crystals, scientific question answering, property prediction, and generation tasks.

2 Results

SciReasoner improves protein-function, RNA, and virtual-screening metrics while extending to open-ended biomedical language tasks. Its protein gains are largest in low-homology settings, where structure-grounded traces integrate local motifs, environments, domain composition, and reference proteins.

  • Protein and RNA prediction: 0.88 subcellular-localization accuracy exceeds ESM2’s 0.84, while DeepFRI-GO reaches 0.59 averaged across three aspects versus SaProt’s 0.52.SciReasoner is also on par with or above the specialist for DNA-promoter and transcription-factor detection and substantially outperforms RNA-function specialists.
  • Cross-domain results: 0.59→0.86 Isoform R2 and 0.74→0.81 RNA–protein interaction MCC improve substantially; DUD-E enrichment at 5.0% rises 7.12→7.70 while matching the best AUC of 0.76.SciReasoner also reaches 0.85 BertScore on biomedical QA and 0.77 ROUGE-L on protein general-function description.
  • Gene Ontology prediction: SciReasoner’s largest CAFA-3 Gene Ontology gains occur in low-homology regimes, particularly for Cellular Component prediction.Figure 2 stratifies Molecular Function, Biological Process, and Cellular Component performance by maximum BLAST identity to training sequences.
  • Gene Ontology prediction: Structure-grounded reasoning integrates domain composition, localized motifs, structural environments, and reference proteins, preserving local functional cues when global sequence identity is weak.This design contrasts with BLAST’s local sequence similarity and ESM2’s evolutionary and sequence-context patterns without explicit structural grounding.
  • Retrosynthesis: The retrosynthesis evaluation uses USPTO-50K, giving the model a target SMILES and requiring the reactant SMILES set that produces it.The task is framed as recursive target disconnection into commercially available precursors, with interpretable traces intended for human verification and route reuse.

AUC 5.0% EF

SciReasoner demonstrates structure-grounded performance across retrosynthesis, materials prediction, and cross-domain evaluations. Its analyses indicate that explicit structural evidence improves predictions and supports chemically and scientifically interpretable reasoning.

  • Retrosynthesis: 0.72 retrosynthesis accuracy exceeded RSGPT by +0.09 points, while Opus-4.7 five-shot scored 0.48 on USPTO-50K.The comparison excluded pretraining reactions whose product SMILES appeared in the test set.
  • Retrosynthesis: SciReasoner recovered the gold reactant set in Top-3 for 5/5 products, versus 2/5 for both RSGPT and Opus-4.7.Its traces used sub-fragments that were mostly smaller than either reactant and whose union reconstructed the gold answer.
  • Materials prediction: SciReasoner was evaluated on ten materials sub-tasks spanning five databases, including stability, energies, band gaps, spin-orbit coupling, CO2 uptake, and pore geometry.The model was compared with CGCNN, LLM-Prop, and Opus-4.7 across inorganic crystals, semiconductors, and metal-organic frameworks.
  • Structure ablation: Removing structural information consistently weakened performance, whereas structural inputs improved prediction across materials, proteins, and small molecules.For GO molecular-function prediction, structure-aware attention concentrated around functional binding sites, unlike sequence-only reasoning.
  • Expert evaluation: In a double-blinded expert pilot, 177 reliable judgments assessed traces across protein annotation, crystalline-material prediction, and retrosynthesis using evidence, plausibility, alignment, coherence, and anti-hallucination criteria.Experts also provided five-point pairwise preferences between SciReasoner and DeepSeek-V4-Pro.

3 Discussion

SciReasoner treats structures as primary objects of inference for native structural reasoning across proteins, small molecules, and inorganic crystals. Human evaluation remains ongoing, leaving the judgment evidence incomplete.

  • Contribution: SciReasoner introduces native structural reasoning across proteins, small molecules, and inorganic crystals through a multimodal scientific foundation model.The model is presented as a unified approach spanning three scientific domains.
  • Rationale: The approach argues that structure–property relationships require structures to be primary inference objects rather than text strings, low-dimensional descriptors, or black-box predictor inputs.This premise motivates a unified structure-aware vocabulary for representing scientific structures.
  • Limitation: Human judgments of SciReasoner’s reasoning are still being collected, so the evaluation record remains incomplete.The discussion explicitly identifies ongoing collection of human judgments.

4 Method

SciReasoner combines domain-specific scientific corpora with a unified causal language-model architecture that represents protein, molecular, and crystal structures as discrete, addressable tokens. Its training uses staged autoregressive curriculum learning followed by task-aware cold-start supervision and reinforcement learning for native structural reasoning.

  • Data construction: SciReasoner integrates protein, small-molecule, materials, RNA, DNA, general web, and reasoning-instruction data into a broad multimodal pretraining corpus.Protein data link structures with UniProt, literature, and functional annotations; molecular data combine text, representations, properties, and conformations; materials data combine compositions, CIF structures, descriptions, and properties.
  • Architecture: SciReasoner uses modality-specific offline structural compressors, a structure-aware vocabulary embedding layer, and a Qwen3-14B-initialized unified causal language model.The architecture processes interleaved structural and textual modalities through a discrete cross-modal projection rather than relying exclusively on continuous structural encoders.
  • Structure-language integration: Discrete structural tokens preserve physical topologies that text subword tokenizers can fragment, while their embeddings are concatenated with native language embeddings to condition autoregressive generation.The projected structural and language representations form a unified prompt consumed by the language-model backbone.
  • Pretraining curriculum: Training minimizes a unified next-token prediction objective and uses a three-stage curriculum that progressively aligns structural tokens, modalities, and the full model.The first stage freezes the transformer backbone while training structural vocabulary and related parameters; later stages unfreeze all parameters and use a shared Warmup-Stable-Decay scheduler.
  • Post-training: Post-training converts the pretrained next-token continuator into a native structural reasoner through task-specific cold-start supervision followed by reinforcement learning.The reinforcement-learning framework calibrates reward magnitudes across tasks, using distance-based, matching-based, and tool-verified rewards for different scientific task groups.

Appendix A Detailed experimental results

Appendix A reports task-level benchmark results for SciReasoner across Chemistry, Material Science, and Biology, comparing it with four frontier general-purpose models. Results are organized by discipline and task type, with best and second-best outcomes highlighted.

  • Experimental results: SciReasoner is evaluated against Opus-4.7, GPT-5.5, Kimi-K2.6, and DeepSeek-V4-Pro across benchmark tasks in Chemistry, Material Science, and Biology.The tables include tasks with available model measurements and group results by Scientific QA, Property Prediction, Property Classification, or Generation and Design.

A.1 Task and metric descriptions

This section defines the tasks and evaluation metrics used in the result tables, clarifying expected model behavior and measurement.

  • A.1 Task and metric descriptions: Task descriptions follow the result-table organization and specify expected model behavior alongside the evaluation metric.

Chemistry tasks.

The chemistry evaluation spans text-based chemical information extraction, molecular property and virtual-screening prediction, biomedical classification, and reaction-planning tasks. These tasks use F1, RMSE, enrichment, MAE, accuracy, and exact-match metrics across complementary chemistry capabilities.

  • Text understanding: The chemistry suite evaluates chemical entity recognition and chemical–protein or chemical–disease interaction extraction from scientific and biomedical text using F1.These tasks require recovering chemical mentions and identifying paired entities with asserted interactions.
  • Molecular property prediction: Molecular modeling tasks assess aqueous solubility and lipophilicity with RMSE, physicochemical properties with MAE, and DUD-E virtual-screening enrichment at 5.0%.The tasks predict continuous molecular properties or rank compounds for early enrichment.
  • Biomedical classification: Biomedical molecular classification covers blood-brain barrier permeability, clinical toxicity, HIV activity, and side-effect associations, evaluated with accuracy.The suite includes BBBP, ClinTox, HIV, and SIDER prediction tasks.
  • Reaction planning: Reaction-planning tasks include forward synthesis, forward reaction, reagent, and retrosynthesis prediction, all evaluated by exact match.The tasks predict products, reaction outcomes, required reagents, or precursor reactants.

Material science tasks.

The material-science evaluation covers heterogeneous property prediction, discrete attribute classification, and constrained composition generation across several benchmark datasets. Regression uses normalized MAD MAE, while classification uses AUC and generation is assessed for chemical validity or plausibility.

  • Evaluation metrics: Database-level heterogeneous regression reports normalized MAD MAE, where larger values indicate lower error relative to target dispersion.
  • Property prediction: Material property prediction spans continuous targets from Materials Project, SNUMAT, JARVIS-DFT, and JARVIS-QETB datasets.Targets include band gap, density, volume, formation energy, stability-related, spin-orbit-related, structural, electronic, elastic, dielectric, and thermodynamic properties.
  • Classification: The evaluation classifies discrete Materials Project and SNUMAT attributes, including direct-gap status, indirect band-gap status, and thermodynamic stability, using AUC.
  • Composition generation: Composition-generation tasks use SMACT to generate chemically valid materials under elemental constraints or target bulk-modulus conditions.The tasks evaluate chemical validity for unconstrained elemental compositions and chemical plausibility when conditioning on a target bulk modulus.

Biology tasks.

The biology tasks evaluate SciReasoner across protein function, phenotype prediction, sequence interactions, gene-expression annotation, and function-guided design. Metrics include text-generation overlap, rank correlation, classification scores, and regression performance.

  • Function and design: Biological function tasks generate protein-function descriptions, broader functional annotations, and enzyme-catalyzed reaction descriptions against reference text.Function and general-function tasks use ROUGE-L; catalytic activity also evaluates reaction descriptions with ROUGE-L.
  • Protein properties: Protein-property tasks predict mutant fluorescence, stability, solubility, and function-guided sequence similarity using Spearman correlation, accuracy, and normalized sequence similarity.Fluorescence and stability use Spearman correlation, solubility uses accuracy, and function-guided protein design uses normalized sequence similarity.
  • Interactions and regulation: Molecular and genomic interaction tasks assess antibody–antigen and RNA–protein interactions, enhancer activity, isoform usage, and mean ribosome loading.The supplied tasks use MCC for interaction prediction, HK-PCC for enhancer activity, and R2 for isoform usage and ribosome loading.
  • Biological annotation: Gene- and tissue-level annotation tasks map gene symbols or names to tissue-expression and cancer-type labels.gSymbol2Tissue, gName2Cancer, and gSymbol2Cancer are evaluated with F1.

Metric definitions.

The section defines classification, ranking, multilabel annotation, imbalanced-class, and regression metrics, with arrows indicating whether higher or lower values are preferred.

  • Metric definitions.: ACC (↑) is the fraction of samples whose predicted label exactly matches the reference label.
  • Metric definitions.: AUC (↑) measures ROC area, F1 (↑) balances precision and recall, and Fmax (↑) is the maximum F1 across candidate thresholds for multilabel annotation.
  • Metric definitions.: MCC (↑) remains informative for imbalanced binary classification, while RMSE (↓) is the root mean squared error for regression.

A.2 Detailed results

Across its full benchmark suite, SciReasoner leads on 67 of 86 tasks, with detailed comparisons organized by discipline and task type against specialist and frontier general-purpose models.

  • Overall results: 67 of 86 tasks favor SciReasoner across the full benchmark suite.The comparisons include specialist baselines and frontier general-purpose models.
  • Cross-discipline comparisons: The appendix provides complete task-level comparisons with frontier general-purpose models across Chemistry, Material Science, and Biology.Results are grouped into Scientific QA, Property Prediction, Property Classification, and Generation and Design.
  • Specialist baselines: Table A1 reports per-task comparisons between SciReasoner and specialist baselines, marking best and second-best performance.Bold denotes the best performance, while underline denotes the second best.

Appendix B Human Evaluation Form · B.1 General Evaluation Instructions · B.2 Blank Scoring Sheet Used for Each Sample

Appendix B defines a double-blind evaluation of anonymized reasoning traces and outputs across crystal-material prediction, Gene Ontology prediction, and retrosynthesis. Evaluators score five trace-quality axes, overall comparisons with expert expectation, direct model preference, and confidence using a standardized scoring sheet.

  • Appendix B Human Evaluation Form: The evaluation samples one item from each task category and presents evaluators with the input, two anonymized traces and outputs, and a read-only ground-truth fact sheet.The categories are crystal-material property prediction, Gene Ontology prediction, and retrosynthesis.
  • B.1 General Evaluation Instructions: Evaluators judge reasoning quality rather than final-answer closeness, emphasizing evidence grounding, domain plausibility, target alignment, coherence, and hallucination risk.These are the five primary trace-quality axes.
  • B.1 General Evaluation Instructions: Each axis is scored from 1 to 10 or N.A., with verdicts distinguishing correct, minor, major, and critical defects according to their severity and evidentiary basis.N.A. applies when an axis has no checkable claim.
  • B.1 General Evaluation Instructions: Task-specific grounding checks cover crystal formulas and periodic structure, protein sequences and 3Di features, and chemical products, atom maps, connectivity, and cited facts.The evaluation verifies that referenced entities, tokens, positions, indices, groups, and facts actually appear in the input.
  • B.1 General Evaluation Instructions: Target alignment evaluates the correct materials regime, GO region, or retrosynthetic reaction class and formed bond, including boundary, neighboring-family, and no-commit cases.Chemical scoring also considers gold-route plausibility, atom-map balance, oxidation or protection state, and feasibility.
  • B.1 General Evaluation Instructions: Coherence requires an evidence-to-conclusion chain, while hallucination scoring penalizes fabricated or unsupported biological, materials, and chemical details, mechanisms, entities, or citations.The expected chains run from structural or sequence evidence to the committed prediction or from product parsing to retrosynthetic proposals.
  • B.2 Blank Scoring Sheet Used for Each Sample: The blank scoring sheet records Model A and B scores, evidence and claim notes, expert-expectation ratings, direct preference, and evaluator confidence from 1 to 10.Direct comparison options range from A much better to B much better, including a tie; expert ratings range from significantly falls short to significantly exceeds.

B.3 Materials: Ag2HgI4, shear modulus · B.4 Gene Ontology: 1bd8 A-P55273, biological process

The evaluated traces connect native structural evidence to materials-property and Gene Ontology predictions, with Model A closer to the Ag2HgI4 ground truth and substantially stronger biological-process performance than Model B. Materials reasoning links coordination and heavy iodide chemistry to softness, while GO reasoning distinguishes committed terms aligned with the ground-truth regulatory region from transcription-centered alternatives.

  • B.3 Materials: Ag2HgI4, shear modulus: 5.62 GPa is closer to the Ag2HgI4 shear-modulus ground truth of 5.77 GPa than Model B’s 8.00 GPa.Both predictions fall in the soft regime, but Model A’s committed value is described as close to ground truth.
  • Input prompt.: The Ag2HgI4 task supplies a chemical formula and structure encoding and requests a precise JSON property prediction.The target is shear modulus, described as resistance to shear deformation related to directional bonding, framework rigidity, and elastic anisotropy.
  • B.3 Materials: Ag2HgI4, shear modulus: The materials trace parses Ag, Hg, and I atoms, lists metal–iodine edges, and invokes tetrahedral coordination to support the shear-modulus prediction.The reasoning treats heavy, polarizable iodide ions as implying a compliant lattice with low shear stiffness.
  • Example claim prompts shown to the evaluator.: Both evaluation traces require checking whether committed conclusions follow from decoded structural or sequence evidence without unsupported identity, specificity, or fabrication claims.For the protein, bZIP, WRKY, leucine-zipper, and transcription-factor claims are explicitly treated as potentially real-but-misassigned or fabricated; materials evaluators similarly inspect named phases and literature-like statements.
  • B.4 Gene Ontology: 1bd8 A-P55273, biological process: Model A achieves F1 = 0.967, precision = 0.993, and recall = 0.943 for 1bd8 A-P55273 biological-process prediction, versus Model B’s F1 = 0.209, precision = 0.314, and recall = 0.156.Model A predicts 139 BP terms and Model B predicts 70, against 145 true BP terms.
  • Model outputs shown in the questionnaire.: Model B’s 70-term GO prediction centers on transcription and gene expression, outside the ground-truth biological-process region centered on cell-cycle, apoptosis, DNA-damage, and stress responses.The ground truth contains 145 terms, including negative regulation of cell cycle, G1/S transition, CDK regulation, and DNA-damage response and repair.
  • B.4 Gene Ontology: 1bd8 A-P55273, biological process: Model A’s GO trace uses sequence, 3Di runs, loop-like segments, and basic clusters such as RRLLHRE to infer regulatory and stress-response biology.Its committed terms largely overlap the ground-truth CDK-inhibitor biological-process region, including cell-cycle, apoptosis, kinase-regulation, and stress-response terms.

B.5 Retrosynthesis: USPTO-50K sample 4, other

For USPTO-50K sample 4 in the Other reaction class, Model A matches the gold reactants, while Model B proposes a related but non-gold set. The traces show amide-bond disconnection and reactant selection, with Model B using a less activated acyl source than the gold route.

  • Retrosynthesis: Model A matches the gold reactants and the gold amide-forming reaction family, whereas Model B proposes a related but non-gold reactant set.Both models identify the same C–N disconnection, but only Model A matches the gold formed bond and reaction family.
  • Retrosynthesis: The trace parses a trifluoroacetamide linked to an ortho-substituted aryl sulfone and cyclopropyl group before disconnecting the C1–N7 amide bond.It identifies the trifluoroacetyl group, amide N, benzyl group, sulfone, and cyclopropyl group as structural components.
  • Retrosynthesis: Model A proposes TFAA plus the primary amine, while Model B uses trifluoroacetic acid plus the amine as a less activated acyl source.The traces follow product parsing to amide disconnection and then reactant selection; evaluators check for unsupported reaction claims and invented reagents.
Loading 2607.07708v1…