Source-linked AI summary
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang, Pengze Li, Encheng Su, Jun Yao, Jiabei Xiao, Yuqi Shi, Jielan Li, Hongxia Hao, Zhangyang Gao, Fang Wu, Ben Fei, Xiangyu Yue, Pan Tan, Bozitao Zhong, Jinouwen Zhang, Aoran Wang, Yan Lu, Jiaheng Liu, Xinzhu Ma, Liang Hong, Mingyue Zheng, Phil Torr, Bowen Zhou, Wanli Ouyang, Lei Bai
TL;DR
Scientific AI는 단백질, 분자, 결정 전반에서 native structural representation과 evidence-linked reasoning을 분리하는 경우가 많아 기계론적 해석이 제한된다. SciReasoner는 domain-native structural token과 언어 추론을 통합하고, 86개 benchmark 중 67개에서 state-of-the-art 성능을 달성하는 동시에 검토 가능한 scientific inference trace를 생성한다.
문제
Scientific AI 시스템은 생물학적·화학적·결정학적 조직을 텍스트로 압축하거나, structural representation을 evidence-linked reasoning과 분리하는 경우가 많다.
방법
SciReasoner는 좌표, topology, 결정학적 lattice를 위한 domain-native token을 하나의 multimodal autoregressive model에서 언어 지시와 통합한다.
결과
86개 benchmark 중 67개에서 state-of-the-art 성능을 달성했으며, low-homology 및 orphan-like protein의 Cellular Component Fmax는 0.42에서 0.55로 증가했다.
시사점 및 한계
SciReasoner는 단백질, small molecule, inorganic crystal 전반에서 정확한 예측과 검토 가능한 scientific inference를 연결한다.
시사점 및 한계
저자들은 여전히 더 많은 인간 판단을 수집하고 있다.
Abstract
from arXiv · showhide
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through scientific principles and physical constraints, from stereochemistry and bonding to symmetry, energetics and periodic order. However, applying artificial intelligence to this process presents a joint challenge of representation and reasoning: models must preserve domain-native structural information while showing how specific evidence supports predictions under these constraints. Here we introduce SciReasoner, a multimodal scientific foundation model for native structural reasoning across proteins, small molecules and inorganic crystals. SciReasoner discretizes coordinates, topologies and periodic connectivities into a unified structure-aware vocabulary, treating structural tokens as addressable evidence units during reasoning. In homology-controlled Gene Ontology prediction, SciReasoner improves Cellular Component annotation for low-homology and orphan-like proteins, increasing $F_{\max}$ from 0.42 to 0.55. In chemistry, it raises single-step retrosynthesis accuracy from 0.63 to 0.72 while generating fragment-level disconnection and precursor-verification traces. In materials science, its representations separate elemental and compound phases and resolve high- and low-band-gap regimes. Across 86 benchmarks, SciReasoner achieves state-of-the-art performance on 67 tasks. Double-blind expert evaluation rates its reasoning traces as preferred or at least comparable to those of a frontier large language model in 98% of cases. By making structure an inspectable substrate for reasoning under scientific constraints, SciReasoner connects accurate prediction with interpretable scientific inference.
1 서론
SciReasoner는 단백질, small molecule, periodic crystal 전반에서 구조–물성 예측과 검증 가능한 증거 연계 추론을 연결하는 native structural reasoning을 도입한다. 통합된 structure-aware vocabulary로 text-centered scientific AI의 한계를 보완하고, 낮은 homology 단백질 annotation과 광범위한 benchmark 평가에서 높은 성능을 보인다.
- Mechanistic structure–property reasoning에는 scientific constraints 아래에서 local motif, non-local contact, chemical environment, conformational geometry, long-range periodic order를 통합하는 과정이 필요하다.이러한 이질적 단서는 단백질, 화학물질, crystalline material 전반에서 기능, 반응, material property를 뒷받침한다.
- Text-centered scientific AI는 구조적 조직을 string이나 description으로 압축할 수 있어, 설명이 직접 addressable한 physical evidence보다 언어적 연관성에 주로 의존하게 만든다.이는 structure–property analysis를 위한 foundation-model paradigm으로서 native structural reasoning의 필요성을 뒷받침한다.
- SciReasoner는 단백질, small molecule, periodic crystal을 unified structure-aware vocabulary로 표현하며, 이 vocabulary의 token은 보조적 language descriptor가 아니라 addressable evidence unit으로 기능한다.이를 통해 중간 주장을 명시적 structural evidence로 추적할 수 있는 inspectable reasoning chain이 가능해진다.
- SciReasoner는 low-homology 및 orphan-like protein의 Cellular Component Gene Ontology prediction을 개선하여 Fmax를 0.42에서 0.55로 높인다.이 모델의 attention은 contact로 정의된 DNA-binding residue에서 강화되고 protein–DNA interface에 국소화되어, 예측을 분자 상호작용을 물리적으로 매개하는 residue와 연결한다.
- 86개 benchmark에서 SciReasoner는 67개 task에서 state-of-the-art performance를 달성하여, structure-grounded modeling strategy의 일반성을 뒷받침한다.Benchmark에는 protein, DNA, RNA, small molecule, inorganic crystal, scientific question answering, property prediction, generation task가 포함된다.
2 결과
SciReasoner는 단백질 기능, RNA, virtual screening 지표를 향상시키는 동시에 개방형 biomedical language task로 확장된다. 단백질 성능 향상은 low-homology 조건에서 가장 크며, structure-grounded trace가 local motif, environment, domain composition, reference protein을 통합한다.
- Protein 및 RNA prediction: 0.88의 subcellular-localization accuracy는 ESM2의 0.84를 상회하며, DeepFRI-GO는 세 가지 aspect 평균 0.59로 SaProt의 0.52에 앞선다.SciReasoner는 DNA-promoter 및 transcription-factor detection에서도 specialist와 대등하거나 더 우수하며, RNA-function specialist를 크게 능가한다.
- Cross-domain results: Isoform R2는 0.59→0.86, RNA–protein interaction MCC는 0.74→0.81로 크게 향상된다. DUD-E enrichment at 5.0%는 7.12→7.70으로 상승하면서 최고 AUC 0.76을 유지한다.SciReasoner는 biomedical QA에서 0.85 BertScore, protein general-function description에서 0.77 ROUGE-L에도 도달한다.
- Gene Ontology prediction: SciReasoner의 가장 큰 CAFA-3 Gene Ontology 향상은 low-homology regime에서 나타나며, 특히 Cellular Component prediction에서 두드러진다.Figure 2는 training sequence에 대한 maximum BLAST identity에 따라 Molecular Function, Biological Process, Cellular Component 성능을 층화한다.
- Gene Ontology prediction: Structure-grounded reasoning은 domain composition, localized motif, structural environment, reference protein을 통합하여 global sequence identity가 약할 때에도 local functional cue를 보존한다.이 설계는 명시적인 structural grounding 없이 local sequence similarity에 의존하는 BLAST 및 evolutionary pattern과 sequence-context pattern을 사용하는 ESM2와 대조된다.
- Retrosynthesis: Retrosynthesis 평가는 USPTO-50K를 사용하며, 모델에 target SMILES를 제공하고 이를 생성하는 reactant SMILES set을 요구한다.이 task는 target을 commercially available precursor로 재귀적으로 disconnection하는 방식으로 정의되며, human verification과 route reuse를 위한 interpretable trace를 목표로 한다.
AUC 5.0% EF
SciReasoner는 retrosynthesis, materials prediction, cross-domain evaluation에서 구조에 근거한 성능을 보인다. 분석 결과, 명시적 구조 증거가 예측을 개선하고 화학적·과학적으로 해석 가능한 reasoning을 뒷받침하는 것으로 나타났다.
- Retrosynthesis: 0.72 retrosynthesis accuracy는 RSGPT보다 +0.09 points 높았으며, Opus-4.7 five-shot은 USPTO-50K에서 0.48을 기록했다.비교에서는 테스트 세트에 product SMILES가 등장한 pretraining reaction을 제외했다.
- Retrosynthesis: SciReasoner는 5/5개 product에서 Top-3 안에 gold reactant set을 복원했으며, RSGPT와 Opus-4.7은 각각 2/5였다.해당 trace는 대부분 어느 reactant보다 작은 sub-fragment를 사용했으며, 이들의 union으로 gold answer를 재구성했다.
- Materials prediction: SciReasoner는 다섯 개 database에 걸친 ten materials sub-tasks에서 평가되었으며, stability, energies, band gaps, spin-orbit coupling, CO2 uptake, pore geometry를 포함했다.모델은 inorganic crystal, semiconductor, metal-organic framework 전반에서 CGCNN, LLM-Prop, Opus-4.7과 비교되었다.
- Structure ablation: 구조 정보를 제거하면 성능이 일관되게 약화된 반면, structural input은 materials, proteins, small molecules 전반에서 prediction을 향상시켰다.GO molecular-function prediction에서 structure-aware attention은 sequence-only reasoning과 달리 functional binding site 주변에 집중되었다.
- Expert evaluation: Double-blinded expert pilot에서 177 reliable judgments가 evidence, plausibility, alignment, coherence, anti-hallucination criteria를 사용해 protein annotation, crystalline-material prediction, retrosynthesis 전반의 trace를 평가했다.전문가들은 SciReasoner와 DeepSeek-V4-Pro 사이의 five-point pairwise preference도 제시했다.
3 논의
SciReasoner는 단백질, 소분자, 무기 결정 전반의 네이티브 구조 추론을 위해 구조를 추론의 일차적 대상으로 다룬다. 인간 평가는 여전히 진행 중이므로 판단 근거는 불완전한 상태다.
- 기여: SciReasoner는 multimodal scientific foundation model을 통해 단백질, 소분자, 무기 결정 전반에 네이티브 구조 추론을 도입한다.이 모델은 세 과학 분야를 아우르는 통합 접근법으로 제시된다.
- 근거: 이 접근법은 structure–property relationship을 텍스트 문자열, 저차원 descriptor, 또는 black-box predictor의 입력이 아니라 구조 자체를 일차적 추론 대상으로 삼아야 한다고 주장한다.이 전제는 과학적 구조를 표현하기 위한 통합 structure-aware vocabulary를 도입하는 근거가 된다.
- 한계: SciReasoner의 추론에 대한 인간 판단은 여전히 수집 중이므로 평가 기록은 아직 불완전하다.논의에서는 인간 판단의 수집이 진행 중이라는 점을 명시적으로 한계로 지적한다.
4 방법
SciReasoner는 분야별 scientific corpora와 통합 causal language-model architecture를 결합해 protein, molecular, crystal structure를 discrete하고 addressable한 token으로 표현한다. 학습에는 staged autoregressive curriculum learning을 적용한 뒤, native structural reasoning을 위한 task-aware cold-start supervision과 reinforcement learning을 수행한다.
- Data construction: SciReasoner는 protein, small-molecule, materials, RNA, DNA, general web 및 reasoning-instruction data를 광범위한 multimodal pretraining corpus로 통합한다.Protein data는 structure를 UniProt, literature 및 functional annotation과 연결하고, molecular data는 text, representation, property 및 conformation을 결합하며, materials data는 composition, CIF structure, description 및 property를 결합한다.
- Architecture: SciReasoner는 modality-specific offline structural compressor, structure-aware vocabulary embedding layer 및 Qwen3-14B-initialized unified causal language model을 사용한다.Architecture는 continuous structural encoder에 전적으로 의존하지 않고, discrete cross-modal projection을 통해 interleaved structural 및 textual modality를 처리한다.
- Structure-language integration: Discrete structural token은 text subword tokenizer가 분할할 수 있는 physical topology를 보존하며, 해당 embedding은 native language embedding과 결합되어 autoregressive generation을 조건화한다.Projected structural 및 language representation은 language-model backbone이 처리하는 unified prompt를 형성한다.
- Pretraining curriculum: Training은 unified next-token prediction objective를 최소화하고, structural token, modality 및 전체 model을 점진적으로 정렬하는 three-stage curriculum을 사용한다.첫 번째 stage에서는 transformer backbone을 고정한 채 structural vocabulary와 관련 parameter를 학습하고, 이후 stage에서는 모든 parameter를 unfreeze하며 shared Warmup-Stable-Decay scheduler를 사용한다.
- Post-training: Post-training은 pretrained next-token continuator를 task-specific cold-start supervision에 이어 reinforcement learning을 수행하는 native structural reasoner로 전환한다.Reinforcement-learning framework는 task 간 reward magnitude를 보정하며, 서로 다른 scientific task group에 distance-based, matching-based 및 tool-verified reward를 사용한다.
부록 A 상세 실험 결과
부록 A는 Chemistry, Material Science, Biology 전반에서 SciReasoner의 task-level benchmark 결과를 제시하고, 이를 네 개의 frontier general-purpose model과 비교한다. 결과는 discipline과 task type별로 정리했으며, best와 second-best 결과를 강조 표시했다.
- 실험 결과: SciReasoner는 Chemistry, Material Science, Biology의 benchmark task 전반에서 Opus-4.7, GPT-5.5, Kimi-K2.6, DeepSeek-V4-Pro와 비교 평가된다.표에는 model measurement를 확인할 수 있는 task가 포함되며, 결과는 Scientific QA, Property Prediction, Property Classification, Generation and Design별로 묶었다.
A.1 과제 및 metric 설명
이 절에서는 결과 표에 사용된 과제와 평가 지표를 정의하고, 기대되는 모델 동작과 측정 방식을 명확히 한다.
- A.1 과제 및 metric 설명: 과제 설명은 결과 표의 구성을 따르며, 평가 지표와 함께 기대되는 모델 동작을 명시한다.
화학 과제
화학 평가는 텍스트 기반 화학 정보 추출, 분자 특성 및 virtual screening 예측, 생의학 분류, 반응 계획 과제를 아우른다. 이 과제들은 상호 보완적인 화학 역량 전반에서 F1, RMSE, enrichment, MAE, accuracy, exact-match metric을 사용한다.
- 텍스트 이해: 화학 과제군은 과학 및 생의학 텍스트에서 화학 개체 인식과 화학–단백질 또는 화학–질병 상호작용 추출을 F1로 평가한다.이 과제들은 화학물질 언급을 복원하고, 상호작용이 명시된 쌍을 이루는 개체를 식별해야 한다.
- 분자 특성 예측: 분자 모델링 과제는 RMSE를 사용해 수용해도와 lipophilicity를, MAE를 사용해 물리화학적 특성을 평가하며, DUD-E virtual-screening enrichment는 5.0%에서 평가한다.이 과제들은 연속적인 분자 특성을 예측하거나 초기 enrichment를 위해 화합물의 순위를 매긴다.
- 생의학 분류: 생의학 분자 분류는 혈액–뇌 장벽 투과성, 임상 독성, HIV 활성, 부작용 연관성을 다루며, accuracy로 평가한다.과제군에는 BBBP, ClinTox, HIV, SIDER 예측 과제가 포함된다.
- 반응 계획: 반응 계획 과제에는 순방향 합성, 순방향 반응, 시약, retrosynthesis 예측이 포함되며, 모두 exact match로 평가한다.이 과제들은 생성물, 반응 결과, 필요한 시약 또는 전구체 반응물을 예측한다.
재료과학 과제
재료과학 평가는 여러 benchmark dataset에 걸쳐 이질적 물성 예측, 이산 속성 분류, 제약 조성 생성을 다룬다. 회귀에는 normalized MAD MAE를 사용하고, 분류에는 AUC를 사용하며, 생성은 화학적 유효성 또는 개연성으로 평가한다.
- 평가 지표: Database 수준의 이질적 회귀는 normalized MAD MAE를 보고하며, 값이 클수록 목표 분산에 비해 오차가 낮음을 의미한다.
- 물성 예측: 재료 물성 예측은 Materials Project, SNUMAT, JARVIS-DFT, JARVIS-QETB dataset의 연속형 목표를 다룬다.목표에는 band gap, density, volume, formation energy, stability 관련, spin-orbit 관련, 구조적, 전자적, 탄성, 유전, 열역학적 물성이 포함된다.
- 분류: 평가는 AUC를 사용해 Materials Project와 SNUMAT의 이산 속성을 분류하며, direct-gap 여부, indirect band-gap 여부, 열역학적 안정성을 포함한다.
- 조성 생성: 조성 생성 과제는 SMACT를 사용해 원소 제약 또는 목표 bulk modulus 조건하에서 화학적으로 유효한 재료를 생성한다.과제는 제약 없는 원소 조성에 대해서는 화학적 유효성을, 목표 bulk modulus를 조건으로 할 때는 화학적 개연성을 평가한다.
생물학 태스크
생물학 태스크는 단백질 기능, 표현형 예측, 서열 상호작용, 유전자 발현 주석, 기능 기반 설계 전반에서 SciReasoner를 평가한다. 평가지표에는 텍스트 생성 중첩, 순위 상관, 분류 점수, 회귀 성능이 포함된다.
- 기능 및 설계: 생물학적 기능 태스크는 단백질 기능 설명, 더 넓은 범주의 기능 주석, 효소 촉매 반응 설명을 기준 텍스트에 맞춰 생성한다.기능 및 일반 기능 태스크는 ROUGE-L을 사용하며, 촉매 활성은 ROUGE-L로 반응 설명도 평가한다.
- 단백질 특성: 단백질 특성 태스크는 변이체 형광, 안정성, 용해도, 기능 기반 서열 유사도를 Spearman 상관, 정확도, 정규화 서열 유사도로 예측한다.형광과 안정성에는 Spearman 상관을 사용하고, 용해도에는 정확도를 사용하며, 기능 기반 단백질 설계에는 정규화 서열 유사도를 사용한다.
- 상호작용 및 조절: 분자 및 유전체 상호작용 태스크는 항체–항원 및 RNA–단백질 상호작용, 인핸서 활성, 아이소폼 사용, 평균 리보솜 로딩을 평가한다.제공된 태스크는 상호작용 예측에 MCC, 인핸서 활성에 HK-PCC, 아이소폼 사용과 리보솜 로딩에 R2를 사용한다.
- 생물학적 주석: 유전자 및 조직 수준 주석 태스크는 유전자 기호 또는 이름을 조직 발현 및 암 유형 레이블에 매핑한다.gSymbol2Tissue, gName2Cancer, gSymbol2Cancer는 F1로 평가한다.
평가지표 정의
이 절에서는 분류, 순위화, multilabel annotation, 불균형 클래스, 회귀 평가지표를 정의하며, 화살표는 값이 높거나 낮을수록 선호되는지를 나타낸다.
- 평가지표 정의: ACC (↑)는 예측 label이 reference label과 정확히 일치하는 샘플의 비율이다.
- 평가지표 정의: AUC (↑)는 ROC 면적을 측정하고, F1 (↑)은 precision과 recall의 균형을 나타내며, Fmax (↑)는 multilabel annotation에서 후보 threshold 전반의 최대 F1이다.
- 평가지표 정의: MCC (↑)는 불균형 이진 분류에서도 유용한 정보를 유지하며, RMSE (↓)는 회귀에서 root mean squared error다.
A.2 상세 결과
전체 benchmark suite에서 SciReasoner는 86개 task 중 67개에서 앞서며, specialist 및 frontier general-purpose model과의 상세 비교를 분야와 task 유형별로 정리한다.
- 종합 결과: 전체 benchmark suite에서 86개 task 중 67개가 SciReasoner의 우세를 보인다.비교에는 specialist baseline과 frontier general-purpose model이 포함된다.
- 분야 간 비교: 부록에서는 Chemistry, Material Science, Biology 전반에 걸쳐 frontier general-purpose model과의 완전한 task-level 비교를 제시한다.결과는 Scientific QA, Property Prediction, Property Classification, Generation and Design으로 묶어 제시한다.
- 전문가 기준선: Table A1은 SciReasoner와 전문가 기준선 간 task별 비교를 제시하고 최고 및 차상위 성능을 표시한다.굵은 글씨는 최고 성능을, 밑줄은 차상위 성능을 나타낸다.
부록 B Human Evaluation Form · B.1 General Evaluation Instructions · B.2 Blank Scoring Sheet Used for Each Sample
부록 B는 crystal-material prediction, Gene Ontology prediction, retrosynthesis 전반에서 익명화된 reasoning trace와 output을 double-blind로 평가하는 절차를 정의한다. 평가자는 다섯 가지 trace-quality 축, expert expectation과의 전반적 비교, 직접적인 model preference, 그리고 standardized scoring sheet를 사용한 confidence를 점수화한다.
- Appendix B Human Evaluation Form: 평가는 각 task category에서 하나의 item을 sampling하고, evaluator에게 input, 두 개의 anonymized trace와 output, 그리고 read-only ground-truth fact sheet를 제시한다.category는 crystal-material property prediction, Gene Ontology prediction, retrosynthesis다.
- B.1 General Evaluation Instructions: 평가자는 final-answer closeness가 아니라 reasoning quality를 판단하며, evidence grounding, domain plausibility, target alignment, coherence, hallucination risk를 중점적으로 본다.이는 다섯 가지 primary trace-quality axis다.
- B.1 General Evaluation Instructions: 각 axis는 1 to 10 or N.A.로 점수화하며, verdict는 심각도와 evidentiary basis에 따라 correct, minor, major, critical defect를 구분한다.N.A.는 검증 가능한 claim이 없는 axis에 적용한다.
- B.1 General Evaluation Instructions: Task-specific grounding checks는 crystal formula와 periodic structure, protein sequence와 3Di feature, chemical product, atom map, connectivity, cited fact를 대상으로 한다.평가는 참조된 entity, token, position, index, group, fact가 실제로 input에 나타나는지 검증한다.
- B.1 General Evaluation Instructions: Target alignment는 올바른 materials regime, GO region, 또는 retrosynthetic reaction class와 formed bond를 평가하며, boundary, neighboring-family, no-commit case도 포함한다.Chemical scoring에서는 gold-route plausibility, atom-map balance, oxidation 또는 protection state, feasibility도 고려한다.
- B.1 General Evaluation Instructions: Coherence는 evidence-to-conclusion chain을 요구하며, hallucination scoring은 날조되었거나 뒷받침되지 않은 biological, materials, chemical detail, mechanism, entity, citation에 불이익을 준다.expected chain은 structural 또는 sequence evidence에서 committed prediction으로, 또는 product parsing에서 retrosynthetic proposal로 이어진다.
- B.2 Blank Scoring Sheet Used for Each Sample: Blank scoring sheet에는 Model A와 B의 score, evidence 및 claim note, expert-expectation rating, direct preference, evaluator confidence from 1 to 10을 기록한다.Direct comparison option은 A much better부터 B much better까지이며 tie를 포함한다. Expert rating은 significantly falls short부터 significantly exceeds까지다.
B.3 Materials: Ag2HgI4, 전단 탄성률 · B.4 Gene Ontology: 1bd8 A-P55273, biological process
평가된 추론 trace는 native structural evidence를 materials-property 및 Gene Ontology 예측과 연결하며, Model A는 Ag2HgI4 ground truth에 더 가깝고 Model B보다 biological-process 성능이 substantially 우수하다. Materials reasoning은 coordination과 heavy iodide chemistry를 softness와 연결하고, GO reasoning은 ground-truth regulatory region과 부합하는 committed terms를 transcription-centered alternatives와 구분한다.
- B.3 Materials: Ag2HgI4, shear modulus: 5.62 GPa는 Ag2HgI4 shear-modulus ground truth인 5.77 GPa에 Model B의 8.00 GPa보다 더 가깝다.두 예측 모두 soft regime에 속하지만, Model A의 committed value는 ground truth에 가깝다고 기술된다.
- Input prompt.: Ag2HgI4 task는 chemical formula와 structure encoding을 제공하고 정밀한 JSON property prediction을 요구한다.Target은 shear modulus이며, 이는 directional bonding, framework rigidity, elastic anisotropy와 관련된 shear deformation에 대한 저항으로 설명된다.
- B.3 Materials: Ag2HgI4, shear modulus: Materials trace는 Ag, Hg, I 원자를 파싱하고 metal–iodine edge를 나열하며, shear-modulus prediction을 뒷받침하기 위해 tetrahedral coordination을 호출한다.이 reasoning은 heavy하고 polarizable한 iodide ion이 낮은 shear stiffness를 갖는 compliant lattice를 시사한다고 해석한다.
- Example claim prompts shown to the evaluator.: 두 evaluation trace 모두 committed conclusion이 decoded structural 또는 sequence evidence에서 뒷받침 없이 도출된 identity, specificity 또는 fabrication claim을 포함하지 않는지 확인해야 한다.Protein의 경우 bZIP, WRKY, leucine-zipper, transcription-factor claim은 실제일 수 있지만 잘못 할당되었거나 fabricated일 가능성이 있는 것으로 명시적으로 다뤄진다. Materials evaluator도 마찬가지로 명시된 phase와 literature-like statement를 점검한다.
- B.4 Gene Ontology: 1bd8 A-P55273, biological process: Model A는 1bd8 A-P55273 biological-process prediction에서 F1 = 0.967, precision = 0.993, recall = 0.943을 달성하는 반면, Model B는 F1 = 0.209, precision = 0.314, recall = 0.156을 기록한다.Model A는 139개의 BP term을, Model B는 70개를 예측하며, 실제 BP term은 145개다.
- Model outputs shown in the questionnaire.: Model B의 70-term GO prediction은 transcription과 gene expression을 중심으로 하며, cell-cycle, apoptosis, DNA-damage, stress response를 중심으로 하는 ground-truth biological-process region 밖에 있다.Ground truth에는 negative regulation of cell cycle, G1/S transition, CDK regulation, DNA-damage response and repair를 포함해 145개 term이 들어 있다.
- B.4 Gene Ontology: 1bd8 A-P55273, biological process: Model A의 GO trace는 sequence, 3Di runs, loop-like segments, 그리고 RRLLHRE와 같은 basic cluster를 사용해 regulatory 및 stress-response biology를 추론한다.Committed term은 대체로 ground-truth CDK-inhibitor biological-process region과 겹치며, cell-cycle, apoptosis, kinase-regulation, stress-response term을 포함한다.
B.5 역합성: USPTO-50K sample 4, 기타
Other reaction class의 USPTO-50K sample 4에서 Model A는 gold reactants와 일치하는 반면, Model B는 관련되지만 gold가 아닌 reactant set을 제안한다. 두 trace는 amide-bond disconnection과 reactant selection을 보여주며, Model B는 gold route보다 덜 활성화된 acyl source를 사용한다.
- 역합성: Model A는 gold reactants 및 gold amide-forming reaction family와 일치하지만, Model B는 관련되나 gold가 아닌 reactant set을 제안한다.두 모델 모두 동일한 C–N disconnection을 식별하지만, gold formed bond와 reaction family에 일치하는 것은 Model A뿐이다.
- 역합성: 이 trace는 ortho-substituted aryl sulfone 및 cyclopropyl group에 연결된 trifluoroacetamide를 파싱한 뒤 C1–N7 amide bond를 절단한다.trifluoroacetyl group, amide N, benzyl group, sulfone 및 cyclopropyl group을 구조적 구성요소로 식별한다.
- 역합성: Model A는 TFAA와 primary amine을 제안하는 반면, Model B는 amine과 함께 덜 활성화된 acyl source인 trifluoroacetic acid를 사용한다.trace는 product parsing에서 amide disconnection, 이어서 reactant selection으로 진행되며, 평가자는 뒷받침되지 않은 reaction claim과 지어낸 reagent가 있는지 확인한다.