Source-linked AI summary

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

arXiv:2608.13558v1cs.AIcs.CL

TL;DR

기존 AI scientist 시스템은 과학적 증거를 텍스트나 scalar summary로 축약하는 경우가 많아 공간적·시간적·절차적 관계를 잃는다. OmniScientist는 아이디어 구상, 실험, 집필 전 과정에서 autonomous agent와 직접적인 multimodal perception을 사용하며, 평가한 36개 사례 모두를 평균 논문 점수 6.3으로 완성한다.

  • 문제

    증거를 텍스트나 scalar summary로 축약하는 AI scientist 시스템은 과학적으로 중요한 공간적·시간적·절차적 관계를 활용할 수 없게 만든다.

  • 방법

    OmniScientist는 아이디어 구상, 실험, 집필을 위한 autonomous agent와 perception layer를 결합하며, 코드로 강제되는 검사가 novelty, validity, provenance, claim traceability를 뒷받침한다.

  • 결과

    36개 사례 모두에서 compiled paper가 생성되었으며, 7개 차원 평가 기준에 따른 mean overall score는 6.3이었다.

  • 시사점 및 한계

    Direct multimodal perception은 여러 분야, evidence family, modality에 걸쳐 evidence-grounded scientific discovery를 지원한다.

Abstract

from arXiv · show

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

1 서론

OmniScientist는 워크플로 자동화와 과학적 증거에 포함된 이질적인 공간적·시간적·통계적·절차적 관계에 대한 접근 사이의 간극을 해소한다. 직접 지각을 아이디어 구상·실험·작성용 자율 agent와 결합하고, 광범위한 36개 사례 모음에서 end-to-end 연구를 완료한다.

  • 1 서론: 기존 멀티모달 벤치마크는 대개 관찰과 질문을 사전에 고정하는 반면, 과학 agent는 일반적으로 워크플로의 국소 단계에서만 지각을 사용한다.이로 인해 전체 워크플로에 걸쳐 연구 질문, 실험, 과학적 작성을 형성할 수 있는 증거에 대한 접근이 제한된다.
  • 1 서론: 이 시스템은 지각 레이어를 아이디어 구상·실험·작성용 자율 agent 3개와 결합하고, 관찰을 사용해 연구 과정을 이끈다.각 단계에서 ReAct 루프는 관찰·추론·행동을 교차 수행한다.
  • 1 서론: OmniScientist는 5개 학문 분야군, 4개 증거군, 다양한 과학적 modality를 아우르는 전체 36개 사례에서 원시 데이터부터 컴파일된 논문까지의 전 과정을 완료했다.사례 모음에는 이미지, 파형, 오디오, 비디오, 3-D 구조, 궤적, 표, 수식, 그래프가 포함되었다.
  • 1 서론: OmniScientist는 이질적인 과학적 증거에서 직접 작동하므로, 발견 과정 전반에서 공간적·시간적·통계적·절차적 관계를 활용할 수 있다.이러한 관계는 현미경 이미지, 스펙트럼, 파형, 오디오, 비디오, 3-D 구조, 분포, 궤적과 같은 modality에 담겨 있다.

2 관련 연구

기존 연구는 자율 과학 에이전트, 범용 멀티모달 인식, 에이전트 기반 추론을 결합하지만, 프롬프트로 매개되는 오케스트레이션에서는 단계 전환과 검증이 오류 가능성이 있는 모델 출력에 좌우된다. OmniScientist는 그 대신 코드로 제어되는 전환을 사용해 관찰을 근거로 삼고, 단계를 검증하며, 필요할 때 백트래킹한다.

  • 자율 과학적 발견: 과학 에이전트는 단일 분야·단일 기기를 사용하는 실험실 절차에서 후속 실험실 검증을 위한 가설을 제안하는 학제 간 시스템으로 발전했다.초기 사례로는 회절 유도 합성과 실험실 자동화가 있으며, 이후 시스템은 여러 분야에 걸친 실제 과학 데이터를 사용했다(Mitchener et al., 2025; Villaescusa-Navarro et al., 2025; Swanson et al., 2025; Gottweis et al., 202
  • 멀티모달 과학적 인식: 범용 멀티모달 모델은 강력한 image-text 및 vision-language 인식을 제공하며, 과학 벤치마크는 차트, 도표, 현미경 이미지, 재료 관찰을 평가한다.인용된 foundation method는 contrastive image-text 사전학습(Radford et al., 2021)과 instruction-following vision-language model(Alayrac et al., 2022; Liu et al., 2023)을 사용한다.
  • 에이전트 기반 추론: 에이전트 기반 추론은 도구 사용, 언어적 성찰, 멀티모달 관찰, 그리고 과학 작업을 에이전트 간에 분담하는 전문화된 역할을 통해 자율적 발견을 지원한다.이 문단은 단일 에이전트 추론과 도구 사용(Wei et al., 2022; Yao et al., 2023; Schick et al., 2023), 성찰(Li et al., 2026a; Shinn et al., 2023), 멀티모달 행동(Yang et al., 2023b), 공유 지식과 검토(Shao et al., 2025; Li et al., 2026b)를 식별한다.
  • OmniScientist의 위치: OmniScientist는 프롬프트와 메시지에 의존하는 오케스트레이션을, 실제 관찰에 대한 근거화를 검증하고 단계가 뒷받침되지 않을 때 백트래킹하는 코드로 제어되는 에이전트 파이프라인으로 대체한다.각 단계 내부의 추론은 여전히 개방형으로 유지되지만, 코드가 그 주변의 전환과 검증을 제어한다.

3 문제 설정

OmniScientist는 heterogeneous raw artifact를 대상으로 하는 evidence-grounded research task로 정식화되며, 주로 text와 numerical data를 처리하는 시스템의 한계를 다룬다. domain-agnostic pipeline은 단일 specification을 입력으로 받아 5개 discipline category와 36개 second-level case에 동일한 perception-to-writeup loop를 적용한다.

  • 동기: 현재 AI-scientist system은 주로 text와 numerical data를 처리하므로 perceptual and procedural evidence를 검토하지 못하고, 질문의 범위와 disciplinary reach를 좁힌다.Token serialization만으로는 local spatial, temporal, cross-channel, procedural relation의 보존을 보장할 수 없으며, textual caption은 이러한 구조를 잃을 수 있다.
  • Evidence 범위: OmniScientist는 image, symbolic structure, numerical result, experimental-process record와 같은 artifact를 아우르는 4 families of scientific evidence를 인식하도록 설계되었다.
  • Task 정식화: Task에는 dataset, scientific subject, target property, corresponding raw data를 기술한 하나의 specification file이 제공되며, methodology는 agent에 맡겨진다.필수 output은 evidence-grounded paper이며, 그 외에는 아무것도 제공되지 않는다.
  • 평가 범위: Demonstration suite는 5 top-level discipline categories와 36 second-level cases를 포괄하여 engine의 domain agnosticism을 평가할 수 있게 한다.
  • Domain 일반성: 새 discipline을 추가하려면 추가 specification file 하나만 필요하다. 변경되지 않은 engine이 domain-specific code 없이 동일한 perception, ideation, experimentation, write-up loop를 재사용하기 때문이다.이 loop는 seismogram, CAD mesh, knowledge graph 전반에서 작동하는 것으로 기술된다.

4 OmniScientist 프레임워크

OmniScientist는 raw evidence perception과 자율적인 ideation, experimentation, writeup agent를 결합한 end-to-end AI scientist다. 결정론적 pipeline은 code-enforced check를 사용해 hypothesis, experiment, result, manuscript claim을 실행 가능한 evidence에 근거시킨다.

  • Perception layer: Perception layer는 raw artifact를 evidence family와 modality별로 구성하고, native numeric analysis를 우선하며 spatial 또는 structural pattern이 중요할 때만 visual rendering을 호출한다.Visual inspection은 철저한 분석과 computational efficiency의 균형을 위해 예산이 제한된다.
  • Framework architecture: 이 framework는 perception layer와 ideation, experiment, writeup을 담당하는 three autonomous agents를 결합해 four evidence families를 twelve modalities에 걸쳐 처리한다.Raw evidence는 sequential pipeline으로 유입되며, ideation이 hypothesis를 수립하고 experimentation이 이를 검증하며 writeup이 manuscript를 편집한다.
  • Ideation: Ideation은 ReAct loop를 사용해 material을 목록화하고 observation을 조사하며 literature를 검색하고, code-enforced validation 전에 계산으로 답할 수 있고 novel하며 falsifiable한 question을 수립한다.Validation에는 research question, hypothesis, experiment sketch, falsification criterion, self-filtered candidate five 개, focused literature search at least three 회가 필요하다.
  • Experiment: Experimentation은 code를 반복적으로 생성하고 debug하며, perception을 사용해 input과 plot을 조사하고, baseline, ablation, mechanism probe, sensitivity sweep 같은 control을 포함한 at least four analyses를 요구한다.Exit check는 모든 attempted test에 대해 dataset access, execution provenance, figure correspondence, multiple-comparison correction을 검증한다.
  • Writeup: Writeup 단계는 venue와 discipline에 맞춰 manuscript structure와 length를 조정하고, OpenAlex를 통해 reference를 검색하며, experimental record에 비추어 claim을 audit하고, 최종 PDF를 compile한다.다섯 가지 structural specification은 machine learning, biomedical, chemistry paper의 section을 포함해 discipline-specific convention을 지원한다.

5 평가

OmniScientist는 36개 사례 suite 전반에서 end-to-end manuscript generation을 완료하며, backbone·discipline·evidence modality에 걸쳐 높고 전반적으로 일관된 품질을 달성한다. 직접적인 multimodal perception은 grounding과 significance를 향상시키며, prior-art search와 반복적 agentic reasoning은 전반적 품질에 핵심적으로 기여한다.

  • 5.1 End-to-end 평가: 6.3 overall score: Claude 기반 OmniScientist는 전체 36개 사례 suite에서 완성된 manuscript를 생성하며, 평가된 subset에서는 성능이 우수한 대체 backbone도 유사한 범위의 결과를 보인다.비교에서는 세 역할을 분리한 채 reasoning backbone만 교체하며, Table 4는 suite 전반의 dispatch, completion, mean composite score를 보고한다.
  • 5.1 End-to-end 평가: 6.1–7.1 median composite scores: discipline family와 evidence modality 전반에서 generation quality가 안정적으로 유지되며, 최고 점수 manuscript는 모든 domain category에 걸쳐 나타난다.보고된 범위는 discipline과 evidence modality 양쪽에 따른 grouping을 포괄하며, 평가된 사례 전반에서의 폭넓은 generalization을 뒷받침한다.
  • 5.3 Perception ablation: +2.8 multimodal grounding and +1.8 significance: 직접적인 perception은 text-only scalar-feature baseline 대비 가장 큰 향상을 내며, factual accuracy는 동일하게 높은 수준을 유지한다.grounding 향상은 raw observation을 직접 처리하는 데서 비롯되며, text-only baseline은 이에 접근할 수 없다.
  • 5.4 Component ablation: 6.9 to 5.7: prior-art search를 제거하면 leave-one-out 성능이 가장 크게 하락하며, iterative agentic loop를 한 pass로 줄여도 manuscript quality가 유의하게 저하된다.prior-art를 누락하면 중복되거나 이미 출판된 아이디어가 생성될 위험이 커지며, search와 iterative reasoning이 핵심 구성요소임을 보여준다.
  • 5.5 정성적 분석: 70%–87%: 동률을 제외하면, perception-driven system이 7개 평가 지표 모두에서 승리했으며, grounding·significance·novelty에서는 동률이 거의 없었지만 accuracy와 reproducibility에서는 대략 사분의 일에 달했다.paired qualitative analysis에 따르면, perception-driven research questions는 raw multimodal records에만 존재하는 속성에 의존했다.
  • 5.6 Robustness and scaling: 2.9-point gaps: factual accuracy와 soundness는 backbone scale에 가장 민감한 반면, multimodal grounding은 동일한 최대–최소 model 범위에서 1.2점만 변한다.결과는 더 강한 backbone이 일부 evaluation dimension을 다른 차원보다 더 크게 향상시키지만, scaling만으로는 perceptual competence를 제공하지 못함을 보여준다.

6 결론

OmniScientist는 연구 수명주기 전반에 multimodal perception을 통합하는 end-to-end, omni-modal, discipline-agnostic AI scientist로서, raw observation이 연구를 이끌고 manuscript claim을 뒷받침하도록 한다.

  • OmniScientist는 multimodal perception을 연구 수명주기에 직접 통합해 raw observation을 아이디어 도출, 실험 실행, manuscript claim과 연결한다.
  • 이 framework는 5개 discipline family를 아우르는 36-case demonstration suite를 성공적으로 완료해 학제 간 적용 가능성을 입증한다.
  • 이 system은 end-to-end, omni-modal, discipline-agnostic AI scientist로 제시된다.

A 추가 평가 세부사항

이 절에서는 계산 비용, backbone 성능, perception 향상, judge 검증을 다루는 보조 평가 분석을 제시한다. 또한 모든 case와 backbone에 일관되게 적용한 고정 review rubric을 명시한다.

  • 보조 분석: 보충 평가에서는 backbone별 논문당 비용, backbone-score heatmap, 차원별 perception 향상, judge-panel 검증 지표를 다룬다.이러한 분석은 Section 5.2의 주요 평가를 뒷받침한다.
  • 계산 비용: 평가한 모든 시스템에서 experiment stage가 전체 계산 비용의 대부분을 차지한다.Table 11은 backbone별 논문당 비용을 보고하며, open-weight backbone은 로컬에서 실행했다.
  • Perception 향상: Table 12는 5개의 blind pair와 1개의 vision-off pair에 대해 각 평가 차원별 perception-enabled minus blind-baseline panel-score 차이를 보고한다.점수는 원래의 0–10 scale을 사용한다. 양수 값은 visual perception을 사용했을 때 성능이 향상되었음을 의미하며, Cardiology에서는 더 약한 vision-off run을 사용한다.
  • Judge 검증: Judge는 novelty, soundness, clarity, significance, reproducibility, multimodal grounding, factual accuracy를 강조하는 고정 rubric을 사용해 1–10으로 일곱 차원을 평가한다.각 judge는 모든 case와 backbone에 대해 동일한 rubric, manuscript source, figure caption, authors’ result ledger를 받으며, validation metric과 사전 설정된 acceptance threshold는 별도로 보고한다.
  • Backbone 성능: Backbone heatmap은 reasoning strength의 순서를 보여준다. 상단에는 더 강한 model이, 하단에는 더 작은 open-weight model이 배치되며, factual accuracy는 일관되게 가장 높다.더 어두운 cell은 더 높은 panel score를 나타내며, 평가한 model 전반에서 factual accuracy가 가장 어두운 열을 이룬다.

B 코드에서 강제되는 검사

OmniScientist는 각 단계의 finalize payload와 누적 loop state에 대해 idea, rigour, claim check를 결정론적 Python predicate로 적용한다. 거부된 결과는 새로운 observation이 되어 승인되거나 step budget이 소진될 때까지 계속 진행하도록 강제하며, 결과 선택이 지배적인 실패 모드임을 드러낸다.

  • B 코드에서 강제되는 검사: 각 check는 해당 단계의 finalize payload와 누적 loop state를 입력으로 받아, 모델의 self-assessment가 아니라 승인 여부 또는 이유 문자열을 반환하는 Python predicate다.claim check는 작성된 manuscript 자체를 평가하는 반면, 다른 predicate는 단계 종료 기록에 적용된다.
  • B 코드에서 강제되는 검사: 거부된 finalization은 실패 이유를 새로운 observation으로 추가하므로 agent는 계속 reasoning을 수행하며, 승인되거나 step budget이 소진될 때까지 단계가 종료될 수 없다.gated-stage algorithm은 실행에서 실제로 일어난 일에 predicate를 적용하고, check가 성공한 뒤에야 payload를 허용한다.
  • B 코드로 강제된 검사: 전체 거부의 삼분의 이는 산술이 아니라 결과 선택과 관련되었으며, 36회 실행에서 provenance 관련 거부가 한 건 발생하는 데 그쳐 노골적인 조작은 드물었다.나머지 거부의 대부분은 빈 판정, 설정되지 않은 주요 선택, 누락된 핵심 수치, 검정 횟수의 과소 집계 등 schema 규율과 관련되었다.
  • B 코드에서 강제되는 검사: 36회 primary run에서 115회의 finalize attempt가 거부되었고, 단 4회 실행만 되돌려 보내지지 않고 두 stage exit에 모두 도달했다.결과 선택이 지배적이었다. 26회 실행은 유의하지 않은 분석을 finding으로 제시하려 했고, 이를 낮은 수준의 결론으로 내리도록 강제되었다.

C 실행 수준 실행 통계

36개 primary run에서 ideation은 search와 perception에 집중된 반면 experiment는 execution 중심이었으며, rigour check는 부정적이거나 약한 결과를 더 중요한 결과로 승격하지 않고 보존했다.

  • 실행 구성: Ideation은 code를 실행하지 않고 literature search와 raw-item inspection에 대부분의 call을 사용한 반면, experiment는 run당 평균 31.8 run_python calls를 수행했고 raw evidence를 재검토한 경우는 드물었다.질문을 형성하는 observation은 대개 experiment 단계에 이르기 전에 이미 이루어졌다.
  • 실행 구성: 어떤 run도 24-step ideation 또는 50-step experiment budget을 소진하지 않았다.이 budget과 미소진 결과는 36개 primary run에 대해 보고되었다.
  • 실행 결과: 67개 분석은 실행된 뒤 타당성이 부족한 것으로 판단되어 원고에서 의도적으로 제외되었으며, 시스템이 계산한 전체 항목의 대략 오분의 일을 차지했다.제외된 분석은 실행 trace에 남아 있었지만 claim check에 의해 원고에서는 제외되었다.

D 인식 계층

인식 계층은 modality-native numeric reader와 visual reader를 제공해, 모든 modality를 이미지로 변환하지 않고도 agent가 이질적인 증거를 직접 검사하게 한다. 각 사례의 specification file에서 tool이 자동으로 해제되며, 36개의 primary run 전반에서 사용되었다.

  • 인식 아키텍처: Modality-native reader는 각 modality 고유의 표현으로 수치를 제공하고 visual reader와 함께 작동해, agent가 numeric analysis와 artifact inspection 중에서 선택할 수 있게 한다.Signals, audio, video, 3-D structures, trajectories에는 각각 native 및 visual reading capability가 제공된다.
  • 인식 아키텍처: Tool은 specification file에서 자동으로 해제되므로 perception suite에서 discipline별 등록이 필요하지 않다.Table 18은 36개의 primary run 전반에서 tool availability와 invocation count를 보고한다.
  • 인식 도구: 이 suite에는 signal, audio, 3-D, trace, trajectory, video analyzer가 포함되며, FFT frequency, RMS, PCA axis, path length, frame-difference motion과 같은 modality-specific measurement를 제공한다.나열된 tool은 보고된 run에서 각각 40, 39, 21, 20, 19, 3회 호출되었다.

E 새로운 discipline의 task specification

OmniScientist는 engine을 수정하지 않고 단일 specification file을 통해 새로운 discipline을 추가한다. 이 specification은 scientific role, data subject, measured property, open request, dataset members를 정의하며 method와 hypothesis는 agents에 맡긴다.

  • E 새로운 discipline의 task specification: 새로운 discipline에는 specification file 하나만 필요하며 engine 변경은 필요하지 않고, engine은 해당 discipline에 관해 그 밖의 어떤 내용도 읽지 않는다.seismology 예시에서 이 file에는 대략 1,500개의 dataset members가 포함되며, members list에는 file paths와 함께 제공된 metadata만 들어 있다.
  • E 새로운 discipline의 task specification: specification의 네 가지 scientific fields는 agent의 role, data subject, measured property, open research request다.이 file에는 method, hypothesis, analysis가 명시되지 않으며, 이러한 선택은 research workflow에 맡긴다.
  • E 새로운 discipline의 task specification: seismology 사례에서 agent는 STEAD catalogue에서 100 Hz로 sampling된 60초 3성분 broadband waveforms를 분석한다.각 record는 직교하는 E, N, Z components에서 시간에 따른 ground-motion amplitude를 측정하며, earthquake/noise labels와 함께 station 및 source metadata를 제공한다.
  • E 새로운 discipline의 task specification: seismology request는 system에 novel하고 testable한 question을 찾고, method를 선택하고, 실제 waveform code를 실행하며, 짧은 publishable paper를 작성하도록 요구한다.member records는 signal files, labels, station metadata, 그리고 해당하는 경우 earthquake magnitude, distance, depth를 제공한다.

F 분야별 작성 사양

작성 단계는 사례 사양 또는 추론된 주제로부터 분야별 구조 사양 5가지를 확정한다. 각 섹션에는 word budget, 문단 수, 문단 수준 개요가 할당되며, 시스템은 할당된 experiment-record slice만 사용해 이를 확장한다.

  • F 분야별 작성 사양: 5가지 구조 사양이 작성 단계의 분야별 섹션 골격과 abstract의 word range를 정의한다.Table 19는 5가지 사양을 제시하고 단일 문단 abstract의 word range를 제공한다.
  • F 분야별 작성 사양: Chemistry 작성물은 Introduction, Experimental Section, Results and Discussion, Conclusions를 구조 섹션으로 사용한다.제시된 Chemistry 구조의 word range는 제공된 사양에서 150–250이다.
  • F 분야별 작성 사양: 시스템은 사례 사양 또는 추론된 주제로부터 사양을 확정한 뒤, 각 섹션에 word budget, 문단 수, 문단 수준 개요를 할당한다.초안 작성 단계는 해당 섹션에 할당된 experiment-record slice만 사용해 한 번에 한 섹션씩 개요를 확장한다.

G Stage 프롬프트

아이디어 도출과 실험은 실행 시 조립되는 분야 불문형 프롬프트를 사용하며, 각 단계 종료 시 요구사항을 predicate로 강제한다. 프롬프트는 검토한 데이터와 제약된 주장에 아이디어를 근거시키고, 포괄적이고 정직하며 통계적으로 엄밀한 실험을 요구한다.

  • 프롬프트 설계: 실행 시 조립되는 프롬프트는 사용 사례 명세와 탐지된 modality를 사용하면서도 특정 분야, modality, method의 명칭에는 의존하지 않는다.아이디어 도출과 실험은 분야 전반에서 동일한 템플릿 구조를 공유하며, writeup 단계에는 이에 상응하는 단일 텍스트가 없다.
  • 프롬프트 설계: 모든 프롬프트 요구사항은 exit-check predicate이기도 하므로, 단계 진행은 모델이 자발적으로 따르기를 선택하는 데만 의존하지 않는다.필수 요구사항이 누락되면 exit check는 해당 단계의 통과를 거부한다.
  • 아이디어 도출 단계: 아이디어 도출은 대표 자료를 검토하고, 집중적인 문헌 검색을 수행하며, 관찰된 데이터에서 질문을 도출하고, novelty와 주장을 보수적으로 근거화할 것을 요구한다.근거화 규칙은 시각적 해석을 주어진 label과 조정하고, 뒷받침되지 않는 novelty 주장을 피하며, 직접 증거와 더 광범위한 mechanism을 구분하고, 실제 데이터에서 통계적 power를 추정하며, label leakage를 방지하도록 요구한다.
  • 실험 단계: 실험 프롬프트는 데이터가 뒷받침하는 경우 primary, baseline, ablation, mechanism, breakdown, sensitivity, grouping 분석을 포괄하는 full-study battery를 지정한다.grouping identifier가 존재하면 grouping은 random-split metric과 함께 group-disjoint 또는 leave-one-group-out 평가를 요구한다.
  • 실험 단계: 실험 단계는 조작된 결과를 금지하고, multiplicity correction, independence check, 정직한 power 표현, 그리고 가장 강하게 뒷받침되는 lead를 선택하는 anti-HARKing을 강제한다.데이터가 검정을 뒷받침할 수 없으면 프롬프트는 이를 실행 불가능한 것으로 보고하도록 요구하며, 약하거나 실패한 분석은 lead로 승격하지 않는다.
Loading 2608.13558v1…