Source-linked AI summary
Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models
Valentin Romanov, Monique Bax, Steven Niederer
TL;DR
Extracting nuanced data from research articles is laborious and difficult to evaluate reproducibly. This paper tests increasingly autonomous LLM workflows and finds strong extraction with fixed sources and expert prompts, but weaker literature discovery and a continuing need for human review.
Problem
Extracting nuanced data from unstructured research articles remains laborious, while reproducible evaluation against clear reference standards is difficult.
Method
The study evaluates four workflows that progressively delegate expert prompting, extraction, literature discovery, and dataset creation to frontier LLMs.
Results
Fixed article sets and expert prompts yielded over 90% recovery of ground-truth values and conditions, while deep-research agents recovered only 53.7% of primary references on average.
Takeaways & Limitations
An auditable workflow assigns evidence standards to experts, repeated cross-model extraction to LLMs, and contested or contextual cases to human review.
Takeaways & Limitations
Browser-based model behavior and availability changed during the study, limiting consistent access to some systems and comparisons over time.
Abstract
from arXiv · showhide
Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We demonstrate four escalating workflows, 1) given an expert curated prompt and research articles, most frontier LLMs perform well at data extraction, however can struggle with interpreting scientific context and nuance, 2) given simple instructions, LLMs can author their own prompts which were almost as eNective as expert-written prompts, 3) autonomous discovery of research literature was diNicult, agents either missed or hallucinated references, and 4) LLMs can create new datasets from published guidelines that closely match human-expert judges, but still require a human-in-the-loop. Together, these findings define an auditable division of labour in which experts specify the evidence standard, models cross-check repeated extractions and researchers resolve disputed cases, providing a practical route to scaling scientific data curation without relinquishing expert oversight.
Introduction
The introduction frames scientific data extraction as laborious and error-prone, while emphasizing that accurate, reproducible, and evaluable LLM-assisted extraction requires clear standards and human oversight. The study therefore evaluates four progressively autonomous browser-based strategies for extracting nuanced information from research articles.
- Motivation: Extracting data from unstructured research articles is time- and labour-intensive, and errors, miscalculations, and misunderstandings of nuance can pollute results.Researchers may spend months tracing the provenance of published parameters.
- Motivation: LLM-assisted extraction requires attention to the browser-based interface, where deployed products may combine hidden instructions, document processing, retrieval, memory, moderation, and provider-controlled tools.Browser sessions differ from isolated API calls because users cannot necessarily control system prompts, temperature, or other settings.
- Related work: LLM performance depends on both prompting and context, with prior studies reporting that prompt length, key-information position, and prompt structure can affect results inconsistently.Longer prompts improved some financial tasks but reduced performance in another study.
- Motivation: Automated literature extraction must be accurate, reproducible, and evaluated against a clear reference standard, but papers often omit key methods and reporting information.Preclinical studies reported fewer than half of recommended ARRIVE items; model consensus may flag disagreements for human review.
- Study aim: The study evaluates four strategies that progressively increase model autonomy in extracting nuanced information from unstructured research articles through the browser interface.The authors present tactics and techniques intended to improve LLM output across challenging data-extraction tasks.
Results & Discussion · Data extraction and curation using LLMs
The study examines LLM-assisted extraction of contextualized measurements from difficult scientific literature, using a cardiac troponin C case study and seven frontier models across four increasingly delegated workflows. It emphasizes that extraction quality depends on prompting, model and access choices, while manual curation still requires locating measurements, recording context, and reconciling reporting conventions.
- Data extraction and curation using LLMs: The initial case study reconstructs Niederer et al.’s review of Ca2+ binding affinity measurements for a cardiac-contraction model.The evidence base spans multiple experimental studies, including older sources with incomplete Methods and era-specific scientific language.
- Data extraction and curation using LLMs: Reconstructing the review is challenging because its older sources often provide incomplete Methods sections and nuanced, era-specific scientific descriptions.These limitations complicate interpretation of the experimental context behind reported measurements.
- Data extraction and curation using LLMs: LLM performance varies with prompt detail, model appropriateness, access mode, model tier, context limits, and non-deterministic responses.The same query can produce different answers, and free and paid models differ in ability and context limits.
- Data extraction and curation using LLMs: Seven frontier LLMs were tested: GPT-5.5, Opus 4.7, Gemini 3.1 Pro, Qwen3.7, Kimi K2.6, GLM-5.1, and DeepSeek v4.The study compares four strategies that progressively delegate more of the extraction task to the model.
- Data extraction and curation using LLMs: The first strategy gives every model the same human-expert-curated prompt and scores retrieval against a gold standard using the median or best of five.This establishes the expert-prompt baseline for the escalating workflow design.
- Data extraction and curation using LLMs: Traditional manual extraction requires locating measurements across article text, tables, and figures, recording experimental context, and reconciling units and reporting conventions.The workflow produces a structured dataset from the full article.
Expert curated prompt vs frontier-models
Frontier LLMs generally extracted numerical Ca2+ binding parameters accurately from research articles, but struggled with contextual interpretation and consistency across repeated runs. Cross-model and repeated-run checking exposed errors and supported more reliable estimates of model performance.
- Human oversight: Repeated extraction and comparison with ground truth identified errors in the original dataset, including the Kd-to-K conversion factor, temperature, and associated magnesium values.Using multiple LLMs across multiple runs supported double-checking extraction quality.
- Contextual accuracy: GPT-5.5 reached 99.5% affinity accuracy, while top models averaged 97.1 ± 2.6% and the remaining models averaged 93.0 ± 4.1%.Models were generally capable of finding or calculating an affinity value but were less reliable at assigning it to the correct preparation and experimental conditions.
- Run-to-run consistency: 76.4% of extracted parameters were correct in all five runs, 4.4% were wrong in all five, and 19.2% changed between correct and incorrect.Consistency remained an issue even among frontier models.
- Run-to-run consistency: The three leading models gave the same correct answer in all five runs for 87.7% of parameters, compared with 67.9% for the other four.A single run could misrepresent model capabilities; Kimi K2.6 scored 95.2% in one run and only 78.4–80.2% in the other four.
- Contextual accuracy: 99.0% and 99.8% of numerical K and Kd parameters were extracted correctly, respectively, versus 80.7% for identifying the correct troponin preparation.Identifying the preparation required deeper reasoning about biological context than finding or calculating the numerical values.
LLM self-authored prompts and the master prompt · Deep Research agents
LLM-generated master prompts nearly matched expert-designed prompts in extraction performance, although prompt-rubric similarity weakly predicted accuracy. Deep Research agents inconsistently recovered the underlying literature, sometimes hallucinated citations, and achieved limited data-extraction accuracy.
- LLM self-authored prompts and the master prompt: Each model searched for model-specific prompt-engineering practices and field terminology, generated three candidate extraction prompts, and combined them into one master prompt.The workflow specified the target data, output-table format, and justifications for extracted findings.
- LLM self-authored prompts and the master prompt: Aggregating prompts improved the 38-component rubric score for every model; GPT-5.5 rose from 82.9% across individual prompts to 88.2% for its master prompt.The rubric covered task definition, source prioritization, required columns and quotations, calculations, special cases, examples, and conflict resolution.
- LLM self-authored prompts and the master prompt: Prompt similarity and extraction accuracy were weakly and non-significantly associated across seven models (Spearman r=0.36, P=0.43).DeepSeek V4 ranked second for master-prompt score but fifth for extraction accuracy, whereas Opus 4.8 ranked fourth and second, respectively.
- LLM self-authored prompts and the master prompt: Affinity extraction was robust: GPT-5.5’s combined accuracy fell 1.1 pp, from 99.2% with the expert prompt to 98.1% with the LLM master prompt.Measurement method showed the largest mean accuracy reduction across models, at 27.5 pp.
- Deep Research agents: Agents ran five independent deep-research searches using the same LLM-authored prompt, seeking up to 20 primary research articles published before 2005.Opus 4.8 returned the most citations, averaging 20.6 ± 1.4 per run.
- Deep Research agents: Opus 4.8 recovered the largest share of the 19 reference articles, at 53.7% ± 14.7 percentage points, but still missed approximately half of the expert-curated references.GPT-5.5 recovered 38.9% ± 2.6 pp and Edison QA3 recovered 29.5% ± 2.6 pp.
- Deep Research agents: Qwen3.7-Plus hallucinated 45% of unique citations (27 of 60), and GLM-5.2 hallucinated 29% (20 of 69), while five other agents produced none.The five agents without hallucinated citations were GPT-5.5, Opus 4.8, Gemini 3.1 Pro, Mistral 3.5, and Edison QA3.
Creating new data sets using the techniques developed in prior sessions
The study combined self-prompting and repeated extraction to create reporting-practice datasets from 10 research articles, comparing seven frontier LLMs with trained human reviewers. Repeated-run consensus provided a reliability signal, while voting across different models achieved the strongest reviewer agreement.
- Dataset creation: Seven frontier LLMs and PhD-trained human reviewers assessed 10 brain-organoid articles against 23 minimum and ideal reporting criteria.Each model generated candidate prompts, aggregated them into a model-specific master prompt, and applied it in five independent runs.
- Human agreement: Human reviewers agreed on 79.1% of assessments and differed on 20.9%, reflecting uncertainty in borderline Yes/No classifications.The study therefore treated reviewer agreement as an imperfect benchmark for contextual reporting judgments.
- Reproducibility: GPT-5.5 was most reproducible at 92.6% ± 7.7 percentage points, followed by Kimi K2.6 at 87.8% and Opus 4.7 at 87.0%.Gemini 3.1 Pro was least reproducible at 81.7%.
- Consensus reliability: A unanimous 5–0 model answer matched reviewer consensus in 93.2% of cases, versus 65.0% for 4–1 splits and 55.2% for 3–2 splits.Agreement among reruns therefore indicated how likely the majority answer was to match reviewer consensus.
- Section-level disagreement: Reviewer agreement varied across guideline sections, reaching 86.7% for “Morphology and Lineage” but falling to 50.0% for “Functional criteria”.“Transcriptomics” had the highest proportion jointly classified as reported at 60.0%, while Metabolic criteria had the highest jointly classified as not reported at 66.0%.
- Cross-model consensus: Majority voting across different models reached 92.1% agreement with reviewers after three runs and 93.3% after five, compared with 89.2% for five same-model runs.The gains from different-model voting were 3.0 and 4.1 percentage points, respectively.
Conclusion
Reliable LLM-assisted literature curation combines expert prompting, curated sources, model consensus, and human validation. The resulting division of labour supports transparent, auditable dataset construction while reserving contested or contextual cases for expert review.
- Frontier models recovered more than 90% of ground-truth values and experimental conditions with fixed articles and an expert prompt, while numerical affinity constants approached 99% accuracy.The main difficulty was interpreting the nuance and context surrounding numerical data.
- LLMs generated effective master prompts from informed-user briefs, but these remained below expert-written prompt performance and required iterative user refinement.Prompt self-generation was a strong starting point rather than a replacement for expert prompting.
- 53.7% of expert-curated primary references were recovered on average by the strongest deep-research agent when locating and interpreting source literature.Greater autonomy therefore exposed a substantial constraint in autonomous literature discovery.
- Model consensus closely reproduced agreed human assessments, while disagreements between models or reviewers identified cases for expert review.Repeated and cross-model assessments provided quality control and helped surface uncertainty.
- The proposed framework assigns experts to define questions, sources, and evidence standards; models to perform repeated cross-system extractions; and humans to review contested, contextual, or high-impact cases.This division of labour is intended to produce high-quality scientific datasets while retaining human verification, feedback, and scientific judgement.
Methods … Generating initial LLM dataset
The initial LLM dataset was generated by running frontier models in web browsers on a corpus of 18 primary articles and 39 experimental conditions. Extraction included validation aids and predefined rules for missing, extra, and majority-of-five results, while browser-model behavior varied over time.
- Models: The study used GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, DeepSeek V4, GLM-5.1, Kimi K2.6, and Qwen3.7-Plus, with internet access enabled where possible.Deep Research analyses were identified by model-era because the served model was unclear.
- Prompts: Prompts for data extraction, LLM-authored prompts, and master prompts were provided in the Supporting Information.
- Browser access: Browser-model behavior changed during the study: Gemini 3.0 Pro searched and generated prompts effectively, whereas Gemini 3.1 Pro typically refused to search and hallucinated that it had searched.The researchers also stopped benchmarking MiniMax 3 after its platform reduced compute for free accounts.
- Generating initial LLM dataset: All models were accessed through their web browsers using five conversations per model, with one expert-prompted PDF added to each conversation.Outputs were collected for subsequent processing.
- Generating initial LLM dataset: The corpus comprised 18 accessible primary articles and 39 experimental conditions drawn from 19 articles and 40 conditions in Niederer et al. (2006).Reference data covered species, temperature, troponin preparation, Ca²⁺-measurement method, magnesium concentration, association constant Kd (M⁻¹), and dissociation constant K (µM).
- Generating initial LLM dataset: Each extracted parameter included a quote for rapid validation against the original article and a nuance column.
- Generating initial LLM dataset: Extraction was anchored on K and Kd values: failure to locate either made the row missing and incorrect for every field, while unassignable rows were extra.A majority-of-five result was positive when at least three of five runs were positive.
LLM self-authored prompts and master prompt · Deep Research Literature-Discovery Benchmark
The study evaluated LLM-authored prompts against an expert prompt and tested seven Deep Research agents on repeated literature-discovery runs. Prompts emphasized structured extraction with quotations and calculations, while reports underwent compliance and reference verification.
- LLM self-authored prompts and master prompt: Models researched provider-specific prompting guidance and scientific terminology relevant to cTnC affinity extraction.This preparation preceded generation of candidate extraction prompts.
- LLM self-authored prompts and master prompt: Each model generated three candidate prompts for extracting a structured table from one uploaded article.The prompts were intended to preserve supporting quotations and calculations.
- LLM self-authored prompts and master prompt: For Figures 3 and 5, models could search the internet using provider-specific research configurations.Configurations included web search, Deep think, Max, Advanced Search, DeepThink, and Search.
- LLM self-authored prompts and master prompt: LLM-designed master prompts were compared with a single expert-curated prompt across 38 components.Components covered task definition, source prioritization, required evidence, calculations, special cases, examples, and conflict resolution.
- LLM self-authored prompts and master prompt: Each prompt component received a score of 0, 1, or 2, producing a maximum total score of 76.Initial scoring was performed by Opus 4.8 (xhigh), followed by human validation, verification, and confirmation.
- Deep Research Literature-Discovery Benchmark: Seven Deep Research agents independently received the same master prompt in five separate runs.Each report was compared against the prompt for compliance, and every reported reference was manually checked for reality and open-access status.
Brain-organoid reporting benchmark · Use of generative AI · Data Availability
The benchmark translated brain-organoid reporting recommendations into a 23-item evaluation applied by seven models to ten articles across five runs. The study also used generative AI for manuscript support and shared prompts, data, and workflow materials.
- Brain-organoid reporting benchmark: Sandoval et al. (2024) recommendations became 23 paper-level items across five sections: methods, morphology and lineage, transcriptomics, functional assessment, and metabolism.The sections contained 7, 6, 2, 3, and 5 items, respectively.
- Brain-organoid reporting benchmark: Seven models generated three candidate prompts each, combined them into model-specific master prompts, and applied fixed prompts to ten brain-organoid articles in five runs.The ten papers were selected because they contained most guideline-requested features.
- Brain-organoid reporting benchmark: LLM outputs were recorded in Excel with raw answers, supporting quotes, and binary Yes/No judgments for each guideline item.This preserved both the model responses and their supporting evidence.
- Brain-organoid reporting benchmark: Two PhD-level reviewers independently scored the same ten papers and 23 items as Yes or No, while majority-of-five agreement required at least three Yes runs.A Yes score meant that the article addressed the guideline.
- Use of generative AI: Generative AI systems GPT-5.6 Sol, Fable 5, and Opus 5 wrote Python scripts for data processing and copy-edited the manuscript.The listed configurations were GPT-5.6 Sol (xhigh), Fable 5 (high), and Opus 5 (med).
- Data Availability: The study made available all prompts, raw and curated data for every figure, and a tutorial for trying the Figure 2 and Figure 3 workflow.These materials support inspection and reuse of the reported workflow.