Source-linked AI summary

LAB-Bench: Measuring Capabilities of Language Models for Biology Research

Jon M. Laurent, Joseph D. Janizek, Michael Ruzo, Michaela M. Hinks, Michael J. Hammerling, Siddharth Narayanan, Manvitha Ponnapati, Andrew D. White, Samuel G. Rodriques

arXiv:2407.10362v3cs.AI

TL;DR

Existing science benchmarks mainly test textbook knowledge rather than practical research capabilities. LAB-Bench addresses this gap with a broad biology benchmark, finding wide variation across tasks and generally stronger human expert performance than models.

  • Problem

    Most science benchmarks test rote knowledge and textbook-style questions rather than practical biology research tasks, limiting evaluation of systems intended for scientific work.

  • Method

    The authors construct LAB-Bench, a dataset of over 2,400 multiple-choice questions spanning literature, figures, tables, databases, protocols, and DNA and protein sequence tasks.

  • Results

    Models show wide performance disparities across tasks, generally underperform human experts, and struggle especially with complicated DNA and protein sequence manipulation.

  • Takeaways & Limitations

    LAB-Bench provides an initial benchmark for practical biology research capabilities, including difficult cloning tasks relevant to evaluating potential AI assistance for molecular biologists.

  • Takeaways & Limitations

    LAB-Bench is not comprehensive because the breadth of biology required tradeoffs between topic coverage and task quality.

Abstract

from arXiv · show

There is widespread optimism that frontier Large Language Models (LLMs) and LLM-augmented systems have the potential to rapidly accelerate scientific discovery across disciplines. Today, many benchmarks exist to measure LLM knowledge and reasoning on textbook-style science questions, but few if any benchmarks are designed to evaluate language model performance on practical tasks required for scientific research, such as literature search, protocol planning, and data analysis. As a step toward building such benchmarks, we introduce the Language Agent Biology Benchmark (LAB-Bench), a broad dataset of over 2,400 multiple choice questions for evaluating AI systems on a range of practical biology research capabilities, including recall and reasoning over literature, interpretation of figures, access and navigation of databases, and comprehension and manipulation of DNA and protein sequences. Importantly, in contrast to previous scientific benchmarks, we expect that an AI system that can achieve consistently high scores on the more difficult LAB-Bench tasks would serve as a useful assistant for researchers in areas such as literature search and molecular cloning. As an initial assessment of the emergent scientific task capabilities of frontier language models, we measure performance of several against our benchmark and report results compared to human expert biology researchers. We will continue to update and expand LAB-Bench over time, and expect it to serve as a useful tool in the development of automated research systems going forward. A public subset of LAB-Bench is available for use at the following URL: https://huggingface.co/datasets/futurehouse/lab-bench

1 Introduction

LAB-Bench addresses the limited evaluation of practical biology research capabilities by assembling benchmarks spanning literature, figures, databases, protocols, and sequence manipulation. It also introduces human-hard cloning tasks and evaluates frontier models against expert biology researchers.

  • Existing science benchmarks mostly test rote knowledge and textbook-style questions rather than practical research tasks.
  • LAB-Bench contains over 2,400 multiple-choice questions covering literature reasoning, figure and table interpretation, database access, protocol writing, and DNA or protein sequence manipulation.
  • The benchmark includes 41 human-hard Cloning Scenarios designed to represent multi-step challenges in molecular cloning workflows.
  • The authors evaluate frontier commercial and open-source models and compare their performance with PhD-level biology researchers.
  • Approximately 80% of each benchmark subtask is publicly released while 20% is retained to monitor contamination.
  • The authors identify model-generated question screening and high-quality distractor design as important considerations for constructing difficult evaluations.

2 Results

LAB-Bench combines programmatic and manual dataset construction to evaluate practical biology research capabilities. Results show substantial variation across tasks: models perform relatively well on some structured or simpler tasks but remain weak on literature lookup, long-sequence manipulation, cloning, and protocol troubleshooting, generally trailing human experts.

  • Dataset construction: LAB-Bench used programmatic generation for scalable tasks and manual expert generation for more difficult categories.The benchmark spans literature, database, figure, table, protocol, sequence, and cloning tasks.
  • Literature and database reasoning: Retrieval-oriented performance was uneven: LitQA2 reached above-random precision, whereas SuppQA and DbQA elicited frequent refusals and low coverage.SuppQA requires supplemental-material lookup, while DbQA requires access to biological databases.
  • Figure and table interpretation: FigQA was near-random for most models, while Claude 3.5 Sonnet performed well above the others on technical image interpretation.TableQA was easier overall, and Claude 3.5 Sonnet narrowly surpassed human precision while matching human accuracy.
  • Sequence comprehension and manipulation: SeqQA precision generally fell around 40% to 50%, with some simpler primer-selection subtasks exceeding 90%.Performance dropped when questions required recalling information from a gene name or manipulating long sequences and subsequences.
  • Sequence comprehension and manipulation: Restriction-digestion subtasks requiring fragment counts or lengths were among those with the poorest model performance.These tasks require accurate internal mapping of enzyme digestion across sequences.
  • Protocols and cloning: Models clustered around 50-60% precision on ProtocolQA and remained well below humans on Cloning Scenarios.On cloning questions, low coverage did not improve precision, and reported successes often appeared to rely on distractor-elimination heuristics.
  • Human comparison: Human experts generally outperformed models in accuracy and precision while answering most questions, with TableQA the closest comparison.Claude 3.5 Sonnet equaled human accuracy and exceeded human precision on TableQA.

3 Discussion and Limitations

LAB-Bench reveals uneven model capabilities across practical biology tasks, with models generally behind human experts and weaker on sequence manipulation. The benchmark’s interpretation is constrained by distractor quality, incomplete human baselines, unequal tool access, and limited breadth.

  • Models show wide performance disparities across LAB-Bench tasks, often refusing information-lookup questions and struggling especially with complicated manipulation of DNA and protein sequences.
  • Human experts generally outperform models across most practical research categories, although the gap varies by subtask.
  • LAB-Bench is not comprehensive because biology’s breadth required trading topic coverage for task quality while focusing initially on foundational areas.
  • Multiple-choice scores can overstate reasoning ability because models may eliminate implausible distractors and guess, making high-quality distractors and open-answer validation important.
  • Human baselines for Cloning Scenarios may be unreasonably low because questions require molecular-cloning expertise and 10–60 minutes each.
  • The initial evaluation excluded model tools and agents, while human evaluators could use tools, so comparisons are not entirely equitable.

A.1.1 DbQA

DbQA includes database-oriented biology questions spanning gene-disease associations, genomic locations, gene sets, interaction predictions, regulatory sites, and clinical variants.

  • DbQA asks which genes are associated with diseases, located at genomic coordinates, or contained in database-derived gene sets.
  • Database questions may require distinguishing answers across resources such as DisGeNet, OMIM, Ensembl, miRDB, MouseMine, and the Gene Transcription Regulation Database.
  • Several questions require identifying benign or pathogenic variants using ClinVar and protein sequence or variant information.
  • The task also covers predicted protein interactions and transcription-factor binding sites from specialized biological databases.

A.1.2 SeqQA

SeqQA covers sequence-oriented biology research tasks, including primer design, restriction digestion, open reading frames, translation, and DNA composition. The examples span both cloning workflows and direct sequence analysis.

  • Cloning and primer design: Restriction-ligation and Gibson-assembly questions test selecting primers or enzymes that match cloning designs.Examples cover pUC19 cloning with SalI/PstI, EcoRI/SacI, SalI/TspMI, and Gibson assembly into HindII- or SmaI-linearized vectors.
  • Cloning and primer design: Primer-design tasks also ask whether candidate pairs generate amplicons with specified lengths or sequences.The examples include expected amplicon length, a requested 422 bp product, and matching a target amplicon sequence.
  • Sequence analysis: Restriction-enzyme questions require predicting both fragment lengths and the number of fragments after digestion.One example asks for fragment lengths after MaeI and Cfr10I digestion, while another asks for the resulting fragment count.
  • Sequence analysis: ORF tasks test identifying translation-related sequence properties, including high-efficiency RNA, amino-acid translation, ORF counts, and encoded residues.The examples ask about translation efficiency, the longest ORF’s amino-acid sequence, ORFs longer than 70 amino acids, and the amino acid at a specified position.

A.1.3 CloningScenarios

CloningScenarios questions apply sequence and protocol reasoning to plasmid assembly, screening, molecular biology protocols, and experimental systems. The examples include construct verification, transfection troubleshooting, and interpretation of biological effects.

  • Cloning and construct validation: Other scenarios connect plasmid or fragment sequences to identifying genetic elements and selecting appropriate assembly designs.One question asks which element is present in a DNA fragment, while the Golden Gate example combines four plasmids using BsaI.
  • Cloning and construct validation: Plasmid-assembly scenarios ask researchers to infer fragment patterns that identify a correct Golden Gate clone.A screening example uses NotI and PvuI digestion and specifies the fragment-length pattern indicating correctness.
  • Experimental interpretation: Additional questions test interpreting molecular-biology measurements and optical-system arrangements.Examples ask about the effect of an ATPase-deficient Spindle E mutant, the ratio of modified protein sites, and the half-waveplate position relative to the SLM.
  • Protocols and experimental procedures: The benchmark includes literature- and protocol-based questions requiring retrieval of experimental conditions or troubleshooting steps.Examples ask for the temperature and duration that denature transposase in a sequencing protocol and for a step that may improve iPSC transfection efficiency.

B Supplemental Results

The supplemental results report accuracy, precision, and coverage for LAB-Bench tasks and subtasks. These measurements use the entire dataset and average results across three runs per model.

  • Accuracy: Accuracy is tabulated for all LAB-Bench tasks and subtasks over the combined public and private dataset splits.Results are averaged over three runs per model.
  • Precision: Precision is tabulated for all LAB-Bench tasks and subtasks using the combined public and private dataset splits.Results are averaged over three runs per model.
  • Coverage: Coverage is tabulated for all LAB-Bench tasks and subtasks over the combined public and private dataset splits.The table caption specifies the same three-run averaging procedure used for the reported model results.

B.1 Public and Private split set performance

LAB-Bench reserves 20% of each subtask as a private split and releases 80% publicly to monitor contamination. Before public release, model performance differed by less than five percentage points in mean accuracy between the splits.

  • Split construction: 80% of each LAB-Bench subtask is public, while 20% remains private for future contamination monitoring.The main results use the full dataset, with split-specific results provided in supplemental material.
  • Split performance: Less than five percentage points separated mean LAB-Bench accuracy on the private and public splits.The comparison weights subtasks equally and is reported in Supplemental Figure 6.
  • Split performance: Before public release, models performed equivalently well on the private and public sets according to the reported mean-accuracy comparison.This comparison is presented as a contamination-monitoring check rather than the main full-dataset result.

B.2 Open Answer results

The authors evaluated a small open-response subset spanning CloningScenario, ProtocolQA, and FigQA, with expert grading and secondary review; results are reported in Supplemental Table 5.

  • A small-scale open-response evaluation covered CloningScenario, ProtocolQA, and FigQA questions.Question wording was modified where necessary to remove multiple-choice-specific phrasing.
  • An expert biologist paper author graded model completions against ideal answers, and a second reviewer reviewed the results.
  • The open-response results are collected in Supplemental Table 5.

C Task generation

Task construction combined programmatic scaling for some categories with predominantly manual, domain-expert question generation, supported by interfaces for collecting and reviewing tasks.

  • SeqQA and DbQA could be scaled programmatically by generating or extracting data and crafting questions formulaically.
  • Most other categories required manual generation and specific domain expertise.
  • An Airtable-based workflow supported task drafting, metadata maintenance, and review, revision, rejection, or approval.

C.1 LitQA

LAB-Bench tasks were constructed to test literature, supplementary-material, figure, table, protocol, cloning, and sequence reasoning through a mix of manual, expert-driven, and programmatic procedures.

  • LitQA: LitQA questions require full-text literature rather than abstracts or titles and were manually generated by authors and contracted experts.Question authors were instructed to identify recent papers published within the last 36 months.
  • SuppQA: SuppQA questions require information from supplementary text or tables, often involving paper-specific methods or reagents.
  • FigQA and TableQA: FigQA and TableQA assess interpretation of figures and tables extracted from biology papers as images.
  • ProtocolQA: ProtocolQA introduces errors into published protocols to test whether models can interpret negative results and propose protocol changes.
  • CloningScenarios: Cloning scenarios model complex workflows involving multiple plasmids, DNA fragments, restriction enzymes, and cloning methods.Questions may share a scenario but are designed to be independently answerable because each includes the full scenario.
  • SeqQA: SeqQA covers PCR design, restriction enzymes, open reading frame translation, molecular cloning, and basic sequence properties such as GC percentage.
  • SeqQA: Sequence-based cloning questions generate ideal primers by selecting ORF-end substrings and appending recognition sequences for randomly chosen eligible enzymes.Distractors include shifted primers, an alternative gene, and a non-overlapping enzyme pair.

D.3 Evaluation strategy

The evaluation used a custom harness to extract multiple-choice answers from natural-language completions, with model-specific coverage constraints for Llama 3.

  • A custom evaluation harness used chem-bench components and regex parsing to extract answer tags from model completions.
  • Llama 3 was excluded from FigQA and TableQA because it lacks vision capabilities.
  • Only 25 of 41 Cloning Scenarios questions were run for Llama 3 because of prompt-token limits.

D.4 Analysis

The analysis establishes human baselines using quizzes, controlled task assignment, flexible research tools, and incentive structures. Human and model performance are compared with accuracy, precision, and coverage metrics.

  • Performance metrics: Accuracy, precision, and coverage were reported for models and human participants.Accuracy counts correct answers; precision measures correctness among non-“Insufficient information” responses; coverage measures the fraction of questions answered without that option.
  • Human baseline: Human quizzes ranged from 20–140 questions and combined multiple categories or focused on a single category.Experts were recruited from the question-drafting group but were not assigned questions they had drafted.
  • Human baseline: Task assignment avoided previously seen questions and prioritized approximately random coverage of selected categories.Some categories did not receive 1X coverage because the tasks were judged repetitive enough for random sampling to support comparison.
  • Administration: Quiz-takers had 5–8 days to complete quizzes without an additional time limit and could use web search, code, or DNA sequence software.Quizzes were administered through a custom Airtable interface.
  • Incentives: Compensation combined per-question payment, performance bonuses, and occasional completion bonuses to encourage completion and quality.Base values ranged from $2 to $12 per question, while completion bonuses ranged from $50 to $200 per quiz.
  • Incentives: The incentive schemes varied bonuses by accuracy or relative performance, while “unsure” responses received partial payment to discourage random guessing.Incorrect answers received no performance bonus, and suspected bad-faith responses could lead to withheld payment or termination.

E.4 Human performance analysis

Human performance was evaluated with the same three metrics as model performance, adapting them to sure versus unsure responses. Accuracy and precision therefore focus on completed answers, while coverage measures answer completion.

  • Human metrics: Human accuracy counts correct sure answers out of all questions.A sure answer is one that the evaluator did not flag as “unsure.”
  • Human metrics: Human precision counts correct sure answers among questions attempted with certainty.The denominator excludes questions marked “unsure.”
  • Human metrics: Human coverage is the fraction of questions for which the evaluator provided a sure answer.This metric captures how often human participants committed to an answer rather than flagging uncertainty.
Loading 2407.10362v3…