Source-linked AI summary

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor, Francisco Guzmán, Nicholas Magazine, Jonas Mueller

arXiv:2608.21357v1cs.AI

TL;DR

Professional life-sciences workflows rely on visual artifacts whose interpretation is not well captured by existing evaluations. The paper introduces VIALS, a 161-task benchmark built from expert-created workflow artifacts, and finds that leading multimodal models achieve limited accuracy, with tool assistance improving performance but leaving a residual interpretation gap.

  • Problem

    Visual artifacts central to professional life-sciences workflows require domain-specific localization, measurement, comparison, and interpretation that existing benchmarks poorly cover.

  • Method

    VIALS evaluates multimodal models on 161 visual question-answering tasks created and reviewed by domain experts from professional life-sciences workflows.

  • Results

    Leading models achieve only 26.5% accuracy, while tool-assisted inspection improves accuracy by up to 43 points but the best agent solves only 65% of tasks.

  • Takeaways & Limitations

    Reliable interpretation of scientific images remains a substantial bottleneck for deploying vision-language models in life-sciences workflows.

  • Takeaways & Limitations

    VIALS is limited in scale and scope, emphasizing relatively straightforward high-frequency tasks rather than the full range of ambiguous interpretations and subsequent research-plan decisions.

Abstract

from arXiv · show

In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual interpretation tasks straightforward. AI that cannot similarly interpret such images will have limited utility in professional life sciences workflows, where such artifacts are central to how scientists reason, communicate, and make decisions.

1. Introduction

Life-science research depends on interpreting visual artifacts, but current vision-language models can describe these images without reliably extracting and reasoning over their critical evidence. VIALS introduces 161 professionally grounded visual question-answering tasks and finds severe accuracy limitations.

  • Motivation: Scientific workflows encode critical information in gel blots, microscopy images, plasmid maps, flow cytometry plots, docking views, and molecular structures.Accurate interpretation requires localizing evidence, extracting labels or values, distinguishing signal from noise, comparing elements, and applying scientific knowledge.
  • Failure modes: Fluent scientific answers can remain wrong when models select the wrong visual evidence or misinterpret spatial and structural relationships.Such errors may be difficult to detect because the final answer can sound scientifically plausible.
  • Gap: Existing benchmarks poorly cover visual artifacts and interpretation tasks encountered across professional life-sciences work.Prior evaluations largely target general academic knowledge or specialized biomedical modalities.
  • Benchmark: 161 VIALS tasks pair scientific images with questions requiring visual-evidence extraction and interpretation central to professional research workflows.PhD-level professional scientists with relevant domain and work expertise create and review the tasks.
  • Results: 26.5% accuracy was achieved by both top-performing models, GPT-5.6 Sol and Gemini 3.7 Flash.Around 90% of unsuccessful tasks involved miscounting, misreading measurements, or missing spatial and structural relationships.

2. Related Work

Related benchmarks evaluate academic or narrowly specialized scientific visual reasoning, whereas VIALS targets artifacts and tasks from professional life-sciences workflows. Expert assessments indicate that VIALS tasks are more professionally relevant than comparable HLE tasks.

  • Existing benchmarks: MMMU, MMMU-Pro, SciVQR, and SciVQA largely use academic materials such as exams, textbooks, publications, or conceptual questions.Other benchmarks focus on particular biomedical settings, including medical imaging, clinical tasks, or microscopy.
  • Task character: Figure 3 contrasts a VIALS flow-cytometry gating workflow with an HLE textbook-like conceptual question.The VIALS task supports determining whether a treatment changed immune-cell composition through direct interpretation of experimental output.
  • Coverage gap: Specialized biomedical benchmarks do not target the visual artifacts and tasks encountered across professional life-sciences work.Their coverage centers on selected modalities rather than the broader range of workflow artifacts.

3. The VIALS Benchmark

VIALS is a 161-task benchmark built from professional life-sciences artifacts and questions that require both visual evidence extraction and scientific reasoning. Multiple expert review stages validate task realism, correctness, answerability, and professional relevance.

  • Task design: Each VIALS task contains one or more images, a question, and a short held-out ground-truth answer.Tasks use open-domain text answers and semantic grading rather than exact matching.
  • Scoring: 99.9% agreement with human judgments was observed for the LLM judge across 4,685 model responses.The judge evaluated the question, model response, and reference answer using a grading prompt.
  • Artifact coverage: VIALS covers artifact domains including phylogenetic trees, flow cytometry plots, blots and gels, plasmid maps, cell-counting images, protein structures, and small-molecule structures.Domains were selected for criticality in high-value biotechnology and pharmaceutical workflows.
  • Artifact sourcing: Artifacts include unpublished experimental data, procedurally generated scientific-software outputs, and open-access publication sources, with emphasis on messy real-world data.Procedural generation was used for selected phylogenetic trees, plasmid maps, and blots, which experts verified as realistic.
  • Expert authorship: 31 contributors with professional domain expertise developed VIALS, and 87% had doctoral-level training.Contributors included researchers with experience at biotechnology, pharmaceutical, clinical, academic, and regulatory organizations.
  • Quality control: Each task undergoes independent expert answering, realism and answerability review, and consensus-based ground-truth validation.Independent expert attempts took 16 minutes on average, and tasks were retained only when answers agreed or consensus was reached.
  • Professional relevance: Experts rated VIALS artifacts as resembling professional work more often than HLE artifacts, at 63% versus 22%.They also reported higher weekly-or-daily exposure and commercial or economic value for VIALS tasks.

4. Evaluation Setup

The evaluation compares frontier multimodal models under a standardized single-turn setup. Each model receives only the image(s) and question, and accuracy is averaged across three independent rollouts.

  • Evaluation protocol: All evaluated multimodal models receive the same prompt containing only the image(s) and the question.The full prompt and model configurations are provided in Appendix C.1.
  • Evaluation protocol: Each model is independently run three times per task, with results reported as average accuracy across the three rollouts.This setup measures performance under repeated single-turn inference calls.

5. Results

VIALS reveals low and uneven vision-language-model performance on professional life-science artifacts, with failures concentrated in reading and interpreting visual evidence. Tool-assisted inspection improves many models but requires substantially greater computational resources.

  • Overall model performance: 26.5% accuracy is achieved by the top-performing models, GPT-5.6 Sol and Gemini 3.7 Flash, on VIALS.Gemini 3.7 Flash incurs lower costs than GPT-5.6 Sol.
  • Overall model performance: Under Pass^3, these models handle under 17% of tasks when all three rollouts must be correct.This evaluates consistency across repeated attempts rather than single-rollout accuracy.
  • Domain performance: Performance differs substantially across domains, with phylogenetics strongest for leading models at 43.5% and 34.8% accuracy.GPT-5.6 Sol reaches 43.5%, while Muse Spark 1.2 reaches 34.8% in phylogenetics.
  • Domain performance: 21.7% is the best reported accuracy for cell counting/quantification, the weakest domain alongside protein structure tasks.Gemini 3.7 Flash achieves 21.7% on cell counting/quantification.
  • Failure mode analysis: 83–93% of errors fall into the four artifact-reading categories, with visual quantification accounting for 28.2–41.8% across models.Quantitative reasoning errors account for 3.8–10.7%, while fabricated findings account for at most 4.3%.
  • Tool-assisted agents: Agentic tool access markedly improves performance for most models, but tool-assisted runs use 10–432× more tokens per attempt and GLM-4.6V declines.Models can iteratively crop, zoom, transform, measure, and otherwise inspect the supplied image programmatically.

6. Conclusion

VIALS shows that leading multimodal models remain unreliable at interpreting scientific artifacts, with failures often occurring in basic visual reading before higher-level reasoning. Tool-assisted inspection improves accuracy but introduces substantial cost and still leaves a residual interpretation gap.

  • Conclusion: VIALS evaluates visual-artifact interpretation in professional life sciences workflows, where models often misread labels, values, regions, or relationships needed for correct answers.The benchmark’s error analysis identifies failures before higher-level reasoning, despite models often recognizing the general artifact type.
  • Conclusion: Up to 43 points: code-assisted iterative inspection improves accuracy, indicating a perception gap in direct visual inference.Agents can crop, zoom, and measure artifacts, but the improvement reflects the additional inspection tools rather than direct visual inference alone.
  • Conclusion: 10–432× more tokens: tool-assisted runs cost substantially more than direct evaluation, while the best agent still solves only 65% of tasks.The residual interpretation gap remains even with unrestricted iterative inspection and programmatic tool use.
  • Additional evaluation findings: Pass@3 reaches around 40%, whereas Pass^3 remains below 17%, showing that models frequently fail to reproduce correct answers across repeated attempts.Pass@3 counts tasks solved in at least one of three attempts; Pass^3 counts tasks solved in all three.
  • Additional evaluation findings: High-confidence responses are only 6–20 percentage points more accurate than low-confidence responses, and high-confidence accuracy reaches at most 30%.Seven of nine models provide enough lower-confidence responses for comparison, but the overall confidence estimates remain severely miscalibrated.

A.3. Additional failure examples

The additional examples illustrate recurring VIALS errors in counting, feature selection, ranking, molecular-mass calculation, and structural interpretation. These examples show that models can apply a stated rule or perform arithmetic correctly while still extracting the wrong visual evidence.

  • Incorrect value or feature selection: Four of ten models call lanes 1 and 3 tied even though densitometry places lane 3 5% above lane 1.This example shows a ranking error caused by treating a real difference as a tie.
  • Visual quantification error: Models miscount objects even when applying an edge-exclusion rule correctly, enumerating five of six objects.The example concerns fully contained green outlined cell objects, with additional objects touching the image edge.
  • Additional failure examples: The examples include structural interpretation errors in which a hydrophobic arc is counted as an electrostatic contact and nonexistent hydrogen bonds are included.These errors concern selecting and interpreting spatial or chemical relationships in molecular structures.
  • Quantitative reasoning error: A model’s molecular-mass arithmetic is internally correct but omits the C3′ methyl, producing 2571.6 Da instead of the correct 2585.7 Da.The difference is 14.1 Da, corresponding to CH₂.
  • Incorrect value or feature selection: Models miss relevant visual features: eight of nine miss panel B, although every model finds the immunostained vessels in panel I.The correct answer includes both B and I.

B. Further comparison with Humanity’s Last Exam

VIALS is compared with HLE using semantic coverage and task construction, revealing that the benchmarks differ substantially in both question content and use of visual evidence.

  • Semantic coverage: VIALS and HLE questions show stark semantic differences in a two-dimensional UMAP projection of text embeddings.The embeddings use OpenAI’s text-embedding-3-small model and cosine distance before UMAP projection.
  • Task construction: HLE examples in overlapping scientific domains often test knowledge or entity identification rather than extracting evidence from the image for a research question.The image may function primarily as a cue for identifying an entity instead of supplying the scientific evidence needed to answer the question.
  • Task construction: The depicted HLE examples do not stem from professional life sciences work.Examples include identifying a protein length from an organism image, inferring a collection locality from morphology, and identifying a chemical reaction.

C.1. VLM evaluation

VLM responses are evaluated through standardized multimodal prompts, answer extraction, and semantic grading, with a rare image-based fallback for labeling ambiguities.

  • Direct evaluation: Each model receives one or more images and a question, then returns concise reasoning followed by a single final answer in a fixed format.The required response ends with an ANSWER: marker containing only the final answer.
  • Answer grading: Candidate answers are extracted after the final ANSWER: marker and judged for semantic equivalence to the ground-truth answer.Formatting, capitalization, unambiguous units, terminology, ordering, and duplicates may vary under the grading rules.
  • Answer grading: Numeric answers are strict: exact values, discrete counts, and stated reporting precision must match unless an inclusive lead-bound range applies.The grading rules reject invented tolerances, same-order-of-magnitude matches, and different coefficients at the same power of ten.
  • Fallback handling: Unverifiable image-dependent labeling correspondences are marked equivalent=false and routed for multimodal regrading or human review.This applies when the mapping between labels such as sample conditions, lane numbers, positional indices, or colors is absent from the question text.
  • Fallback handling: Well under 1% of judge calls escalate to the multimodal fallback across 10 models, 161 tasks, and 3 rollouts.The typical trigger is a mismatch between a reference sample condition and a model-reported lane number requiring image-based mapping.
  • Evaluation configuration: 483 total responses include retries for failed, refused, or empty model outputs, which are counted as incorrect if unsuccessful after three additional attempts.Muse Spark 1.2 has 13 such responses, Kimi K3 has 6, Grok 4.6 has 2, and Mistral Medium 3.5 has 1.

C.2. Failure classification methodology

Failure analysis assigns one primary substantive failure mode to each task–model pair with an incorrect rollout, using an LLM classifier that compares the attempt with the ground truth without seeing the image.

  • Analysis unit: Each task–model pair with at least one incorrect rollout contributes one observation to the failure-mode analysis.Multiple incorrect rollouts for the same pair do not create multiple observations.
  • Classifier inputs: The classifier receives the question, ground-truth final answer, and failing attempt’s final answer and reasoning trace, but not the task image.It identifies why the attempt diverged from the ground truth rather than solving the task independently.
  • Failure labels: The classifier assigns exactly one primary failure mode and must name the substantive scientific or reasoning cause of the error.Reasons should identify the broken scientific concept or reasoning step and what a competent expert would have done instead.

C.3. Tool-assisted agent evaluation

The tool-assisted evaluation gives agents a sandboxed Linux environment for image analysis while retaining the same tasks and answer-grading procedure as direct evaluation.

  • Execution environment: Tool-assisted agents use the same VIALS tasks as direct evaluation but can inspect and analyze images with shell commands and scientific Python packages.Available tools include image-processing, OCR, numerical, plotting, and graph-analysis libraries.
  • Inference cost: Mean tool-assisted cost per attempt ranges from 0.011 for GLM-4.6V to 1.047 for Claude Opus 5 with Claude Code.Table A1 reports direct cost, tool-assisted cost, and their ratio for the evaluated model–harness configurations.
  • Execution environment: Agents are instructed to use only the task files, without internet, external databases, or outside information.The image is available at /task/image.png, and scratch files may be written under /task/.
  • Output and grading: Tool-assisted responses are written to /task/answer.txt and scored using the same semantic grading procedure as direct evaluation.The final answer follows the ANSWER: marker format.
  • Professional relevance assessment: The professional relevance assessment asks experts to classify whether artifacts resemble professional work, how often they encounter them, and whether interpretation has commercial value.The response options cover appearance, frequency of exposure, and agreement about real commercial or economic value in life sciences.
Loading 2608.21357v1…