Source-linked AI summary
VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology
Luca L. Weishaupt, Simone de Brot, Javier Asin, Llorenç Grau-Roma, Nic G. Reitsam, Andrew H. Song, Dongmin Bang, Stefan T. Kaluziak, Long Phi Le, Jakob Nikolas Kather, Faisal Mahmood, Guillaume Jaume
TL;DR
Existing pathology vision-language benchmarks provide limited coverage of veterinary and toxicologic pathology, despite animal-tissue review being central to preclinical safety assessment. The paper introduces VIPER, an expert-curated rat histology benchmark with morphology-grounded questions, and evaluates 16 models. Results show limited transfer from human and general-purpose models, while veterinary-specific instruction tuning yields consistent gains across backbones.
Problem
Veterinary and toxicologic pathology lack broad public vision-language benchmarks despite their role in preclinical drug safety assessment.
Method
VIPER uses expert-validated rat histology questions across seven organ systems and three formats, with standardized evaluation across 16 models.
Results
Veterinary-specific instruction tuning yields consistent gains across backbones, while human pathology-specialized and general-purpose frontier models show limited transfer to veterinary pathology.
Takeaways & Limitations
VIPER benchmarks progress in non-human pathology VLMs and provides a template for visually grounded biomedical evaluation datasets.
Takeaways & Limitations
VIPER is ROI-level, omits some organs and species, and remains modest in size at 419 images and 1,251 questions.
Abstract
from arXiv · showhide
Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non-human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert-curated benchmark for vision-language model evaluation in toxicologic pathology. VIPER contains 1,251 questions associated with 419 H&E-stained rat histology images across seven organ systems, covering multiple-choice, KPrim, and free-text formats. All questions were curated and validated by board-certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary-pathology models, seven human pathology-specialized models, and seven general-purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over-diagnosis of normal tissue in frontier models, and show that domain-specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.
1 Introduction
AI pathology evaluation has matured mainly around human oncology, while veterinary and toxicologic pathology remain underbenchmarked despite their role in preclinical drug safety. VIPER addresses this gap with expert-curated, morphology-grounded vision-language evaluation and veterinary-specialized models.
- Human pathology benchmarks have enabled systematic evaluation across diagnosis, grading, molecular prediction, and segmentation.
- Preclinical toxicologic pathologists examine animal tissues across organs and dose groups to characterize drug toxicity and determine whether compounds can advance to human testing.
- Veterinary pathology lacks public benchmarks, while existing datasets emphasize isolated tasks and provide limited vision-language reasoning evaluation.
- VIPER contains 1,251 questions linked to 419 H&E-stained rat histology images across seven organ systems and three question formats.
- Questions were authored and validated by ECVP board-certified veterinary pathologists, then refined through human-in-the-loop generation and adversarial filtering against mirage predictions.
- The study also introduces TOXSCRIBE models and benchmarks veterinary-specialized, human pathology-specialized, and general-purpose multimodal systems.
3 VIPER: Veterinary Pathology Image Evaluation and Reasoning
VIPER combines diverse rat histology regions of interest with expert-authored, morphology-grounded questions, structured evaluation, reader validation, and comparisons across 16 multimodal models.
- Benchmark composition: VIPER comprises 1,251 question-answer pairs associated with 419 rat H&E regions of interest across seven organ systems and 23 named anatomic structures.
- Benchmark composition: Board-certified veterinary pathologists validated annotations, while questions span MCQ, KPrim, and free-text formats to test multiple forms of multimodal reasoning.
- Curation pipeline: The curation workflow uses randomly extracted ROIs, morphology-based seed questions, LLM-generated variants, expert refinement, and adversarial filtering.
- Evaluation protocol: MCQs use exact-match accuracy averaged across cyclic answer-order permutations, KPrim uses the ETH half-point rule, and free text uses a calibrated judge weighting accuracy 70% and completeness 30%.
- Validation and models: A reader study sampled 100 image-question pairs and compared three ECVP-certified veterinary pathologists, while 16 models were evaluated under the same prompting and generation protocol.
- Validation and models: The study developed two veterinary-specialized models, TOXSCRIBEQwen and TOXSCRIBEGemma, alongside human pathology-specialized and general-purpose multimodal models.
5 Results
VIPER results show that veterinary pathology specialization improves overall and organ-level performance, while model behavior varies substantially by question category and visual access. The benchmark also reveals transfer gaps, instruction biases, complementary errors, and measurable gains from domain-specific fine-tuning.
- Overall performance: 62.4% overall score for TOXSCRIBEQwen and 61.2% for TOXSCRIBEGemma lead VIPER, significantly exceeding every non-veterinary-specialized model.The two TOXSCRIBE variants are statistically similar (P = 0.13), suggesting their domain-adaptation benefit transfers across backbones.
- Performance by question category: 80.5% over-reading-probe performance for TOXSCRIBEGemma and 77.2% for TOXSCRIBEQwen exceed PathChat+ at 62.1% and GPT-5.4 at 55.1%.The category tests whether models resist hallucinating lesions in normal tissue; pathology identification remains challenging, with TOXSCRIBEQwen at 50.3%.
- Performance by organ system: TOXSCRIBE ranks first in six of seven organ systems, with gains of 9.5% in urinary and 6.0% in respiratory tissues.The narrowest reported margin is 2.0% in digestive tissue, and TOXSCRIBE outperforms PathChat+ in every organ system.
- Veterinary vs. human pathology: TOXSCRIBEQwen exceeds PathChat+ by 11.5% overall, while veterinary reader VP2 scores 93.1% versus 72.4% for human-trained physician reader PP1 on MCQ.KPrim and free-text reader gaps are not significant, with differences of 7.6% and −0.8%, respectively.
- The importance of visual grounding: Removing images reduces scores by 41.7% for TOXSCRIBEQwen, 28.0% for PathChat+, and 29.0% for GPT-5.4 across VIPER.Reader MCQ performance also drops by 31.6% for VP3 and 35.7% for PP1 when images are removed.
- Gain from domain-specific fine-tuning: 10.0% improvement over base Qwen3.5-27B and 6.8% over base Gemma 4 result from veterinary pathology-specific fine-tuning.Both gains are statistically significant at P < 0.001, and the training data were independent of VIPER.
- Challenges in understanding instructions: Patho-R1-3B and Patho-R1-7B answer True on 90.2% and 87.6% of KPrim statements, while LLaVA-Med selects option A on 67.5% of MCQs.The biases occur despite competitive free-text scores for the Patho-R1 models and uniformly rotated correct-answer positions for LLaVA-Med.
- Oracle performance: An oracle selecting the highest per-question score among TOXSCRIBEQwen, PathChat+, and GPT-5.4 reaches 78.2% overall.The models fail on substantially different questions, indicating complementarity across their predictions.
6 Discussion
VIPER addresses a missing evaluation resource in non-human pathology through expert-curated, visually grounded ROI-level questions. The benchmark shows limited transfer from human pathology and frontier models, while veterinary-specific instruction tuning improves performance across backbones.
- Benchmark contribution: VIPER provides an expert-curated benchmark for ROI-level vision-language evaluation in rat toxicologic pathology.Its design uses expert-authored questions, human-in-the-loop refinement, and adversarial filtering to emphasize visual grounding.
- Benchmark contribution: The benchmark extends pathology evaluation beyond human tissue by focusing on non-human toxicologic pathology.
- Empirical findings: Human pathology-specialized and general-purpose frontier models show limited transfer to veterinary pathology.
- Empirical findings: Veterinary-specific instruction tuning improves performance across different model backbones.
- Reliability implications: Normal-tissue performance provides a complementary reliability test because safety assessment requires detecting lesions while avoiding false positives in normal tissue.
2. Limitations
The paper identifies scope and evaluation limitations that constrain how VIPER should be interpreted. These include ROI-level rather than whole-slide assessment, incomplete organ and species coverage, modest size, and subjective free-text scoring.
- Scope limitations: VIPER evaluates selected regions of interest rather than complete slide-level toxicologic pathology workflows.Slide-level review and dose-group comparison remain outside the benchmark’s current scope.
- Interpretation boundary: VIPER performance should not be interpreted as evidence that models are ready for deployment in toxicologic pathology.
6. Experimental setting/details
The paper reports evaluation and release practices intended to support reproducibility and responsible interpretation. It describes training details, statistical procedures, compute requirements, societal risks, and a research-only framing.
- Experimental details: The paper reports benchmark composition, curation, evaluation, prompting, scoring, and statistical-analysis details needed to understand the main results.
- Statistical analysis: Key comparisons use uncertainty and statistical-significance reporting, including agreement metrics and paired bootstrap testing.
- Compute requirements: TOXSCRIBEQwen was trained on 8× NVIDIA H200 GPUs for approximately 90 hours, while TOXSCRIBEGemma required approximately 30 hours on the same hardware.
- Responsible use: The paper discusses positive impacts such as improved evaluation and safer AI development, alongside risks from overinterpreting benchmark performance and inappropriate deployment.
- Responsible use: VIPER is framed as a research evaluation resource rather than a clinical or regulatory decision system.
12. Licenses for existing assets
The paper documents asset provenance, licensing, and benchmark documentation while positioning VIPER as a research resource. It also describes expert collaborators and the absence of human-subject research.
- Existing assets: The paper credits original sources for preclinical data, benchmarks, baseline models, and software assets.
- Existing assets: Images are redistributed under their source-repository licenses, including CC BY-SA 2.1 JP for Open TG-GATEs and CC BY-NC 4.0 for MMO.
- New assets: VIPER documentation covers image sources, organ composition, question types, annotation, validation, intended use, licensing, and limitations.
- Human involvement: The reader study involved three board-certified veterinary pathologists and one board-certified physician pathologist as research collaborators, without crowd workers or monetary compensation.
- Human involvement: The paper states that human-subject IRB approval was not applicable because the study used preclinical animal pathology images and expert-authored or expert-validated questions.
A Additional information about VIPER
VIPER combines expert-curated questions with veterinary pathology model development, structured coverage of rat tissues, and multiple evaluation formats. Its design also specifies category definitions and reader-agreement procedures for assessing benchmark quality.
- Model development: TOXSCRIBEQwen and TOXSCRIBEGemma are LoRA fine-tuned versions of Qwen/Qwen3.5-27B and google/gemma-4-31B-it, using the same instruction-tuning corpus.
- Optimization: The training pipeline uses AdamW, cosine scheduling, LoRA adapters, a frozen vision backbone, multi-magnification images, and 62,201 optimization steps.LoRA rank is 64, with α=128 and dropout 0.05.
- Benchmark composition: VIPER covers seven organ systems and 23 named anatomic structures, with question counts organized by organ system and format.The benchmark includes MCQ, KPrim, and free-text formats.
- Question categories: The question categories include anatomy identification, hallucination resistance, spatial localization, pathology identification, feature description, artifact recognition, and numeric extent estimation.Examples include identifying renal cortex, rejecting a false myositis claim, locating tubular necrosis, and estimating gland atrophy.
- Reader-study analysis: MCQ and KPrim agreement are analyzed at choice and statement levels, while free-text reliability uses ICC(2,1) alongside Pearson correlation.The reliability summaries distinguish discrete-category agreement from continuous-score agreement.
A.4 Free-text LLM judge robustness
The free-text evaluation is robust to judge repetition and judge family, although absolute scores depend on the selected judge. TOXSCRIBEQwen’s advantage over PathChat+ remains significant across both independent judging conditions.
- Score calibration: TOXSCRIBEQwen scores 58.3 under GPT-5.4, 59.1 under replicate GPT-5.4, and 56.0 under Claude Opus 4.7.The absolute score changes with the judge while the model ordering is broadly retained.
- Judge agreement: Pearson r = 0.98 and Spearman ρ = 0.95 show close item-level agreement between GPT-5.4 and Claude Opus 4.7 judges.Their disagreement is concentrated in absolute score levels rather than response ranking.
- Judge agreement: A replicate GPT-5.4 judge shifts model means by only 0.8–1.1 %, indicating small within-judge stochasticity.The replicate correlation with the headline pipeline is Pearson r = 0.99.
- Model comparison: TOXSCRIBEQwen’s free-text advantage over PathChat+ is 5.4–5.7 % and statistically significant under both GPT-5.4 and Claude Opus 4.7.Holm-adjusted p ≤0.008 across the two-judge family.
A.7 Performance by organ system
The supplied appendix material defines VIPER’s organ-system evaluation and question-authoring protocols while documenting scope boundaries for reader comparisons and image-ablation conditions. It also reports limited transfer to PathMMU, where PathChat+ remains strongest.
- Statistical comparison: Only one physician reader participated, limiting the generalizability of the KPrim and free-text null results.
- Image ablation: The image-ablation matrix distinguishes Normal, Black, Random, and No image conditions and reports the Normal −No-image gap.
- Image ablation: Random-image ablations are reported only for MCQ because KPrim and free-text random-image conditions used an earlier protocol.
- Cross-benchmark generalization: PathChat+ leads PathMMU by 2.1% over ToxScribe (Qwen), while ToxScribe (Gemma) is statistically similar to GPT-5.4 at ∆= −0.7 pp.The cross-benchmark evaluation uses n = 8,468 questions.
- Generation prompt: VIPER’s question-generation prompt targets rat H&E histopathology from TG-GATEs or MMO studies for veterinary pathology students and residents.The prompt supplies an image, seed question-answer pair, and organ, source, and magnification metadata.
2. Generation, user prompt.
The generation prompt converts expert seed questions and histopathology images into three structured exam formats. It emphasizes image dependence, text-only anti-leakage, plausible distractors, and explicit scoring structures.
- 2. Generation, user prompt.: The metadata fields identify the source, rat organ, and magnification for each generated question.
- 2. Generation, user prompt.: Each generation request supplies a histopathology image, an expert seed question-answer pair, and instructions to produce MCQ, KPrim, and free-text variants.
- 2. Generation, user prompt.: Every question must require image examination, and question text must not name or describe the visible finding that determines the answer.
- 2. Generation, user prompt.: MCQ options must be plausible for the rat organ, similarly long, and checked for text-only answer leakage.
- 3. Structured output schema.: The structured schema defines MCQ stems, answers, explanations, and four distractors, plus KPrim stems with four true-or-false statements.
- 3. Structured output schema.: The schema defines KPrim questions with exactly four statements and free-text questions with a stem, expected answer, synonyms, scoring rubric, and maximum points.
- 4. Adversarial text-only checker, MCQ, system prompt.: The text-only MCQ checker must select exactly one option and return an answer index, confidence level, and reasoning as JSON.
5. Adversarial text-only checker, MCQ, user prompt.
The MCQ and KPrim prompts are designed to preserve seed-topic fidelity while forcing image-grounded judgments and reducing answerability from textual cues alone.
- 5. Adversarial text-only checker, MCQ, user prompt.: MCQs require selecting one correct answer from five options and returning JSON only.
- 6. Adversarial text-only checker, KPrim, system prompt.: KPrim prompts require true/false judgments for all four statements, even when the image cannot be seen.
- 6. Adversarial text-only checker, KPrim, system prompt.: KPrim outputs must contain four Boolean answers, a confidence label, and brief reasoning in JSON format.
- 7. Adversarial text-only checker, KPrim, user prompt.: New questions must test the same seed topic and remain grounded in the seed answer rather than changing subjects.
- 7. Adversarial text-only checker, KPrim, user prompt.: To reduce guessability, questions emphasize image-specific appearance, location, distribution, counts, extent, and relationships.
- 7. Adversarial text-only checker, KPrim, user prompt.: Distractors and statements are balanced to avoid textbook-answer, stem-hint, detail-level, or generally true-or-false cues.MCQ options are constrained to similar lengths, and failed items are regenerated after text-only readers guess correctly.
10. Regeneration after a pathologist rejection, system prompt.
The regeneration prompt uses expert-authored rat histopathology questions and reviewer feedback to produce improved, non-duplicative veterinary pathology assessments.
- 10. Regeneration after a pathologist rejection, system prompt.: The regeneration system positions the model as an expert veterinary pathologist and exam-question author.
- 10. Regeneration after a pathologist rejection, system prompt.: The questions concern H&E-stained rat histopathology sections from TG-GATEs or MMO studies.
- 10. Regeneration after a pathologist rejection, system prompt.: The intended audience is veterinary pathology students and residents.
- 10. Regeneration after a pathologist rejection, system prompt.: Each regeneration receives a histopathology image, an expert-authored seed question and answer, metadata, and a reviewer-rejected question with feedback.
- 11. Regeneration after a pathologist rejection, user prompt.: The prompt also provides other questions for the image so the replacement tests a different aspect.
- 11. Regeneration after a pathologist rejection, user prompt.: The regenerated item must address reviewer feedback, avoid repeating the rejected approach, and reference the rat-organ image.