Source-linked AI summary
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, Qi Liu
TL;DR
LVLMs remain limited in understanding abstract scientific figures because scientific training data are scarce. Multimodal ArXiv introduces ArXivCap and ArXivQA and uses them for scientific comprehension evaluation; ArXivQA improves mathematical reasoning, while in-domain training substantially improves vision-to-text performance despite persistent captioning challenges.
Problem
LVLMs struggle with abstract scientific figures because training datasets for scientific domains requiring complex reasoning are inadequate.
Method
The paper constructs ArXivCap from 6.4M images and 3.9M captions across 572K ArXiv papers, generates 100K GPT-4V QA pairs as ArXivQA, and defines four vision-to-text evaluation tasks.
Results
Across evaluations, ArXivQA improves open-source LVLM mathematical reasoning, including a 10.4% absolute MathVista accuracy gain, while in-domain training substantially improves all four vision-to-text tasks.
Takeaways & Limitations
Multimodal ArXiv provides broad scientific figure data and benchmarks, but current LVLMs still misinterpret visual context, make recognition errors, and produce overly simplified captions.
Takeaways & Limitations
The dataset draws from ArXiv papers, which may overlook the disciplinary and modality diversity of the broader scientific literature.
Abstract
from arXiv · showhide
Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes. However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a scarcity of training datasets in scientific domains. To fill this gap, we introduce Multimodal ArXiv, consisting of ArXivCap and ArXivQA, for enhancing LVLMs scientific comprehension. ArXivCap is a figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers spanning various scientific domains. Drawing from ArXivCap, we introduce ArXivQA, a question-answering dataset generated by prompting GPT-4V based on scientific figures. ArXivQA greatly enhances open-sourced LVLMs' mathematical reasoning capabilities, achieving a 10.4\% absolute accuracy gain on a multimodal mathematical reasoning benchmark. Furthermore, employing ArXivCap, we devise four vision-to-text tasks for benchmarking LVLMs. Evaluation results with state-of-the-art LVLMs underscore their struggle with the nuanced semantics of academic figures, while domain-specific training yields substantial performance gains. Our error analysis uncovers misinterpretations of visual context, recognition errors, and the production of overly simplified captions by current LVLMs, shedding light on future improvements.
1 Introduction
Multimodal ArXiv addresses LVLMs’ difficulty with abstract scientific figures by providing large-scale figure-caption and figure-based QA data, then evaluates scientific comprehension across reasoning and vision-to-text tasks.
- Dataset construction: 6.4M images and 3.9M captions from 572K papers form ArXivCap, a diverse scientific figure-caption dataset spanning academic domains.The dataset preserves subfigure structure and original-paper titles to support diverse evaluation tasks.
- Dataset construction: 100K GPT-4V-generated multiple-choice QA pairs derived from ArXivCap form ArXivQA for scientific figure reasoning.The QA resource is designed to improve scientific reasoning abilities in LVLMs.
- Evaluation: 10.4% absolute accuracy gain on MathVista shows that ArXivQA improves Qwen-VL-Chat’s multimodal mathematical reasoning.The evaluation measures reasoning through QA accuracy and generation through vision-to-text tasks.
- Evaluation: Four vision-to-text tasks benchmark single-figure captioning, multi-figure summarization, in-context captioning, and paper-title generation.These tasks test increasingly varied forms of scientific-figure comprehension.
- Findings: Current LVLMs struggle with nuanced scientific-figure semantics, while in-domain training produces substantial gains across the four tasks.Error analysis identifies visual-context misinterpretation, recognition errors, and overly simplified captions.
2 Related Work
Recent LVLM research advances along three fronts: model architecture, training paradigms, and dataset curation. These efforts emphasize modular design, scaling and alignment strategies, and increasingly high-quality multimodal data.
- 2 Related Work: Recent LVLM progress spans model architecture, training paradigms, and dataset creation.The related-work overview characterizes these as major areas of advancement.
- Model Architecture: LVLMs typically combine a vision encoder, modality alignment module, and LLM backbone to process and decode multimodal context.CLIP is commonly used for image encoding, while LLaMA and Vicuna are popular language-model backbones.
- Training Paradigms: Training research examines scaling vision and language components, increasing image resolution, and varying which modules are unfrozen.PaLI-X studies component scaling, while Qwen-VL explores resolution and unfreezing strategies.
- Training Paradigms: Alignment methods including RLHF and AI-feedback preference optimization improve LVLM alignment with human preferences.The cited examples include LLaVA-RLHF and preference optimization through AI feedback.
- Dataset Curation: Dataset quality substantially affects LVLM performance, motivating cleaned image captions and high-quality multimodal instruction-tuning datasets.Web-scale image-caption pairs support modality alignment, while instruction fine-tuning helps models respond to user queries.
3 Multimodal ArXiv
Multimodal ArXiv comprises ArXivCap and ArXivQA, built from curated scientific paper figures and captions to support LVLM training and evaluation. Its datasets span broad scientific domains and improve multimodal mathematical reasoning.
- ArXivCap: ArXivCap is constructed by filtering papers using publication records, extracting figure-caption pairs, cleaning captions, and filtering problematic images.The pipeline removes short captions and images with extreme aspect ratios, small edges, or excessive pixel counts.
- ArXivCap: ArXivCap contains 572K papers and 6.4M high-quality images across 32 scientific domains, making it a broad real-paper figure-caption resource.The dataset includes 193M words and covers fields including computer science, mathematics, physics, and economics.
- Dataset comparison: Table 2 positions ArXivCap as the largest captioning dataset and ArXivQA as the only QA dataset covering broad domains from real papers.The comparison emphasizes scale, domain breadth, and the use of figures from real scientific papers.
- ArXivQA: ArXivQA contains 100K question-answer pairs generated from sampled ArXivCap figures using GPT-4V prompting, after invalid samples were filtered.The questions average 16.98 words, with 4.20 options per question on average.
- ArXivQA: 10.4% absolute accuracy boost for Qwen-VL-Chat on MathVista demonstrates the benefit of fine-tuning with ArXivCap and ArXivQA.Table 4 reports that the two datasets together enhance Qwen-VL-Chat’s overall MathVista performance.
4 Experiments
Experiments evaluate ArXivQA for multimodal mathematical reasoning and ArXivCap across four scientific vision-to-text tasks. Domain-specific training improves several reasoning and generation outcomes, while current LVLMs remain challenged by academic figures and contextual captioning.
- 4.1 Results: Domain effects varied by task: computer-science QA produced a 27.09% relative reasoning improvement, while most domains harmed Figure QA.Astrophysics helped geometry problem solving, and condensed-matter data improved math word problems.
- 4.2 Evaluated Tasks: ArXivCap benchmarks four tasks: single-figure captioning, multiple-figure captioning, contextualized captioning, and paper-title generation.The tasks increase demands on scientific detail, cross-figure reasoning, context use, and semantic generation.
- 4.2 Results: Fine-tuning improved single-figure captioning, raising Qwen-VL-Chat’s BLEU-2 from 4.4 to 8.9, while title generation BLEU-2 rose from 2.6 to 6.7.Adding paper titles improved captioning scores, whereas abstracts provided negligible gains.
- 4.2 Results: Qwen-VL-Chat initially achieved only 3.0 BLEU-2 and 7.2 ROUGE-L on multiple-figure captioning, but ArXivCap training eventually surpassed GPT-4V.The result indicates stronger performance after training on summaries requiring reasoning across multiple images.
- 4.2 Results: Contextualized captioning favored IDEFICS-Instruct-9B, while shuffled contexts caused a 31% ROUGE-L drop for the original model versus 8% after fine-tuning.The findings support the importance of contextual cues and show greater robustness after ArXivCap training.
- 4.3 Analysis: Manual analysis identified oversimplified captions and recognition problems, motivating OCR, metadata, improved perception, and external information.These observations indicate remaining weaknesses in scientific figure comprehension.
5 Conclusion
The paper introduces Multimodal ArXiv through ArXivCap and ArXivQA to advance scientific comprehension in LVLMs. Experiments show improved mathematical reasoning and substantial gains from in-domain training, while error analysis identifies persistent challenges in scientific-figure understanding.
- 5 Conclusion: Multimodal ArXiv combines ArXivCap and ArXivQA to improve LVLM comprehension of scientific figures and literature.The paper evaluates both mathematical reasoning and four vision-to-text tasks.
- 5 Conclusion: Fine-tuning on ArXivQA enhances LVLM mathematical reasoning, while ArXivCap evaluations expose challenges and gains in scientific figure understanding.The conclusion links reasoning improvements with the broader benchmark of scientific vision-to-text capabilities.
Limitations
The paper documents dataset curation and examples supporting Multimodal ArXiv. These materials include caption cleaning, representative figure-caption pairs, domain labels, and visualizations of caption vocabulary.
- Dataset curation: Caption cleaning uses pylatexenc, with Table 9 showing captions before and after processing.The cleaning step is documented as part of dataset preparation.
- Dataset examples: The dataset includes illustrated single-figure, multiple-figure, and ArXivQA examples spanning different figure types and question sets.Figures 7, 8, and 10–13 provide representative visualized cases.
- Dataset analysis: Caption word clouds indicate diverse vocabulary for describing academic figures.The visualization is presented in Figure 9.
- Dataset composition: Table 10 lists the names of the domains represented in the dataset.The supplied passage identifies the table’s purpose but not the individual domain names.
A.5 Quality Analysis of ArXivQA
ArXivQA was manually evaluated across multiple quality dimensions using an annotator grading protocol. Most samples were clear and high quality, although some questions were unanswerable because figures were misread.
- The manual assessment evaluated ArXivQA across six quality aspects using the study’s grading protocol.
- 79 of 100 samples scored at least 5 from both annotators, meeting the study’s stringent quality threshold.Most samples had clear images and unambiguous question and option descriptions.
- Some generated questions were unanswerable because models mis-recognized elements in the figures, lowering factual-alignment scores.
B Evaluation Details
The evaluation compares multiple open-source and proprietary LVLMs on four vision-to-text tasks using standardized prompts. The tasks span single- and multiple-figure captioning, contextualized captioning, and title generation.
- All models receive the same task prompts to support consistent comparison across the evaluation.
- Four tasks test single-figure captioning, multiple-figure summarization, contextualized captioning, and title generation from figure-caption pairs.
B.3 GPT-4 Evaluation of Caption
GPT-4 was used to score sampled single-figure captions alongside automatic metrics. Caption quality remained low overall, while ArXivCap training improved Qwen-VL-Chat’s GPT-4 score.
- The GPT-4 evaluation sampled 500 generated captions and mapped GPT-4’s quality assessment to a 1-to-5 scale.
- 12% improvement in GPT-4 score was achieved by ArXivCap-trained Qwen-VL-Chat over the original model.This model obtained the most favorable GPT-4 evaluation among the compared systems.
- ROUGE-L correlated most strongly with GPT-4 scores at Pearson r = 0.91, followed by BLEU-2 at 0.64 and BERT-S at 0.39.
- Uniformly low GPT-4 scores indicate that the evaluated models struggled to produce satisfactory scientific captions.
C.2 Failure Sample of MathVista
A challenging MathVista geometry problem exposed a failure case in which both evaluated models produced incorrect answers. The authors suggest more focused corpora for geometry and mathematical reasoning.
- Both evaluated models failed on a challenging geometry mathematical-reasoning problem from MathVista.
- The authors suggest incorporating more focused corpora to improve LVLMs’ geometry and mathematical reasoning abilities.
D Results with LLaVA Backbone
Combining ArXivQA with the original SFT data improves LLaVA-v1.5-7B across scientific reasoning benchmarks and overall multimodal evaluation. These results support benefits beyond a single benchmark or model backbone.
- D Results with LLaVA Backbone: The LLaVA experiment mixes ArXivQA with the LLaVA SFT 665K-instruction dataset and follows the original training recipe.LLaVA-v1.5-7B is used as the backbone.
- D Results with LLaVA Backbone: Together with Qwen-VL-Chat results, the findings indicate that ArXivQA benefits different model backbones across multiple benchmarks.The reported gains extend across scientific reasoning and general multimodal evaluation.