Source-linked AI summary
ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures
Fahad Ahmed, Sören Auer, Jennifer D'Souza
TL;DR
Scientific figure understanding lacks comprehensive, domain-specific benchmarks that capture end-to-end reasoning over complex ALD/E figures. Sci-ImageMiner introduces an expert-annotated four-task benchmark and competition, showing that structured, task-specific approaches work well overall while data extraction and visual question-answering remain challenging.
Problem
Existing scientific-figure datasets and competitions provide limited domain coverage, quantitative figures, and end-to-end comprehension and reasoning tasks.
Method
The paper introduces Sci-ImageMiner, an expert-annotated ALD/E figure benchmark and competition spanning four complementary end-to-end scientific understanding tasks.
Results
Across four tasks, top-performing approaches favored problem decomposition, data-centric optimization, contextual grounding, and structured task-specific reasoning.
Takeaways & Limitations
The benchmark establishes a rigorous platform for studying domain-specific scientific figure comprehension and identifying challenges that require further research.
Takeaways & Limitations
Significant challenges remain in data extraction and visual question-answering, which require robust multimodal reasoning and precise scientific understanding.
Abstract
from arXiv · showhide
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.
1 Introduction
The introduction frames scientific figure comprehension as multimodal interpretation requiring both visual perception and domain-specific reasoning, and presents Sci-ImageMiner as an expert-annotated ALD/E benchmark addressing current limitations. It motivates the competition by highlighting multimodal models’ difficulty with complex scientific figures and domain-specific reasoning.
- Scientific figure comprehension requires visual perception and domain-specific reasoning to interpret trends, relationships, and meaningful scientific knowledge from research figures.
- ALD and ALE are foundational semiconductor-manufacturing technologies for advanced electronic and functional materials and next-generation nanoelectronics at atomic scale.
- Multimodal models perform strongly on natural images but struggle with complex scientific-figure semantics and abstract representations because domain-specific datasets and reasoning capabilities remain limited.Previous ICDAR competitions contributed significantly but did not fully capture the complexity and end-to-end needs of real-world scientific figure understanding.
- Sci-ImageMiner is introduced as a comprehensive, expert-annotated, domain-specific dataset of ALD/E scientific figures designed to investigate multimodal understanding, improvement strategies, and task difficulty.
2 Sci-ImageMiner Benchmark Dataset
Sci-ImageMiner comprises 205 ALD/E research publications organized by process type and study modality, with figures and structured content preserved for multimodal analysis. A custom platform supports coordinated annotation across four progressive comprehension and reasoning tasks, including subfigure-level VQA guided by Bloom’s taxonomy.
- Dataset construction: The benchmark collects 205 experimental and simulation-based ALD/E publications and develops a 49-figure taxonomy with visualizations and descriptions.The taxonomy is shared through the project’s Github repository.
- Curation workflow: MinerU extracts structured textual content as JSON and high-resolution figures as JPEGs, preserving semantic structure and visual fidelity for downstream multimodal analysis.The end-to-end curation workflow is illustrated in Figure 1.
- Dataset organization: The dataset is organized into ALD and ALE categories, each divided into experimental and simulation-based studies, with standard train/dev/test splits.Dataset statistics are reported in Table 1, while Figure 2 shows the hierarchy and associated PDF, content.json, figure, and annotation files.
- Annotation framework: Because the large, complex dataset requires coordinated annotation across four progressive tasks, the authors developed a custom multi-user web platform for fine-grained end-to-end labeling.The platform is hosted on institutional infrastructure because existing annotation tools were insufficient for the domain-specific workflow.
- Annotation framework: For VQA, annotators assign four subfigure-level question–answer pairs per figure, using Bloom’s revised taxonomy to assess progressively deeper understanding across scientific question types.The question types cover process-oriented mechanisms and workflows, comparative or trend-based quantitative relationships, and structure–property connections.
3 Competition
The competition was designed to evaluate comprehensive, end-to-end scientific figure understanding and reasoning across complementary tasks. It used separate development and evaluation phases with structured submission and leaderboard management.
- Competition organization: The competition progressed from a development phase to final evaluation and rankings, using CodaBench tracks for scalable submission and evaluation management.
- Competition purpose: The competition collectively evaluated end-to-end comprehensive scientific figure understanding and reasoning capabilities.
- Competition tasks: The four tasks covered figure classification, chart-data extraction into Markdown tables, concise factual summarization, and domain-specific reasoning over expert-annotated question-answer pairs.
- Evaluation protocols: Classification used Accuracy, Precision, Recall, and F1-score with F1-score as the primary ranking metric, while data extraction used RMS and TEDS.
- Evaluation protocols: Summarization used ROUGE-1/2/L and BERTScore-F1, while reasoning was evaluated separately for paragraph, factoid, list, and yes/no answers with task-specific metrics.
4 Results and Discussion · 4.1 Overview · 4.2 Baseline Results
The competition drew substantial international engagement, while baseline evaluations showed that scientific figure comprehension remains difficult for current LVLMs across tasks. No single baseline performed consistently well, underscoring the need for domain-specific multimodal reasoning.
- 4.1 Overview: 68 active participants contributed 1,263 public and private submissions to the ICDAR Sci-ImageMiner competition.These participation and submission statistics are reported in Table 3.
- 4 Results and Discussion: The competition attracted substantial engagement from participants around the globe.
- 4.1 Overview: The full competition leaderboards were made available on the competition website.
- 4.2 Baseline Results: Baselines evaluated a range of LVLMs, including Gemma 4 E4B 8b, Qwen3-VL-8B-Instruct [14], GLM-4.6V-Flash [15], and Intern VL 3.5 8b [17].
- 4.2 Baseline Results: The baseline evaluations used prompts 9 consistently across tasks.
- 4.2 Baseline Results: No single baseline performed well across all tasks, demonstrating the intrinsic difficulty of scientific figure comprehension and the need for advanced domain-specific multimodal reasoning.
4.3 Competition Results · Task 1: Classification Results
Section 4.3 summarizes the methodologies of the top-five leaderboard teams, while Task 1 reports their classification leaderboard and presents the best baseline alongside the top-five teams. The described approaches include hierarchical multi-model fine-tuning, contextual augmentation with iterative prompting and multi-agent consensus, and other systematic experimental strategies.
- 4.3 Competition Results: The competition results summarize methodologies proposed by the top-five leaderboard teams, with an overview of their approaches provided in Table 6.The summary is based on the availability of corresponding method reports.
- Task 1: Classification Results: Task 1 presents the top-five submissions and their classification leaderboard in Table 4.Table 4 covers classification scores for the best baseline and top-five teams, but the supplied passage does not provide the numerical scores.
- Task 1: Classification Results: Ricoh_SRCB uses hierarchical fine-tuning from a coarse 6-class model to a refined 49-class model, supplemented by auxiliary binary and 4-class classifiers.Final predictions are obtained by fusing outputs from these models.
- Task 1: Classification Results: DocMiner enriches inputs with category-specific sample documents, image captions, and source text to improve semantic understanding.Its method also iteratively optimizes prompts using development-set feedback and refines category descriptions to resolve ambiguities between visually similar classes.
- Task 1: Classification Results: DocMiner further introduces multi-agent inference that reconciles independent predictions through consensus or re-evaluation to mitigate individual-model bias.The supplied passage identifies consensus and re-evaluation as the mechanisms for reconciling predictions.
- Task 1: Classification Results: VLMinators’ approach is described as adopting a systematic experimental strategy, although the supplied passage does not provide further methodological details.This team is identified among the Task 1 top-five approaches.
Task 2: Data Extraction Results
Task 2 evaluates data-extraction systems through a leaderboard covering the best baseline and top-five teams. The leading approaches combine structure-aware training, parameter-efficient fine-tuning, multimodal input design, architectural enhancements, and auxiliary textual context.
- Task 2: Data Extraction Results: TeleOCR-VL combines structure-aware training, synthetic data generation, and ensemble inference, separating structural learning from content recognition.Separate training on structure-only and full-content samples targets layout understanding and semantic extraction, while synthetic data addresses annotation scarcity and generalization.
- Task 2: Data Extraction Results: VLMinators uses QLoRA fine-tuning of Qwen2.5-VL-7B-Instruct with prompt constraints, chart semantics, and chart-type context injection for structured table generation.The method focuses on parameter-efficient adaptation and uses axis labels, legends, and metadata to guide structural predictions.
- Task 2: Data Extraction Results: Ricoh_SRCB combines full charts with cropped subgraphs in multi-image prompts and applies category-wise augmentation to improve coverage across chart types.The representation captures global and local visual patterns while data balancing addresses class imbalance and supports generalization.
- Task 2: Data Extraction Results: Vassilis Sioros uses Qwen3.5-9B with early vision-language fusion, Gated Delta Networks, sparse Mixture-of-Experts, and document-context inputs parsed into structured Markdown.Its inputs combine cropped subfigures with surrounding source text, and dense table strings are subsequently parsed into structured Markdown.
- Task 2: Data Extraction Results: DocMiner integrates captions and surrounding text with a multi-agent architecture that dynamically routes different question types to specialized agents.The auxiliary textual information is intended to enhance semantic grounding, while routing enables task-specific processing.
- Task 2: Data Extraction Results: Table 5 reports Task 2 data-extraction scores for the best baseline and top-five teams.The supplied passage identifies the leaderboard scope but does not provide individual scores or establish a ranking comparison.
Task 3: Summarization Results
Task 3 presents the approaches used by the top-five submissions and reports their ranking in the task leaderboard, including DeepVitminC’s modular summarization and visual question-answering framework.
- Task 3: Summarization Results: The task leaderboard in Table 7 presents the rankings of the top-five submissions.
- Task 3: Summarization Results: DeepVitminC combines context extraction from document information, primarily figure captions, with consistent visual-context alignment for task-specific outputs.It uses retrieval-augmented generation to inject contextual information into the model input and jointly processes visual and contextual signals.
Tasks 4: Visual Question-Answering Results
The VQA leaderboard compares the best baseline with the top-five teams, whose approaches combine contextual retrieval, supervised multimodal fine-tuning, structured prompting, and cross-task information flow.
- Top-submission approaches: DeepVitminC retrieves figure captions and surrounding document text as complementary context, then uses enriched prompts and supervised fine-tuning to train a vision-language model with a Consistent Alignment Module.
- Top-submission approaches: Ricoh_SRCB constrains answer type, output structure, and length through prompt design and adaptive decoding limits.
- Top-submission approaches: Sequential context chaining incorporates upstream classification, table extraction, and summarization outputs into VQA prompts, improving numerical precision and contextual reasoning for factoid and paragraph-based queries.
- Leaderboard: The VQA leaderboard presents scores for the best baseline and top-five teams.
4.4 Discussion
The discussion compares participating teams across all four competition tasks and identifies methodological trends behind strong performance. Overall, robust results arise from combining data design, contextual input engineering, model ensembling, and task-specific pipeline or training strategies.
- Comparative analysis: Across all tasks, the analysis highlights consistently strong approaches and finds that combining data design, contextual input engineering, and model ensembling yields robust performance.The comparison focuses on effective models and strategies across the competition rather than a single task.
- Task 1: Classification: Ricoh_SRCB leads classification with hierarchical multi-stage, multi-granularity decomposition and dual-image inference, while IIT_Patna_CV_1 uses Qwen2.5-VL-7B-Instruct, prompt design, and test-time augmentation.The approaches contrast structured task decomposition with generative reformulation and robustness-oriented inference.
- Task 2: Data Extraction: TeleOCR-VL leads data extraction through structure-aware training, synthetic data augmentation, and ensemble inference, while VLMinators remains competitive with lightweight fine-tuning and prompt-based context injection.These results emphasize data-centric and system-level optimization alongside the strength of pretrained models.
- Task 3: Summarization: Summarization leaders combine structured pipelines, retrieval-augmented context, alignment methods such as DPO with hard negatives, and robust multi-candidate consensus re-ranking.DeepVitminC, Ricoh_SRCB, and TeleOCR-VL exemplify complementary advances in training and inference strategies.
- Task 4: Visual Question-Answering: Visual question-answering balances pipeline engineering and training optimization, with retrieval-grounded context, iterative prompt refinement, agent-based reasoning, and sophisticated pipeline design supporting strong results.DeepVitminC emphasizes grounding, DocMiner adaptability, and Ricoh_SRCB shows limited gains from purely training-based improvements.
Comparative Analysis and Key Takeaways
Across all tasks, top-performing approaches combine strong vision-language backbones with pipelines integrating context, structure, and task-specific constraints, while task-specific strategies influence performance.
- Overall Findings: Top-performing approaches combine Qwen-based vision-language models with pipelines that integrate context, structure, and task-specific constraints.The passage identifies Qwen2.5-VL-7B-Instruct and Qwen3.5-9B as predominant examples.
- Task 1: Task 1 favors explicit decomposition and ambiguity handling.
- Task 2: Task 2 highlights data-centric design and ensembling.
5 Conclusions
The ICDAR Sci-ImageMiner competition establishes the first benchmark dataset and competition for scientific comprehension and reasoning over ALD/E figures, using expert annotations to capture domain complexity. Strong global participation and novel methods underscore interest, while results show persistent challenges in data extraction and visual question-answering.
- Sci-ImageMiner is the first benchmark dataset and competition dedicated to scientific comprehension and reasoning over ALD/E figures.
- Domain-expert annotations reflect the complexity and specificity of ALD/E scientific content.
- The competition attracted substantial global participation across all four task categories, and teams proposed novel methodologies and approaches.
- Results reveal significant remaining challenges, particularly in data extraction and visual question-answering.