Source-linked AI summary
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Yida Yin, Harish Krishnakumar, Chung Peng Lee, Boya Zeng, Wenhao Chai, Shengbang Tong, Wenhu Chen, Hu Xu, Xingyu Fu, Gabriel Sarch, Aleksandra Korolova, Zhuang Liu
TL;DR
Existing multimodal benchmarks often prioritize task categories over visual diversity, limiting coverage of open-ended visual inputs. WorldBench instead builds a taxonomy-guided, visually diverse benchmark with challenging questions, and evaluations show greater visual diversity than existing benchmarks while the strongest model reaches only 64.0% accuracy.
Problem
Task-centric multimodal benchmarks often overlook image diversity, limiting which visual abilities can be evaluated on open-ended inputs.
Method
WorldBench uses a taxonomy of thousands of visual concepts to curate diverse images and iteratively design questions that challenge frontier MLLMs.
Results
WorldBench exhibits greater visual diversity than existing diverse benchmarks and challenges current MLLMs, with the strongest model reaching 64.0% accuracy.
Takeaways & Limitations
The evaluations highlight visual diversity as an important consideration when building multimodal benchmarks.
Takeaways & Limitations
Pre-trained vision encoders provide inconsistent rankings across diversity metrics, limiting their use as general visual-diversity measures.
Abstract
from arXiv · showhide
In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended visual inputs. We present WorldBench, a challenging and visually diverse reasoning benchmark to evaluate Multimodal Large Language Models (MLLMs). We build a taxonomy of thousands of visual concepts across multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. On quantitative and human evaluations, WorldBench achieves higher visual diversity than any existing diverse benchmark. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance-level. We hope our work highlights the importance of visual diversity in building multimodal benchmarks.
1 Introduction
WorldBench evaluates multimodal reasoning through visual diversity rather than task diversity, using a taxonomy-guided image collection paired with human-intuitive but model-challenging questions. Its evaluations show stronger visual diversity than existing benchmarks and substantial weaknesses in current MLLMs.
- Motivation and benchmark design: WorldBench shifts benchmark construction from task diversity to visual diversity, because image variety determines the questions and abilities that can be evaluated.It pairs a diverse image set with questions that are intuitive for humans yet challenging for models.
- Benchmark construction: A taxonomy of thousands of fine-grained visual concepts across 7 visual domains guides curation of high-quality, diverse images from web searches and existing datasets.The taxonomy is semi-automated with an LLM and involves light human effort; curation prioritizes non-iconic images with rich contextual scenes.
- Visual-diversity evaluation: WorldBench consistently ranks among the strongest benchmarks on embedding-based diversity measures and receives the highest overall human-rated diversity score.Quantitative evaluation uses the effective rank and participation ratio of image-embedding feature covariance matrices, while human comparisons are aggregated with the Bradley–Terry model.
- Model evaluation: 64.0%: Gemini-3.1-Pro achieves the highest accuracy averaged across all domains, while 56.6%: Qwen3.5-VL-27B is the best-performing open-source model.WorldBench evaluation covers 15 MLLMs, and several models perform only moderately above chance-level.
- Model limitations: No model exceeds 75% accuracy in any domain, with failures involving fine-grained perception and ungrounded inferences despite reasoning models outperforming non-reasoning variants.Increasing reasoning-token counts does not always improve performance, a pattern observed across all domains.
2 WorldBench
WorldBench is a 2,000-question reasoning benchmark designed to evaluate MLLMs on visually diverse images spanning seven domains. It builds a 2,000-concept taxonomy, curates representative images, and manually designs natural, challenging questions that frontier models often answer incorrectly.
- Benchmark overview: WorldBench spans 7 domains and contains 2,000 questions that challenge frontier MLLMs while remaining natural and intuitive for humans.The domains include Living Things, Objects, Scenes, Digital World, Academics, Documents, Charts & Tables, and Agents.
- Construction procedure: The benchmark construction uses three steps: building a visual-concept taxonomy, curating one high-quality image per concept, and manually proposing difficult questions.Questions are designed so that frontier MLLMs answer them incorrectly.
- Visual taxonomy: Existing taxonomies often emphasize tangible entities or narrow practical domains, whereas WorldBench explicitly includes digital imagery and broader visual content.ImageNet allocates >120 of 1,000 categories to fine-grained dog breeds but covers no digital imagery such as webpages, screenshots, or games.
- Visual taxonomy: WorldBench’s taxonomy covers 2,000 visual concepts across 7 domains, including both real-world content and web-native concepts such as shopping interfaces, games, and Web Agents.Examples range from animals and events to product reviews, supply-and-demand curves, booking a flight, and installing an app.
- Image curation: WorldBench consists of 2,000 images and offers greater visual diversity than comparison benchmarks, which often rely on textbook-style charts, diagrams, or single centered objects.Its images range from real-world objects and scenes to web content and games.
- Question design: The questions are designed to be more natural and closer to real-world settings than questions that chain several independent sub-questions into one complex task.This contrasts WorldBench with the example construction strategy described for ZeroBench.
3 Measuring Visual Diversity of Images
WorldBench’s visual diversity is evaluated quantitatively through embedding covariance and qualitatively through pairwise human comparisons. It achieves strong diversity results across vision encoders and the highest human-rated diversity, though rankings depend on the encoder and human judgments reflect multiple criteria.
- Quantitative evaluation: WorldBench often achieves the highest or second-highest effective rank and participation ratio across SigLIP 2, Perception Encoder, and DINO v3.These metrics measure how broadly embedding variance is distributed across feature-space directions.
- Quantitative evaluation: Diversity rankings vary by encoder: MMBench has the highest participation ratio under SigLIP 2 but ranks fourth under Perception Encoder.This indicates that different encoders capture distinctive visual representations and limits the universality of a single encoder-based metric.
- Human evaluation: WorldBench receives the highest human-rated visual diversity score from pairwise comparisons of benchmark image panels.The study collected 360 comparisons from 12 volunteers, with each user completing 30 rounds.
- Human evaluation: Users judged diversity using semantic and category breadth, chart/photo balance, visual variety, and the absence of duplicate or near-duplicate images.Some criteria, such as fewer duplicates, were considered objective, while semantic content and category judgments depended more on individual interpretation.
4 Model Performance
WorldBench is challenging for both proprietary and open-source MLLMs, with the strongest model reaching 64.0% average accuracy. Reasoning helps initially but can saturate or regress at higher budgets, while WorldBench’s 0.94 average correlation with other benchmarks is the lowest reported.
- Overall performance: 64.0% average accuracy is achieved by the strongest proprietary model, Gemini-3.1-Pro, while Qwen3.5-VL-27B reaches 56.6% as the best open-source model.Qwen3.5-VL-27B also outperforms proprietary models such as Claude-Opus-4.7 and Grok-4.2.
- Overall performance: 73.8% accuracy is reached by Gemini-3.1-Pro on Documents, Charts, & Tables, versus 39.0% for Gemma-4-E4B; no model exceeds 75% on any domain.The 39.0% result is only moderately above chance-level performance, and within-domain performance gaps are substantial.
- Benefits of reasoning: Accuracy consistently improves from no reasoning to a low budget across all domains, but higher budgets cause domain-dependent saturation, regression, or continued gains.Digital World, Objects, Scenes, and Agents continue climbing through the high budget, whereas Documents, Charts, & Tables, Academics, and Living Things saturate or regress.
- Correlation with other benchmarks: 0.94 is WorldBench’s lowest average correlation with other diverse benchmarks, indicating that it evaluates a different dimension of MLLM capability.The correlation is computed from Pearson correlations between model accuracies on WorldBench and MMMU, MMStar, and MMT-Bench.
5 Related Work
Related work covers rapid advances in multimodal large language models and benchmarks targeting specific image-understanding capabilities. MLLMs have improved across visual-understanding benchmarks, while multimodal benchmarks address tasks including perception, OCR, hallucination, GUI navigation, and coding.
- Multimodal large language models: MLLM capabilities have drastically improved over the past years, with proprietary models achieving unprecedented performance across visual-understanding benchmarks.The passage names OpenAI, Google, Anthropic, and xAI as examples of proprietary models.
- Multimodal large language models: Open-source MLLM progress has advanced rapidly from early pioneers to more recent models.The passage cites early pioneers including Dai et al., Awadalla et al., and Liu et al., alongside newer Qwen and GLM models.
- Multimodal benchmarks: Many multimodal benchmarks target specific MLLM image-understanding capabilities, including visual perception, OCR, object hallucination, GUI navigation, and coding.The passage organizes these benchmark areas around specialized capabilities rather than a single general task.
6 Conclusion · A Question Annotation Interface
WorldBench combines a taxonomy of thousands of visual concepts, diverse image curation, and structured trial-and-error question design to challenge MLLMs. Its annotation interface supports image-fact extraction, model evaluation, and LLM-assisted refinement of questions and explanations.
- 6 Conclusion: WorldBench introduces a visually diverse reasoning benchmark designed to challenge MLLMs by representing the visual world broadly.The benchmark is built around seeing the whole world rather than narrowly expanding task types.
- 6 Conclusion: Thousands of visual concepts guide the curation of a diverse image set for WorldBench.The taxonomy spans visual concepts across multiple domains and supports comprehensive image collection.
- 6 Conclusion: A structured trial-and-error process designs questions targeting frontier MLLMs’ failure modes.Questions are manually developed to expose weaknesses that current frontier models fail to handle.
- 6 Conclusion: Quantitative and human evaluations show that WorldBench has greater visual diversity than existing diverse benchmarks.The conclusion reports this advantage across both automated and human assessments.
- 6 Conclusion: WorldBench evaluations reveal that current models find the benchmark very challenging.The authors frame these results as motivation for datasets that prioritize visual diversity.
- A Question Annotation Interface: Annotators can query an LLM to extract image facts or devise questions independently.The interface provides a “Generate Image Facts” button using an LLM such as GPT-5.
- A Question Annotation Interface: For each question, annotators use “Run Model Evaluation” to test it against frontier models.The interface is designed to examine images, propose questions, and evaluate them against frontier models.
- A Question Annotation Interface: An “Improve via GPT-5” function refines question wording, answer choices, and explanations with an LLM.The interface combines automatic refinement with annotator-created questions and model-based evaluation.
B Additional Results
WorldBench’s additional analysis evaluates visual diversity across six widely used pre-training datasets using multiple vision encoders. Encoder rankings differ, likely because their learned representations reflect different pre-training data distributions.
- Pre-training dataset evaluation: The analysis randomly samples 1,000 images from six pre-training datasets and computes their visual diversity with different vision encoders.The datasets are YFCC, CC, DataComp, WIT, LAION, and ImageNet.
- Encoder-dependent rankings: SigLIP2 ranks DataComp as most diverse, whereas DINO v3 ranks ImageNet highest.The differing rankings are likely influenced by differences in the encoders’ pre-training data distributions.
- Pre-training distribution effects: SigLIP2’s WebLI training may align its representations with web-scale datasets such as DataComp, while DINO v3 uses Instagram images and ImageNet.ImageNet comprises 1/10 of DINO v3’s pre-training data, according to the supplied passage.
C User Study Interface for Human Evaluation
The human evaluation interface compares the visual diversity of two benchmarks through side-by-side image panels, allowing users to inspect images and select the more diverse panel.
- Evaluation interface: The interface randomly samples 100 images from each of two benchmarks and displays them side-by-side in a 20x5 image grid.Users can click any image to view it in full resolution.
- Evaluation interface: After examining all images on both sides, users choose the panel they judge to be visually more diverse.
D Implementation Details · D.1 Benchmarks Used in Diversity Evaluation
The diversity evaluation uses carefully selected benchmark images to avoid penalizing datasets for repeated or overly similar content. WorldBench curation also applies licensing and safety reviews before release.
- D.1 Benchmarks Used in Diversity Evaluation: Repeated or highly similar images are removed because they would reduce benchmark visual diversity and disadvantage WorldBench.Some evaluated benchmarks reuse identical images across questions or include multiple similar images within one question.
- D.1 Benchmarks Used in Diversity Evaluation: One image is randomly sampled per question, followed by removal of exact duplicates within each benchmark’s sampled set.
- D.1 Benchmarks Used in Diversity Evaluation: MEGA-Bench excludes in-context-example images and samples only images on which models are actually evaluated.
- D.1 Benchmarks Used in Diversity Evaluation: The evaluation uses VQAv2 validation, MMBench and MMMU test splits, MME, MMStar, SEED-Bench-2, and MEGA-Bench full sets.It also uses the MMT-Bench test split.
- D.1 Benchmarks Used in Diversity Evaluation: WorldBench filters web-search results for Creative Commons-licensed images during curation.
- D.1 Benchmarks Used in Diversity Evaluation: Manual review removes unsafe, private, or otherwise inappropriate images and questions before WorldBench release.
D.2 Model Performance · E Model Responses
WorldBench evaluates proprietary and open-source MLLMs with standardized generation and answer-parsing procedures, then presents representative responses from all 15 evaluated models. The benchmark’s prompt ordering slightly improves accuracy, while its regex parser achieves over 99% average extraction success.
- D.2 Model Performance: The evaluation covers six proprietary model groups and Qwen3.5-VL-Plus-Thinking/Instruct using default sampling, except when a model exposes a reasoning-effort setting.The listed proprietary evaluations include GPT-5.4-Thinking, Gemini-3.1-Pro, Gemini-3-Flash, Claude-Opus-4.7, Grok-4.2, and Qwen3.5-VL-Plus-Thinking/Instruct.
- D.2 Model Performance: Open-source models are evaluated locally with vLLM using default or officially recommended generation parameters, except Kimi-K2.5, which is evaluated through its API.Each local evaluation uses approximately 4–8 NVIDIA A100 GPUs and takes roughly one hour, depending on model size.
- D.2 Model Performance: Asking models to explain briefly before giving the final answer slightly improves accuracy compared with requesting the final answer first.When the answer is requested first, models sometimes change to a different answer during the explanation.
- D.2 Model Performance: The answer parser applies three regex rules sequentially: explicit “Answer: X,” a leading option letter, and then a standalone option letter anywhere in the response.Responses matching none of these patterns are considered incorrect.
- D.2 Model Performance: >99% average parsing success is achieved across all WorldBench questions, eliminating reliance on advanced methods such as LLM-as-a-judge.The response is normalized to uppercase before parsing.
- E Model Responses: Representative responses are presented for all 15 evaluated models across Figures 13–27, with exposed reasoning traces included when available.Some thinking models do not return private reasoning traces through the API.
- E Model Responses: The examples span questions about objects, scenes, digital worlds, documents, charts, tables, and academics, covering visual attributes, spatial relations, text, cell types, and scene interpretation.Representative prompts include identifying scratch orientation, light-source direction, bowling-ball color, text styling, and the lightest-blue cell type.