Source-linked AI summary

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Mohamed Hefeeda, Ehsaneddin Asgari

arXiv:2606.05531v1cs.CVcs.AIcs.CLcs.LG

TL;DR

Existing VLM benchmarks often fragment multimodal evaluation and provide limited coverage of cognitive depth. BloomBench addresses this gap with a bilingual, Bloom-grounded benchmark and hybrid-validated generation pipeline, finding strong semantic understanding but weaker factual recall, creative synthesis, and Arabic performance.

  • Problem

    Existing evaluations use disconnected tasks and do not comprehensively diagnose multimodal cognitive abilities across reasoning levels.

  • Method

    BloomBench defines a hierarchical six-level multimodal taxonomy and constructs bilingual tasks through a scalable semi-automated pipeline with hybrid validation.

  • Results

    Models show near-ceiling Understand and Evaluate performance above 0.88 RAE but substantially weaker Apply and Create performance, with additional English-Arabic degradation.

  • Takeaways & Limitations

    Cognition-driven bilingual evaluation exposes reasoning asymmetries and cross-lingual gaps that aggregate multimodal performance can mask.

  • Takeaways & Limitations

    GPU infrastructure and proprietary API costs restricted model coverage, while manual validation was limited to a stratified subset rather than every dataset item.

Abstract

from arXiv · show

Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence. Most existing evaluations focus on piecemeal or disconnected tasks, obscuring critical cognitive weaknesses and providing little insight for targeted improvement. To address this gap, we introduce BloomBench, part of the Almieyar benchmarking series, the first cognitively human-grounded, bilingual (English-Arabic) multimodal benchmark for VLMs. Grounded in Bloom's Taxonomy, BloomBench systematically evaluates six levels of cognition (Remember, Understand, Apply, Analyze, Evaluate, Create) through carefully designed image-question-answer tasks. Built with a semi-automated pipeline and validated through a stratified hybrid quality assurance protocol, it ensures scalability, cultural inclusivity, and linguistic fidelity. Leveraging this framework, we conduct a comprehensive study of state-of-the-art VLMs to diagnose their cognitive profiles. Our analysis reveals a sharp cognitive asymmetry: while state-of-the-art models achieve strong performance ceilings in semantic understanding, they struggle substantially with factual recall and creative synthesis. This demonstrates that current general multimodal proficiency masks deeper limitations in specific cognitive layers. Furthermore, our study highlights a critical performance gap between Arabic and English, exposing limitations in current cross-lingual multimodal reasoning. These findings establish a foundation for developing more cognitively aligned and inclusive VLMs. The benchmark framework and dataset is available at: https://github.com/qcri/Almieyar-Oryx-BloomBench.

1 Introduction

Existing VLM evaluations often fragment multimodal abilities across disconnected tasks, limiting diagnosis of cognitive depth. BloomBench addresses this gap with a bilingual benchmark grounded in Bloom’s Taxonomy and designed for comprehensive, interpretable reasoning assessment.

  • Conventional benchmarks may not capture human-level task complexity or transfer across distinct cognitive skills.
  • Bloom’s Taxonomy organizes cognition from Remember through Create, enabling multimodal evaluation beyond surface-level accuracy.
  • BloomBench introduces a bilingual multimodal benchmark explicitly grounded in Bloom’s Taxonomy to evaluate multiple cognitive-complexity levels.
  • The benchmark develops tasks that measure reasoning depth rather than performance on disconnected tasks.
  • BloomBench combines scalable generation, LLM-as-a-judge and human validation, answer-based and likelihood-based evaluation, and English-Arabic assessment.

2 Related Works

Prior multimodal benchmarks address important specialized capabilities but remain fragmented across tasks and modalities. BloomBench extends cognitively structured evaluation to bilingual vision-language reasoning across the full Bloom spectrum.

  • Widely used benchmarks provide scalable headline scores across heterogeneous tasks, often using multiple-choice or short-answer formats.
  • VLM benchmarks have expanded beyond VQA and captioning to visual arithmetic, geometry, spatial reasoning, and broader evaluation settings.
  • Arabic multimodal resources have grown through image-text alignment and multilingual datasets, alongside culturally and dialectally aware benchmark efforts.
  • Text-only Bloom analyses find that evaluations concentrate on Apply and Analyze while under-representing Remember, Evaluate, and Create.
  • Existing Bloom-oriented efforts remain restricted to text and code modalities, leaving cognitive-depth evaluation in vision-language processing insufficiently covered.

3 BloomBench Methodology

BloomBench translates Bloom’s cognitive hierarchy into a fine-grained, bilingual multimodal benchmark spanning six reasoning levels. Its semi-automated pipeline combines cognitively grounded generation with hybrid validation to support scalable, reliable coverage.

  • 3.1 Design Principles: BloomBench decomposes six cognitive levels into sub-levels and task types for comprehensive, fine-grained VLM capability assessment.
  • 3.1 Design Principles: The benchmark uses cognitive completeness, hierarchical dependence, real-world generalization, scalability, and transparency as design principles.
  • Data Generation Pipeline: The pipeline combines scenario ideation, VQA generation, multiple-choice conversion, translation, and hybrid quality validation.
  • Data Generation Pipeline: The data-generation pipeline integrates LLMs with human oversight to produce cognitively grounded, visually relevant, bilingually validated items.
  • Data Generation Pipeline: Each open-ended VQA pair becomes a four-option Arabic-translated multiple-choice question with plausible and trap distractors.
  • Quality Validation: Hybrid validation filters samples with an LLM judge and statistically validates a stratified subset covering all 106 taxonomy leaf nodes.
  • Dataset Statistics: 7,747 bilingual image-question-answer pairs span 106 task types across Remember, Understand, Apply, Analyze, Evaluate, and Create.

4 Evaluation

BloomBench evaluates representative open- and closed-source VLMs using answer extraction and likelihood-based scoring. It reports accuracy overall, with Micro and Macro measures accounting for class imbalance across cognitive levels.

  • Setup: The evaluation covers representative open-source and closed-source VLMs, including multiple Gemma 3 sizes, Qwen models, and GPT-4o mini.Multiple Gemma 3 sizes support analysis of model-scale effects on cognitive performance.
  • Evaluation methods: Regex-based Answer Extraction parses free-form outputs for answer choices and assigns an incorrect choice when no valid format is produced.This reflects direct interpretation of generated responses while accounting for instruction-following failures.
  • Evaluation methods: Likelihood-based Scoring normalizes each choice’s conditional log-probability by its token count and selects the choice with the highest normalized score.The normalization makes comparisons fair across choices of different lengths.
  • Evaluation methods: Likelihood-based scoring measures model confidence more directly by avoiding parsing and formatting errors in final answers.The method is intended to expose reasoning signals that conventional accuracy evaluations may overlook.
  • Metrics: Accuracy is the primary metric, calculated against ground-truth answers, and results are reported using both Micro and Macro accuracy.Macro accuracy accounts for class imbalance across cognitive levels; closed-source models are marked N/A for likelihood-based scoring when unsupported.

5 Discussion

BloomBench reveals that current VLM performance is uneven across evaluation methods, languages, and cognitive levels. Models show strong semantic and discriminative abilities but persistent weaknesses in likelihood calibration, factual recall, procedural reasoning, and creative synthesis.

  • Coverage Comparison with Existing Benchmarks: 66.4% of MMMU samples map to Analyze, while Create and Evaluate together represent under 1.1%.Forty-five BloomBench taxonomy leaf nodes receive no MMMU representation, including Ambiguity Resolution, Toxicity Detection, and Dialogue Generation.
  • Model and Cross-Linguistic Effects: 89.8% English and 87.6% Arabic Regex-based accuracy make Gemma 4 31B state of the art, yet it struggles under LBS.The Gemma 3 family shows stronger cross-lingual consistency, while the 4B variant loses substantially more performance than the 12B and 27B variants.
  • Comparison Between RAE and LBS Evaluation: 0.883 RAE accuracy for Gemma 3 27B falls to 0.336 under LBS, whereas Qwen2.5-VL declines from 0.869 to 0.654.The divergence suggests that RAE and LBS capture complementary aspects of model performance: answer behavior and underlying reasoning confidence.
  • Model Size Effect on Performance: RAE performance scales as 27B > 12B > 4B, but Gemma 3 shows inverse scaling under LBS.The largest Gemma 3 model yields the lowest likelihood scores, suggesting a trade-off between deterministic answer generation and raw probabilistic calibration.
  • Cognitive-Level Analysis: Understand and Evaluate exceed 0.88 RAE, while Apply and Create degrade significantly and Remember performs poorly under LBS.This cognitive asymmetry indicates that strong discriminative visual reasoning does not transfer reliably to procedural, recall, or generative tasks.
  • Cross-Linguistic Performance: English performance generally exceeds Arabic performance, with Create showing consistent cross-lingual degradation across evaluated model families.Understand transfers most effectively, while high-order generative synthesis remains a pronounced Arabic bottleneck; Arabic tokenization also disproportionately penalizes LBS.

6 Conclusion

BloomBench is a cognitively informed bilingual benchmark that evaluates VLM reasoning across Bloom’s six cognitive levels. Its semi-automated, hybrid-verified design supports systematic assessment while exposing persistent weaknesses in current models.

  • 6 Conclusion: BloomBench evaluates VLMs across Remember, Understand, Apply, Analyze, Evaluate, and Create using Bloom’s hierarchical taxonomy.The benchmark integrates cognitive theory with multimodal task evaluation to support systematic and interpretable assessment.
  • 6 Conclusion: A semi-automated, hybrid-verified pipeline enables scalable and quality-controlled bilingual benchmark construction.The framework combines automated generation with validation to support reliable evaluation across English and Arabic.
  • 6 Conclusion: BloomBench analyses reveal both strengths and persistent weaknesses in current VLMs across cognitive levels.The findings support cognition-driven evaluation as a way to make multimodal reasoning performance more interpretable.

7 Limitations

The benchmark has scope, format, and validation limitations that constrain model coverage, open-ended reasoning assessment, and full-dataset verification.

  • Model coverage: Resource constraints restricted the evaluated model set, excluding several noteworthy recent VLMs.The authors selected models across architectural families and scales but recommend broader future coverage.
  • Future expansion: Future benchmark iterations may improve as stronger agents, newer models, and more computing resources become available.The pipeline relies on agent-driven generation over rich public web data.
  • Question format: BloomBench uses only multiple-choice questions, which may not capture open-ended and real-world reasoning abilities.Future versions could add fill-in-the-blank, short-answer, and multi-step reasoning formats.
  • Validation scope: Validation covered a stratified subset of approximately 1,000 items rather than every item in the full dataset.The authors report over 97% agreement across all 106 taxonomy nodes.

8 Ethical Considerations.

BloomBench addresses data provenance and content safety by linking to publicly available images rather than redistributing them and filtering unsafe web-crawled content.

  • Data provenance: The dataset releases image URLs and a download script instead of hosting or redistributing image files.This approach retrieves content from original public repositories to respect copyright and distribution restrictions.
  • Content safety: Automated safety classifiers and manual inspection remove images containing offensive or violent content.Users are still advised to exercise standard caution with web-crawled data.

A Full BloomBench Taxonomy

BloomBench organizes multimodal evaluation through Bloom’s six cognitive levels and decomposes them into fine-grained task categories, including activities and attributes.

  • Task categories: Activity Recognition includes individual activities, interactions, and professions.
  • Task categories: Attribute Recognition includes artistic style, color, shape, size, and texture.
  • Hierarchical structure: The taxonomy spans Remember, Understand, Apply, Analyze, Evaluate, and Create, each decomposed into sub-levels and specific task types.This hierarchical structure is intended to provide comprehensive and fine-grained assessment of VLM capabilities.
  • Benchmark statistics: The benchmark distributes items across Bloom’s taxonomy categories and provides separate hierarchical statistics for each cognitive level.The reported tables cover Remember, Understand, Apply, Analyze, Evaluate, and Create.

C Cross-Lingual LBS Ablation Study

A Spanish ablation separates likelihood-based scoring effects from genuine reasoning performance by comparing deterministic accuracy with likelihood scores under different tokenization conditions.

  • Deterministic accuracy: 84.91% Spanish RAE remained close to 87.22% English RAE, a gap of 2.31 percentage points.The comparison uses Qwen2.5-VL-7B on a stratified subset of 908 samples.
  • Likelihood-based scoring: Spanish LBS fell 8.26 points below English despite near-parity in RAE.The dissociation supports sensitivity of LBS to lower non-English probability priors rather than genuine reasoning differences.

D MMMU Taxonomy Coverage Analysis

Mapping MMMU samples onto BloomBench reveals substantial gaps in cognitive coverage, especially for higher-order and socially grounded tasks. The comparison illustrates the kinds of visually grounded, Bloom-aligned tasks used to characterize those gaps.

  • Coverage gaps: 45 taxonomy leaf nodes received zero coverage in MMMU, while Create and Evaluate together comprise only ≈1.0% (11/1,080) of samples.The analysis excludes 94 unmatched samples from leaf-level counts.
  • Coverage gaps: MMMU lacks task types for contextual inference, harm and safety evaluation, and creative generation.Examples include ambiguity resolution, toxicity detection, dialogue generation, and experiment design.
  • Benchmark comparison: MMMU provides stronger coverage of expert domain knowledge and analytical reasoning than of creative, evaluative, and socially grounded multimodal cognition.The comparison positions BloomBench as broader rather than as a replacement for MMMU’s existing strengths.
  • BloomBench framework: BloomBench organizes multimodal assessment across Remember, Understand, Apply, Analyze, Evaluate, and Create, with tasks aligned to specific cognitive abilities.The framework includes leaves and hierarchical paths such as Analyzing → Differentiating → Discriminating → Pattern Recognition.
  • Task design: The generation process targets visually rich, minimally textual scenes and requires image-grounded, Bloom-aligned, deterministic question-answer pairs.Examples include identifying a modern smartphone as an anachronism and inferring intended contents from contrasting storage designs.
  • Quality assurance: Human reviewers verify image alignment, question integrity, choice validity, and English-Arabic translation fidelity before samples are retained.Flagged samples are removed or revised, and concurrency controls ensure each question is reviewed once.

H Detailed Results of Benchmarking

The detailed results are organized as model-performance comparisons using two extraction and scoring procedures across BloomBench’s six cognitive categories. The supplied tables cover Analyze, Apply, Create, Evaluate, Remember, and Understand for both LBS and RAE.

  • Likelihood-based Scoring: Tables 12–17 report Likelihood-based Scoring performance for Analyze, Apply, Create, Evaluate, Remember, and Understand.Each table corresponds to one BloomBench cognitive category.
Loading 2606.05531v1…