Source-linked AI summary

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

Chuer Chen, Zichen Wang, Yi He, Zhengxi Yu, Nan Cao

arXiv:2609.02502v1cs.CVcs.AI

TL;DR

T2I models perform well on literal rendering but remain insufficiently evaluated for visual metaphor generation, which requires blending distinct conceptual domains. The paper introduces VMetaphor-Bench and a hybrid MLLM-as-judge framework, then evaluates 11 models and finds persistent difficulty with compositional structuring and cross-domain mapping, even among the strongest proprietary systems.

  • Problem

    Existing T2I benchmarks largely test literal fidelity, leaving the cross-domain conceptual blending required for visual metaphor generation largely unexamined.

  • Method

    VMetaphor-Bench contains 1,500 structured visual metaphors with two prompt granularities and combines 9,594 MCQs across four fidelity levels with three-dimension scoring.

  • Results

    Evaluation of 11 representative T2I models shows proprietary models substantially outperform open-source counterparts, yet even the strongest models struggle with compositional structuring and cross-domain mapping.

  • Takeaways & Limitations

    Visual metaphor generation remains an important frontier for T2I research because compositional structuring and cross-domain mapping remain unresolved aspects of metaphorical expression.

  • Takeaways & Limitations

    The benchmark has uneven thematic-category coverage, relies on MLLM judges for subjective judgments, and is predominantly English and rooted in Western creative traditions.

Abstract

from arXiv · show

Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, remains largely unexamined. To bridge this gap, we introduce VMetaphor-Bench, the first benchmark for evaluating visual metaphor generation in T2I models. It comprises 1,500 visual metaphors curated from real-world creative imagery, organized into three levels and ten categories, with each sample paired with two prompts of differing specificity. For evaluation, we develop a hybrid framework within an MLLM-as-judge paradigm, combining a multiple-choice question (MCQ) based protocol of 9,594 questions across four levels of metaphorical fidelity with a dimension-based scoring protocol along three perceptual dimensions. Extensive evaluation of 11 representative T2I models reveals that even the strongest proprietary models struggle with compositional structuring and cross-domain mapping, key aspects of metaphorical expression, highlighting visual metaphor generation as an important frontier for future T2I research.

1 Introduction

Visual metaphor generation requires coherent cross-domain blending, meaningful element mappings, and communication of intended meaning, but existing T2I benchmarks largely test literal prompt following. VMetaphor-Bench addresses this gap and evaluation shows that even leading models struggle with compositional structuring and cross-domain mapping.

  • Motivation: Visual metaphors combine source and target domains through visual analogy, requiring coherent composition, element-level mappings, and intended meaning.Models may omit a domain, use the wrong structure, blend domains awkwardly, or fail to convey the intended meaning.
  • Motivation: Existing T2I benchmarks emphasize object, attribute, spatial, factual, or physical fidelity, leaving cross-domain conceptual blending largely unexamined.Prior metaphor work focuses mainly on understanding, while a small generation probe used only 300 samples and limited metrics.
  • Benchmark: VMetaphor-Bench contains 1,500 visual metaphors organized into three levels and ten categories, with structured annotations and conceptual and descriptive prompts.Annotations include source and target domains, structure type, metaphorical meaning, and element-level mappings.
  • Evaluation: The hybrid evaluation uses 9,594 MCQs across four metaphorical-fidelity levels and dimension scoring across three perceptual dimensions.The framework is implemented with MLLM-as-judge evaluation for fine-grained diagnosis.
  • Findings: Evaluation of 11 representative T2I models finds that proprietary systems outperform open-source counterparts, while even the strongest models struggle with structure and mapping.These difficulties concern key aspects of metaphorical expression and motivate further T2I research.

2 Related Works

T2I research has progressed through diffusion, autoregressive, and unified multimodal architectures, alongside increasingly semantic evaluation methods. Visual metaphor research has mainly addressed understanding rather than generation, leaving systematic generation evaluation limited.

  • T2I Models: T2I generation has advanced through diffusion, autoregressive, and unified multimodal architectures.These directions differ in their generation mechanisms and compositional-control properties.
  • Evaluation: Early T2I evaluation used FID and CLIPScore, which capture perceptual quality but not fine-grained semantic correctness.MLLM-as-judge evaluation allows a vision-language model to assess generated images against prompts more flexibly.
  • Visual Metaphor Research: Visual metaphor benchmarks have predominantly targeted classification, localization, interpretation, captioning, and visual question answering.MetaCLUE includes only a small generation probe, while ImageMet and MetaphorStar focus on understanding or reasoning.

3 VMetaphor-Bench

VMetaphor-Bench is built from curated and manually verified creative imagery with structured metaphor annotations and prompts at two granularities. Its hybrid MLLM-as-judge framework evaluates both metaphorical fidelity and expressive quality through complementary MCQ and dimension-based protocols.

  • Benchmark Overview: VMetaphor-Bench evaluates cross-domain conceptual blending in T2I models using a hierarchical benchmark and multi-dimensional assessments.The complete pipeline covers dataset construction and evaluation framework stages.
  • Dataset Construction: Approximately 5,000 candidate images are collected from Pinterest and filtered through review to construct the real-world visual-metaphor source pool.The collection targets creative content such as advertisements, editorial illustrations, and conceptual art.
  • Dataset Construction: A two-stage annotation pipeline extracts structured metaphor annotations before generating descriptive and conceptual prompts at different granularities.The conceptual prompt states domains and intended meaning without specific visual details or element mappings, while the descriptive prompt details visual realization.
  • Dataset Construction: Manual inspection and cross-validation remove ambiguous images and correct inaccurate annotations and prompts, yielding 1,500 final images.Three trained annotators inspect the images, annotations, and prompts.
  • Evaluation Framework: The hybrid evaluation combines MCQ-based fidelity assessment with 1–5 dimension-based scoring of expressive quality.MCQs assess metaphorical correctness, while dimension scoring captures overall perceptual quality.
  • Evaluation Framework: The MCQ protocol produces 9,594 manually verified questions covering domain presence, structure type, meaning, and element mapping.Per-level accuracies are computed and averaged into an overall MCQ score.
  • Evaluation Framework: Dimension scoring rates metaphoric efficacy, metaphor logic, and perceptual harmony on a 1–5 Likert scale, with the overall score averaged across dimensions.Together with MCQ evaluation, these dimensions assess metaphorical fidelity and expressive quality.

4 Experiments

Experiments evaluate 11 representative T2I models on VMetaphor-Bench, compare conceptual and descriptive prompts, and test evaluator reliability. Proprietary models lead overall, but structure and mapping remain major bottlenecks, while descriptive prompts especially help weaker models.

  • Main results: 85.4% MCQ accuracy and a 4.30 dimension score make GPT Image 1.5 the strongest model, ahead of Nano Banana 2 at 84.8% and 4.11.FLUX.2-dev, the best open-source model, scores 76.0% MCQ accuracy and 3.62 dimension score.
  • Main results: Structure and Mapping are the weakest fidelity levels, with accuracies of 48.1–78.1% and 49.1–76.5%, respectively.Seedream 5.0 Lite illustrates the gap by reaching 90.6% Domain and 96.0% Meaning but only 65.6% Structure and 66.5% Mapping; replacement is especially difficult.
  • Main results: Perceptual Harmony scores highest across models at 3.43–4.38, while Metaphoric Efficacy and Metaphor Logic generally trail behind.OmniGen2 has a Harmony–Efficacy gap of 0.92, compared with 0.11 for GPT Image 1.5.
  • Effect of prompt granularity: Descriptive prompts improve every model’s MCQ accuracy, while most also improve dimension scores, with weaker models gaining more.GPT Image 1.5 gains +3.1 percentage points in MCQ overall and −0.06 in dimension score, whereas Z-Image gains +16.6 and +0.69.
  • Effect of prompt granularity: Descriptive prompts narrow the dimension-score gap between GPT Image 1.5 and Z-Image from 1.09 under conceptual prompts to 0.34.The corresponding scores are 4.30 versus 3.21, and 4.24 versus 3.90.
  • Evaluator reliability: GPT-5.4 and Qwen3.5-27B produce highly consistent rankings, with Spearman’s ρ of 0.991 for MCQ and 0.936 for dimension scores across all 11 models.Human alignment exceeds 80% balanced accuracy overall, while Qwen3.5-27B’s dimension-score MAE remains below 0.8 on a 1–5 scale.

5 Conclusion

The paper introduces VMetaphor-Bench and uses it to show that proprietary T2I models outperform open-source counterparts, while even the strongest models struggle with compositional structuring and cross-domain mapping. The benchmark is intended to support research bridging literal rendering and meaning-rich visual generation.

  • Conclusion: VMetaphor-Bench is the first benchmark for visual metaphor generation in T2I models, containing 1,500 structured visual metaphors across three levels and ten categories.Its hybrid evaluation combines MCQ-based fidelity assessment with dimension-based scoring.
  • Conclusion: Across 11 evaluated models, proprietary systems substantially outperform open-source counterparts, but compositional structuring and cross-domain mapping remain difficult even for the strongest models.These capabilities are identified as key aspects of metaphorical expression.
  • Conclusion: The benchmark is positioned as a foundation for future research on bridging literal rendering and creative, meaning-rich visual generation.

A.1 Benchmark Statistics

The benchmark contains 1,500 images organized into ten thematic categories across three levels, with conceptual and descriptive prompts differing substantially in length. Table 5 summarizes prompt-length and category-distribution statistics.

  • Dataset statistics: 1,500 images are organized into 10 thematic categories across three levels: Individual contains 495 images and External contains 499.Individual covers personal cognition, emotion, health, and growth; External covers societal, environmental, technological, and interpersonal themes.
  • Prompt statistics: Conceptual prompts average 25.6 words, while descriptive prompts average 65.0 words; adding a style suffix roughly doubles both lengths.

A.2 Experimental Setup Details

All experiments run on a single server equipped with four NVIDIA H200 GPUs, with model and evaluator details provided separately in Table 6.

  • Hardware: All experiments are conducted on a single server equipped with 4 × NVIDIA H200 GPUs.Detailed information about evaluated models and MLLM evaluators appears in Table 6.

A.3 Per-Structure Analysis

The structure analysis shows that models handle fusion and juxtaposition considerably better than replacement, which is especially difficult because it requires inferring a missing element.

  • Analysis setup: The analysis constructs a confusion matrix by comparing MCQ-predicted structure types with ground-truth annotations across all 11 evaluated models.The L2 questions distinguish fusion, replacement, juxtaposition, and cannot be determined.
  • Structure-type performance: 67.9% fusion accuracy and 69.2% juxtaposition accuracy contrast with only 18.9% replacement accuracy.For replacement metaphors, models instead predicted fusion 37.8% of the time or separate-object juxtaposition 36.6% of the time.
  • Structure-type performance: Replacement errors are concentrated in predicting fusion or juxtaposition rather than substituting one domain for the other.This pattern indicates a specific difficulty with the replacement structure type.
  • Interpretation: The observed difficulty hierarchy mirrors visual rhetoric theory, which places replacement above fusion and juxtaposition in cognitive complexity.Replacement requires viewers to infer the missing element, whereas fusion and juxtaposition retain both domains visibly.

A.4 Full Results with Descriptive Prompts

Descriptive prompts generally improve evaluation results by providing more explicit visual guidance, although the strongest proprietary model changes little and model rankings remain largely stable.

  • Overall results: Most models show substantial improvements with descriptive prompts because detailed visual descriptions provide more explicit generation guidance.The comparison is against conceptual-prompt results reported in Table 1.
  • Overall results: GPT Image 1.5 has an essentially unchanged dimension score under descriptive prompts, with a change of −0.06.This suggests the model can already infer suitable visual realizations from abstract conceptual prompts.
  • Ranking stability: Despite differing score magnitudes, the overall model ranking remains largely consistent across conceptual and descriptive prompts.The prompt type changes performance levels more than the relative ordering of models.

A.5 Cross-Judge Results with GPT-5.4

Using GPT-5.4 as an alternative judge yields model rankings highly consistent with those from Qwen3.5-27B. The judges differ somewhat in absolute dimension scores, but the paper's main conclusions remain unchanged.

  • Cross-judge consistency: 0.991 Spearman correlation for MCQ and 0.936 for dimension scoring show highly consistent rankings across the two judges.The comparison covers all 11 evaluated models.
  • Cross-judge consistency: GPT-5.4 assigns somewhat more lenient absolute dimension scores than Qwen3.5-27B.The relative performance ordering is nevertheless preserved.
  • Implication: The preserved model ordering confirms that the main experimental conclusions do not depend on the specific evaluator.GPT-5.4 results are reported as a complement to the Qwen3.5-27B results in Table 1.
  • Evaluation materials: The study's evaluation materials include MCQ and dimension-scoring protocols, alongside prompts and interfaces for annotation, generation, and human alignment checks.These materials support the benchmark's structured evaluation workflow.
  • Benchmark examples: The benchmark includes visual metaphors spanning ten categories, with representative generated samples shown across evaluated models under conceptual prompts.The examples illustrate thematic diversity and variation in model capability.
  • Release scope: The benchmark's samples are released through source-image URLs, structured annotations, prompts, and 9,594 multiple-choice questions rather than redistributed source images.Image retrieval is subject to Pinterest's terms of service.
Loading 2609.02502v1…