Source-linked AI summary

Evaluating Numerical Reasoning in Text-to-Image Models

Ivana Kajić, Olivia Wiles, Isabela Albuquerque, Matthias Bauer, Su Wang, Jordi Pont-Tuset, Aida Nematzadeh

arXiv:2406.14774v3cs.LGcs.CLcs.CV

TL;DR

Text-to-image models lack comprehensive evaluation of numerical reasoning despite often producing high-quality images. The paper introduces GECKONUM, a controlled benchmark spanning exact, approximate, and partial-quantity reasoning, and finds that models have rudimentary skills concentrated on small exact quantities. The benchmark also exposes limitations in evaluation protocols and supports analysis of automatic metrics and vision–language counting.

  • Problem

    Existing text-to-image benchmarks lack comprehensive coverage of numerical reasoning across exact counts, approximate quantities, and partial quantities.

  • Method

    GECKONUM evaluates twelve text-to-image models with controlled prompts, generated images, and human annotations across three numerical reasoning tasks.

  • Results

    Models show rudimentary numerical reasoning, perform best on small exact quantities, and DALL·E 3 has the highest overall accuracy on exact and approximate number generation.

  • Takeaways & Limitations

    GECKONUM discriminates between models and supports research on automatic evaluation metrics and pretrained vision–language models for counting.

  • Takeaways & Limitations

    The benchmark relies on laborious, costly human annotation and covers only some of the broader numerical capabilities in human cognition.

Abstract

from arXiv · show

Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical reasoning tasks of varying difficulty, and show that even the most advanced models have only rudimentary numerical skills. Specifically, their ability to correctly generate an exact number of objects in an image is limited to small numbers, it is highly dependent on the context the number term appears in, and it deteriorates quickly with each successive number. We also demonstrate that models have poor understanding of linguistic quantifiers (such as "a few" or "as many as"), the concept of zero, and struggle with more advanced concepts such as partial quantities and fractional representations. We bundle prompts, generated images and human annotations into GeckoNum, a novel benchmark for evaluation of numerical reasoning.

1 Introduction

The paper introduces GECKONUM to address the lack of comprehensive benchmarks for numerical reasoning in text-to-image models. Across twelve models, results show rudimentary numerical skills, with strongest performance on small exact quantities.

  • 1 Introduction: Existing text-to-image systems can produce high-quality images yet fail to accurately depict specified quantities such as “7 pistachios” [20].The paper frames numerical reasoning as an unresolved evaluation gap alongside capabilities such as alignment and compositionality.
  • 1 Introduction: GECKONUM evaluates numerical reasoning through exact number generation, approximate number generation, and reasoning about partial quantities.Its prompt templates control sentence structure, number context, and the number of attributes or entities.
  • 1 Introduction: The benchmark evaluates twelve models from five families using 1386 prompts, 52,721 generated images, and 479,570 human annotations.The evaluated families include DALL·E 3, Midjourney, Imagen, Muse, and Stable Diffusion.
  • 1 Introduction: Models show rudimentary numerical reasoning and are most accurate when generating small exact quantities.GECKONUM can distinguish powerful models with similar image quality, including Imagen-D [31] and DALL·E 3.

2 Related Work

Existing text-to-image benchmarks target broad capabilities or specific skills but provide limited coverage of numerical reasoning. GECKONUM expands evaluation across number ranges, representations, quantifiers, and partial quantities.

  • 2 Related Work: Existing benchmarks often contain too few or overly complex numerical prompts to isolate numerical reasoning from other skills.Interpreting prompts may also require object generation, relations, and attribute binding.
  • 2 Related Work: GECKONUM systematically covers number ranges, noun types, number representations, linguistic approximations, and partial quantities missing from other datasets.The benchmark is designed to evaluate these dimensions comprehensively rather than relying on naturally harvested prompts.
  • 2 Related Work: Prior number-generation work used a small prompt set, while GECKONUM evaluates more model families and adds estimation and conceptual quantitative reasoning.The cited prior set contains N=59 prompts and focuses on the simplest prompt structure [31].
  • 2 Related Work: Counting benchmarks for image-to-text models include CountBench and TallyQA, but CountBench is small and TallyQA is skewed toward small numbers with mixed image and label quality.CountBench contains 540 images, whereas TallyQA contains approximately 20K evaluation images.

3 Tasks to Examine Numerical Reasoning

The paper decomposes numerical reasoning into three tasks and uses controlled prompt templates to vary difficulty and context. These tasks span exact counts, approximate quantities and zero, and conceptual reasoning about parts and fractions.

  • 3 Tasks to Examine Numerical Reasoning: The benchmark uses 12 prompt templates across three tasks, generating 1386 prompts while varying numbers, nouns, contexts, and sentence structures.This design targets the abstraction principle: representing the same set size independently of object identity.
  • 3.1 Task 1: Exact Number Generation: Task 1 tests whether models depict the exact number of objects specified in a prompt.Prompt variations examine number magnitude from 1 to 10, Arabic numerals versus number words, noun frequency, and compositional structure.
  • 3 Tasks to Examine Numerical Reasoning: Table 1 organizes twelve prompt types with examples and templates for probing different aspects of numerical reasoning.The templates vary prompt structure and the context in which numerical terms appear.
  • 3.2 Task 2: Approximate Number Generation and Zero: Task 2 tests approximate quantities expressed by linguistic quantifiers such as “many,” “a few,” and “more,” and separately evaluates zero.The benchmark includes prompts with one entity and prompts relating quantities between two entities.
  • 3.3 Task 3: Conceptual Quantitative Reasoning: Task 3 examines conceptual understanding of objects and parts through whole, fractional, relational-fraction, and part-whole prompts.Examples include cakes divided into quarters, unequal pieces, and objects split into multiple pieces.

4 Human Annotations of Images

Human annotations measure whether generated images match prompts across the three numerical tasks. The study withholds original prompts from participants and aggregates their responses into task-specific accuracy labels.

  • 4 Human Annotations of Images: Participants count prompted object categories in Task 1 using automatically generated questions, entering counts, ranges, or “10+” responses.Each noun in a source prompt receives one counting question.
  • 4 Human Annotations of Images: Task 2 presents three or five candidate descriptions, and participants select the line that best describes each generated image.The number of candidate lines depends on whether the prompt contains one or two approximate entities.
  • 4 Human Annotations of Images: Task 3 uses DSG-generated, prompt-specific questions whose answers should all be “yes” when the image accurately depicts the prompt.Numerically irrelevant questions are excluded from analysis.
  • 4 Human Annotations of Images: Participants do not see the original generation prompts, supporting unbiased estimates of depicted quantities.Twenty-five crowd-sourced participants annotated each image five times under an ethically reviewed study design.
  • 4 Human Annotations of Images: Task 1 labels use the mode of five participant counts, while Tasks 2 and 3 compare encoded responses with ground-truth or binary accuracy values.Task 2 matches encoded quantifier values; Task 3 averages binary question responses.

5 Evaluating Text-to-Image Models

Across five model families, numerical accuracy depends strongly on number magnitude, representation, prompt structure, and task difficulty. Models handle small exact quantities best, but struggle with linguistic quantifiers, zero, and partial quantities.

  • 5.1 Task 1: Exact Number Generation: Accuracy drops sharply as exact quantities increase: DALL·E 3 decreases 18 percentage points from 1→2, 9 from 2→3, and 23 from 3→4.Models generally overestimate counts; except DALL·E 3, they also sometimes underestimate or omit the requested entity.
  • 5.1 Task 1: Exact Number Generation: Ten of twelve models are more accurate when numbers appear as words rather than digits, while nine of twelve favor frequent over rare nouns.DALL·E 3 and SD1.5 show no significant digit-versus-word difference; DALL·E 3, Imagen-D, and SDXL show no significant noun-frequency difference.
  • 5.1 Task 1: Exact Number Generation: Prompts combining multiple number–noun pairs or spatial relationships are harder than simple single-number prompts, with spatial prompts significantly degrading eleven models’ performance.These findings show that numerical generation is sensitive not only to magnitude but also to compositional context.
  • 5.2 Task 2: Approximate Number Generation and Zero: Approximate-number prompts with one entity are easier than those with two, while zero-object prompts are hardest for every model in the one-entity setting.For two entities, eight of twelve models perform best on “more X than Y” and worst on “as many X as Y.”
  • 5.3 Task 3: Conceptual Quantitative Reasoning: Task 3 is hardest overall, with most models near or below random chance; fractional-simple prompts are easiest, followed by part-whole and fractional-complex prompts.The benchmark’s question-based evaluation can overstate performance when a presence question yields a 50% baseline without distinguishing whole, sliced, or fractional objects.

6 Measuring What Counts: Challenges in Evaluation of Numerical Reasoning

The benchmark is also used to examine whether automatic metrics and vision–language models capture numerical reasoning. These evaluations show useful metric discrimination and gains from counting-specific fine-tuning and synthetic data, alongside task-specific evaluation challenges.

  • Evaluating auto-evaluation metrics: Gecko, DSG, and VNLI reliably distinguish correct from incorrect image generations across all evaluated models on small-number exact-count prompts.The study compares CLIPscore, TIFA, Gecko, DSG, and VNLI using statistical tests; TIFA sometimes asks about concepts absent from the prompt.
  • Evaluating counting in vision-language models: Adding Imagen-generated synthetic data does not significantly change TallyQA test performance but improves Muse-B performance by more than 20 percentage points in some cases, including held-out classes.The results motivate broader public datasets and benchmarks for counting and numerical reasoning.

7 Discussion and Conclusion

GECKONUM fills a gap in numerical-reasoning evaluation for text-to-image models and shows that current models remain weak on several numerical capabilities. The benchmark also supports evaluation of automatic metrics and counting improvements.

  • Existing text-to-image benchmarks lack comprehensive evaluation of numerical reasoning, motivating GECKONUM’s focus on controlled numerical prompts.
  • GECKONUM evaluates exact and approximate number generation plus conceptual reasoning about quantities across controlled prompt types and twelve models.
  • DALL·E 3 achieves the highest overall accuracy on exact and approximate number generation, but remains close to or below 50%.
  • Approximate quantities, zero, and parts or fractions remain difficult across models, with fraction reasoning performing close to baseline.
  • Only VNLI [35] and Gecko [33] among tested automatic metrics reliably distinguish correct from incorrect images and rank models pairwise on simple numerical prompts.
  • The benchmark relies on laborious, costly human annotation and covers only three manually designed numerical capabilities, leaving broader numerical cognition and more complex evaluation protocols open.

Checklist

The checklist records that the paper addresses its claims, limitations, ethics, reproducibility, and asset documentation. It reports releasing the materials and documenting participant protections and experimental details.

  • The authors report that the paper discusses its contributions, scope, limitations, societal impacts, and conformity with ethics-review guidelines.
  • The paper reports providing code, data, and reproduction instructions for the main experimental results through supplemental material or a URL.
  • The checklist states that training details, error bars, compute resources, asset citations, licenses, and newly released assets are documented.
  • For crowdsourced human-subject research, the authors report including participant instructions, screenshots, risk and review information, compensation, and consent-related documentation.

A Additional Information on Prompts

The appendix documents the benchmark’s prompt vocabulary, answer encoding, distributions, statistical procedures, and analyses of number and prompt-structure effects. It also describes comparisons involving colors and spatial relationships.

  • A.1 Words and Word Frequencies: Task 1 uses 40 nouns spanning everyday objects, food, nature, and animals, divided into 21 frequent and 19 rare words.
  • A.2 The Distribution of Prompt Types in the Benchmark: The benchmark contains 1,386 prompts, while appendix figures summarize prompt-number distributions and confusion matrices for numeric-simple prompts.
  • A.3 The Encoding Scheme for Task 2 Answers: Task 2 encodes quantity comparisons with five labels ranging from zero to more, including relations such as fewer, as many, and many.
  • B Additional and Detailed Experimental Results: The appendix reports that statistical significance was tested with α = .05 and Chi-squared comparisons based on binary prompt–image accuracy counts.
  • B.1.2 Number representation and word frequency: Appendix analyses compare accuracy for digits versus word numerals, rare versus frequent words, and specific numbers across prompt types.
  • B.1.3 Prompt structure: Additive prompts: Additive-prompt results show accuracy drops when specific numbers occur in different prompt types, with all models showing a substantial and significant drop for the reported higher number condition.
  • B.1.4 Prompt structure: Colors and spatial relationships: Color and spatial analyses compare numeric-simple with attribute-color prompts, two-additive with colored variants, and two-additive with spatial prompts.
  • B.1.4 Prompt structure: Colors and spatial relationships: Most color-term differences were not significant, whereas spatial relationships produced significant differences for all models except DALL·E 3.

B.2 Qualitative Analysis of Model Failures

The qualitative analysis identifies object omissions, overgeneration, ambiguity, and failures involving attributes, entities, and spatial arrangements. It also documents annotation procedures, agreement, and disagreement sources.

  • Models failed on exact-one prompts by omitting objects, generating multiple objects, or producing ambiguous depictions requiring interpretation.
  • Counting edge cases involved background or partial objects, and foreground-background boundaries were sometimes unclear even under explicit annotation instructions.
  • DALL·E 3 and Muse-B often represented the correct number of entities in additive prompts but miscounted those entities, while smaller models sometimes omitted an entity.
  • Attribute-spatial prompts were generally hardest: DALL·E 3 often produced near-correct counts without the required arrangement, while other models omitted objects.
  • The study used trained annotators, task-specific instructions, pilot testing, and five annotations per image-question pair, aggregating Tasks 1 and 2 with the mode.
  • Annotator agreement was high, but disagreement arose from differing interpretations of partial objects, uncertain identities, salience, distorted or morphed entities, and color perception.

C.3.2 Task 2: Approximate Number Generation and Zero

Approximate-quantity judgments are subjective, but annotators were broadly consistent about “no,” “few,” and “many.” Evaluation is further complicated by ambiguity in generated objects, images, and questions.

  • Annotation ambiguity: Ambiguous or morphed objects often prevented annotators from recognizing the entities needed to judge quantity.Examples included morphed trowel-manatee and fish-seahorse objects, plus guard rails that were not clearly part of a crib.
  • Approximate quantities: Individual interpretations of linguistic quantifiers varied, with some annotators labeling the same flower quantity “many” and others “few.”
  • Approximate quantities: Annotators usually selected “no” for zero objects, while “few” spanned roughly 1–10 objects and “many” was associated with counts above 10.The corresponding ground-truth-prompt analysis found that “no” labels were linked to fewer generated items, while “many” labels shifted toward larger counts.
  • Conceptual quantities: Task 3 annotations also disagreed about broken, sliced, or partial objects because their visual status and the questions were sometimes unclear or not meaningful.The authors attribute disagreement to ambiguity in images, questions, or both.
  • Evaluation method: For Task 3, evaluation used several prompt-grounded questions per image because alignment depends on depicting operations such as cutting or slicing, not only counting visible objects.The questions were generated with an automatic method, but some were nondiscriminative, confusing, or uninformative.

D.2.1 GECKONUM as a VQA benchmark

GECKONUM can serve as a VQA benchmark for counting, with simpler scenes than TallyQA and improvements from fine-tuning that are strongest at higher counts. Adding synthetic GECKONUM images preserves TallyQA performance while substantially improving transfer to Muse-B.

  • GECKONUM as a VQA benchmark: GECKONUM appears easier than TallyQA because its images contain one or very few object classes rather than complex, cluttered scenes.The base model already answers many GECKONUM questions correctly, especially at lower counts.
  • GECKONUM as a VQA benchmark: Fine-tuning PaLIGemma on TallyQA particularly improves counting at higher counts (≥5) on both TallyQA and GECKONUM.The authors note that TallyQA has few examples with counts ≥9, increasing uncertainty there.
  • Fine-tuning data mixtures: Adding Imagen-derived GECKONUM images to TallyQA training does not significantly change TallyQA test performance, while replacing TallyQA data can reduce it by a few percentage points.
  • Fine-tuning data mixtures: Including Imagen images markedly improves Muse-B performance by more than 25 percentage points, including when frequent classes train evaluation on rare classes.The improvement is slightly smaller for the frequent-to-rare split, at about 20 percentage points.
  • Fine-tuning data mixtures: With the full TallyQA training set, adding 53k Imagen images changes TallyQA test performance by less than 1% while drastically improving Muse-B, even in the held-out case.
Loading 2406.14774v3…