Source-linked AI summary

Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing

Shaoan Zhao, Fang Zhao, Xueqiang Guo, Xinpei Su, Huanlin Gao, Qiang Hui, Ting Lu, Fuyuan Shi, Chao Tan, Bikun Yang, Kai Wang, Shiguo Lian

arXiv:2608.24112v1cs.AI

TL;DR

Modern T2I models can have similar overall scores while differing substantially in capability, and existing fine-grained evaluations often obscure attribution, complexity, and compositional failure. QC-T2I-Bench uses attributed atomic questions, dependency graphs, hierarchy-constrained aggregation, and reusable records for ranking, diagnosis, and routing. Across bilingual evaluations, joint completion declines from 80.7% to 37.2% as component complexity increases, while routing matches ERNIE’s estimate with 21.3% lower GPU-s/MP.

  • Problem

    Similar aggregate scores can mask different T2I capability profiles, while prompt-level aggregation and fixed categories limit attribution, complexity awareness, and cross-prompt comparison.

  • Method

    QC-T2I-Bench converts open prompts into attributed atomic questions, organizes dependencies with DSGs, and applies HCQ for validity-aware, complexity-aware aggregation.

  • Results

    Joint completion falls from 80.7% for two-tag components to 37.2% for components with seven or more tags, while the router matches ERNIE’s point estimate with 21.3% lower GPU-s/MP.

  • Takeaways & Limitations

    The same auditable question records support reliable ranking, fine-grained diagnosis, and training-free cost-aware model selection.

  • Takeaways & Limitations

    The closest model pair is not deterministically separated, and the router’s interval precludes a lossless-routing claim.

Abstract

from arXiv · show

Modern text-to-image (T2I) models often have similar total scores but different strengths, making practical selection difficult. Fine-grained benchmarks decompose prompts into questions, yet often return them to prompt scores and fixed categories, weakening attribution and ignoring complexity. Related requirements are also scored separately or as one total, obscuring basic versus compositional failure. We present QC-T2I-Bench, a question-centric framework that converts open prompts into attributed atomic questions and organizes their dependencies with Davidsonian Scene Graphs (DSGs). We use hierarchy-constrained question aggregation to exclude downstream questions after a prerequisite fails and to prevent simple and complex prompts from receiving the same total weight. We then use the DSG structure to measure joint success within prompts and compare repeated entities across prompts, separating basic realization failures from failures under additional requirements. We evaluate multiple open-source T2I models on English and Chinese prompts. The resulting question-level evidence supports reliable ranking and fine-grained diagnosis: joint completion falls from 80.7\% for components with two capabilities to 37.2\% for those with seven or more. Finally, we reuse the same records for training-free routing; our cost-aware router matches ERNIE's 89.51-point estimate with 21.3\% less GPU-s/MP.

Introduction

QC-T2I-Bench addresses the limits of prompt-level evaluation by treating attributed atomic questions and their dependencies as reusable evidence. It supports ranking, diagnosis, and cost-aware routing across bilingual evaluations of 13 T2I models.

  • Introduction: Similar aggregate scores can conceal different model strengths, so practical selection requires evidence about capability-specific successes and failures.
  • Introduction: Question-generation methods decompose prompts into local checks, but evidence attribution, aggregation, and reuse remain central challenges.
  • Introduction: Prompt-centered pipelines make category scores difficult to interpret for multi-capability prompts and may weight simple and complex prompts equally.
  • Introduction: Existing benchmarks variably support atomic attribution, open prompts, dependency awareness, complexity awareness, and cross-prompt evidence, but none combines all five.
  • Introduction: QC-T2I-Bench changes the evaluation unit from prompts to attributed atomic questions and uses HCQ to remove invalid evidence and balance capabilities.
  • Introduction: DSGs test joint satisfaction of related requirements, while repeated-entity comparisons distinguish basic generation failures from failures under added attributes or relations.
  • Introduction: Across 13 models with bilingual evidence, the framework supports reliable ranking, fine-grained diagnosis, and routing; the router reduces GPU-s/MP by 21.3% at the matched ERNIE point estimate.

Related Work

Prior T2I evaluation work spans atomic question verification, capability-specific benchmarks, and model-selection methods. These lines of work provide important specialized evidence but differ in coverage and routing scope.

  • Fine-grained evaluation benchmarks: GenEval and GenEval 2 evaluate predefined object properties, with GenEval 2 adding atom-level questions and atomicity analysis.
  • Fine-grained evaluation benchmarks: TIFA decomposes prompts into VQA pairs, while DSG organizes atomic questions through valid dependencies.
  • Fine-grained evaluation benchmarks: Other benchmarks improve supervision, calibration, holistic alignment, hierarchical capability rubrics, or dependency-aware checklists.
  • Capability-specific and closed-set benchmarks: Capability-specific benchmarks target world knowledge, physical rules, compositional instructions, bilingual scenarios, creativity, reasoning, long text, or multilingual evaluation.
  • Evidence reuse and generator routing: Routing methods learn or plan model selection for prompt-conditioned quality–cost decisions, APIs, stateful orchestration, or multi-step generation.

Method

QC-T2I-Bench evaluates atomic, attributed questions while preserving dependency and prompt context, then uses HCQ and DSG-based diagnostics for balanced scoring and compositional analysis.

  • Downstream uses: The same question records support ranking, diagnosis, bilingual evidence views, and training-free generator routing.The scoring and structural paths are reused for ranking, diagnosis, and routing rather than requiring separate records.
  • Question construction: Each prompt is converted into independently judgeable atomic questions with one capability label, covering salient constraints under a fixed construction contract.The contract requires Target Yes, Atomicity, Coverage, and Tag Validity, with auditing and cleanup.
  • Question construction: Evaluation records preserve prompt, question, model, language, reporting-group, capability, outcome, and dependency-validity coordinates.The taxonomy has five first-level groups and 21 secondary capabilities; Text Content uses a dedicated transcription evaluator.
  • Dependency-aware scoring: Davidsonian Scene Graph prerequisite edges score downstream questions only when prerequisites succeed, preventing repeated penalties from one missing entity.Invalid records are excluded from capability-level means rather than entering the score.
  • Hierarchy-constrained aggregation: HCQ micro-averages valid questions within capabilities, then macro-averages capabilities within groups and across five reporting groups.The groups contain 5, 1, 6, 4, and 5 capabilities respectively, balancing capabilities instead of prompt counts.
  • Compositional analysis: DSG diagnostics measure complete-component exactness, cross-prompt root matching, and topology-localized coupling beyond marginal success multiplication.Components are maximal weakly connected structures scored once for their full tag set; edge contrasts compare parent–child pairs with matched disconnected pairs.

Experiments

Experiments show that QC-T2I-Bench supports stable, diagnostically traceable rankings across bilingual evaluations while exposing compositional difficulty, topology-localized coupling, and quality–cost trade-offs.

  • Experimental setup: 13 open-source T2I models are evaluated on 6,573 prompts with 94,547 English and 94,555 Chinese questions, analyzed separately by language.The benchmark reports bilingual leaderboards, dependency analyses, and routing summaries.
  • Reliable ranking: Equal atomic-question weighting minimizes adjudication-error variance, whereas prompt-first weighting amplifies errors from relatively sparse prompts.Under the stated independent, zero-mean, equal-variance error model, equality holds only when prompts contribute equal question counts.
  • Reliable ranking: FLUX.2-dev leads in English and ERNIE-Image in Chinese, but no model is best across every capability view.The scalar ordering remains traceable to diagnostic capability coordinates.
  • Reliable ranking: 68 of 78 model-pair bootstrap intervals exclude zero, while FLUX.2-dev and ERNIE-Image have top-1 probabilities of 55.2% and 44.8%, respectively.Both models have 95% rank intervals of [1, 2], so their 0.042-point gap does not establish deterministic separation.
  • Fine-grained diagnosis: Joint exact completion falls from 80.7% for two-tag components to 37.2% for components with seven or more tags.This measures increasing joint difficulty as requirements accumulate, not interaction by itself.
  • Fine-grained diagnosis: The direct-edge coupling contrast is +1.10 points [ +0.69, +1.50 ], versus an unresolved non-ancestral contrast of +0.38 points [−0.17, +1.01].The excess coupling localizes to explicit DSG topology; cross-prompt root controls show local roots trail matched baselines by 0.14–0.88 points.
  • Scalable cost-aware routing: Q-Profile-C matches ERNIE’s point estimate while reducing GPU-s/MP cost by 21.3% relative to ERNIE and 65.0% relative to FLUX2.Its interval crosses the preregistered 0.20-point non-inferiority margin, supporting a quality–cost trade-off rather than lossless routing.

Conclusion

QC-T2I-Bench reuses question-centric evidence for ranking, diagnosis, and cost-aware routing. Across models, compositional completion declines with added capabilities, while routing matches ERNIE’s point estimate at lower cost.

  • Conclusion: Question decomposition becomes reusable evidence rather than merely an intermediate step toward a prompt score.The same evidence base supports ranking, diagnosis, and cost-aware selection without treating any scalar score as universal.
  • Conclusion: 21.3% lower GPUs/MP accompanies routing that matches ERNIE’s point estimate, although the interval precludes a lossless-routing claim.The records support cost-aware selection while preserving uncertainty about non-inferiority.
  • Conclusion: 80.7% exact completion for two-tag components falls to 37.2% for components with seven or more tags.After marginal control, excess coupling is resolved on direct Davidsonian Scene Graph edges but not matched non-ancestral pairs.
Loading 2608.24112v1…