Source-linked AI summary

Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

Yichen Feng, Yuetai Li, Chunjiang Liu, Yuanyuan Chen, Fengqing Jiang, Yue Huang, Hang Hua, Zhengqing Yuan, Kaiyuan Zheng, Luyao Niu, Bhaskar Ramasubramanian, Basel Alomair, Xiangliang Zhang, Misha Sra, Zichen Chen, Radha Poovendran, Zhangchen Xu

arXiv:2605.12684v1cs.CVcs.AIcs.HC

TL;DR

Existing aesthetic benchmarks usually score images independently, leaving comparative preference and execution quality difficult to measure. This paper introduces VAB, a set-based benchmark with matched subjects and expert-grounded rankings, finding that the strongest system achieves 26.5% accuracy versus 68.9% for human experts.

  • Problem

    Existing aesthetic methods infer preference from independent scalar image scores, which may not faithfully measure comparative execution quality across matched visual works.

  • Method

    VAB evaluates comparative selection over matched-subject candidate sets using labels consensually derived from 10 independent expert judges across 400 tasks and 1,195 images.

  • Results

    26.5% of tasks were solved by the strongest system on TB-1 pass^3, versus 68.9% for human experts, with models also showing positional instability across permutations.

  • Takeaways & Limitations

    VAB exposes a persistent gap between current multimodal models and expert aesthetic judgment on permutation-robust comparative evaluation.

  • Takeaways & Limitations

    The human evaluation protocol was determined exempt human-subjects research and involved voluntary online participation without collecting sensitive personal information.

Abstract

from arXiv · show

Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an explicit aesthetic judgment. Most existing solutions reduce this judgment to predicting a scalar score for a single image. We first ask whether such scores faithfully capture comparative preference: in a controlled study with eight expert annotators, score-derived rankings align poorly with the same annotators' direct comparisons, while direct ranking yields substantially higher inter-annotator agreement on best- and worst-image labels. Motivated by this finding, we introduce the Visual Aesthetic Benchmark (VAB), which casts aesthetic evaluation as comparative selection over candidate sets with matched subject matter. VAB contains 400 tasks and 1,195 images across fine art, photography, and illustration, with labels derived from the consensus of 10 independent expert judges per task. Evaluating 20 frontier MLLMs and six dedicated visual-quality reward models, we find that the strongest system identifies both the best and the worst image correctly across three random permutations of the candidate order in only 26.5% of tasks, far below the 68.9% achieved by human experts. Fine-tuning a 35B-parameter model on 2,000 expert examples brings its accuracy close to that of a 397B-parameter open-weight model, suggesting that the comparative signal in VAB is transferable. Together, these results expose a clear and measurable gap between current multimodal models and expert aesthetic judgment, and VAB provides the first set-based, expert-grounded testbed on which that gap can be tracked and closed.

1 Introduction

The paper argues that direct comparative ranking captures aesthetic preference more faithfully than independent image scoring and introduces VAB to evaluate set-based expert aesthetic judgment. Across 20 frontier MLLMs and six reward models, performance remains far below human experts, while expert-data fine-tuning substantially improves transfer.

  • Existing aesthetics methods infer preference from independent scalar image scores, although aesthetic applications require comparing execution quality across composition, color, technical execution, and cultural convention.The paper frames aesthetic judgment as a demanding higher-order multimodal reasoning task.
  • Comparative ranking produced 42 percentage points higher inter-annotator agreement on best-image selection than score-derived rankings among eight expert annotators.The study found that scoring and direct comparison elicited different preference signals from the same annotators.
  • VAB contains 400 tasks and 1,195 images across fine art, photography, and illustration, spanning 24 topics with two to six matched-subject candidates per task.The benchmark asks models to select the best image and, under the stricter setting, also the worst image.
  • 26.5% TB-1 pass^3 was achieved by Claude Sonnet 4.6, 42.4 points below the 68.9% human baseline, with severe sensitivity to candidate order.Models retained only 34–65% of ap@1 accuracy when required to succeed across all three permutations.
  • Fine-tuning a 35B-parameter model on 2,000 expert examples brought it close to a 397B-parameter open-weight model, indicating transferable comparative judgment.This result suggests lightweight adaptation can effectively transfer expert comparative judgments.

2 Human Annotation Study: Ranking vs. Scoring

A controlled study of eight experts finds that absolute scoring poorly preserves the preferences expressed through direct comparison, whereas ranking produces more reproducible and concentrated aesthetic judgments. Accordingly, VAB adopts ranking for annotation.

  • Setup: The study compares scoring and ranking across 119 homogeneous-content and 107 heterogeneous-content tasks, using eight experts and candidate groups of two to five images.Scoring assigns independent values in [0, 10], whereas ranking orders images directly within each group.
  • Results: Score-derived rankings often diverge from direct comparisons, with mean Kendall’s τ of 0.188 and 0.351 and top-1 self-consistency of 45.5% and 52.0% for homogeneous and heterogeneous tasks.Nearly half of numeric score annotations fail to recover the same annotator’s directly expressed order.
  • Results: Ranking improves inter-annotator reproducibility of best/worst labels over scoring on homogeneous-content tasks, with gains of 42.0, 37.0, and 51.3 points for best, worst, and both labels.All reported improvements are statistically significant at p < 0.001; the same pattern holds for heterogeneous-content tasks.
  • Results: Ranking concentrates top-choice preferences more strongly, reducing entropy from 1.213 to 0.729 and increasing unanimous choices from 0.8% to 10.1%.The difference is statistically significant (Wilcoxon p = 1.7 × 10−15).
  • Interpretation: The authors attribute ranking’s advantage to relational aesthetic judgment and the cognitive primacy of comparative evaluation, then adopt ranking for VAB annotation.These interpretations are presented as complementary aesthetic and cognitive explanations.

3 The VAB Benchmark

VAB evaluates expert-aligned aesthetic judgment through comparative selection among matched-subject candidate sets, rather than scalar ratings aggregated from crowd annotations (Murray et al., 2012; Kong et al., 2016; He et al., 2022). It includes best-image and stricter best-and-worst-image tasks, with labels from direct expert comparisons.

  • Task formulation: VAB contains 400 tasks and 1,195 images spanning fine art, photography, and illustration across 24 topics, with candidate sets of 2–6 matched-subject images.Top-1 selects the best image, while TB-1 selects both the best and worst image.
  • Evaluation protocol: Each task is evaluated under three independent random candidate-order permutations using direct comparative expert judgments rather than aggregated absolute scores.The benchmark reports ap@1 as average accuracy across evaluations and pass^3 when all three evaluations are correct.
  • Data construction: VAB organizes data around matched image groups whose members share subject matter but vary in execution quality, enabling comparison of aesthetic execution rather than depicted content.The construction pipeline is summarized in Figure 3, with additional raw-pool and domain-level details in Appendices B.2 and B.3.
  • Data construction: Fine-art sets use constrained artist prompts to preserve subject matter while varying composition, value structure, color handling, and finish.The commissioned collection comprises 426 painting sets and 1,142 images.
  • Data construction: Photography sets originate from single source photographs and are expanded through expert-retouching and scalable agent-edit pipelines that generate improved, degraded, or multi-expert variants.The collection comprises 670 photography sets and 1,809 images, including source material from professional photographers.

Step 1: Data Collection by Domain

VAB organizes candidate images into matched-subject sets spanning artwork, photography, and illustration, with varied aesthetic execution. Illustration contributes 271 sets and 908 images generated from text prompts or controlled renderings of CC0 3D assets.

  • Step 1: Data Collection by Domain: The collected domains include artwork, photography, and illustration, covering subjects such as architecture, food and product, landscape, portrait, sports, street, city, and wildlife.Illustration categories also include anime/manga, comic, concept art, digital/AI art, pixel art, and stylized 3D.
  • Step 1: Data Collection by Domain: The data-construction pipeline organizes images into candidate sets with matched subject matter and varying aesthetic execution before expert review and consensus filtering.The pipeline overview is summarized in Figure 3.
  • Step 1: Data Collection by Domain: Illustration contributes 271 candidate sets comprising 908 images with matched subject matter but varying aesthetic execution.Sets are produced either through prompt-based text-to-image generation or by rendering CC0 3D assets with controlled viewpoint and presentation changes.
  • Step 1: Data Collection by Domain: All collected sets are deduplicated and reviewed for semantic consistency before annotation.

4 Experiments

Experiments show that frontier MLLMs and reward models remain substantially below expert aesthetic judgment on comparative VAB evaluation. Performance also degrades with larger candidate sets and varies across model families, domains, and decoding conditions.

  • Progress across families: Model progress is uneven: Claude Sonnet 4.6 improves over its predecessor from 14.5% to 26.5% on TB-1 pass^3, while GPT-5.2 reaches 15.5%.A closed/open gap also persists, with Qwen3.5-397B-A17B at 17.2% versus Claude Sonnet 4.6 at 26.5%.
  • Domain difficulty and stability: GLM-4.6V drops from 34.1% to 11.5% on TB-1 across ap@1 and pass^3, indicating sensitivity to candidate order and response instability.Accuracy also varies by domain: Claude Sonnet 4.6 reaches 19.0% on illustration, while o4-mini reaches 30.2% on photography.
  • Reward models: 52.2%: Q-Align’s Top-1 accuracy remains below the human baseline of 77.7%, while its TB-1 accuracy is 38.2% versus 68.9% for humans.Reward models can be competitive with MLLMs on Top-1, but their strongest overall results remain below experts.
  • Temperature ablation: Greedy decoding improves Top-1 pass^3 by 1.1 points and TB-1 pass^3 by 2.3 points on average, so stochasticity adds noise but does not explain the main gap.The evaluation covers 20 MLLMs and six reward models, using model defaults unless otherwise noted.
  • Candidate set size: Candidate-set difficulty is substantial: the best model’s TB-1 pass^3 falls from 47.3% on two-image tasks to 6.7% on four-image tasks, while human accuracy drops from 87.1% to 43.6%.Larger candidate sets amplify comparative difficulty and instability for both models and experts.

5 Fine-Tuning on Expert Aesthetic Judgments

LoRA fine-tuning QWEN3.5-35B-A3B on 2,000 disjoint expert-annotated examples improves all four VAB metrics. The resulting KALLISTI-35B-A3B nearly matches the much larger Qwen3.5-397B-A17B, indicating that comparative supervision transfers under lightweight adaptation.

  • Training Setup: KALLISTI-35B-A3B uses rank-64 LoRA with α = 128, freezes the vision tower, updates the multimodal projector, and trains for three epochs on 2,000 disjoint expert examples.Training uses a learning rate of 1 × 10−4 with cosine scheduling.
  • Improvement over the base model: 29.5% versus 25.5% Top-1 pass^3, 17.5% versus 15.8% TB-1 pass^3, 51.3% versus 47.5% Top-1 ap@1, and 38.8% versus 36.5% TB-1 ap@1 show consistent gains over the base model.These results compare KALLISTI-35B-A3B with its base model across all four reported metrics.
  • Scaling implication: 29.5% versus 29.8% on Top-1 pass^3 places KALLISTI-35B-A3B near Qwen3.5-397B-A17B, with the other metrics within roughly one point.This provides preliminary evidence that VAB’s comparative supervision can compensate for an order-of-magnitude reduction in parameter count under lightweight adaptation.

6 Conclusion

VAB reframes aesthetic evaluation as expert-grounded comparative selection over matched candidate sets rather than scalar scoring. It includes 400 tasks and 1,195 images spanning fine art, photography, and illustration.

  • 6 Conclusion: VAB evaluates aesthetics through comparative selection among matched-subject candidate sets, using expert rankings instead of scalar scores.The benchmark contains 400 tasks and 1,195 images across fine art, photography, and illustration.
  • 6 Conclusion: Direct ranking produced 42 percentage points higher inter-annotator agreement than score-derived rankings in a controlled human study.

Ethics Statement … B.2 Raw Collection Statistics Before Expert Filtering

The paper establishes ethical safeguards and consent procedures for collecting aesthetic judgments and image data, while positioning VAB against prior single-image aesthetic resources. Its appendices document annotation examples, and benchmark details report both final filtered statistics and the larger pre-filtering collection pool.

  • Ethics Statement: The study was classified as exempt minimal-risk human-subjects research, with voluntary participation, withdrawal allowed, and no sensitive personal information collected.The University of Washington Human Subjects Division reviewed the protocol under Categories 2 and 3; only aggregated benchmark results were collected.
  • Ethics Statement: Fine art, photography, and illustration data were obtained through commissioned work, licensing agreements, and structured text-to-image prompts.Artists consented to research use; source photographs were appropriately licensed, and agent-edited variants remained within those agreements.
  • Appendix: The appendix provides topic-level annotation examples organized into fine art, photography, and illustration sections.The listed sections begin at pages 62, 71, and 79, respectively.
  • A Related Work: Existing aesthetic resources largely frame assessment as single-image score prediction, whereas newer MLLM efforts add richer critique but remain mostly single-image, score-centric, or domain-specific.The cited landscape includes AVA and AADB, alongside AesBench, AesExpert, UNIAA, ArtiMuse, and The Photographer’s Eye.
  • B. Benchmark Details: The benchmark-details appendix reports final statistics after expert annotation and consensus filtering, including domain/topic breakdowns and candidate-set sizes.Table 7 covers final domain and topic statistics, while Tables 8 and 9 cover candidate-set and topic-level distributions.
  • B.2 Raw Collection Statistics Before Expert Filtering: Before expert filtering, the collected pool contained 1,367 tasks and 3,859 images.The raw collection is further summarized by domain, topic, and photography collection pipeline in Tables 10–12.

B.3 Data Construction Details … D.1 Selected Thresholds

The benchmark constructs matched-subject candidate sets across fine art, photography, and illustration, then evaluates them through expert comparative judgments guided by domain-specific rubrics. Consensus filtering and documented thresholds determine the benchmark’s ground-truth selections.

  • B.3 Data Construction Details: The dataset spans fine art, photography, and illustration, with controlled candidate sets designed to preserve subject matter while varying composition and aesthetic quality.Fine art includes 426 sets and 1,142 images; photography includes 670 sets and 1,809 images; illustration includes 271 sets and 908 images.
  • B.3 Data Construction Details: Photography candidate sets are produced through expert retouching and agent-based editing, then deduplicated and reviewed for semantic consistency before annotation.The agent-edit pipeline uses ArtiMuse reviews, GPT-5 prompts, and Gemini-3-Pro-Image generation, while expert-edit sets compare original and improved variants or multiple independent edits.
  • B.3 Data Construction Details: Illustration sets come from prompt-based generation or 3D rendering, with fixed semantic and scene conditions creating matched-subject comparisons with controlled visual variation.Generative sets vary subject, composition, viewpoint, lighting, and style through structured prompts; rendered sets vary camera viewpoint while fixing object identity, scene semantics, lighting, and background.
  • C.1 Annotation Rubrics: Judges use structured rubrics as comparison guides rather than numerical score sheets, assessing domain-specific criteria before making final selections.The rubrics emphasize visible properties and explicitly exclude factors such as personal preference, contextual information, assumed difficulty, or software used.
  • Fine Art Rubric: Fine-art judgments compare artistic completeness and mature aesthetic expression through subject definition, composition, color, mark-making, resolution, atmosphere, and overall persuasiveness.For each set, judges identify the strongest work and, when there are at least three images, the weakest work.
  • Photography Rubric: Photography judgments evaluate photographic language and visual expression through emphasis, framing, background, light, exposure, focus, color, technical quality, and narrative impact.For each set, judges identify the strongest photograph and, when there are at least three images, the weakest photograph.
  • Digital Illustration Rubric: Digital-illustration judgments assess visual design and professional quality through focal hierarchy, flow, structure, palette, lighting, detail control, stylistic alignment, and use-case suitability.For each set, judges identify the strongest illustration and, when there are at least three images, the weakest illustration.
  • C.2 Annotation Interface / D Ground Truth Mathematical Details / D.1 Selected Thresholds: The annotation interface presents all candidates side by side: judges select the best image for k = 2 and both the best and worst images for k ≥3.Consensus filtering is documented mathematically, with selected thresholds and corresponding null pass probabilities reported by candidate-set size k for n = 10 judges.

D.2 Derivation of Null Pass Probabilities … E.2 Comprehensive Reward-Model Results

The appendix derives and verifies null pass probabilities, handles binary-task thresholds separately, and reports comprehensive MLLM and reward-model results across metrics and domains. These results expose order sensitivity, model-family variation, domain difficulty, and asymmetric strengths in selecting best versus worst images.

  • D.2 Derivation of Null Pass Probabilities: For k ≥3, null probabilities model each judge as uniformly selecting an ordered best–worst pair and sum multinomial masses over configurations meeting both consensus thresholds.The acceptance event is conjunctive: both best-vote and worst-vote thresholds must be met.
  • D.3 Monte Carlo Verification: Exact null probabilities were verified by Monte Carlo simulation using 107 samples per (k, m) configuration and 95% Wald confidence intervals.Table 14 compares the exact values against the Monte Carlo estimates.
  • E Detailed Experimental Results: The detailed experimental appendix preserves the main evaluation focus while expanding coverage of expert-relative performance, domain variation, candidate-order sensitivity, decoding, and task size.It supplies full tables that are too large for the main text.
  • E.1 Comprehensive MLLM Results: Comprehensive MLLM results show pervasive gaps between ap@1 and pass^3, with predictions often sensitive to candidate order and response instability.Table 15 reports Top-1, Bot-1, and TB-1 accuracy under both metrics across overall and domain-level evaluations.
  • E.1 Comprehensive MLLM Results: Illustration is the most difficult domain, photography is generally more tractable, and fine art lies between them across the comprehensive evaluations.The domain columns reinforce the main-text headline while also clarifying cross-domain variation.
  • E.2 Comprehensive Reward-Model Results: Reward models can be materially stronger at rejecting poor images than identifying the best, especially Q-Instruct on photography.The comprehensive table makes this Bot-1 failure mode directly visible.

E.3 Topic-Level MLLM Results … F Human Study Details

Topic-level analyses show that domain averages conceal substantial variation across subjects, while decoding and candidate-set size affect performance without changing the central comparative-judgment limitation. Illustration is consistently hardest, photography generally strongest but uneven, and larger candidate sets widen the gap from human experts.

  • E.3 Topic-Level MLLM Results: Illustration remains the hardest MLLM domain across Top-1 ap@1, Top-1 pass^3, TB-1 ap@1, and TB-1 pass^3, while photography is more tractable but highly uneven.These domain-level trends are not driven by a single topic or metric.
  • E.3 Topic-Level MLLM Results: Within MLLM topics, remaining illustration signal concentrates in anime-and-manga and comic tasks, while concept art, digital-and-AI art, and stylized 3D often collapse under pass^3.Chinese painting and portrait-color tasks are often relatively tractable in fine art, whereas landscape-color and some sketch-based tasks are less stable.
  • E.4 Topic-Level Reward-Model Results: Reward-model domain averages also mask topic variation: photography’s stronger results come from a subset of topics, whereas illustration remains difficult across most topics for all six models.Fine-art performance is mixed, with several models competitive on Chinese painting and portrait-color tasks.
  • E.4 Topic-Level Reward-Model Results: Photography is the strongest reward-model domain, but its aggregate advantage is driven by a subset of topics, while most illustration topics remain near random, especially under TB-1.Outside anime-and-manga and occasional comic tasks, concept art, digital-and-AI art, pixel art, and stylized 3D are especially difficult.
  • E.5 Temperature Ablation: Greedy decoding modestly improves stability overall, but variable model-specific effects leave comparative judgment—not decoding noise—as the dominant limitation.Gemini 3.1 Pro and Claude Sonnet 4.5 benefit most, especially on TB-1 pass^3, while several other models are flat or slightly worse on at least one metric.
  • E.6 Results by Candidate Set Size: Performance drops sharply as candidate-set size k increases, especially under TB-1 pass^3, widening the gap between models and human experts on larger sets.Table 35 provides the full breakdown by candidate set size, reporting pass^3 and ap@1 (%).

F.1 Detailed Setup and Metric Definitions … G.1 Hyperparameters

The study formalizes aesthetic evaluation through paired scoring and set-based ranking protocols, with domain-specific criteria, explicit agreement and entropy metrics, and statistical tests across homogeneous and heterogeneous tasks. Its appendices detail task construction, annotation rubrics, notation, and fine-tuning hyperparameters.

  • F.1 Detailed Setup and Metric Definitions; F.3 Task Examples; Heterogeneous-Content Tasks (107 Tasks); Homogeneous-Content Tasks (119 Tasks): The human study spans 119 homogeneous-content tasks and 107 heterogeneous-content regroupings, while Figure 6 illustrates that homogeneous tasks share underlying content and heterogeneous tasks regroup distinct source tasks.The heterogeneous set excludes 12 near-saturated regrouped tasks, retaining 107 for analysis.
  • F.1 Detailed Setup and Metric Definitions; F.5 Notation: The evaluation defines fidelity with Kendall’s τ and top-1 self-consistency, reproducibility through five-of-eight majority accuracy for best, worst, and joint labels, and distinguishability through top-choice entropy.For entropy, only the 119 homogeneous-content tasks are used, and lower unnormalized entropy indicates stronger concentration of judgments.
  • F.2 Extended Interpretation: The interpretation argues that direct comparison is more natural and avoids the calibration variance, scale-use heterogeneity, and added measurement noise introduced when absolute judgments are translated into numeric scores.This framing connects comparative aesthetic judgment to a relational view in which preferences become more reliable when alternatives are juxtaposed.
  • F.1 Detailed Setup and Metric Definitions; F.4 Annotation Protocol; F.5 Notation: Ranking presents all n images jointly for a complete best-to-worst ordering, whereas scoring rates each image independently on a 0–10 scale before inducing rankings.The protocols use groups of n ∈{2, 3, 4, 5}; scoring covered 400 images, while induced rankings were computed for 119 homogeneous-content and 107 heterogeneous-content tasks.
  • Fine Art Single-Image Scoring Rubric; Photography Single-Image Scoring Rubric; Digital Illustration Single-Image Scoring Rubric; Fine Art Ranking Rubric; Photography Ranking Rubric; Digital Illustration Ranking Rubic: Ranking rubrics compare matched fine art, photography, and digital illustrations using composition, subject emphasis, tonal or color control, technical execution, coherence, and overall expressive quality.The rubrics exclude personal subject preference and contextual information, while the illustration rubric also excludes software, tools, and hypothetical client requirements.
  • F.6 Statistical Tests: Statistical comparisons use exact two-sided McNemar tests for majority-agreement metrics and a two-sided Wilcoxon signed-rank test for entropy; after Bonferroni correction, all six majority-agreement tests remain significant at p < 0.001.For homogeneous tasks, McNemar p-values are 1.0 × 10−11, 2.4 × 10−9, and 1.2 × 10−14 for accbest, accworst, and accbest&worst; heterogeneous values are 1.2 × 10−7, 2.4 × 10−5, and 1.0 × 10−9, while entropy yields p = 1.7 × 10−15.
  • F.7 Task Lists; Homogeneous-Content Tasks (119 Tasks); Heterogeneous-Content Tasks (107 Tasks): Task lists document the 119 homogeneous tasks and their topic-specific candidate counts, while heterogeneous tasks preserve topic consistency by regrouping images from distinct source tasks without reusing an image.The homogeneous tasks span four topics across three domains.
  • G Fine-Tuning Details; G.1 Hyperparameters: The fine-tuning appendix records the hyperparameters used to fine-tune KALLISTI-35B-A3B and points to the associated topic-level fine-art Top-1 ap@1 results table.The supplied passages identify these implementation tables but do not provide their hyperparameter values or result cells.

G.2 Detailed Results

Topic-level fine-tuning results across four metrics show uneven gains for KALLISTI-35B-A3B: strongest improvements occur in several fine-art topics, while photography gains are inconsistent. The model remains competitive with larger open-weight baselines on illustration.

  • Topic-level fine-tuning results: Fine-tuning gains are uneven across topics: KALLISTI-35B-A3B improves most clearly on several fine-art topics, remains competitive with larger open-weight baselines on illustration, and does not uniformly surpass its base model in photography.The breakdown covers Top-1 ap@1, Top-1 pass^3, TB-1 ap@1, and TB-1 pass^3.
  • Evaluation scope: The detailed breakdown reports topic-level results for fine art, illustration, and photography under Top-1 ap@1, Top-1 pass^3, TB-1 ap@1, and TB-1 pass^3.Tables 40–51 provide the domain-specific topic-level breakdowns across these four metrics.

H Prompt Templates and Scoring Interfaces … Stylized 3D

The appendix specifies the prompt templates and scoring interfaces used to construct and evaluate VAB across photography, illustration, MLLM, and reward-model pipelines. It also documents representative topic-level annotation examples.

  • H.1 Photography Pipeline Prompts: The photography pipeline uses ArtiMuse for eight aspect-wise critiques, GPT via LiteLLM for edit planning, and Gemini-3-Pro-Image for image editing.The eight responses are concatenated into one review string, while editing generates better and worse variants in separate rounds.
  • Round 2: Worse-Image Editing Prompt with Context: The worse-image editing prompt references the improvement made to the original, then instructs editing the original image using the supplied edit instruction.The runtime template identifies the improved result as the second reference image and the original as the first image.
  • H.2 Illustration Pipeline Prompts: Illustration prompts use a structured subject-to-style order for digital, AI, concept, and pixel art, while stylized 3D is rendered from assets rather than text prompts.The structured order includes subject, motion, environment, lighting, atmosphere, camera or composition, and art style.
  • H.2.2 Manga / Comic Storyboard Generation: Anime, manga, and comic generation uses storyboard prompts with consistent characters and environments, specifying shot type, actions, props, and scene details for each panel.Each prompt describes a concise multi-panel micro-story.
  • H.3 MLLM Evaluation Prompts: MLLM evaluation prompts ask models to select the best image, the worst image, or both within each candidate set.These prompts are generic and vary according to the task setting.
  • H.4 Reward-Model Scoring Interfaces: Reward models score images independently using their released interfaces, with num_trials scores averaged and converted into within-task rankings for predicted best and worst images.Unlike MLLMs, reward models do not share a single comparative prompt.
  • H.4.1–H.4.3 Reward-Model Interfaces: Reward-model interfaces vary: ArtiMuse decodes two-letter outputs to [0, 100], Q-Align weights five token categories, and Q-Instruct uses the good–poor logit difference.ArtiMuse and Q-Align use implicit prompts, whereas Q-Instruct uses an explicit image-token-wrapped instruction.
  • H.4.4–H.4.5 and I Topic-Level Annotation Examples: RealQA parses generated scores in [0.0, 10.0], PEAS models predict expected values from 10-bin distributions, and the appendix includes one annotation example for each of 24 VAB topics.PEAS-aes and PEAS-ava use no language prompt, while RealQA can issue a constrained retry when parsing fails.
Loading 2605.12684v1…