Source-linked AI summary

AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability

Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu, Jianjun Gao

arXiv:2608.26713v1cs.CVcs.AI

TL;DR

Existing aesthetic benchmarks leave unclear whether visually appealing images fit particular purposes, audiences, cultures, and conventions. AesCanvas addresses this gap with large-scale critique supervision and expert-reviewed suitability scenarios, finding that critique capability does not reliably transfer to contextual judgment. Aesthetic specialists remain competitive on selected critique metrics but substantially lag strong general-purpose MLLMs on ContextCanvas.

  • Problem

    Existing benchmarks mainly assess intrinsic visual quality or fixed domain criteria, leaving contextual suitability for a stated purpose, audience, cultural setting, or convention insufficiently evaluated.

  • Method

    AesCanvas combines CritiqueCanvas, with 519,136 instruction–response pairs from 54,300 images, and ContextCanvas, with 301 expert-reviewed use scenarios.

  • Results

    Aesthetic specialists achieve below 30% accuracy on ContextCanvas, while Claude Opus 5 reaches 91.69%, showing critique performance and contextual suitability are distinct capabilities.

  • Takeaways & Limitations

    Culturally situated suitability should be treated as a distinct objective requiring integration of visual evidence, cultural knowledge, intended use, and domain-specific constraints.

Abstract

from arXiv · show

Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward interpretable critique and guidance. Yet existing benchmarks mainly assess intrinsic visual quality or fixed domain criteria, leaving open whether an appealing image is appropriate for a specific purpose, audience, cultural setting, or domain convention. We introduce AesCanvas, a unified suite with two complementary components: CritiqueCanvas with 519,136 instruction-response pairs from 54,300 images supports long-form, multi-dimensional critique across photography, painting, and virtual imagery, whereas ContextCanvas with 301 expert-reviewed use scenarios evaluates contextual aesthetic suitability in realistic use scenarios. Under a unified protocol, we evaluate closed-source frontier, open-weight general, and aesthetic-specific MLLMs. Results reveal a clear separation between critique generation and context-sensitive judgment: reference-based lexical and semantic metrics only partially capture critique quality, while aesthetic specialists remain competitive on selected critique metrics yet substantially lag strong general-purpose MLLMs on ContextCanvas. Further analyses show that aesthetic specialization does not reliably transfer to contextual suitability and that model decisions may fail to track or ground themselves in decisive contextual visual cues. These findings establish culturally situated, evidence-grounded suitability as a distinct objective for aesthetic modeling.

I Introduction

Aesthetic assessment has expanded from scalar quality prediction to language-based critique, but practical judgment also requires contextual suitability grounded in purpose, audience, and convention. AesCanvas evaluates these complementary capabilities and finds that fluent critique performance does not reliably ensure context-sensitive decisions.

  • Image Aesthetic Assessment has progressed from scalar prediction and preference modeling toward language-based aesthetic perception, critique, diagnosis, and guidance.
  • Visual appeal alone is insufficient because symbols, styles, dress, and genre conventions can change whether an image suits a particular purpose, audience, or convention.
  • Models may produce plausible critiques while mistaking visually appealing attributes for appropriateness in a mismatched use context.
  • Four recurring failure modes include visually ungrounded critiques, inappropriate domain priors, and suitability judgments based on surface aesthetics or stylistic plausibility.
  • AesCanvas combines large-scale multi-domain critique supervision with expert-reviewed use scenarios to evaluate critique grounding and contextual suitability.
  • Aesthetic specialists remain competitive on selected critique metrics, yet all score below 30% on ContextCanvas versus 91.69% for the strongest evaluated general-purpose model.

II Related Work

Related work has extended Image Aesthetic Assessment from scalar prediction toward language-based perception, scoring, and critique. AesCanvas instead examines whether aesthetic articulation transfers to contextual decisions involving intended use.

  • Early Image Aesthetic Assessment modeled classification, ranking, or score-distribution prediction from crowd preferences and incorporated photographic attributes and taste variation.
  • Aesthetic multimodal models now support language-based recognition, description, interpretation, scoring, critique, professional critique, and multi-attribute analysis.
  • Figure 3 depicts a shared multi-domain image pool branching into long-form aesthetic critique construction and expert-reviewed contextual suitability construction.
  • Unlike prior work centered on aesthetic perception, scoring, or reference-aligned critique, AesCanvas tests whether attractive or stylistically plausible images conflict with intended use.

C Contextual Multimodal Reasoning

Contextual multimodal reasoning combines visual evidence with task conditions, external knowledge, cultural conventions, and criterion-specific constraints. ContextCanvas targets whether contextual evidence should outweigh surface appeal in concrete visual-use decisions.

  • Contextual multimodal reasoning requires combining visual evidence with task conditions, external knowledge, cultural conventions, and criterion-specific constraints.
  • Existing benchmarks primarily assess factual correctness, grounding, functional inference, or criterion adherence rather than context-sensitive aesthetic suitability.
  • ContextCanvas evaluates use-oriented aesthetic suitability judgments grounded in both visual and contextual evidence.

III Dataset

AesCanvas uses a shared multi-domain image pool to construct CritiqueCanvas for structured critique and ContextCanvas for expert-reviewed suitability judgments. Its curation combines dimension-aware annotation, verification, and prespecified criteria for context-dependent, visually answerable cases.

  • Dataset Overview: AesCanvas contains 519,136 instruction–response pairs from 54,300 images and 301 expert-designed, reviewed cases spanning multiple visual domains.
  • Dataset Construction: The pipeline branches from a shared multi-domain image pool into CritiqueCanvas for structured aesthetic critique and ContextCanvas for closed-form contextual suitability judgment.
  • CritiqueCanvas: CritiqueCanvas uses shared aesthetic dimensions, domain-specific criteria, prompt-conditioned annotation, and multiple MLLMs to create multi-turn critique conversations.
  • CritiqueCanvas: An independent Claude Opus 5 verifier screens generated pairs for instruction consistency, visual grounding, and response quality; a 1,000-pair audit reports 95.3% acceptance and 97% raw agreement.
  • ContextCanvas: Four PhD-level researchers developed and reviewed 1,060 candidate ContextCanvas cases in plausible exhibition, publication, advertising, education, commemoration, and communication scenarios.
  • ContextCanvas: Retained cases must satisfy criteria including context dependence, visual necessity, decisive evidence, scenario naturalness, answerability, and anti-shortcut validity.
  • ContextCanvas: The final benchmark contains 301 cases covering 302 images, 291 binary questions, and 10 three-way questions across painting, sculpture, illustration, animation, photography, and film.

C Evaluation Protocol

The evaluation uses a unified, tool-free protocol to assess long-form critique generation and contextual suitability with complementary metrics and diagnostics. Critique metrics capture lexical, semantic, and image–text relevance, while coverage statistics measure breadth rather than correctness.

  • Models receive identical inputs and instructions without web search, retrieval, or auxiliary recognition tools for both evaluation tasks.
  • 6,000 open-ended samples balanced across photography, painting, and virtual imagery evaluate aesthetic analysis and image-grounded improvement suggestions.Image-level splits keep related conversations together.
  • BLEU, ROUGE-L, and METEOR measure lexical overlap, while BERT-F1, SBERT-Cos, and CLIPScore measure semantic similarity and image–text relevance.These metric families are interpreted jointly rather than as complete measures of aesthetic reasoning.
  • ContextCanvas reports exact-match accuracy on all 301 cases, with MacroF1 and predicted Yes Rate computed on 291 binary cases against a 38.14% gold Yes Rate.The remaining 10 three-way cases contribute only to accuracy.
  • Average dimensions addressed per response measure explicit aesthetic coverage breadth, not whether the selected dimensions are relevant, accurate, or grounded.A higher coverage value therefore does not necessarily indicate better critique quality.
  • Contextual predictions are correct only when a unique parsed option matches the gold answer; missing, conflicting, or unparsable responses count as incorrect.

IV Experiments

Experiments compare closed-source frontier, open-weight general, and aesthetic-specific MLLMs on contextual suitability and long-form critique generation. Results show a gap between context-sensitive judgment and reference-aligned critique metrics, with specialization competitive on selected critique measures but weak on contextual suitability.

  • Contextual Aesthetic Suitability Judgment: ContextCanvas accuracy ranges from 44.19% to 91.69% among closed-source MLLMs, while Qwen3-VL-Instruct reaches 42.19% as the strongest evaluated open-weight model.
  • Contextual Aesthetic Suitability Judgment: All evaluated aesthetic-specific models score below 30% accuracy on ContextCanvas, with ArtQuant-APDD highest at 29.57% versus 42.19% for Qwen3-VL-Instruct and 91.69% for Claude Opus 5.ArtiMuse, Q-SiT, and AesExpert obtain 19.60%, 19.60%, and 18.27%, respectively.
  • Contextual Aesthetic Suitability Judgment: Aesthetic-specific models predict Yes in 59.11%–76.98% of binary cases against a 38.14% gold rate, indicating pronounced over-acceptance despite contextual mismatch.Frontier models reach up to 91.69% accuracy on the full benchmark, suggesting performance is not explained by majority-label guessing.
  • Long-form Aesthetic Critique Generation: No model dominates all CritiqueCanvas metric families: Gemini 3.1 Pro leads SBERT-Cos at 0.848, LLaVA-OneVision leads BERT-F1 at 0.724 and CLIPScore at 0.319, and Qwen3-VLFT leads lexical scores.Qwen3-VLFT obtains 12.23 BLEU, 0.297 ROUGE-L, and 0.296 METEOR.
  • Long-form Aesthetic Critique Generation: Qwen3-VLFT improves over Qwen3-VL by 7.11 BLEU, 0.065 ROUGE-L, and 0.022 METEOR, but SBERT-Cos falls from 0.827 to 0.787 and CLIPScore from 0.313 to 0.285.Adaptation improves reference-style matching without uniform gains in semantic similarity or image grounding.
  • Long-form Aesthetic Critique Generation: Aesthetic-specific models remain competitive on selected critique metrics, while overall critique performance is heterogeneous across lexical, semantic, and image-grounding criteria.ArtiMuse reaches 0.718 BERT-F1 and UniPercept reaches 0.318 CLIPScore, near the respective best scores.
  • Together, the tasks reveal a gap between reference-aligned critique generation and contextually valid aesthetic judgment.

C Analysis

The analysis tests whether critique metrics, aesthetic specialization, and contextual judgments align. Results show that metric overlap is incomplete, specialization transfers inconsistently, and decisive contextual cues are often missed.

  • Validity of Critique Metrics: Human quality ratings do not consistently track reference-similarity metrics, so critique evaluation should include grounding, specificity, and usefulness.Valid critiques can emphasize different evidence and interpretations, limiting lexical and semantic overlap as a complete quality measure.
  • Counterfactual Audit: The 36-pair counterfactual audit holds scenarios and overall presentation fixed while changing the decisive cue to reverse the intended judgment.The variants are diagnostic matched pairs, not a revised annotation set or additional accuracy split.
  • Counterfactual Audit: 25.00–30.56 NCU values indicate positive net updating for all three general-purpose MLLMs despite substantially different benchmark accuracies.Absolute contextual competence and responsiveness to decisive visual evidence are therefore related but separable.
  • Base–Specialist Transfer: Aesthetic specialization improves ArtQuant and AesExpert over their base models, but produces no measurable gain for ArtiMuse or Q-SiT.Strong CritiqueCanvas metric alignment also does not consistently translate into ContextCanvas accuracy.
  • Decisive-Cue Grounding: ArtiMuse falls from 21% accuracy to 3% Evidence-Grounded Accuracy, showing that correct answers can lack decisive-cue grounding.Evidence-Grounded Accuracy requires correctness, decisive-cue identification, and contextual explanation.

V Conclusion

AesCanvas evaluates aesthetic critique alongside contextual suitability and finds that critique strength does not reliably transfer to context-sensitive judgment. The findings support culturally situated suitability as a distinct objective requiring integrated visual, cultural, functional, and domain-specific reasoning.

  • AesCanvas combines CritiqueCanvas and ContextCanvas to evaluate aesthetic critique and contextual suitability across multiple MLLM classes.
  • Strong critique performance does not reliably translate into context-sensitive judgment: aesthetic specialists can remain competitive on critique metrics yet show low ContextCanvas accuracy and over-acceptance.
  • Culturally situated suitability should be treated as a distinct training and evaluation objective integrating visual evidence, cultural knowledge, intended use, and domain-specific constraints.

Supplementary Material

The supplementary material identifies the paper and its focus on aesthetic critique and contextual suitability.

  • The supplementary material presents AesCanvas as a benchmark for aesthetic critique and contextual suitability.

S1 Supplement Overview and Positioning

The supplement documents the dataset, evaluation, and implementation materials, and positions AesCanvas against existing aesthetic resources. Its distinguishing affordance is auditable suitability judgment for concrete use contexts alongside long-form critique.

  • Supplement Overview: Sections S2 and S3 cover dataset, prompt, evaluation, implementation, model-output, robustness, and transfer analyses.
  • Positioning: Use-context judgment requires an auditable decision about whether an image suits a stated purpose, audience, or convention, rather than generic context reasoning or intrinsic scoring.
  • Positioning: The comparison spans photographic-aesthetics datasets, AI-generated-image quality datasets, and multimodal aesthetic-critique resources.
  • Positioning: Existing resources provide substantial intrinsic aesthetic supervision, while AesCanvas couples long-form critique with suitability judgments that can change under cultural, narrative, functional, or domain-specific evidence.

S2 Dataset and Evaluation Materials

AesCanvas combines large-scale multi-domain critique supervision with expert-reviewed contextual suitability cases, alongside evaluation and governance procedures designed to assess grounding, quality, and scope. Its materials distinguish valid but diverse critiques from context-sensitive visual-use judgments.

  • Dataset Overview: AesCanvas contains CritiqueCanvas and ContextCanvas as complementary components with distinct data units, outputs, and evaluation roles.CritiqueCanvas supports long-form critique, while ContextCanvas evaluates contextual aesthetic suitability in use scenarios.
  • CritiqueCanvas: CritiqueCanvas covers photography, painting, and virtual imagery using ten shared aesthetic dimensions plus domain-conditioned criteria.The dimensions include content, composition, color, lighting, style, emotion, technique, symbolism, and visual appeal, with image-relevant subsets selected per record.
  • ContextCanvas: ContextCanvas retains 301 cases covering 302 unique images, with model inputs restricted to images, questions, and fixed options while labels and rationales remain withheld.Each case stores a gold option, source-grounded rationale, source basis, and provenance metadata for evaluation and governance.
  • ContextCanvas: ContextCanvas cases are organized around cultural, historical, narrative, functional, and medium-specific mechanisms that make aesthetic-use decisions non-trivial.Examples separate visible cues, activated contextual knowledge, and resulting visual-use decisions rather than testing title recall or context-free cultural knowledge.
  • Quality Control: Expert review rejected cases whose answers could be inferred from coarse emotion, depended on off-image events, or admitted defensible alternative labels.Professional wording and plausible intended answers were insufficient without visual necessity and a uniquely compelled gold label.
  • Scope and Governance: The datasets reflect their source pool and contributor expertise, and ContextCanvas is an expert diagnostic benchmark rather than exhaustive cultural or design coverage.CritiqueCanvas references are supervision targets, not the only legitimate analyses of an image.
  • Quality Control: Five reviewers accepted 953 of 1,000 audited CritiqueCanvas pairs after response-quality review and a separate asset-clarity validation.The reported 97.0% raw agreement is unanimous absence of blocking issues, while 17 otherwise acceptable pairs were removed for failing the final clarity threshold.
  • Evaluation Materials: Critique evaluation combines lexical, semantic, image–text, and human rubric-based measures, while lexical diagnostics do not establish correctness, relevance, or grounding.The raw CLIP implementation uses mean cosine similarity without max truncation or 2.5× scaling, and density counts repeated mentions rather than distinct categories.

S3 Additional Analyses and Insights

Additional analyses show that contextual suitability errors reflect systematic acceptance biases, weak grounding in decisive cues, and limited transfer from aesthetic specialization. They also show that critique breadth and reference similarity do not reliably indicate critique usefulness or contextual competence.

  • Behavioral patterns: Aesthetic-specific models predict Yes for 59.11–76.98% of cases, compared with 20.96–30.24% for the strongest closed-source models, despite a 38.14% gold Yes rate.The differing rates reveal systematic decision tendencies rather than equivalent contextual judgment.
  • Behavioral patterns: Q-SiT’s 76.98% Yes rate coincides with only 4.05 No-class F1, whereas GPT-5.2’s conservative 20.96% Yes rate produces a different error profile.Class-specific F1 exposes failures that aggregate scores can conceal.
  • Grounding and evidence use: Grounding audits separate label accuracy from evidence use: Gemini retains 79 of 83 correct decisions under EGA, while ArtiMuse retains only 3 grounded correct answers from 21.ArtiMuse’s substantive rationales make the drop difficult to attribute to missing explanations.
  • Failure modes: Reviewed examples show cultural-neighbor substitution, aesthetic-halo errors, and salient-composition myopia, but these subtypes are not corpus-wide frequency estimates.The labels summarize illustrative casebook patterns and would require independent coding to support frequency claims.
  • Critique behavior: Qwen3-VL produces 199.07-word critiques, while UniPercept reaches 13.26 dimension density at 102.12 words; breadth and prompt alignment likewise vary independently.Gemini covers 6.53 dimensions, whereas Qwen3-VL reaches 95.2% focused-prompt alignment; longer or broader critiques are not thereby shown correct or grounded.
  • Critique evaluation: Lexical and semantic reference alignment only partially reflects grounding, image specificity, criterion choice, and usefulness, so lower BLEU, SBERT, or CLIPScore need not imply poorer critique.Human direct-quality ratings favored GPT-5.2 and Gemini over adapted or aesthetic-specific models, despite those models not dominating every reference-similarity metric.
  • Counterfactual and grounding analyses: Counterfactual and rationale analyses reveal source-scene anchoring, where responses repeat a canonical narrative after its supporting visual cue is removed.This diagnostic complements grounding audits by testing whether decisions track changed visual evidence.
  • Visual information and specialization: Text-only performance stays at or below the 61% majority baseline, while neutral descriptions improve general-purpose models by 28–34 points; specialization still yields heterogeneous ContextCanvas effects.ArtQuant and AesExpert improve over their bases after correction, but ArtiMuse and Q-SiT show no measurable gain.
Loading 2608.26713v1…