Source-linked AI summary
On Semiotic-Grounded Interpretive Evaluation of Generative Art
Ruixiang Jiang, Changwen Chen
TL;DR
GenArt evaluation often focuses on appearance or literal prompt matching, leaving symbolic and abstract meaning underassessed. The paper models HGI as cascaded Peircean semiosis and introduces SemJudge with Hierarchical Semiosis Graphs to reconstruct meaning from prompts to artifacts. SemJudge aligns more closely with human judgments and yields more informative interpretations, while its benchmark remains culturally and artistically incomplete.
Problem
GenArt evaluators often focus on realism, prompt-image alignment, and visual appeal rather than deeper artistic meaning.
Method
SemJudge reconstructs prompt-to-artifact meaning conveyance using an interpretation-centric evaluator and Hierarchical Semiosis Graphs.
Results
SemJudge aligns more closely with human judgments and produces more informative, auditable interpretations of artistic meaning.
Takeaways & Limitations
The framework shifts GenArt evaluation toward recognizing the deeper ideas and intentions embedded in visual art.
Takeaways & Limitations
SemiosisArt may underrepresent cultural minorities and contemporary conceptual art because stable shared human judgments are harder to obtain for these categories.
Abstract
from arXiv · showhide
Interpretation is essential to deciphering the language of art: audiences communicate with artists by recovering meaning from visual artifacts. However, current Generative Art (GenArt) evaluators remain fixated on surface-level image quality or literal prompt adherence, failing to assess the deeper symbolic or abstract meaning intended by the creator. We address this gap by formalizing a Peircean computational semiotic theory that models Human-GenArt Interaction (HGI) as cascaded semiosis. This framework reveals that artistic meaning is conveyed through three modes - iconic, symbolic, and indexical - yet existing evaluators operate heavily within the iconic mode, remaining structurally blind to the latter two. To overcome this structural blindness, we propose SemJudge. This evaluator explicitly assesses symbolic and indexical meaning in HGI via a Hierarchical Semiosis Graph (HSG) that reconstructs the meaning-making process from prompt to generated artifact. Extensive quantitative experiments show that SemJudge aligns more closely with human judgments than prior evaluators on an interpretation-intensive fine-art benchmark. User studies further demonstrate that SemJudge produces deeper, more insightful artistic interpretations, thereby paving the way for GenArt to move beyond the generation of "pretty" images toward a medium capable of expressing complex human experience. Project page: https://github.com/songrise/SemJudge.
1 Introduction
The paper argues that GenArt evaluation must assess artistic meaning, not only appearance or literal prompt adherence. It introduces SemJudge, which reconstructs meaning conveyance from prompts to artifacts using semiotic structure.
- Motivation: Existing GenArt evaluators emphasize realism, prompt-image alignment, or visual appeal, while often misaligning with trained human judgments.These metrics largely assess what is visually observable rather than deeper artistic meaning.
- Motivation: Artistic meaning often relies on juxtaposition, abstraction, and metaphor, so surface fidelity can diverge from the intended message.The paper uses Guernica to illustrate how distortion and fragmentation convey moral outrage and an anti-war stance.
- Motivation: Prompts frequently express artistic directions about vibe, theme, or motif rather than fully specified visual layouts.A prompt such as “in the spirit of Guernica” requires interpretation rather than literal rendering.
- Approach: Semiotics models HGI as meaning communication from creator intention through prompts and generated artifacts, exposing conventional metrics’ iconicity bias.The framework addresses meaning conveyed through metaphor, symbolism, or convention rather than literal resemblance.
- Approach: SemJudge uses Hierarchical Semiosis Graphs to link interpretive claims with prompt spans and image regions, extending evaluation beyond surface alignment.The resulting representation supports interpretation-based as well as resemblance-based criteria.
- Findings: SemJudge aligns more closely with human judgments and produces more informative, auditable interpretations of artistic meaning.The reported validation uses the SemiosisArt dataset.
2 Related Work
Related work has progressed from realism and text-image alignment toward preference and structured evaluation, while computational art interpretation and semiotics provide complementary foundations. The paper identifies remaining shortcomings for evaluating meaning in generated art.
- GenArt Evaluation: Early GenArt metrics evaluated realism through distances between generated and real image distributions, including Inception Score, FID, and ArtFID.These metrics focus on distributional similarity rather than interpretation.
- GenArt Evaluation: Text-conditional generation shifted evaluation toward text-image alignment, later supplemented by preference models for generic visual appeal.PickScore and HPS provide global preference scores but remain black boxes.
- GenArt Evaluation: Question Generation and Answering models make evaluation more interpretable and structured, but existing approaches still leave deep artistic meaning insufficiently addressed.The related-work discussion contrasts structured evaluation with the unresolved meaning-level problem.
- Art Interpretation: Computational art-interpretation methods use retrieval or curated-dataset tuning, but canonical artworks can make performance difficult to separate from memorization.This concern limits their direct use for GenArt evaluation.
- Computational Semiotics: Computational semiotics formalizes meaning-related concepts for intelligent systems and HCI, while showing that contemporary AI often manipulates surface patterns without genuine semiotic grounding.This provides theoretical motivation for modeling sign relations in GenArt evaluation.
3 Human-GenArt Interaction as Semiosis
The paper formalizes HGI as cascaded Peircean semiosis, in which signs, objects, and interpretants are repeatedly related through interpretation and reification. In generation, prompts are interpreted into representations, artifacts are synthesized, and those artifacts are interpreted again.
- 3.1 Formulating Peircean Triadic Semiosis: Peircean semiosis represents meaning-making as a triadic relation among a sign, an object, and an interpretant.The sign is the perceptible form, the object is the referent or intended content, and the interpretant is constructed meaning.
- 3.1 Formulating Peircean Triadic Semiosis: Interpretation is modeled as interpreter-dependent: an interpreter maps a sign to an interpretant.The interpreter may be human or computational.
- 3.1 Formulating Peircean Triadic Semiosis: Signs can be iconic, symbolic, or indexical, and a single sign may involve all three grounds to different degrees.Because art often combines resemblance, convention, allegory, and contextual reference, resemblance-only evaluation is unreliable.
- 3.1 Formulating Peircean Triadic Semiosis: The semiotic ground is treated as a computational evidence layer that maps, under an interpreter, to the immediate object represented by a sign.This separates the external intent or reality from the object as represented within the sign.
- 3.2 Human-GenArt Interaction as Semiosis: Cascaded semiosis occurs when an interpretant from one stage becomes the next sign and is interpreted again.The cascade includes successive atomic semiosis units, interpreters, and reification processes such as image generation.
- 3.2 Human-GenArt Interaction as Semiosis: In a basic HGI workflow, a user’s intended goal is expressed as a prompt, interpreted by the generator, reified as an artifact, and interpreted by another evaluator.Even single-round generation therefore forms at least a two-round cascade.
4 Semiotics-Grounded GenArt Evaluation
This section frames HGI evaluation as cascaded semiosis and argues that conventional ground-space metrics can fail when artistic meaning is indirect. It introduces HSG-based SemJudge to reconstruct interpretable meaning from prompts to artifacts.
- 4.1 Semiosis Quality Measure: Semiosis-grounded evaluation treats HGI quality as the quality of meaning-making induced by human–GenArt interaction.
- 4.1 Semiosis Quality Measure: An N-round semiosis is theoretically evaluated by the distance between its initial and final dynamic objects, with smaller distance indicating higher quality.Because dynamic objects are latent, empirical quality uses interpreter-reconstructed immediate objects.
- 4.2 Demystifying Conventional GenArt Metrics: The Interpretive Principle states that increasing mismatch between intended and interpreted iconicity decreases semiosis quality.The framework illustrates this with symbolic art that an iconicity-biased evaluator may misread as poor depiction.
- 4.2 Demystifying Conventional GenArt Metrics: Conventional GenArt metrics operate in ground space, comparing prompt-image grounds or artifacts against idealized priors rather than recovering interpreted objects.Prompt-aware metrics include CLIP, PickScore, and MLLM-based scoring; prompt-agnostic metrics include FID and aesthetic predictors.
- 4.2 Demystifying Conventional GenArt Metrics: Their shared limitation is treating ground-space comparison as a universal proxy for semiosis quality, which can yield high scores despite low human satisfaction when symbolic meaning is missed.The mismatch can arise during generation or evaluation when intended iconicity diverges from interpreted iconicity.
- 4.3 The SemJudge: SemJudge introduces Hierarchical Semiosis Graphs whose nodes encode atomic semioses and whose edges represent relations such as support, elaboration, and contrast.Localizable sub-semioses can be grounded in prompt spans and image bounding boxes for fine-grained, auditable analysis.
5 The SemiosisArt
SemiosisArt addresses the difficulty of evaluating symbolic and indexical artistic meaning by grounding tasks in canonical motifs and combining comparative judgment with fine-grained interpretation.
- Challenge: Existing GenArt and art-interpretation benchmarks are poorly aligned with meaning-level evaluation because most GenArt sets emphasize iconic prompts and appearance quality.
- Dataset design: SemiosisArt uses canonical motifs rooted in traditions such as iconology, culture, theology, and literature to reduce interpretive arbitrariness.The construction process collaborates with 12 experts and uses quality control for comparative tasks.
- Tasks: Figure 3 depicts prompts constructed from canonical motifs, images generated by multiple models, and two evaluation formats: 2AFC and VQA.
- Tasks: The dataset contains 1,870 2AFC comparative judgment tasks and 600 VQA questions.2AFC captures relative quality at the instance level, while VQA probes fine-grained semiotic interpretation.
6 Experiment and Analysis
The experiments evaluate SemJudge against conventional scorers and structured or interpretive baselines using human-alignment metrics, art-interpretation accuracy, subjective ratings, iconicity-bias tests, and ablations. Across these analyses, SemJudge shows stronger alignment and interpretation quality, while HSG structure is identified as a major contributor; the benchmark remains culturally bounded.
- Experiment Settings: SemJudge is compared with scoring models, structured-rationale evaluators, and art-interpretation or aesthetic models across correlation and VQA tasks.The evaluation uses KRCC, SRCC, CCC, and VQA accuracy to measure alignment with human judgments.
- Quantitative Correlation Experiment: Conventional low-level scorers show near-zero or weak correlation with expert judgments, indicating that appearance and generic preference signals are insufficient for symbolic and indexical meaning.The comparison includes image-quality, prompt-alignment, and preference-based scoring methods.
- Quantitative Correlation Experiment: SemJudge achieves the strongest overall alignment with expert judgments across instance concordance, discrete rank correlation, and continuous Elo correlation.The advantage persists across Qwen-9B and Gemini-Flash backbones.
- Quantitative Correlation Experiment: 92.4% VQA accuracy for Gemini-3.1-Flash-lite approaches the 93.2% expert-human performance, while SemJudge achieves the best overall interpretation performance.The result indicates strong performance on explicit, fine-grained art interpretation.
- Subjective Interpretation Quality: SemJudge is significantly preferred across all four subjective interpretation dimensions, including causal agreement and depth.Figure 4 reports feedback from 70 users on 5-point Likert ratings, and the study collected 4,943 responses overall.
- Iconicity Bias: Conventional evaluators exhibit significant iconicity bias, whereas SemJudge’s agreement with humans is not concentrated on highly iconic artworks and extends to symbolic and indexical works.The iconicity-bias test uses Δ, bootstrap confidence intervals, Cohen’s d, and one-sided permutation-test significance.
- Ablation Study: Ablations show that HSG structure improves fixed-judge performance, strong transferred HSGs elevate weak judges, and these gains are especially pronounced for VQA interpretation.The controlled design separates contributions from HSG construction and final-judge scaling.
- Limitations: SemiosisArt is culturally grounded in Christian, East Asian, Hindu, and Islamic traditions and modern motifs but may underrepresent cultural minority and contemporary conceptual art.These categories are harder to evaluate through stable shared human judgments.
7 Conclusion
The study identifies a critical gap in GenArt evaluation: conventional metrics emphasize appearance, while SemJudge reconstructs meaning-making and improves human correlation for judging and interpreting GenArt.
- Conventional metrics struggle to grasp the symbolic and indexical depth of visual art.
- SemJudge shifts evaluation from surface-level appearance to the mechanics of meaning-making.
- SemJudge reconstructs the interpretive process connecting creator intention and generated artifacts.
- The study reports a significant improvement in human correlation for judging and interpreting GenArt.
A Details on the SemiosisArt
SemiosisArt is a meaning-oriented GenArt benchmark designed to evaluate symbolic and indexical interpretation rather than appearance alone. It combines expert-designed prompts, iterative crowd quality control, generated images, comparative judgments, and fine-grained VQA annotations.
- SemiosisArt focuses on meaning and interpretation, rather than appearance, by using prompts whose intended meanings rely substantially on symbolic or indexical interpretation.
- 187 HSG initiatives produce 935 images, 1,870 pairwise comparative judgments, and 600 VQA questions for interpretation benchmarking.
- Experts propose challenging low-iconicity prompts grounded in canonical motifs, then retain prompts whose generated images achieve sufficient inter-subjective agreement through crowdsourced 2AFC judgments.
- 0.58 Cohen’s κ agreement was obtained for non-expert annotators after the iterative process, despite the interpretive nature of the judgments.
- VQA questions are generated semiautomatically from expert image and region-of-interest annotations, then regenerated when they fail automatic or expert quality checks.
- High-iconicity tasks emphasize low-level features or identity preservation, whereas low-iconicity tasks involve history, convention, stories, or causality.
B.1 Bounding Box Grounding Quality
SemJudge evaluates bounding boxes by their interpretive usefulness for semiotic analysis rather than exact overlap with fixed ground-truth boxes. A user study collected satisfaction judgments and identified differences across models, while the authors note a limitation in zero-shot localization.
- Because SemJudge generates open-vocabulary sub-sign descriptions, fixed-ontology metrics such as mIoU lack static ground truth for evaluating its predicted concepts.
- Bounding boxes are evaluated by whether they provide valid visual evidence for the model’s semiotic argument, using human satisfaction rate instead of exact box matching.
- 450 satisfaction annotations were collected by asking annotators whether image-and-box visualizations were satisfactory.
- 74.7% satisfaction was achieved by Gemini-3.1-Flash-Lite, compared with 56.0% for Qwen-3.5-35B-A3B and 57.8% for Qwen-3.5-9B.
- MLLMs remain limited in precise zero-shot bounding-box prediction, motivating a dedicated grounding module such as GroundingDINO as future work.
B.2 Details of the Iconicity-Bias Analysis
The iconicity-bias analysis tests whether evaluator–human agreement is concentrated on visually resembling cases. It defines semiotic components and net iconicity, then uses aligned-versus-misaligned comparisons with permutation and bootstrap-based inference.
- The analysis quantifies subjective iconicity to test whether conventional GenArt evaluators agree with humans primarily on iconic prompt–artifact relations.
- Net iconicity aggregates six experts’ 7-point ratings of iconicity, indexicality, and symbolism across signs and instances.
- Positive net iconicity indicates iconic-resemblance dominance, while negative values indicate stronger symbolic and indexical contributions.
- The analysis compares average iconicity between evaluator–human aligned and misaligned instances, defining Δ as the difference between these subsets.
- A positive Δ is interpreted as iconicity bias and tested with a one-sided permutation test, a one-sided 95% bootstrap confidence interval, and Cohen’s d.
- Robust evaluators should not show a significantly positive Δ because their human agreement should not concentrate on highly iconic cases.
- An atomic semiosis comprises an object, sign, and interpretant, with an interpreter mapping signs to meanings through evidential grounds and related functions.
- Cascaded semiosis represents a chain of N atomic semioses, and HGI evaluation uses hierarchical semiosis graphs to visualize prompt and output signs.
C Implementation Details
SemJudge is implemented as a staged procedure that reconstructs prompt and artifact semiosis before grounding a judgment in cited evidence. The accompanying HSG examples connect visual elements, literary or religious context, and iconic, symbolic, or indexical grounds.
- SemJudge procedure: Algorithm C.1 organizes SemJudge into prompt-semiosis reconstruction, artifact-semiosis reconstruction, and evidence-grounded judgment stages.The judgment stage compares reconstructed chains, cites evidence, parses a binary judgment, and extracts node-level rationales.
- HSG construction settings: Standard HSG construction permits up to three nodes with succinct descriptions, whereas complex construction permits up to five nodes with more detailed descriptions.These settings are specified for the input prompt, image, and 2AFC summarization system prompts.
- HSG example: Chinese ink-wash painting: The Jiang Jie example represents a triptych of youth, middle age, and old age, with the structure conventionally encoding progression or sequence.The example links poem verses and visual scenes to themes including joy, intimacy, isolation, hardship, and life’s responsibilities.
- HSG example: Chinese ink-wash painting: The Chinese painting example grounds meaning through visual resemblance, conventional symbolic language, and indexical cues such as turbulent water and dark clouds implying hardship.Calligraphy and literary context anchor the visual metaphors to the poem’s meaning.
- HSG visualizations: The visualizations compare SemJudge HSG outputs with analyses from other models across Chinese ink-wash, retro-game Nativity, and modern vector-art examples.The figures present three artifact-sign cases, each pairing a prompt with a top SemJudge HSG and lower compared-model analysis.
- HSG example: Nativity scene: The Nativity example combines pixel-art and FPS-HUD conventions with Christian iconography, using numbers, bars, and text to convey abstract status information.The composition recontextualizes the Biblical Nativity within a retro video-game engine and frames the scene as an interactive digital experience.