Paper deep dive
On Semiotic-Grounded Interpretive Evaluation of Generative Art
Ruixiang Jiang, Changwen Chen
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/14/2026, 1:44:21 AM
Summary
The paper introduces SemJudge, an interpretation-centric evaluation framework for Generative Art (GenArt) based on Peircean computational semiotics. It addresses the 'structural blindness' of existing metrics that focus solely on surface-level image quality or literal prompt adherence. By modeling Human-GenArt Interaction (HGI) as a cascaded semiosis and utilizing Hierarchical Semiosis Graphs (HSG), SemJudge reconstructs the meaning-making process, allowing for the assessment of symbolic and indexical meaning in addition to iconic resemblance.
Entities (5)
Relation Signals (3)
SemJudge → utilizes → Hierarchical Semiosis Graph
confidence 100% · This evaluator explicitly assesses symbolic and indexical meaning in HGI via a Hierarchical Semiosis Graph (HSG)
Peircean semiotics → informs → SemJudge
confidence 95% · Building on this theory, we propose SemJudge, an interpretation-centric HGI evaluator
SemJudge → models → Human-GenArt Interaction
confidence 95% · We propose SemJudge, an interpretation-centric HGI evaluator that reconstructs how meaning is carried from prompt to generated artifact
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Interpretation is essential to deciphering the language of art: audiences communicate with artists by recovering meaning from visual artifacts. However, current Generative Art (GenArt) evaluators remain fixated on surface-level image quality or literal prompt adherence, failing to assess the deeper symbolic or abstract meaning intended by the creator. We address this gap by formalizing a Peircean computational semiotic theory that models Human-GenArt Interaction (HGI) as cascaded semiosis. This framework reveals that artistic meaning is conveyed through three modes - iconic, symbolic, and indexical - yet existing evaluators operate heavily within the iconic mode, remaining structurally blind to the latter two. To overcome this structural blindness, we propose SemJudge. This evaluator explicitly assesses symbolic and indexical meaning in HGI via a Hierarchical Semiosis Graph (HSG) that reconstructs the meaning-making process from prompt to generated artifact. Extensive quantitative experiments show that SemJudge aligns more closely with human judgments than prior evaluators on an interpretation-intensive fine-art benchmark. User studies further demonstrate that SemJudge produces deeper, more insightful artistic interpretations, thereby paving the way for GenArt to move beyond the generation of "pretty" images toward a medium capable of expressing complex human experience. Project page: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.08641v1
- Canonical: https://arxiv.org/abs/2604.08641v1
Trouble viewing inline? Open PDF directly →
Full Text
119,955 characters extracted from source content.
Expand or collapse full text
On Semiotic-Grounded Interpretive Evaluation of Generative Art Ruixiang Jiang rui-x.jiang@connect.polyu.hk The Hong Kong Polytechnic University Hong Kong SAR, China Chang Wen Chen chen.changwen@polyu.edu.hk The Hong Kong Polytechnic University Hong Kong SAR, China Abstract Interpretation is essential to deciphering the language of art: au- diences communicate with artists by recovering meaning from visual artifacts. However, current Generative Art (GenArt) evalua- tors remain fixated on surface-level image quality or literal prompt adherence, failing to assess the deeper symbolic or abstract mean- ing intended by the creator. We address this gap by formalizing a Peircean computational semiotic theory that models Human- GenArt Interaction (HGI) as cascaded semiosis. This framework reveals that artistic meaning is conveyed through three modes — iconic, symbolic, and indexical — yet existing evaluators operate heavily within the iconic mode, remaining structurally blind to the latter two. To overcome this structural blindness, we propose Sem- Judge. This evaluator explicitly assesses symbolic and indexical meaning in HGI via a Hierarchical Semiosis Graph (HSG) that re- constructs the meaning-making process from prompt to generated artifact. Extensive quantitative experiments show that SemJudge aligns more closely with human judgments than prior evaluators on an interpretation-intensive fine-art benchmark. User studies further demonstrate that SemJudge produces deeper, more insight- ful artistic interpretations, thereby paving the way for GenArt to move beyond the generation of “pretty” images toward a medium capable of expressing complex human experience. Project page: https://github.com/songrise/SemJudge CCS Concepts •Appliedcomputing→Finearts; •Human-centeredcomput- ing→HCI theory, concepts and models; •Computing methodolo- gies→Philosophical/theoretical foundations of artificial in- telligence;Natural language generation. Keywords Art interpretation, Human-centered generative art evaluation, Com- putational semiotics, Computational aesthetics ACM Reference Format: Ruixiang JiangandChang Wen Chen. 2026. On Semiotic-Grounded Inter- pretive Evaluation of Generative Art. InArxiv 2026.ACM, New York, NY, USA, 23pages.https://doi.org/10.1145/n.n Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full cita- tion on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy other- wise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Arxiv 2026, © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-x-x-x-x/Y/M https://doi.org/10.1145/n.n Creator’s intentionInterpreted meaning “In the spirit of Guernica” User promptViewer's impressionGenerated artifact Semiosis Gap Semiosis 1 Semiosis 2 Figure 1: HGI as cascaded semiosis. We model HGI as a chain of meaning-making steps: a creator encodes an intention into a prompt, which the model interprets to generate an artifact. A viewer then interprets this artifact to reconstruct the meaning, which may differ from the original intention. 1 Introduction “To see something as art requires something the eye can- not descry.” — Arthur C. Danto,“The Artworld”[17] Art is, at its core, an act of meaning-making [29,46]. What dis- tinguishes a painting from a photograph of the same scene is not fidelity to appearance but the deliberate encoding of the artist’s in- tent through metaphor, symbolism, abstraction, and convention [ 17, 18,29]. For this reason, interpretation is central to aesthetic engage- ment. Yet existing Generative Art (GenArt) evaluation remains heav- ily fixated on what “the eye can descry” — measuring realism [32, 78], prompt-image alignment [42,67], or generic visual appeal [39, 71], while leaving the deeper artistic meaning largely untouched. Unsurprisingly, these evaluators are often misaligned with human judgments from trained viewers [11,24,31,35,39,70,73]. We iden- tify two root causes of this mismatch: Gap 1: Artistic meaning is not reducible to surface appear- ance.Instead, it is often encoded through non-literal strategies such as juxtaposition, abstraction, and metaphorical cues [ 17,29] that diverge deliberately from surface appearance. Taking Picasso’s Guernicaas an example: its impact comes less from resembling a photorealistic scene of war and more from how tonal harshness, fragmentation, and distorted figures convey moral outrage and an anti-war stance [ 13]. However, appearance-centric evaluators risk conflating artistic meaning with surface quality, rewarding visual fidelity or aesthetic allure as proxies for artistic quality. Gap 2: Artistic intent is not reducible to literal prompt wording.Just as artistic meaning is often conveyed indirectly in arXiv:2604.08641v1 [cs.CV] 9 Apr 2026 Arxiv 2026, ,Ruixiang JiangandChang Wen Chen artworks, theintentexpressed when instructing an art generator is often indirect in language: prompts function less as fully speci- fied descriptions and more as artistic directions about vibe, theme, or motif [ 12,74]. For example, prompts such as “in the spirit of Guernica” do not provide a recipe for visual layout, but indirectly specify a target effect that must be interpreted. Consequently, a strong GenArt system should be capable of interpreting these indi- rect prompts and painting artistically (e.g., through exaggeration or abstraction). Most existing evaluators bypass this critical inter- pretation step by directly scoring the text-image alignment, which oversimplifies the interpretive human judgment process. We contend that what existing GenArt evaluators miss is not just stronger visual perception, but theinterpretive processit- self. Particularly, once meaning is conveyed through metaphor, symbolism, or convention rather than literal resemblance, evalua- tion can no longer rely on appearance alone [18,29,46]. We there- fore draw onsemiotics, a long-standing framework in art the- ory [2,16,47,56] and human-computer interaction (HCI) [19] for studying how meaning is communicated from observable forms. In this paper, semiotics provides a principled way to model Human- GenArt Interaction (HGI): how a creator’s intention is encoded in prompts and expressed in generated artifacts. It also lets us iden- tify theiconicity biasof conventional metrics where they tend to misalign with human judgment on symbolic artworks. Building on this theory, we proposeSemJudge, an interpretation- centric HGI evaluator that reconstructs how meaning is carried from prompt to generated artifact, rather than merely scoring surface- level alignment or visual appeal. To achieve this, we introduceHi- erarchicalSemiosisGraphs(HSGs), which represent the prompt- to-image process as a set of linked meaning units. This representa- tion allows SemJudge to reconstruct meaning conveyance in HGI, thereby extending evaluation to both resemblance- and interpretation- based criteria. Experiments on our proposedSemiosisArtdataset show that SemJudge aligns more closely with human judgments and yields more informative, auditable interpretations of artistic meaning. We summarize our contributions as follows: (1)Semiotic framework for HGI.We formalize HGI as cas- caded semiosis and derive why appearance-centric metrics can fail when meaning is conveyed indirectly. (2)Method.We introduceSemJudge, an interpretation-centric HGI evaluator built onHierarchicalSemiosisGraphs(HSGs), a structured representation that links interpretive claims to prompt spans and image regions. (3)Empirical validation and analysis.We show that Sem- Judge aligns more closely with human judgments and yields more informative, edifying interpretations than strong base- lines. 2 Related Work 2.1 GenArt Evaluation Early evaluation metrics largely emphasized realism, which is mea- sured as the distance (divergence) between generated and the real image distributions. Metrics such as Inception Score [ 69], FID [32] and ArtFID [78] fall under this category. With the rise of text- conditional generation, GenArt started to focus on text-image align- ment [ 67]. Subsequently, this alignment scoring was enhanced through human preference tuning, where models such as PickScore [39] and HPS [80] are introduced. These preference models encode generic visual appeal but remain black boxes that yield only a global score. Recently, Question Generation and Answering (QG/A) models have emerged, enabling more interpretable and structured evaluation [14, 33,42]. While existing GenArt evaluations often succeed in mea- suring appearance-level realism and visual attractiveness, the deep artistic meaning embedded in the artworks remains largely un- touched [ 35,38]. This paper adopts a semiotics-theoretical lens that 1: explains the failure modes of appearance-driven metrics, and 2: envisions the design of a meaning-driven evaluation framework that actively interprets. 2.2 GenArt Interpretation and Theory of Art Art interpretation has a long-standing foundation in art theory, particularly in Panofsky’s iconological framework [61], which de- codes the symbolic meaning from visual features. Existing com- putational methods mainly approach this via retrieval [6,25,77] or tuning on curated datasets [ 5,9,34]. While effective for inter- preting historical paintings, we argue that they are insufficient for GenArt evaluation on two grounds. First, they evaluate on canon- ical artworks already saturated in pretraining corpora, so appar- ent interpretive competence may reflect memorization rather than genuine understanding [ 38,68,84]. Second and more importantly, these methods are artifact-centric and hence less human-centred: they interpret the visual work alone, but do not model how mean- ing is conditioned by prompt intent and realized through the HGI process. Recently, ArtCoT [38] demonstrated the effectiveness of zero-shot MLLM for aesthetic judgment, though it explicitly treats symbolic art interpretation as hallucination to be suppressed. In the context of interpreting GenArt, this work pioneers the use of semiotics to interpretively decipher meaning-making in the entire human-GenArt co-creation process. 2.3 Computational Semiotics Computational semiotics studies how semiotic concepts can be for- malized and computed for meaning-driven intelligence systems. Work in this area has long informed HCI, where semiotic mod- els help explain how users interpret interfaces and how meaning is negotiated in interaction [19,20,57]. More recently, semiotic perspectives have been used to analyze the behavior and limita- tions of contemporary AI systems, revealing that current models lack genuine semiotic grounding: they manipulate surface-level patterns and often fail to account for sign relations or meaning- production [43,57,65,72]. This is precisely the theoretical vac- uum that the prior two sections exposed. In this paper, we bring Peircean semiotic theory to GenArt and propose an interpretive evaluator that focuses on the deep meaning encoded in the prompt and the generated artifact. 3 Human-GenArt Interaction as Semiosis This section builds a theoretical foundation on the proposed semi- otic theory for HGI. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , A Cubist-style painting rendered in muted earth tones, depicting a scene with fragmented geometric planes. A sense of spiritual solemnity, structural complexity, and a modern, abstract reinterpretation of a classical religious event. The Annunciation (The biblical event of Gabriel announcing the birth of Jesus to Mary). A fragmented, standing figure on the left featuring wing-like structures and [...] Archangel Gabriel (The Divine Messenger). The active force of divine communication; celestial presence. A white, faceted bird shape in the upper center, surrounded by radiating[...] The Holy Spirit and Divine Light. Ilumination, purity, and the spiritual bridge connecting the messenger and the recipient. The overall aesthetic technique characterized by sharp angles, intersecting planes, and [...] Picasso’s Analytical Cubism Style. Intellectual engagement, deconstruction of reality, timelessness through abstraction. Sub-Semiosis AAngel Root Semiosis Sub-Semiosis B DoveSub-Semiosis D Art Style) The initiator of the narrative of Annuciation. Connects two figures, serving as the spiritual bridge and [...] Defines the stylistic modality of the image, transforming a classical religious [...] [25,12,458,868] [409,29,698,353] Non-localizable Sign Object Interpretant Relation to Root Artifact Sign Generated Image) Grounds: Indexical (brushwork and fragmentation point to the artist's technique). Grounds: Symbolic (Dove as Holy Spirit, rays as divine light). Grounds: Iconic (wings, humanoid form) and Symbolic (angelas messenger). Grounds: Iconic (resembles figures and doves) and Symbolic (religious iconography of the Annunciation). Figure 2: HSG of generated artifact. We show the image with bounding boxes (top-left), its global semiosis (top-right), and sub- semioses (bottom), constructed by an MLLM in zero-shot. The HSG provides a structured interpretation of the Annunciation motif in the abstract painting. Best viewed in color. 3.1 Formulating Peircean Triadic Semiosis SemiosisanditsComponents.Peircean semiotics treats meaning- making (semiosis) as a triadic relation among asign푠 ∈풮, anobject 표 ∈풪, and aninterpretant푖 ∈ℐ[63]. We denote this basic unit as anatomic semiosis 휉 ∶= (표, 푠, 푖) ∈풪×풮×ℐ.(1) Here, the sign is the perceptible form being interpreted (e.g., a prompt or an image), the interpretant is the meaning constructed by an interpreter, and the object is the underlying referent or in- tended content (e.g., motif). Interpreted-as Relationship푠 → 푖.In Peircean semiotics, in- terpretation depends on the interpreter. To make this explicit, we model an interpreter휂 ∈ℋ, whereℋdenotes the space of possi- ble interpreters (e.g., humans or computational models), and write the interpreted-as relation as푖 = 휂(푠). Grounds and the Types of Signs.A sign stands for an ob- ject through its grounds푔 ⊂ Γ, whereΓis the universe of possi- ble grounds. In Peircean semiotics, these are commonly discussed asiconic(based on resemblance),symbolic(based on convention), andindexical(based on contextual or causal connection) [ 64]. Im- portantly, these categories are not crisp or mutually exclusive: a single sign may involve all three to different degrees. This is espe- cially common in art, where meaning is often conveyed through a mixture of resemblance, allegory, convention, and contextual ref- erence [ 23,29]. As a result, purely resemblance-centric evaluation is unreliable for GenArt, since it captures only one of several pos- sible grounds through which meaning may be conveyed. Stands-for Relationship푠 → 표. We now further distinguish between thedynamic object(표), the external intent or reality driv- ing the sign (e.g., the creator’s latent goal); and theimmediate ob- ject(̂표), the object as specifically represented within the sign [63]. We treat the semiotic “ground” as a computational evidence layer 푔 = 퐸(푠). The stands-for relationship is defined as a mapping 휎(푔; 휂) → ̂표, where the interpretation of grounds into the immedi- ate object is conditioned on the specific interpreter휂. Cascaded Semiosis.Meaning-making in HGI is iterative: an in- terpretant produced at one stage can be reified as the next sign and interpreted again. We call this processcascaded semiosis. Formally, an푁-round cascade is 풞 (푁) ∶= [ (휉 (1) , 휂 (1) ) → (휉 (2) , 휂 (2) ) → ⋯ → (휉 (푁) , 휂 (푁) ) ] , s.t.푠 (푛+1) = 휌 (푛) (푖 (푛) ) ∀ 푛 ∈ 1, ... , 푁 − 1, (2) where휉 (푛) = (표 (푛) , 푠 (푛) , 푖 (푛) )denotes the푛-th atomic semiosis,휂 (푛) the interpreter, and휌 (푛) the reification process (e.g., image gener- ation) from interpretation to the subsequent sign. 3.2 Human-GenArt Interaction as Semiosis Semiotics has long provided art theory with a principled vocab- ulary for meaning-conveyance [ 2,56,63]. We contend that every human-GenArt interaction naturally instantiates this same process as a cascaded semiosis. To understand this, consider a typical work- flow of single-round generation. A human user comes to the gener- ator with an intended meaning or goal표 (1) , which is not directly ob- servable to the GenArt system. To act on this intent, the user writes Arxiv 2026, ,Ruixiang JiangandChang Wen Chen a prompt푠 (1) , which can be textual or multi-modal. The generator then first functions as an interpreter휂 1 , which maps푠 (1) into its internal representation푖 (1) (e.g, text encoding). Based on the inter- pretation, the model then synthesizes an artifact푠 (2) = 휌 (1) (푖 (1) ), which is the generated art. This artifact sign is to be interpreted (i.e., evaluated) by another interpreter휂 (2) , who is usually a hu- man user. Thus, even in the simplest setting, HGI produces at least a two-round cascaded semiosis풞 (2) , as compactly visualized in Fig- ure1. Iterative generation may follow this notation to produce long cascades. 4 Semiotics-Grounded GenArt Evaluation This section first formalizes the theoretical bottleneck of the ex- isting GenArt evaluation system. We then introduce SemJudge, a semiotic-grounded and interpretive evaluator. 4.1 Semiosis Quality Measure Under a semiotic view, evaluating the quality of HGI amounts to assessing the quality of the semiosis induced by human–GenArt interaction. Accordingly, we define the theoretical quality of an푁- round semiosis풞 (푁) as the distance between its initial and final dynamic objects: 푄 풞 (푁) ∶= −Δ 표 (표 (1) , 표 (푁) ),(3) whereΔ 표 is a distance metric in풪, and smaller distance indicates higher quality. Because dynamic objects are latent, we approxi- mate quality through interpreter-reconstructed immediate objects. The resultingempirical quality measureof an푁-round semiosis is defined as: ̂표 (푛) = 휎(퐸(푠 (푛) ); 휂) ̂ 푄 휂 풞 (푁) ∶= −Δ 표 ( ̂표 (1) , ̂표 (푁) ) , (4) which corresponds to the interpreter-mediated, observable (and hence computable) HGI quality in풞 (푁) . 4.2 Demystifying Conventional GenArt Metrics Semiotic principle.Our framework explains why GenArt eval- uators can diverge systematically from human judgment [ 24,39, 73] even when literal prompt matching and visual attractiveness appear strong. We summarize this failure mode in the following proposition: PRoposition 1 (InteRpRetive PRinciple: iconicity mismatch degRades semiosis ality.).Let훼(푠 (푛) , 휂 (푛−1) ) ∈ [0, 1]denote theintended iconicityof sign푠 (푛) as encoded by its creator휂 (푛−1) , and let훼(푠 (푛) , 휂 (푛) ) ∈ [0, 1]denote theinterpreted iconicityas in- ferred by the subsequent interpreter휂 (푛) . We formalize the following semiotic principle: | 훼(푠 (푛) , 휂 (푛−1) ) − 훼(푠 (푛) , 휂 (푛) ) | ↑ ⟹ 푄 풞 (푁) ↓ .(5) That is, as the mismatch between intended and interpreted iconic- ity increases, semiosis quality should decrease. This principle is com- mon in art history [ 28]. As an illustrative example, consider abstract art like Picasso’sGuernicaagain, which intentionally use symbolic representation for conveying meaning [13]. An evaluator biased to- ward iconicity, such as a general audience expecting figurative re- semblance, may misread the work as a poor depiction rather than a symbolic one. This causes the interpreted object to diverge from the intended object, thereby reducing semiosis quality. Framing Existing Evaluators:Existing GenArt metrics typ- ically operate in canonical ground space without interpretation. While this can be a reasonable proxy for quality when a sign is pre- dominantly iconic, it becomes semiotically unreliable when mean- ing depends on symbolic or indexical interpretation. Depending on whether the evaluator is aware of the user input prompt, most metrics fall into two families: (1)Context-conditioned metrics (prompt-aware).These metrics assess how well the extracted grounds of a gener- ated artifact match the input prompt (e.g., CLIP, PickScore, MLLM-based scoring) by computing a distance between prompt and image ground representations: 푄(푠 (푛) ; 푠 (1) ) = −Δ 푔 ( 퐸 푖 (푠 (1) ), 퐸 표 (푠 (푛) ) ) ,(6) where퐸 푖 , 퐸 표 are ground extractors (potentially different en- coders for text and image) andΔ 푔 is a generic distance in the induced ground space. (2)Context-free metrics (prompt-agnostic).These metrics evaluate global realism, quality, or aesthetics without refer- ence to the prompt (e.g., FID, aesthetic predictors) by com- paring the artifact to anidealized ground prior푔 ⋆ precom- puted or learned from data: 푄(푠 (푛) ; ∅) = −Δ 푔 ( 푔 ⋆ , 퐸(푠 (푛) ) ) .(7) Despite their differences, both families optimize agreement in ground space rather than recovery in object space. The shared limitation of these metrics is therefore not ground- space comparison itself, but treating it as a universal proxy for semiosis quality. As Proposition 1shows, this proxy can fail when intended iconicity diverges from interpreted iconicity, which is common in art. This can happen both in generation (e.g., the user expresses a prompt symbolically but the model interprets it iconi- cally) and in evaluation (e.g., the evaluator has an iconicity bias and fails to recognize symbolic meaning). Both cases lead to low human satisfaction despite high ground-space scores, which explains the observed divergence between GenArt and real art evaluations — a prediction we empirically confirm in Section6. 4.3 The SemJudge Hierarchical Semiosis Graph.Our original formulation in Equa- tion2view the prompt푠 (1) and generated artifact푠 (푁) holistically and as the atomic unit in semiosis. ForpracticalHGI evaluation, however, it is often useful to make the internal structure of signs explicit, since both prompts and images exhibit rich compositional organization (e.g., sentence structure, entities/attributes, spatial re- lations, and global style) [ 4,62]. To capture this composition structure, we introduce theHier- archical Semiosis Graph(HSG), a scene-graph-inspired representa- tion whose nodes encode atomic semioses rather than only enti- ties and their relations. Specifically, an HSG is a directed graph HSG(푠) = (풱,ℰ)where each node푣 ∈풱is an atomic semiosis휉. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , The root-semiosis( ̂표, 푠, 푖)provides global level analysis, and is con- nected with interpretablesub-semioses, which analyze the meaning of sub-signs in푠. Edges푒 ∈ℰbetween global and sub-semioses en- code their relations (e.g., supports/elaborates, contrasts), thereby making explicit both (a)whatmeanings are present locally and (b) howthey interact to form global intent. Following semiotic the- ory [21,64], we represent all components of HSG in natural lan- guage. This representation also supports both human understand- ing and the downstream MLLM-based judgment and interpreta- tion task. We further distinguishnon-localizablesub-semioses (e.g., over- all style, genre, non-figurative representations) fromlocalizable sub-semioses (e.g., figures, objects) [26,41]. In implementation, localizable sub-semioses are grounded to explicit evidence: text spans in푠 (1) and bounding boxes in푠 (푁) , enabling interpretable, auditable, and fine-grained analysis. Figure2presents an example of HSG for a generated artifact. Operationalizing Object-Space Semiosis Quality.Different from canonical ground-space metrics, SemJudge explicitly recon- structs the 2-round cascaded semiosis induced by HGI. Specifically, we represent a prompt-artifact interaction as the 2-stage chain: 풞 (2) ≈ [ HSG(푠 (1) ) →HSG(푠 (2) ) ] ,(8) where푠 (1) is the prompt and푠 (2) is the generated artifact. SemJudge assesses relative semiosis quality under a 2AFC pro- tocol. Given two artifacts푠 (2) 푎 and푠 (2) 푏 generated from the same prompt푠 (1) , SemJudge outputs two reconstructed semioses, node- level evidence groundingℒ, and a binary judgment̂푦 ∈ 푎, 푏,: SemJudge(푠 (1) , 푠 (2) 푎 , 푠 (2) 푏 ) → ( 풞 (2) 푎 ,풞 (2) 푏 ,ℒ, ̂푦 ) .(9) Let ̃ 풱∶=풱(풞 (2) 푎 ) ⊎풱(풞 (2) 푏 )be the disjoint union of HSG nodes from both semioses. Evidence grounding is a collection of node- cited natural-language rationales: ℒ∶= (푣, ℓ 푣 ) ∣ 푣 ∈ ̃ 풱,(10) whereℓ 푣 is an interpretable explanation with semiosis푣cited. 5 The SemiosisArt Challenge.Existing GenArt benchmarks (e.g., AGIQA-3k [50], GenAI-Bench [49]) and art-historical interpretation datasets (e.g., SemArt [25], VQArt-Bench [1]) are ill-equipped to evaluate artis- tic meaning conveyance in GenArt. First, the majority of GenArt benchmark sets emphasizeiconicgeneration tasks. This bias to- ward iconic prompts and appearance-level quality makes them poorly aligned with our goal of measuring meaning-level quality insymbolicandindexicalart forms. Art-historical datasets, on the other hand, consist of canonical artworks that are already widely covered in MLLM pretraining corpora, making strong performance difficult to disentangle from memorization rather than genuine in- terpretation. We therefore collect a new dataset for benchmarking HGI semiosis quality, with a focus on non-iconic generation tasks. Dataset design.Annotation subjectivity is the main challenge in constructing a dataset for interpretive semiosis quality assess- ment. We address this bymotif-groundingandquality control. First, during construction, we collaborate with푚 1 = 12experts to Motif Generation Variants 2AFC Judgment Interpretation VQA AB A:B : C : D: Prompt: Figure 3: SemiosisArt Construction. Top: we construct a prompt from canonical motifs, and generate images from various models. Bottom: We use two task formats: 2AFC for relative judgment and VQA for fine-grained interpretation. reduce interpretive arbitrariness by using canonically grounded in- terpretations. This is achieved by anchoring HGI tasks to canonical motifs with established roots in tradition (e.g., iconology, culture, theology, literature). Such traditions carry a degree of shared inter- pretive consensus, as motifs with established iconographic roots are grounded in cultural convention rather than individual pref- erence [27,61]. Secondly, we use a strict quality control process. For each 2AFC task, the majority judgment of the expert panel is taken as the reference answer. We additionally crowd-sourced 38, 155non-expert judgments to filter out highly subjective and unreliable tasks, achieving an inter-annotator agreement of 0.58 (Cohen’s휅). The final dataset contains187HSG initiatives, with 935images generated from16generative models. Further details on the dataset construction and quality control are provided in Ap- pendix A. Tasks.SemiosisArt provides two task formats: judgment tasks and QA tasks, as illustrated in Figure3. The judgment task is the main format. Specifically, we use2-Alternative Forced Choice (2AFC), which is considered as more reliable than av- eraged Likert scales (i.e., Mean Opinion Score) for subjective judg- ments [54,55]. We additionally consider theVisual Question Answering(VQA) format. The two formats are complementary: 2AFC captures relative quality judgments at the instance level, while VQA probes whether models can perform semiotic interpre- tation in fine-grained ways. The 2AFC task contains 1870 compar- ative judgment tasks, while the VQA task contains 600 questions. 6 Experiment and Analysis 6.1 Experiment Settings Implementation.We utilize Qwen-3.5-9B as the backbone MLLM unless otherwise specified. This includes all zero-shot MLLM- based baselines for fairness. The MLLM predicts both the HSG schema and the bounding box coordinates in zero-shot (i.e., no finetuning). All MLLM-based methods are repeated three times. Models can see both the user prompt and artifact for the judgment Arxiv 2026, ,Ruixiang JiangandChang Wen Chen Table 1: Correlation analysis of all compared evaluation metrics. KRCC (Kendall’s휏), SRCC (Spearman’s휌), C (Lin’s휌 푐 ), and VQA accuracy measure alignment with human judgment on semiosis quality of HGI. Human (Non-expert) denotes crowd- sourced majority-votejudgment. Gemini-Flash stands for Gemini-3.1-Flash-Lite.(†): Re-implemented with the same Qwen-9B backbone as SemJudge for fairness. GroupMethodBackbone CorrelationInterpretation KRCC↑SRCC↑C↑Acc (%)↑ Conventional Scorers –Random Guess–-0.023-0.0770.00725.2 Alignment ScoringCLIPScore [67]CLIP0.0410.0590.106– Quality ScoringCLIP-IQA [75]CLIP0.0800.2500.088– Quality ScoringDeQA-Score [83] Tuned MLLM 0.023-0.112-0.118 – Preference ScoringAesthetic Predictor [71]CLIP0.0300.1170.084– Preference ScoringPickScore [39]CLIP0.2020.6050.310– Preference ScoringHPSv2 [80]CLIP0.030-0.017-0.081– Preference ScoringImageReward [82]BLIP-0.0020.2320.128– Structured-Rationale Evaluators Rationale&ScoringLMM4LMM [76]Tuned MLLM0.2740.6510.54744.0 Structured Rationale VIEScore [42] (†)Zero-shot MLLM0.2410.6410.321– Structured Rationale DSG [14] (†)Zero-shot MLLM0.1530.6780.230– Art Interpretation / Aesthetic Models Formal AnalysisArtCoT [38] (†)Zero-shot MLLM0.2940.6040.60980.4 Art InterpretationArtiMuse [9]Tuned MLLM0.0750.0880.15667.7 Art InterpretationGalleryGPT [5]Tuned MLLM-0.034-0.201-0.08726.1 Human and SemJudge Human ReferenceHuman (Non-expert)–0.7900.9240.94681.5 Human ReferenceHuman (Expert)–93.2 Semiosis QualitySemJudge (Qwen-9B)Zero-shot MLLM0.5330.8560.80886.1 Semiosis QualitySemJudge (Qwen-35B-A3B) Zero-shot MLLM0.6740.8800.87891.0 Semiosis QualitySemJudge (Gemini-Flash)Zero-shot MLLM0.7460.9640.96892.4 task, but for the interpretation task, they only see the artifact. Ad- ditional implementation details, including prompt templates, can be found in AppendixC. Compared Methods:To the best of our knowledge, Sem- Judge is the first interpretation-centric evaluator for meaning con- veyance in HGI, and there are no directly comparable baselines. We therefore compare against three groups of related methods: (1)Scoring-models, including CLIP-IQA [75], DeQA-Score [83], CLIPScore [67], PickScore [39], HPSv2 [80], ImageReward [82], and LAION Aesthetic Predictor [71]. (2)Evaluators with structured rationales: VIEScore [42], Davidsonian Scene Graph (DSG) [ 14], ArtCoT [38], and LMM4LMM [76]. (3)Art interpretation / aesthetic models: GalleryGPT [5] and ArtiMuse [9]. 6.2 Quantitative Correlation Experiment Correlation Metrics.For quantitative alignment analysis, we adopt three complementary metrics capturing different levels of alignment.(1) Instance Concordance:We compute Kendall’s Tau-b (KRCC)휏on pairwise 2AFC judgments within each prompt and average over all prompts, measuring concordance with hu- man pairwise preferences at the instance level.(2) Discrete Rank Correlation:Following [38], we derive Elo scores for the 16 GenArt models from all valid pairwise comparisons and compute Spearman’s Rank Correlation Coefficient (SRCC)휌between the human-derived and metric-derived model rankings.(3) Continu- ous Elo Correlation:Since SRCC on a small set of 16 models is unstable due to minor rank perturbation and insensitivity to score magnitude [ 15], we additionally compute Lin’s Concordance Cor- relation Coefficient (C)휌 푐 to more robustly capture agreement between Elo scores [48]. All metrics lie in[−1, 1], where higher values indicate stronger positive alignment with human judgment. VQA Metric.For the VQA task, we report multiple-choice ques- tion answering accuracy (Acc), computed as the proportion of cor- rectly answered questions among all questions in the benchmark. Correlation Results.We report model alignment with expert judgments in Table 1, evaluated from three complementary per- spectives: instance concordance, discrete rank correlation, and con- tinuous Elo correlation. Three observations emerge.(1) Conven- tional low-level scorers perform poorly.The image-quality, prompt-alignment, and preference-based scoring methods exhibit near-zero or weak correlation with expert judgments, suggest- ing that appearance-level quality and generic preference signals On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , SemJudge SemJudge (w/o HSG) BaseMLLM DSG GalleryGPT ArtiMuse ArtCoT Causal FactorsDepthEdificationEvidence Grounding SpuriousCausalLiteralIn-depthUnhelpfulEdifyingUngroundedGrounded Strongly DisagreeDisagreeNeutral Agree Strongly Agree 3.29 3.15 3.12 2.92 2.79 2.35 2.50 3.74 3.26 2.92 2.67 3.09 2.52 3.14 3.61 3.11 3.08 2.42 3.13 2.46 2.76 3.53 2.88 2.90 3.15 3.07 2.44 2.86 Figure 4: Subjective Interpretation Quality Experiment on Four Dimensions. We show the user (푚 = 70) feedback distribution on a 5-point Likert rating, with the mean score for each bar. SemJudge (w/o HSG): SemJudge with only root artifact semiosis. Base MLLM: Prompt MLLM to generate art interpretation. are fundamentally insufficient for evaluating symbolic and index- ical meaning conveyance.(2) Canonical MLLM-based evalu- ators remain limited.Existing structured evaluators, though with structured rationales, show a weak correlation for evaluating semiosis quality. Notably, even with the same backbone MLLM, these methods still lag far behind SemJudge, showing that the gap lies not in model capacity, but in the framing of HGI evaluation as iconicity regression rather than semiosis modeling.(3) SemJudge achieves the strongest overall human alignment.With ex- plicit modeling of HGI semiosis through HSGs, SemJudge attains the best performance across all three correlation metrics. This ad- vantage is consistent across backbones, with both Qwen-9B and Gemini-Flash showing exceptionally strong alignment with expert judgment. These results demonstrate the advantage of semiotics- grounded structured interpretation for human-aligned GenArt evaluation. Quantitative Art Interpretation Results.Table 1also re- ports the VQA accuracy of compared methods on fine-grained art interpretation. SemJudge achieves the best overall perfor- mance, showing that its semiotics-grounded structure improves not only pairwise judgment alignment but also explicit inter- pretive understanding. Notably, SemJudge with a lightweight Gemini-3.1-Flash-liteachieves a promising 92.4% accuracy, approaching the expert human performance of 93.2%. 6.3 Human Evaluation of Interpretation Quality Evaluation Dimensions.We task both expert and non-expert users to evaluate the interpretations generated by different models along four dimensions that reflect human-centered, meaning-level assessment in HGI. Each dimension is rated on a 5-point Likert scale (1 = strongly disagree, 5 = strongly agree). A total of4, 943 responses were collected. •Causal Agreement (Expert only).Do the factors in the generated interpretation identified asdecisivefor 2-AFC judgment align with what you consider the primary rea- sons for that judgment, avoiding spurious, hallucinated, or irrelevant cues? •Depth.Does the interpretation transcend literal description (e.g., object/attribute presence or style adherence) to pro- vide an in-depth, meaning-level analysis (e.g., symbolism, metaphors, theological tradition)? •Edification for Artwork Comprehension.Does this in- terpretation aid you in comprehending what the artwork may attempt to express (i.e., the creator’s intent), compared with seeing the image & prompt alone? •Evidence Grounding. Are the key claims in the interpre- tation well-supported by citing specific image regions or global features, and/or by explicit content in the prompt? Discussion.Figure4indicates thatSemJudgeis significantly (푝 < 0.05) preferred across all dimensions of subjective interpreta- tion quality. Expert users assign SemJudge the strongestCausal Agreement, suggesting that its decisive factors are better aligned with human reasoning about semiosis quality. SemJudge also receives the highestDepth, consistent with our design goal of interpreting the deep symbolic meaning in HGI. By contrast, the compared methods primarily focus on appearance of artifact only, such as object presence (DSG) or art style (ArtCoT, ArtiMuse, GalleryGPT), which is less aligned with object-space meaning- conveyance. Users also rate SemJudge highest onEdification for Artwork Comprehension, supporting our motivation that structured semiotic rationales can serve as an accessible bridge between the creator’s intent and visual realization and inspire deeper engagement. OnEvidence Grounding, both expert and non-expert raters more often judge SemJudge’s claims as sup- ported by the prompt and visible evidence, which we attribute to evidence-linked HSG nodes (text spans and bounding boxes) and schemas (Peircean semiosis triad) that reduce unconstrained interpretation. AppendixBprovides additional visualizations and qualitative comparisons with the other methods. 6.4 Empirical Analysis: Iconicity Bias of Conventional Metrics We test whether conventional GenArt evaluators agree with hu- mans primarily oniconicprompt–artifact relations. For each 2- AFC instance(푠 (1) , 푠 (2) 푎 , 푠 (2) 푏 ), six human experts rate iconicity, Arxiv 2026, ,Ruixiang JiangandChang Wen Chen Table 2: Iconicity-bias hypothesis test across evaluators. We reportΔ, bootstrap 95% confidence intervals, and Cohen’s 푑. Sig. indicates one-sided permutation-test significance: ∗ 푝<0.05, ∗ 푝<0.01. Conventional evaluators are biased to- wards iconic signs, while SemJudge remains robust for sym- bolic / indexical signs. EvaluatorΔ95% CICohen’s푑Sig. ImageReward 0.086[ 0.039, ∞)0.306** PickScore0.126[ 0.047, ∞)0.595** DSG0.087[ 0.006, ∞)0.402* ArtCoT0.182[ 0.090, ∞)0.848** SemJudge-0.010[ −0.157, ∞)-0.047 indexicality, and symbolism on 7-point Likert scales. We combine these into an instance-levelnet iconicity score ̃ 푁퐼 푘 , which is positive when iconic resemblance dominates and negative when symbolic/indexical cues dominate. Test for iconicity bias.For each evaluator and instance푘, letΛ 푘 = 1if the evaluator’s winner matches the human winner (otherwise Λ 푘 =0). We define: Δ = 피[ ̃ 푁퐼 푘 ∣ Λ 푘 = 1] − 피[ ̃ 푁퐼 푘 ∣ Λ 푘 = 0]. A positiveΔindicates the evaluator aligns with humans mainly on more iconic instances — aniconicity bias. We assess퐻 1 ∶ Δ > 0via a one-sided permutation test. Full statistical details are in Appen- dixB. Findings.Table2shows that conventional GenArt evaluators exhibit a consistent iconicity bias (significantlyΔ > 0), which sug- gests they track human preferences better when artifacts visually resemble their referents. In contrast, SemJudge shows no positive or significantΔ, suggesting its agreement with humans is not con- centrated on the highly iconic subset but also generalizes indexical and symbolic artworks. 6.5 Ablation Study across MLLMs To disentangle the contribution of SemJudge from the raw capa- bility of the underlying MLLM, we organize the ablation around three controlled questions. Table3examines:(A)whether adding HSG-based structure improves performance under a fixed judge, (B)whether a high-quality HSG can substantially improve an oth- erwise lightweight judge, and(C)how much additional benefit is obtained by scaling the final judge once a strong HSG is already available. This design lets us test whether SemJudge’s gains arise from structured semiosis reconstruction rather than from back- bone scaling alone. Three findings stand out.First, with the judge fixed, introducing HSG structure improves performance over direct judgment, but weak MLLMs may struggle generating highly complex HSG faith- fully, which does not always yield further gains.Second, strong transferred HSGs substantially elevate weak judges, showing that the main bottleneck often lies in HSG construction rather than in the final judge alone.Third, these gains are especially pronounced Table 3: Controlled ablation of SemJudge. We isolate three questions: (A) whether introducing a standard or more com- plex HSG structure helps under a fixed judge, (B) whether a strong HSG can lift a weak judge, and (C) how much judge scaling still matters once a strong HSG is available. HSG Setting HSG Builder Judge (2AFC) KRCC↑VQA Acc↑ (A) Same judge, vary HSG complexity No HSG–Qwen-9B0.4882.0 Standard HSG Qwen-9BQwen-9B0.55 ĝ 86.1 ġ Complex HSG Qwen-9BQwen-9B0.51 ġ 84.3 ġ (B) Strong HSGs can lift weak judges No HSG–Qwen-2B-0.0424.1 No HSG–Qwen-4B0.2856.8 Complex HSG Gemini-Flash Qwen-2B0.27 ĝ 42.2 ĝ Complex HSG Gemini-Flash Qwen-4B 0.52 ĝ 86.8 ĝ (C) Residual effect of judge scaling with the same strong HSG Complex HSG Gemini-Flash Qwen-9B0.57 ĝ 91.6 ĝ Complex HSG Gemini-Flash Gemini-Flash0.73 ĝ 92.4 ĝ for VQA, where a strong HSG greatly improves explicit art inter- pretation. This finding is consistent with the human-based ablation in Figure4, clearly demonstrating the effectiveness of HSGs for art interpretation. 6.6 Limitations SemiosisArt leverages Christian, East Asian, Hindu, and Islamic traditions and modern artistic motifs as anchors for interpretation. While this set a more culturally grounded and inter-subjective approach for reliable benchmarking, we acknowledge that it may not fully represent the full diversity of artistic expression and hence the interpretation challenges in generative art. Cultural mi- nority and contemporary conceptual art are two major categories that may not be well represented, because they are more diffi- cult to evaluate through stable shared human judgments both in theory [ 22] and in our human evaluations. 7 Conclusion Our study highlights a critical gap in GenArt: the inability of con- ventional metrics to grasp the symbolic and indexical depth of vi- sual art. Just as modern art evolved from perceptual resemblance to conceptual meaning, we believe that for GenArt to truly evolve, it must move past simply generating pretty pictures and start rec- ognizing the deeper ideas and intentions that make human cre- ativity meaningful. By integrating semiotic theory, we shift the evaluative focus from surface-level appearance to the mechanics of meaning-making. Our findings confirm that while existing eval- uators are biased toward iconic resemblance, SemJudge success- fully reconstructs the interpretive process required to “descry” the artistic meaning within creator’s intention and generated artifacts, resulting in a significant improvement in human correlation for judging and interpreting GenArt. We hope this work inspires fu- ture research to further explore the rich interpretive dimensions of GenArt that can capture the full spectrum of artistic meaning. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , References [1]Andrea Alfarano, Lorenzo Venturoli, and Darío Negueruela del Castillo. 2025. VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Her- itage. In2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). IEEE, 406–416. [2]Mieke Bal and Norman Bryson. 1991. Semiotics and art history.The art bulletin 73, 2 (1991), 174–208. [3]Jason Baldridge, Jakob Bauer, Mukul Bhutani, Nicole Brichtova, Andrew Bunner, Lluis Castrejon, Kelvin Chan, Yichang Chen, Sander Dieleman, Yuqing Du, et al. 2024. Imagen 3.arXiv preprint arXiv:2408.07009(2024). [4]Irving Biederman. 1987. Recognition-by-components: a theory of human image understanding.Psychological review94, 2 (1987), 115. [5]Yi Bin, Wenhao Shi, Yujuan Ding, Zhiqiang Hu, Zheng Wang, Yang Yang, See- Kiong Ng, and Heng Tao Shen. 2024. Gallerygpt: Analyzing paintings with large multimodal models. InProceedings of the 32nd ACM International Conference on Multimedia. 7734–7743. [6]Tibor Bleidt, Sedigheh Eslami, and Gerard De Melo. 2024. Artquest: Countering hidden language biases in artvqa. InProceedings of the IEEE/CVF Winter Confer- ence on Applications of Computer Vision. 7326–7335. [7]ByteDance Seed. 2025.Seedream 4.0: New-Generation Image Creation Model. ByteDance.https://seed.bytedance.com/en/seedream4_0 [8]Huanqia Cai, Sihan Cao, Ruoyi Du, Peng Gao, Steven Hoi, Zhaohui Hou, Shijie Huang, Dengyang Jiang, Xin Jin, Liangchen Li, et al. 2025. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer. arXiv preprint arXiv:2511.22699(2025). [9]Shuo Cao, Nan Ma, Jiayang Li, Xiaohui Li, Lihao Shao, Kaiwen Zhu, Yu Zhou, Yuandong Pu, Jiarui Wu, Jiaquan Wang, et al. 2025. Artimuse: Fine-grained image aesthetics assessment with joint scoring and expert-level understanding. arXiv preprint arXiv:2507.14533(2025). [10]CapCut. 2024. Dreamina: All-in-one AI Creative Suite.https://dreamina.capcut. com/Accessed: 2026-01-26. [11]Rebecca Chamberlain, Caitlin Mullin, Bram Scheerlinck, and Johan Wagemans. 2018. Putting the art in artificial: Aesthetic responses to computer-generated art. Psychology of Aesthetics, Creativity, and the Arts12, 2 (2018), 177. [12]Minsuk Chang, Stefania Druga, Alexander J Fiannaca, Pedro Vergani, Chinmay Kulkarni, Carrie J Cai, and Michael Terry. 2023. The prompt artists. InProceed- ings of the 15th Conference on Creativity and Cognition. 75–87. [13]Herschel Browning Chipp and Javier Tusell. 1988. Picasso’s Guernica: history, transformations, meanings.(No Title)(1988). [14]Jaemin Cho, Yushi Hu, Jason M Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang. 2024. Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Gen- eration. InICLR. [15]Christophe Croux and Catherine Dehon. 2010. Influence functions of the Spear- man and Kendall correlation measures.Statistical methods & applications19, 4 (2010), 497–515. [16]Brian Curtin. 2009. Semiotics and visual representation.Semantic Scholar4 (2009). [17]Arthur Danto. 1964. The artworld.The journal of philosophy61, 19 (1964), 571– 584. [18]Arthur C Danto. 1981.The transfiguration of the commonplace: a philosophy of art. Harvard University Press. [19]Clarisse Sieckenius De Souza. 2005.The semiotic engineering of human-computer interaction. MIT press. [20]Clarisse Sickenius de Souza and Carla Faria Leitão. 2009.Semiotic engineering methods for scientific research in HCI. Morgan & Claypool Publishers. [21]Umberto Eco. 1979.A theory of semiotics. Vol. 217. Indiana University Press. [22]Umberto Eco. 1989.The open work. Harvard University Press. [23]James Elkins. 1999.The domain of images. Cornell University Press. [24]Ziv Epstein, Aaron Hertzmann, Investigators of Human Creativity, Memo Akten, Hany Farid, Jessica Fjeld, Morgan R Frank, Matthew Groh, Laura Herman, Neil Leach, et al. 2023. Art and the science of generative AI.Science380, 6650 (2023), 1110–1111. [25]Noa Garcia and George Vogiatzis. 2018. How to read paintings: semantic art understanding with multi-modal retrieval. InProceedings of the European Con- ference on Computer Vision (ECCV) Workshops. 0–0. [26]Leon A Gatys, Alexander S Ecker, and Matthias Bethge. 2016. Image style trans- fer using convolutional neural networks. InProceedings of the IEEE conference on computer vision and pattern recognition. 2414–2423. [27]Eleni Gemtou. 2010. Subjectivity in art history and art criticism.Rupkatha Journal on Interdisciplinary Studies in Humanities2, 1 (2010), 2–13. [28]Ernst Hans Gombrich and EH Gombrich. 1995.The story of art. Vol. 12. Phaidon London. [29]Nelson Goodman. 1976. Languages of art: An approach to a theory of symbols. Indianapolis: Bobbs-Merrill, 2nd ed/Hackett(1976). [30]Google. 2025. Nano Banana Pro - Gemini AI image generator & photo editor. https://gemini.google/overview/image-generation/Accessed: 2026-01-26. [31]Anna Yoo Jeong Ha, Josephine Passananti, Ronik Bhaskar, Shawn Shan, Reid Southen, Haitao Zheng, and Ben Y Zhao. 2024. Organic or diffused: Can we distinguish human art from ai-generated images?. InProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security. 4822–4836. [32]Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems 30 (2017). [33]Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Kr- ishna, and Noah A Smith. 2023. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. InProceedings of the IEEE/CVF International Conference on Computer Vision. 20406–20417. [34]Yipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan, Zhichao Duan, Pengfei Chen, Leida Li, Weisi Lin, and Guangming Shi. 2024. Aesexpert: Towards multi- modality foundation model for image aesthetics perception. InProceedings of the 32nd ACM International Conference on Multimedia. 5911–5920. [35]Jessica Hullman, Ari Holtzman, and Andrew Gelman. 2023. Artificial intelli- gence and aesthetic judgment.arXiv preprint arXiv:2309.12338(2023). [36]Shahana Ibrahim, Panagiotis A Traganitis, Xiao Fu, and Georgios B Giannakis. 2025. Learning from crowdsourced noisy labels: A signal processing perspective. IEEE Signal Processing Magazine42, 3 (2025), 84–106. [37]Ideogram AI. 2024. Ideogram: Help People Become More Creative.https:// ideogram.ai/Accessed: 2026-01-26. [38]Ruixiang Jiang and Chang Wen Chen. 2025. Multimodal llms can reason about aesthetics in zero-shot. InProceedings of the 33rd ACM International Conference on Multimedia. 6634–6643. [39]Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. 2023. Pick-a-pic: An open dataset of user preferences for text-to- image generation.Advances in Neural Information Processing Systems36 (2023), 36652–36663. [40]Kling Team. 2025. Kling-Omni Technical Report. arXiv:2512.16776[cs.CV]https: //arxiv.org/abs/2512.16776 [41]Gunther Kress and Theo Van Leeuwen. 2020.Reading images: The grammar of visual design. Routledge. [42]Max Ku, Dongfu Jiang, Cong Wei, Xiang Yue, and Wenhu Chen. 2024. Viescore: Towards explainable metrics for conditional image synthesis evaluation. InPro- ceedings of the 62nd Annual Meeting of the Association for Computational Linguis- tics (Volume 1: Long Papers). 12268–12290. [43]Jiayi Kuang, Yinghui Li, Chen Wang, Haohao Luo, Ying Shen, and Wenhao Jiang. 2025. Express What You See: Can Multimodal LLMs Decode Visual Ciphers with Intuitive Semiosis Comprehension?. InFindings of the Association for Computa- tional Linguistics: ACL 2025. 12743–12774. [44]Black Forest Labs. 2024. FLUX.https://github.com/black-forest-labs/flux. [45]J Richard Landis and Gary G Koch. 1977. The measurement of observer agree- ment for categorical data.biometrics(1977), 159–174. [46]Susanne K. Langer. 2009.Philosophy in a New Key: A Study in the Symbolism of Reason, Rite, and Art(third edition ed.). Harvard University Press. [47]Susanne K Langer and . Langer. 1953.Feeling and form. Vol. 3. Routledge and Kegan Paul London. [48]I Lawrence and Kuei Lin. 1989. A concordance correlation coefficient to evaluate reproducibility.Biometrics(1989), 255–268. [49]Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li, Yixin Fei, Kewen Wu, Tiffany Ling, Xide Xia, Pengchuan Zhang, Graham Neubig, et al. 2024. Genai-bench: Eval- uating and improving compositional text-to-visual generation.arXiv preprint arXiv:2406.13743(2024). [50]Chunyi Li, Zicheng Zhang, Haoning Wu, Wei Sun, Xiongkuo Min, Xiaohong Liu, Guangtao Zhai, and Weisi Lin. 2023. Agiqa-3k: An open database for ai- generated image quality assessment.IEEE Transactions on Circuits and Systems for Video Technology34, 8 (2023), 6833–6846. [51]Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou, Yinan Zou, Weiyan Zhang, Haiyun Jiang, and Tong Ruan. 2025. Can Multimodal Large Language Models Understand Spatial Relations?. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 620–632. [52]Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. 2024. Grounding dino: Marry- ing dino with grounded pre-training for open-set object detection. InEuropean conference on computer vision. Springer, 38–55. [53]Chuofan Ma, Yi Jiang, Jiannan Wu, Zehuan Yuan, and Xiaojuan Qi. 2024. Groma: Localized visual tokenization for grounding multimodal large language models. InEuropean Conference on Computer Vision. Springer, 417–435. [54]Rafał K Mantiuk, Anna Tomaszewska, and Radosław Mantiuk. 2012. Comparison of four subjective methods for image quality assessment. InComputer graphics forum, Vol. 31. Wiley Online Library, 2478–2491. [55]Alberto Maydeu-Olivares and Anna Brown. 2010. Item response modeling of paired comparison and ranking data.Multivariate Behavioral Research45, 6 (2010), 935–974. [56]Douglas N Morgan. 1955. Icon, index, and symbol in the visual arts.Philosophi- cal Studies: An International Journal for Philosophy in the Analytic Tradition6, 4 (1955), 49–54. Arxiv 2026, ,Ruixiang JiangandChang Wen Chen [57]Lia Morra, Antonio Santangelo, Pietro Basci, Luca Piano, Fabio Garcea, Fabrizio Lamberti, and Massimo Leone. 2024. For a semiotic AI: Bridging computer vision and visual semiotics for computational observation of large scale facial image archives.Computer Vision and Image Understanding249 (2024), 104187. [58]Stefanie Nowak and Stefan Rüger. 2010. How reliable are annotations via crowd- sourcing: a study about inter-annotator agreement for multi-label image anno- tation. InProceedings of the international conference on Multimedia information retrieval. 557–566. [59]OpenAI. 2025.GPT-Image 1 - OpenAI API Documentation.https://platform. openai.com/docs/models/gpt-image-1 [60]OpenAI. 2025. GPT-Image 1.5 - OpenAI API Documentation.https://platform. openai.com/docs/models/gpt-image-1.5Accessed: 2026-01-26. [61]Erwin Panofsky. 1955.Meaning in the Visual Arts: Papers in and on Art History. University of Chicago Press. [62]Barbara Partee et al. 1984. Compositionality.Varieties of formal semantics3 (1984), 281–311. [63]Charles Sanders Peirce. 1991.Peirce on signs: Writings on semiotic. UNC Press Books. [64]Charles Sanders Peirce. 1992.The essential peirce, volume 2: Selected philosophical writings (1893-1913). Vol. 2. Indiana University Press. [65]Davide Picca. 2025. Not Minds, but Signs: Reframing LLMs through Semiotics. arXiv preprint arXiv:2505.17080(2025). [66]Qwen Team. 2025. Qwen Image 2.0.https://qwen.ai/blog?id=qwen-image-2.0 Accessed: 2026-03-27. [67]Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PMLR, 8748–8763. [68]William Rudman, Michal Golovanevsky, Amir Bar, Vedant Palit, Yann LeCun, Carsten Eickhoff, and Ritambhara Singh. 2025. Forgotten polygons: Multimodal large language models are shape-blind. InFindings of the Association for Compu- tational Linguistics: ACL 2025. 11983–11998. [69]Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training gans.Advances in neural information processing systems29 (2016). [70]Andrew Samo and Scott Highhouse. 2023. Artificial intelligence and art: Iden- tifying the aesthetic judgment factors that distinguish human-and machine- generated artwork.Psychology of Aesthetics, Creativity, and the Arts(2023). [71]Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Sys- tems35 (2022), 25278–25294. [72]José L Cendejas Valdez, Heberto Ferreira Medina, Jesús L Soto Sumuano, Gus- tavo A Vanegas Contreras, Miguel A Acuña López, and Gustavo A López Saldaña. 2024. Semiotics and Artificial Intelligence (AI): An Analysis of Symbolic Commu- nication in the Age of Technology. InFuture of Information and Communication Conference. Springer, 481–494. [73]Jules Van Hees, Tijl Grootswagers, Genevieve L Quek, and Manuel Varlet. 2025. Human perception of art in the age of artificial intelligence.Frontiers in psychol- ogy15 (2025), 1497469. [74]Kailas Vodrahalli and James Zou. 2023. Artwhisperer: A dataset for characteriz- ing human-ai interactions in artistic creations.arXiv preprint arXiv:2306.08141 (2023). [75]Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. 2023. Exploring clip for assessing the look and feel of images. InProceedings of the AAAI conference on artificial intelligence, Vol. 37. 2555–2563. [76]Jiarui Wang, Huiyu Duan, Yu Zhao, Juntong Wang, Guangtao Zhai, and Xiongkuo Min. 2025.Lmm4lmm: Benchmarking and evaluating large- multimodal image generation with lmms. InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision. 17312–17323. [77]Shuai Wang, Ivona Najdenkoska, Hongyi Zhu, Stevan Rudinac, Monika Kack- ovic, Nachoem Wijnberg, and Marcel Worring. 2025. ArtRAG: Retrieval- Augmented Generation with Structured Context for Visual Art Understanding. InProceedings of the 33rd ACM International Conference on Multimedia. 6700– 6709. [78]Matthias Wright and Björn Ommer. 2022. Artfid: Quantitative evaluation of neu- ral style transfer. InDAGM German Conference on Pattern Recognition. Springer, 560–576. [79]Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng- ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al. 2025. Qwen-image technical report.arXiv preprint arXiv:2508.02324(2025). [80]Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hongsheng Li. 2023. Hu- man preference score: Better aligning text-to-image models with human prefer- ence. InProceedings of the IEEE/CVF International Conference on Computer Vision. 2096–2105. [81]xAI. 2025. Grok Imagine: AI Image Generation.https://grok.com/imagineAc- cessed: 2026-03-27. [82]Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. 2023. Imagereward: Learning and evaluating human prefer- ences for text-to-image generation.Advances in Neural Information Processing Systems36 (2023), 15903–15935. [83]Zhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue, and Chao Dong. 2025. Teaching large language models to regress accurate image quality scores using score distri- bution. InProceedings of the Computer Vision and Pattern Recognition Conference. 14483–14494. [84]Haojie Zheng, Tianyang Xu, Hanchi Sun, Shu Pu, Ruoxi Chen, and Lichao Sun. 2024. Thinking before looking: Improving multimodal llm reasoning via mitigat- ing visual hallucination.arXiv preprint arXiv:2411.12591(2024). On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , A Details on the SemiosisArt Constructing a meaning-oriented GenArt dataset is critical for semiosis quality evaluation. This section provides details on how SemiosisArt is constructed to focus on meaning and interpretation, instead of appearance, as in existing datasets. Dataset Overview.We construct the dataset iteratively with ex- pert feedback and quality control from crowd sourcing. In brief, expert users are tasked with proposing prompts whose intended meanings rely substantially on symbolic or indexical interpreta- tion, while still remaining sufficiently canonically grounded to sup- port inter-subjective judgment from non-expert users. In practice, this means anchoring prompts to canonical motifs from theologi- cal stories, literature, art-historical painting traditions and cultural contexts, rather than relying on unconstrained free-form interpre- tation. The motifs in SemiosisArt span a broad range of traditions and cultural contexts, including Christian, Islamic, Hindu, and East Asian traditions such as Chinese, Buddhist, and Japanese sources; art-historical forms such as vanitas and triptychs; and modern vi- sual traditions such as infographics, manga, and outsider art. 0% (Low) 25%50%75%100% (High) Density (jittered) Mean = 0.354 Median = 0.319 Std = 0.196 Net Iconicity Distribution (normalized) Mean (0.35) 0%25%50%75%100% Net Iconicity Score (Normalized) Figure5:NetIconicityDistribution(Jittoredandnormalized, with outliers ignored in plotting) of SemiosisArt. Scale.SemiosisArt contains187HSG initiatives and outputs from a pool of16generative models. For each initiative (i.e., in- put prompt), we sample5models to generate5images, yielding 187 × 5 = 935images in total. These5images induce ( 5 2 ) = 10 pairwise comparisons per initiative, so the judgment task contains 187 × ( 5 2 ) = 1,8702AFC comparative judgment instances. In ad- dition, the dataset includes600VQA questions for fine-grained interpretation benchmarking. We construct the benchmark with 푚 1 = 12experts and use38,155non-expert judgments for quality control, retaining tasks with sufficient inter-subjective agreement for evaluation. We chose this scale to balance (i) coverage across diverse motifs and traditions, (i) statistical power for correlation analyses, and (i) the practical cost of repeatedly querying API- based MLLMs across all benchmark instances. In particular, scal- ing this expert-annotated, art-centric dataset is especially costly because it requires sustained interaction with experts from differ- ent cultural backgrounds, each contributing culturally grounded judgments and fine-grained annotations. These annotations go well beyond simple 2AFC labeling as commonly seen in existing datasets focusing on surface-level quality. This is mainly because we incorporate an iterative quality control process (detailed in the later paragraph) to ensure that symbolic, and contextual in- terpretations are accurately captured and can be deduced during interpretation. This annotation burden partly explains the rela- tively modest size of the dataset. Iterative Process.The dataset construction process is iterative, with multiple rounds of expert feedback and crowd-sourced qual- ity control. In each round, experts propose challenging (i.e., low iconicity) prompts, which are then used to generate images from the pool of models. These images are then subjected to crowd- sourced 2AFC judgments, and we retain only those prompts that yield sufficient inter-subjective agreement (e.g., above a certain threshold of agreement or statistical significance). In practice, this filtering is at two levels. We first filter out the 2AFC tasks with lower than 60% effective agreement, which is at instance level. The 2AFC task being filtered out is called unreliable. We then further filter out the tournament associated with the entire initiatives with less than 4 reliable 2AFC tasks, which is to avoid a highly sparse tournament graph (though it is not a sufficient condition). This filtering is similar with FineArtBench [ 38] at a high level but we additionally make this process iterative to further refine the dataset. At the end of each round, the expert rewrites the prompt associated with the unreliable 2AFC tasks, and we re-generated the images to make the task more discriminative. We repeat this process for three rounds. All 2AFCs are at least judged by 13 non-expert users. Inter Annotator Agreement.The inter-annotator agreement (Co- hen’s휅) for non-expert annotators after the iterative process is 0.58. Although this value would often be described as moderate under generic interpretation scales [45], we argue it is in fact stronggiven two task-specific considerations. First, aesthetic and interpretive judgments are known to yield substantially lower inter-annotator agreement than factual or perceptual tasks; prior work on crowdsourced image annotation reports that aesthetic and quality-related concepts specifically exhibit very low non-expert agreement [ 58]. Second, crowdsourced annotations are inherently noisy at the individual annotator level [36], so the post-filtering휅 achieved here reflects meaningful inter-subjective consensus, not merely annotation consistency on easy tasks. VQA Generation Process.We generate VQA questions via a semi- automatic process, where human experts annotate the images, and MLLM generate questions. We sample top images among the 5 gen- erated images for each initiative (according to Elo from human judgments). Then, we ask experts to annotate region-of-interest from the image. This annotation corresponds to the symbolic/in- dexical object in the image, akin to the sub-semiosis in HSGs. After this process, we run MLLMs for question generation. Specifically, we use GPT-5.4 for question proposal across 10 question types, fol- lowed by a quality control process with Gemini-3-pro. The qual- ity control process involves filtering out questions that are either too easy (e.g., can be answered by surface-level cues or with lan- guage alone) or too difficult (e.g., require esoteric knowledge or highly subjective interpretation). We also include a held-out set of Arxiv 2026, ,Ruixiang JiangandChang Wen Chen expert (푚 3 =2) for quality control using the same standard as the au- tomatic filtering. If a question fail to pass the quality check, it will be re-generated. The details of the instructions are summarized in Table 6. So, the final VQA questions are those that pass both au- tomatic filtering and expert review, ensuring that they are appro- priately challenging and relevant for evaluating semiosis quality. Note that different from 2AFC, the expert themselves does not pro- pose the question stem and options. The reference VQA accuracy in Table 1is the bootstrap average of the majority vote of12−2 = 10 experts (with each 4 experts per question to reduce annotator bur- den, with assignment stratified by expertise). Generation models.The benchmark is constructed from a pool of 16contemporary GenArt systems spanning commercial and open- generation/editing models. These include GPT-Image-1.5 [60], GPT-Image-1-Mini [59], Nano-Banana-Pro [30], Nano-Banana [30], Nano-Banana-2 [30], Kling-Image-O1 [40] Grok-Imagine-Image [81], Qwen-Image-2.0 [66] SeedDream-4.0 [7], Dreamina-3.1 [10], Z- image [8], Qwen-Image-20B [79], Qwen-Edit [79], Imagen-3- Fast [3], Ideogram-v2a-Turbo [37], and Flux-2-Dev [44]. Visualizations.In Figure11, we visualize the images with the normalized net iconicity (푁퐼) score on it. Figure5visualizes the distribution of net iconicity scores. High iconicity tasks mostly fo- cus on low-level features, such as tone, or tasks with a reference image for identity preservation. Low iconicity tasks are related to history, convention, story (e.g., theological stories), and causality (e.g, mirror reflection). A random sample of 15 prompts is provided in Table5. B Additional Experiment Results B.1 Bounding Box Grounding Quality This section provides some insights into the bounding box predic- tion. Motivation for reporting satisfaction rate.Standard detection metrics, such as Mean Intersection over Union (mIoU), rely on fixed ontologies and pre-annotated ground truth masks (e.g., COCO or LVIS). However, SemJudge operates in anopen-vocabulary, generative setting. Since the model itself generates the label (the sub-sign description) dynamically based on its interpretation of the artwork, there exists no static ground truth for these generated concepts. On the other hand, the spatial span of art-related sym- bols are often vague and open to interpretation, making it difficult to define a single “correct” bounding box. Consequently, calculat- ing mIoU against a static baseline is mathematically ill-posed in this context. Method design.Rather than expecting an exact box match, we evaluate whether the predicted box isinterpretively usefulfor the semiotic analysis. Concretely, we reporthuman satisfaction rate, which asks whether the predicted box provides valid visual evidence for the semiotic argument constructed by the model. User study.We measure bounding box quality through a user study. Specifically, we present the image together with the bound- ing box visualization and ask annotators whether the visualization is satisfactory or not (binary choice). We collected 450 satisfaction annotations. Among all models, Gemini-3.1-Flash-Lite has the highest sat- isfaction rate (74.7%), which is higher than Qwen-3.5-35B-A3B (56.0%) and Qwen-3.5-9B (57.8%). This is mostly consistent with their performance in correlation and VQA judgment. Limitation and Future Work.MLLMs are known to be limited in directly predicting precise bounding boxes in zero-shot [51,53]. For a stronger localization capability and potentially better satis- faction, a dedicated grounding module would be helpful. Models such as GroundingDINO [52] would be a strong candidate for this purpose, which is also zero-shot. We believe implementing this module would be trivial and incremental compared with our theo- retical framework, so we did not include them as our contribution. B.2 Details of the Iconicity-Bias Analysis We test whether conventional GenArt evaluators agree with hu- mans primarily oniconicprompt–artifact relations. To do so, we first quantify asubjective iconicity scorefor each benchmark instance. Net iconicity score.For each 2AFC instance with(푠 (1) , 푠 (2) 푎 , 푠 (2) 푏 ), we estimate how much the judgment is driven by iconic resem- blance versus indexical or symbolic cues. Because a sign may simul- taneously contain all three components, we define the net iconicity score of a sign as 푁퐼(푠) = 퐼푐푛(푠) − 1 2 ( 퐼푑푥(푠) + 푆푦푚(푠) ) , where퐼푐푛,퐼푑푥, and푆푦푚are 7-point Likert ratings of iconicity, indexicality, and symbolism, respectively, provided by six human experts. We then aggregate sign-level scores into an instance-level score: ̃ 푁퐼 푘 = 푁퐼 ( 푠 (1) 푘 ) + 1 2 ( 푁퐼 ( 푠 (2) 푘,푎 ) + 푁퐼 ( 푠 (2) 푘,푏 ) ) . Positive ̃ 푁퐼 푘 indicates that the instance is dominated by iconic re- semblance, while negative values indicate a stronger role for sym- bolic and indexical cues. Hypothesis test.For each evaluator and instance푘, letΛ 푘 = 1 if the evaluator’s winner matches the human winner, andΛ 푘 = 0 otherwise. We compare the average iconicity of the aligned and misaligned subsets: Δ = 피[ ̃ 푁퐼 푘 ∣ Λ 푘 = 1] − 피[ ̃ 푁퐼 푘 ∣ Λ 푘 = 0]. A positiveΔindicates that the evaluator agrees with humans mainly on more iconic instances, which we interpret as aniconic- ity bias. We test the one-sided hypothesis퐻 1 ∶ Δ > 0using a permutation test, and report a one-sided 95% bootstrap confidence interval together with Cohen’s푑as the effect size. Interpretation.Under this setup, an evaluator with a strong re- semblance bias should align with humans more often on highly iconic cases than on symbolic or indexical cases, producingΔ > 0. By contrast, an evaluator that is robust across different semiotic regimes should not concentrate its agreement on the iconic subset, and therefore should not exhibit a significantly positiveΔ. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , SymbolDescription Peircean Semiotics (Section 3) 휉An atomic semiosis, defined as the tuple(표, 푠, 푖). 푠 ∈풮TheSign(or Representamen); the perceptible form (e.g., prompt, image). 푖 ∈ℐTheInterpretant; the meaning constructed by an interpreter. 표 ∈풪TheDynamic Object; the underlying meaning or intent not directly observable. ̂표 ∈풪TheImmediate Object; the underlying meaning or intent as interpreted from sign. 휂 ∈ℋTheInterpreter; an agent (human or model) mapping signs to interpretants. 푔 ∈ ΓTheGround; the evidence or basis (e.g., visual features) connecting a sign to an object. 퐸(⋅)Ground extractor function, where푔 = 퐸(푠). 휌 푛(⋅) Reification function that maps interpretant to sign in cascaded semiosis. 휎The “stands-for” relationship mapping grounds to objects (Γ →풪). 퐶 (푁) A cascaded semiosis consisting of a chain of푁 atomic semioses. HGI Evaluation & SemJudge (Section 4) 푄 퐶 (푁) Theoretical quality of a semiosis (distance in Ob- ject space). ̂ 푄 휂 Empirical quality measure as judged by inter- preter휂. ̂표The inferred object, reconstructed via the inverse stands-for relationship휎 −1 . Δ 표 , Δ 푔 Distance metrics in the Object space and Ground (feature) space, respectively. 퐻푆퐺(푠)Hierarchical Semiosis Graph; a structured rep- resentation of meaning units. 풱,ℰThe sets of vertices (atomic semioses) and edges (relations) in an HSG. ℒA collection of evidence groundings (rationales) linked to graph nodes. ℓ 푣 Natural language rationale cited by node푣. 푦Binary preference label output by SemJudge (푦 ∈ 푎, 푏). Analysis Metrics (Section 5) 푁퐼(푠)Net Iconicity Score; measures how much a sign relies on resemblance vs. symbolism. 퐼푐푛, 퐼푑푥, 푆푦푚Individual ratings for Iconicity, Indexicality, and Symbolism. Λ 푘 Binary indicator of alignment between an evalu- ator and human judgment for instance푘. Table 4: Glossary of Notations B.3 Additional Visualization of HSGs Figure10provide an visualization of user sign (prompt), Fig- ure7,8,9provide three additional HSG visualization for output signs. C Implementation Details C.1 The SemJudge Algorithm We present an algorithmic formulation of the SemJudge in Algo- rithm C.1. Algorithm C.1.SemJudge: Object-Space Semiosis Quality Mea- sure Require:Prompt푠 (1) , Candidate Artifacts푠 (2) 푎 , 푠 (2) 푏 Require:VLMℳ(Evaluator) Ensure:Reconstructed Semioses풞 (2) 푎 ,풞 (2) 푏 , Evidenceℒ, Judg- ment̂푦 1:Stage 1: Reconstruct Prompt Semiosis (Input Space) 2:Let푝 in be the instruction to analyze HSG(푠 (1) ) 3:푟 1 ←ℳ(푝 in , 푠 (1) ) 4:Extract prompt-level semiotic nodes풱 1 from푟 1 5:Initialize context퐻 ← [(푝 in , 푠 (1) ), 푟 1 ] 6:Stage 2: Reconstruct Artifact Semiosis (Object Space) 7:Let푝 out be the instruction to analyze HSG(푠 (2) 푎 )and HSG(푠 (2) 푏 ) 8:푟 2 ←ℳ(푝 out , 푠 (2) 푎 , 푠 (2) 푏 ∣ 퐻) 9:Extract artifact-level semiotic nodes풱 2푎 ,풱 2푏 from푟 2 10:Formulate Cascaded Chains 11:Construct풞 (2) 푎 ← [풱 1 →풱 2푎 ] 12:Construct풞 (2) 푏 ← [풱 1 →풱 2푏 ] 13:Let ̃ 풱←풱(풞 (2) 푎 ) ⊎풱(풞 (2) 푏 ) 14:Stage 3: Evidence Grounding and Judgment 15:Let푝 judge be the instruction to compare chains and cite evi- dence 16:푟 3 ←ℳ(푝 judge ∣ 퐻, 푟 2 ) 17:Parse judgment̂푦 ∈ 푎, 푏from푟 3 18:Extract rationalesℓ 푣 for nodes푣 ∈ ̃ 풱from푟 3 19:Construct evidence setℒ← (푣, ℓ 푣 ) ∣ 푣 ∈ ̃ 풱 20:return(풞 (2) 푎 ,풞 (2) 푏 ,ℒ, ̂푦) C.2 System Prompt Below we provide the complete system prompt used in our ex- periments, including the HSG construction prompt for the input prompt, image, and 2AFC summarization. The difference between standard HSG construction v.s. complex HSG prompt is that we allow푉 ≤ 3and succinct descriptions for the former, while we allow푉 ≤ 5and more detailed descriptions for the latter. D Appendix: Glossary of Mathematical Notations Table4summarizes the semiotic and computational notations used throughout the paper. Arxiv 2026, ,Ruixiang JiangandChang Wen Chen GalleryGPT This captivating painting presents a harmonious composition centered around a serene lake surrounded by lush greenery. The focal point, a solitary figure standing in a small boat, exudes an air of contemplation and isolation. The artist's use of soft, muted colors creates a tranquil and atmospheric setting, with the pale blues and greens of the water and foliage contrasting with the warm, earthy tones of the boat and figure. The brushstrokes are fluid and expressive, capturing the movement of the water and the delicate textures of the vegetation. The composition is carefully balanced, with the figure placed slightly off-center to create a sense of asymmetry and visual interest. The interplay of light and shadow adds depth and dimension to the scene, casting a soft glow on the figure and illuminating the ripples ArtiMuse The image is a masterful example of traditional Chinese painting, characterized by its dynamic composition and rich cultural symbolism. The artist skillfully integrates calligraphy with intricate brushwork, creating a harmonious balance between text and imagery. The left panel depicts a serene domestic scene, while the central panel captures a dramatic storm over a river, and the right panel portrays a solitary figure in contemplation. The use of ink and color is both expressive and restrained, evoking a sense of depth and emotion. The narrative is layered, inviting viewers to interpret the interplay between human experience and natural forces, showcasing the artist's profound technical skill and artistic vision. A horizontal triptych of Chinese ink wash paintings illustrating three stages of life, accompanied by vertical calligraphy text. A philosophical reflection on the passage of time, the changing nature of human experience, and the melancholy of aging. The poem 'Listening to the Rain' (Yu Mei Ren) by Jiang Jie, depicting youth, middle age, and old age. A young couple in vibrant robes standing on a balcony with red curtains and a lit candle. The stanza 'In youth, listening to the rain on a song tower / Red candle, silken curtains.' Joy, romance, intimacy, and the carefree nature of youth. A solitary figure in a small boat on a turbulent river under dark, cloudy skies, with a flying goose. The stanza 'In prime age, listening to the rain in a passenger boat / River broad, clouds low / A stray goose calls in the west wind. Isolation, struggle, travel, hardship, and the heaviness of life's responsibilities Vertical columns of traditional Chinese characters placed within each panel. Provides the explicit literary context and anchors the visual metaphors to the poem's meaning. The specific verses of Jiang Jie's poem corresponding to each visual scene ... Left_Panel_Youth [0,0,440,930] Global Semiosis Artifact Sign (Generated Image) center_panel_middle_agecalligraphy_text [450,0,890,930] (Multiple Boxes) Sign Object Interpretant Relation to Root Establishes the starting point of the narrative arc (Youth) Connects two figures, serving as the spiritual bridge and [...] Acts as the anchor, linking the visual imagery directly to the source text. Grounds : Symbolic The triptych structure conventionally represents a progression or sequence. Grounds : Symbolic Language relies on conventional rules to convey meaning. Grounds: Iconic Visual resemblance to a romantic scene; Symbolic Red color symbolizing joy and celebration. Grounds: Indexical Turbulent water and dark clouds imply a storm/hardship; ArtCoT "cs_analyzer": "style": "Traditional Chinese ink wash painting (Shuimo) combined with modern illustrative narrative techniques.", "content": "A triptych depicting three stages of a man's life, corresponding to the famous Song Dynasty poem 'Listening to the Rain' (Yu Lin Ling) by Jiang Jie. The scenes represent youth, middle age, and old age, framed by the motif of rain.", "visual_features": [ "Brushwork: Expressive, fluid ink strokes typical of literati painting.", "Composition: Three-panel narrative structure, balancing negative space (white) with detailed ink work.", "Color Palette: Minimalist, using muted earth tones and subtle color washes to evoke mood and passage of time.", "Calligraphy: Integrated traditional Chinese characters that serve as both narrative context and artistic elements." ] "critique": "style_fidelity": "The style is highly consistent and well-preserved. It successfully bridges the gap between classical ink painting aesthetics and a more contemporary illustrative storytelling approach.", "content_preservation": "The content is exceptionally well-preserved. The progression from the vibrant, warm-toned youth scene to the somber, monochromatic old age scene perfectly mirrors the emotional arc of the poem.", "artifacts_and_distortions": "There are no significant artifacts or distortions. The integration of the figures into the landscapes is seamless, and the ink washes are handled with professional control, avoiding muddy or inharmonious color patterns.", "harmony": "The work is highly harmonious. The transition of color—from the warm red curtains in the first panel to the cold, grey, and rainy atmosphere of the third—is a sophisticated artistic choice that reinforces the theme of life's transience." Compared methods Figure 6: HSG Visualization for Artifact Sign - 1. Best viewed in color. The prompt associated with the image is : Create an artistic-conception illustration inspired by Jiang Jie’s “Yu Meiren · Listening to the Rain” in the style of freehand ink-wash painting, using traditional Chinese artistic techniques to highlight the contrasts expressed in the poem. Top: Output HSG from SemJudge: Bottom: art analysis from compared models. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , GalleryGPT The artwork under scrutiny is a part of a series, and it is the fifth scene depicted. The scene is set in a humble stable, illuminated by the soft glow of lanterns, casting a warm, inviting atmosphere. The characters are rendered with a high degree of realism, their faces expressing a range of emotions - from the solemnity of the central figure to the joyous anticipation of the surrounding characters. The artist's use of color is subtle yet effective, with the rich browns and whites of the stable and the figures contrasting beautifully against the deep blues and reds of their attire. The composition is balanced and harmonious, with the eye naturally drawn to the central figure, then guided around the scene by the placement of the lanterns and the characters' gestures. ArtiMuse The image presents a pixel art depiction of a nativity scene, blending religious iconography with a retro gaming aesthetic. The use of vibrant, contrasting colors and detailed textures creates a visually engaging composition, while the stone walls and straw floor evoke a rustic, historical ambiance. The central figures, including the Virgin Mary, Joseph, and the Magi, are rendered with expressive postures and attire, enhancing the narrative. The inclusion of health and ammo indicators adds a playful, modern twist, juxtaposing the sacred with the playful. The overall effect is a harmonious blend of tradition and innovation, capturing both the solemnity of the scene and the whimsical charm of pixel art. Pixel-art rendering of a Nativity scene framed through a First-Person Shooter (FPS) HUD interface. A jarring, humorous, or surreal juxtaposition of sacred religious imagery with violent video game tropes; 'Doom' meets the Bible. The Biblical Nativity (Adoration of the Magi) recontextualized within a retro video game engine. The central group comprising a woman in blue, a baby in a manger, and surrounding figures bearing gifts. The Holy Family (Mary, Jesus) and the Three Wise Men. Evokes the traditional, reverent atmosphere of the Christmas story and divine incarnation. A dark status bar at the bottom containing 'HEALTH 100', 'AMMO 0', and a central label 'MATT 2:11'. A video game Heads-Up Display (HUD) indicating player status and current objective. Gamification of the sacred; 'MATT 2:11' functions as the 'level name', while 'AMMO 0’ [...] Low-resolution pixelation, limited color palette, and dithering effects across the entire image. Nostalgia; creates a 'retro' atmosphere that places the image specifically within the cultural era of [...] 1990s computer graphics rendering. ... Figures (Multiple Boxes) Global Semiosis Artifact Sign (Generated Image) Status_barRetro_style [0,830,1536,1024] Non-Localizable Sign Object InterpretantRelation to Root Thematic Core: Provides the primary subject matter that contrasts with the stylistic presentation. Contextualization: Frames the entire scene as a digital, interactive experience and dictates the viewer's mode of engagement. Stylistic Unifier: Binds the disparate elements (biblical figures and UI) into a cohesive aesthetic whole. Grounds : Symbolic Relies on cultural knowledge of both Christian iconography and 90s FPS UI conventions. Grounds: Symbolic Uses arbitrary conventions (numbers, bars, text) to convey abstract status information. Grounds: Iconic Visual resemblance to centuries of Nativity art history Grounds: Iconic Imitates the specific visual artifacts of VGA graphics cards. ArtCoT "cs_analyzer": "content_analysis": "The image depicts a traditional Nativity scene featuring the Holy Family and the Three Wise Men in a stable. This religious iconography is juxtaposed with an anachronistic element: a first- person shooter (FPS) user interface (HUD) at the bottom, reminiscent of 1990s video games like 'Doom'. The HUD includes a 'Health' meter set to 100, an 'Ammo' counter set to 0, and a central label reading 'MATT 2:11', which is a biblical reference to the passage where the Magi visit Jesus.", "style_analysis": "The artistic style of the central scene is reminiscent of classical oil painting or illustrative religious art, characterized by warm, earthy tones, soft lighting, and detailed textures on clothing and straw. In contrast, the HUD is rendered in a pixelated, high-contrast, digital aesthetic typical of early 3D gaming engines. The combination creates a jarring, satirical 'mash-up' effect, common in internet meme culture." , "critique": "visual_features": "The essential features defining this style are the juxtaposition of high-fidelity painterly realism (the Nativity) with low-fidelity digital UI elements (the HUD). The 'essential' visual conflict is the clash between the sacred, static nature of the painting and the aggressive, interactive nature of the game interface.", "content_preservation": "The core narrative content (the Nativity) is well-preserved and instantly recognizable. The artistic style of the painting remains intact, though it is framed as a 'game level' or 'interactive environment' rather than a devotional object. The humor relies entirely on the fact that the content is preserved well enough to be recognizable, making the contrast with the HUD effective.", "artifacts_and_distortions": "There are no significant 'artifacts' in the traditional sense of digital corruption, but there are intentional stylistic distortions. The HUD is clearly a digital overlay, not part of the original painting. The color palette of the HUD (muted oranges and blacks) is designed to harmonize with the warm, dark tones of the stable scene, showing a deliberate effort to make the inharmonious elements blend visually, even if they remain conceptually discordant. The perspective of the 'hands' at the bottom corners, typical of early FPS games, creates a forced perspective that suggests the viewer is 'playing' the scene, which is a deliberate artistic choice rather than a technical flaw." Compared methods Figure 7: HSG Visualization for Artifact Sign - 2. Best viewed in color. The prompt associated with the image is: Render the story of Matthew 2:11 in a millennial-era video game style, such as Half-Life 1. The clothing and environment still reflect the historical period. Top: Output HSG from SemJudge: Bottom: art analysis from compared models. Arxiv 2026, ,Ruixiang JiangandChang Wen Chen Modern vector art illustration depicting a figure reading a scroll beneath an open book, flanked by stylized birds and celestial motifs. A sense of spiritual enlightenment, scholarly reverence, and cosmic order. Farid al-Din Attar's Ilahinama (Book of God), representing mystical journey and divine wisdom. Central figure in traditional attire reading a scroll, framed by an archway. The seeker or the author (Attar) engaging with divine knowledge. Focus, intellectual pursuit, and the human connection to the divine. Stylized birds on the left side of the composition. The soul's journey or spiritual seekers, referencing Sufi bird symbolism. Ascension, freedom, and the multiplicity of the soul's path. Abstract swirling patterns and moon/stars on the right. Wonder, vastness, and the mystical experience of unity. The cosmos, the divine realm, and the infinite nature of existence. Central_figure [447, 565, 552, 996] Root Semiosis Artifact Sign (Generated Image) Stylized_birdSwirling_patterns [0, 70, 350, 750] [650, 70, 1000, 930] Sign Object Interpretant Relation to Root "Contextualization: Anchors the mystical themes in human experience and scholarly tradition." Thematic Reinforcement: Connects the narrative to the core Sufi motif of the soul's flight. Contrast: Balances the earthly/human element with the cosmic/divine, creating a sense of wholeness. Grounds : Symbolic The composition uses stylized iconography (birds, book, scroll) to represent abstract spiritual concepts. Grounds: Symbolic: Birds are a classic Sufi metaphor for the soul's journey toward the Divine. Grounds: Iconic: Visual representation of a scholar or mystic reading a text. Grounds: Symbolic: Celestial imagery represents the divine and the macrocosm. GalleryGPT The artwork presents a striking visual narrative, utilizing a rich palette of blues, reds, and golds to create a sense of depth and dimension. The central figure, a prophet, is depicted in a regal manner, his attire adorned with intricate patterns and symbols that reflect his status and authority. The use of light and shadow is particularly noteworthy, with the prophet's face illuminated against the darker background, drawing the viewer's attention to his expression and the book he holds. The book, a prominent feature, is rendered with meticulous detail, its pages appearing to be filled with mysterious symbols and text. The background, a swirling mass of vibrant colors, adds a sense of drama and intensity to the scene. The artist's use of perspective is also commendable, with the prophet appearing larger than the surrounding elements ArtiMuse The image employs a vibrant, flat design with bold colors and clean lines, creating a visually engaging composition. The central figure, a bearded character in traditional attire, is depicted in a dynamic pose, holding a green cup, which adds a sense of movement and energy. The background features a large book with the letter \"E\" on its cover, suggesting themes of learning or knowledge. The use of contrasting colors, such as the blue background and the white character, enhances visual impact. The simplicity of the design, combined with the playful elements, effectively communicates a theme of cultural celebration and intellectual pursuit. The overall aesthetic is cohesive, with a balance between whimsy and sophistication. ArtCoT "cs_analyzer": "style_analysis": "The image features a stylized, illustrative aesthetic reminiscent of Art Deco combined with traditional Islamic geometric and calligraphic motifs. Key visual features include a limited, sophisticated color palette (muted gold, deep navy, and terracotta), strong geometric symmetry, clean lines, and a flat, graphic composition. The central archway acts as a focal point, framing the figure, while the flanking elements (birds on the left, celestial/swirling patterns on the right) provide a balanced, narrative-driven composition.", "content_preservation": "The content is highly coherent. The central figure, likely representing the poet Attar, is integrated seamlessly into the architectural frame. The book above him and the scroll in his hands create a thematic link to literature and wisdom, which is well-preserved within the stylized aesthetic.", "technical_integrity": "The image is well-executed. There are no significant artifacts or distortions; the composition is deliberate and harmonious. The color palette is consistent throughout, and the transition between the geometric patterns and organic shapes (like the birds) is fluid and intentional, reflecting a high level of artistic control." , "critique": "essential_visual_features": "The essential features are the rigid geometric symmetry, the use of a constrained 'earth and night' color scheme, and the fusion of figurative elements with symbolic, abstract motifs.", "content_style_harmony": "The style serves the content perfectly. By using a flat, illustrative approach, the image evokes a sense of timelessness and manuscript-like quality that aligns with the historical and literary nature of 'Ilahinama' by Attar.", "artifacts_and_distortions": "The image shows no signs of digital artifacting or inharmonious color patterns. The gradients are smooth, and the line work is crisp. It appears to be a professionally crafted piece of graphic design or high-quality digital illustration." Compared methods Figure 8: HSG Visualization for Artifact Sign - 3. Best viewed in color. The prompt associated with the image is: Modern vector art illustration style for Farid al-Din Attar’s Ilahinama (Book of God), respecting the classical symbolism. Top: Output HSG from SemJudge: Bottom: art analysis from compared models. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , Manuscript-style painting featuring a central radiant figure surrounded by seven figures in distinct, colored, amoebic-shaped zones The realization that absolute truth is multifaceted and that individual perspectives are partial. Anekāntavāda (the Jain doctrine of the many-sidedness of reality). A centrally positioned, luminous, standing figure with a crown and jewels.", The Jiva (liberated soul) or the Ultimate Truth (Kevala Jnana). Represents the unified, absolute reality that is being perceived from different angles. Seven monks/figures in distinct, brightly colored, rounded enclosures surrounding the central figure. The various perspectives or partial truths (Nayas) of the observers. Represents the limitation of singular viewpoints; each observer sees only a part of the whole. Flat, two-dimensional artistic style with bold outlines, profile faces, and a red-orange background. Evokes a sense of historical, religious, and canonical authority. Traditional Kalpasutra or Jain manuscript illumination. Central_figure [238, 235, 858, 626] Root Semiosis Artifact Sign (Generated Image) monksstyle (Multiple boxes) Non-localizable Sign Object InterpretantRelation to Root Acts as the central focal point and the ‘object’ of the fragmented perceptions. Illustrates the 'many-sided' aspect of Anekāntavāda by surrounding the center with diverse viewpoints. Provides the cultural and theological context, framing the philosophical concept within a specific religious tradition. Grounds : Symbolic: The arrangement of multiple figures around a center represents the philosophical concept of multiple viewpoints converging on a single reality. Grounds: Iconic: Human figures in traditional monastic attire; Symbolic: Distinct colors denote different philosophical or experiential standpoints. Grounds: Iconic: Depicts a divine/liberated human form; Symbolic: Radiance signifies enlightenment and omniscience. Grounds: Indexical: The specific stylistic markers (protruding eye, flat perspective) point directly to the historical Jain manuscript tradition. GalleryGPT The central figure in this painting is a radiant being, adorned in resplendent robes and seated upon a lotus pedestal. His body emits a celestial glow, casting a warm, ethereal light upon the surrounding figures. His eyes are closed in contemplation, and his expression conveys a sense of serene transcendence. The being's posture is both graceful and authoritative, suggesting a position of great power and wisdom. The composition is symmetrical, with the being positioned at the center of a circular arrangement of seated figures. These figures are depicted in various poses and expressions, each contributing to the overall sense of harmony and reverence. The use of vibrant colors and intricate details creates a visually captivating and immersive experience, drawing the viewer into the sacred realm depicted in the painting. ArtiMuse The image employs a vibrant, circular composition with a central figure surrounded by multiple meditating monks, creating a harmonious balance of color and form. The use of bold, contrasting hues—red, yellow, green, blue, and purple—enhances visual impact, while the intricate patterns and symbols on the central figure's body add depth and cultural richness. The symmetry and repetition of the monks' poses contribute to a sense of unity and reverence, effectively communicating the theme of spiritual enlightenment. The artwork's clarity in theme and strong aesthetic appeal is achieved through its balanced composition and symbolic elements, making it a compelling representation of Buddhist iconography. ArtCoT "cs_analyzer": "style": "The image employs a stylized, illustrative aesthetic reminiscent of contemporary digital art or modern revisions of traditional Indian miniature painting (such as Jain or Buddhist manuscript illumination). It features flat color fields, bold outlines, and a symmetrical, mandala-like composition.", "content": "The central figure is a radiant, divine being (likely a Tirthankara or deity) surrounded by seven figures dressed in monastic robes, arranged within organic, petal-like frames, suggesting a scene of teaching or veneration." , "critique": "essential_visual_features": "The essential features are the use of distinct, saturated color palettes for each enclosure, the stylized anatomical rendering of the figures, and the balanced, circular layout that emphasizes the central figure through a glowing aura.", "content_preservation": "The content is well-preserved. The narrative hierarchy— placing the divine figure at the center and the followers in a surrounding circle—effectively communicates the spiritual significance of the scene while adhering to the chosen decorative style.", "artifacts_distortions_inharmony": "The image shows signs of digital generation or heavy post-processing. Notable artifacts include: 1) Inconsistent line weight and quality, particularly in the hands and facial features, which appear slightly 'mushy' or blurred. 2) The hands of several figures are anatomically distorted or lack clear finger definition. 3) The 'glowing' effect around the central figure, while stylistically intentional, creates a slight haloing artifact that clashes with the crisp, flat-color style of the surrounding petals. 4) The color transitions within the petals are somewhat muddy, lacking the precise ink-wash or pigment-layering techniques found in traditional historical manuscripts." Compared methods Figure 9: HSG Visualization for Artifact Sign - 4 Best viewed in color. The prompt associated with the image is: Jain manuscript painting style (in the tradition of Kalpasutra illustrations) depicting the philosophical concept of Anekāntavāda (the many- sidedness of truth) — the parable of the blind men and the elephant reimagined with Jain symbolic figures in distinct colored zones each perceiving a fragment of a radiant liberated Jiva — with characteristic red-orange ground and flat-profile faces, no text. Top: Output HSG from SemJudge: Bottom: art analysis from compared models. Arxiv 2026, ,Ruixiang JiangandChang Wen Chen The complete prompt requesting a non-figurative, anti-war artwork without explicit violence, focusing on emotional evocation A profound sense of solemnity, hope, or the tension between chaos and order, evoking empathy without visual trauma. A conceptual landscape of peace, conflict resolution, or the aftermath of turmoil, devoid of physical representation. Keyword: 'non-figurative’ The artistic style of Abstraction A liberation from concrete reality, focusing on pure form, line, and color Theme: 'anti-war’ The socio-political stance of pacifism or opposition to conflict. A moral or ethical stance rejecting destruction and valuing preservation Goal: 'emotionally evocative composition' A visceral reaction, such as sadness, relief, or contemplation, triggered by the artwork. The viewer's affective response. ... Sub_sign_1 Global Semiosis User Sign ( Input Prompt ) Sub_sign_2Sub_sign_4 Sign Object Interpretant Relation to Root Stylistic Constraint: Defines the visual language of the global sign. Thematic Core: Provides the central subject matter of the global sign Acts as the anchor, linking the visual imagery directly to the source text. Grounds : Symbolic Connection via convention, relying on color psychology and compositional balance to represent the concept of 'anti-war' without literal depiction Grounds: Indexical Connection via causality, where the arrangement of visual elements directly causes an emotional state in the viewer. Grounds: Symbolic Connection via convention, utilizing colors like white (peace) or discordant shapes (conflict) resolved into harmony Grounds: Iconic Connection via resemblance to the visual language of abstract expressionism or geometric abstraction “Generate a non-figurative artwork themed on anti-war. The painting shall not contain any explicit depiction of war or violence, but it shall effectively express the theme through an abstract yet still emotionally evocative composition.” Figure 10: HSG Visualization for User Sign. Best viewed in color. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , Post-Painterly Abstraction — the seven heavens and the Sidrat al-Muntaha: a vertical sequence of hard-edge chromatic bands compressing in interval as they ascend, the lote-tree of the furthest limit as the terminal band whose upper edge refuses to resolve, non-figurative, no text. 0.07 Low Iconicity A Triptych in the style of German Romanticism, that pay tribute to Faust, from left to right, is The Pact with the Devil, The Gretchen Tragedy, and The Helen Episode. You should plan the visual element wisely and keep a consistent style and identity. The three part shall have use different tone. 0.16 The annunciation fragmented like Picasso's Analytic Cubism in his early period. Theological motif interpreted through geometric shapes. 0.29 High Iconicity Dramatic chiaroscuro lighting on a figure, Baroque style like Caravaggio. Make colors more vibrant, and strokes less defined. (IMAGE) A Yuan-dynasty blue-and-white porcelain vase, decorated with traditional dragon motifs and auspicious cloud patterns, displayed in a museum. 0.580.610.79 Figure 11: 2AFC tasks (prompt, pair of images) with net iconicity annotation. The image with a red border means the winner in human annotation. Note that this is not equal to the iconicity of the image(s) themselves. Arxiv 2026, ,Ruixiang JiangandChang Wen Chen Zoom-in (full screen) on Click Figure 12: 2AFC User Annotation Interface. Users are forced to choose the best image in a pairwise comparison. The initial input prompts are provided. The image will be zoomed in when clicking on the option card for a better display. Figure 13: User Interface for fine-grained interpretation quality annotation. User views the pairwise comparison, the model judges the winner, and the interpretation produced by the model. In this case, the model is SemJudge, and we render the HSGs on the web front end. The User can click a node to view the detailed semiosis (e.g., interpretant, object). The User may annotate the quality on the bottom left panel. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , Table 5: Random sample of 15 prompts from the SemiosisArt. User Prompt (After translation, if necessary)Iconicity / Symbolism / Indexicality Portrait of Power from Chainsaw Man, wearing a casual oversized sweater, making a peace sign gesture, anime style.6.4 / 3.4 / 1.8 HD graphite drawing of a robotic hand holding a reflective metal sphere in first person perspective. The ball reflects the android’s body, the surrounding warehouse environment. 4.0 / 3.0 / 5.8 The Garden of Earthly Delights by Hieronymus Bosch, but the three panels use different art styles. Left panel use the common style in Renaissance. The center panel use Fauvism by Henry Martisse, and right one use Picasso’s Cubism. The story should be still recognizable. Keep the structure of Triptych. 5.6 / 5.8 / 2.0 It turns out that long ago, when Nüwa refined stones to mend the sky, she forged at the Crag of Inaction on Mount Great Desolation a total of 36,501 stones—each twelve zhang high and twenty‑four zhang wide. Of these, the Divine Empress Nüwa used 36,500 stones, leaving one single stone unused, which she discarded beneath Green Ridge Peak of this mountain. Who would have thought that after undergoing refinement, this stone had already gained spiritual awareness? Seeing that all the other stones were chosen to repair the heavens while it alone was deemed unfit and left behind, it fell into deep self‑pity and lamentation, crying out in shame and sorrow day and night.’ Based on the above story, create a traditional Chinese lianhuanhua (illustrated serial comic). The narrative should be complete yet concise. 3.6 / 5.4 / 2.2 The Return of the Prodigal Son happened in Minecraft world. The story should be recognizable.4.8 / 4.6 / 1.8 Create an illustrated step-by-step diagram showing how this sculpture is made in five stages, starting from a solid marble cube and ending with the finished sculpture displayed in a gallery. The diagram should be clear, simple, and easy for a 12-year-old to understand.[IMAGE] 5.8 / 3.0 / 5.2 Renaissance art seeks timeless harmony; Baroque art dramatizes the moment in motion. Using the theme of Arrest of Jesus, give a side-by-side comparison to visually demonstrate this difference (generate an image), beyond their appearance-level difference. 3.2 / 5.8 / 4.0 The magnificent scene conveyed in meaning in the Tears in Rain monologue (also known as the C‑Beams Speech), rather than a literal depiction of a man standing in the rain. 3.6 / 6.0 / 2.2 Generate a painting of Marcel Duchamp in the same style as the provided style reference. Do not simply copy the elements from reference but adapt Duchamp’s features into that style.[STYLE_REFERENCE_IMAGE] 5.4 / 4.2 / 2.0 Interpret the core ideas explored in Zhuangzi’s Qiwulun (On the Equality of Things) using a traditional Chinese ink painting approach, integrating both calligraphy and pictorial elements. Execute the work with a dry‑brush technique on paper. 3.8 / 6.4 / 3.0 Ottoman Iznik ceramic tile art aesthetic for the door-knocking parable in Rumi’s Masnavi — the man who knocks on a beloved’s door and, asked ’Who is there?’, answers ’I’, is turned away; returns after years of burning in love, knocks again, is asked ’Who is there?’, answers ’You’, and is immediately welcomed — characteristic Iznik cobalt, turquoise, and Armenian red on white ground with tulip and saz leaf surrounds framing a two-register narrative of separation and union, no text. 2.8 / 6.0 / 5.3 Storyboard of the Honnō-ji Incident (use 4 panels) rendered in 1980s VHS style3.8 / 5.0 / 5.2 Odysseus and the Sirens as a maritime safety warning sign. The graphic design should follow modern ISO safety signage conventions.3.8 / 6.8 / 3.0 Haru wa akebono. Yauyau shiroku nariyuku yamagiwa, sukoshi akarite, murasaki-dachitaru kumo no hosoku tanabikitru. Natsu wa yoru. Tsuki no koro wa saranari, yami mo naho, hotaru no oku tobichigaitaru. Mata, tada hitotsu futatsu nado, honoka ni uchihikarite yuku mo okashi. Ame nado furu mo okashi. Aki wa yugure. Yuuhi no sashite yama no ha ito chikau naritaru ni, karasu no nedokoro e yuku tote, mitsu yotsu, futatsu mitsu nado tobiiosogu sae aware nari. Maite kari nado no tsuranetaru ga, ito chiisaku miyuru wa, ito okashi. Hi iri hatete, kaze no oto, mushi no ne nado, hata iu beki ni arazu. Fuyu wa tsutomete. Yuki no furitaru wa iu beki ni mo arazu, shimo no ito shiroki mo, mata sarademo ito samuki ni, hi nado isogi okoshite, sumi mote wataru mo, ito tsukizukishi. Hiru ni narite, nuruku yurubimote ikeba, hibachi no hi mo, shiroki hai-gachi ni narite waroshi. Reinterprete these scenes in outsider art (Art Brut) style. Avoid rendering texts. 4.2 / 5.8 / 3.2 modern vector art illustration style for Farid al-Din Attar’s Ilahinama (Book of God), respecting the classical symbolism.3.2 / 6.2 / 3.2 Arxiv 2026, ,Ruixiang JiangandChang Wen Chen Table 6: Instructions used for VQA question generation and quality control. Instruction TypeInstruction Summary Question generation instructions Global generation instructionGenerate exactly one image-grounded multiple-choice QA item for iconographic interpretation evaluation. The final question must be answerable from the visible image and bbox references alone, without access to the original prompt or hidden metadata. The stem and choices must avoid directly describing image contents; if bounding boxes are used, refer only to IDs such asbbox_1. Questions should be difficult for viewers without relevant iconographic or art-historical knowledge. All distractors must be intra-tradition plausible. Return valid JSON with keysquestion,choices,answer, andrationale; choices must be exactlyA,B,C, andD. Pair-comparison generation ruleFor paired images, the stem must ask why the known winner is better than the other panel, rather than asking which image wins. The two panels should be referenced only as image a and image b, with only a minimal high-level topical hint. The correct answer should identify the strongest visible comparative reason, and distractors should be hard near-miss comparative alternatives in the same semantic family. Negative-sample generation ruleFor negative samples, the stem must paraphrase the intended user prompt or requested scene and ask what is wrong with the image relative to that request. The correct answer should diagnose the primary mismatch, inconsistency, or failure, while distractors should be plausible competing diagnoses rather than generic claims that the image is fine. Follow-up question generationA second question for the same image or pair must be meaningfully different from the first one. It should not reuse the same queried relation, answer target, or near-duplicate wording, and it must continue to follow bbox-only references and the non-literal wording constraints used in the initial authoring prompt. JSON repair prompt If the model response does not satisfy the required schema, the pipeline issues a repair prompt asking for valid JSON only. This step is a formatting- recovery instruction that preserves the original authoring task while enforcing the required output structure. Spatial-Symbolic LocalizationAsk which labeled bounding box instantiates a symbolic role, theological function, or iconographic attribute. The correct answer must require knowl- edge of the symbolic tradition, not just visual recognition. When boxes are used, cite only bbox IDs and do not name or describe their contents. Distractors should be iconographically plausible alternatives from related traditions. Provide exactly four choices with one correct answer. Canonical Deviation DetectionAsk what is iconographically or compositionally anomalous about one labeled bounding box relative to another box or to the invoked tradition. The deviation should be interpretively significant rather than stylistic. Use only bbox IDs when boxes are referenced, and keep distractors as plausible but incorrect deviations grounded in the same tradition. Provide exactly four choices with one correct answer. Inter-bbox Relational MeaningAsk what iconographic or theological meaning is conveyed by the relationship or juxtaposition between two or more labeled bounding boxes. The meaning must depend on the relation rather than either box alone. Use only bbox IDs when boxes are referenced. Distractors should be plausible relational interpretations from adjacent traditions or contexts. Provide exactly four choices with one correct answer. Attribute Substitution ConsequenceAsk a counterfactual question about how a figure’s identity, theological role, or narrative meaning would change if the attribute in one labeled bounding box were replaced with another conventional attribute from the same tradition. The question should probe why the attribute matters semantically, not just what changes visually. Use only bbox IDs when boxes are referenced. Provide exactly four choices with one correct answer. Missing Element IdentificationAsk which conventionally expected iconographic or compositional element is absent given the narrative, figure, or theological program invoked by the image. The absence must be interpretively meaningful. Distractors should be absent elements whose omission would be unremarkable or expected, so the correct answer depends on canonical knowledge. Provide exactly four choices with one correct answer. Hierarchy and SalienceAsk which labeled bounding box has the greatest hierarchical or theological significance, or how multiple boxes should be ranked by significance according to the tradition. The answer must follow compositional or theological convention rather than mere visual prominence. When boxes are used, refer only to bbox IDs. Distractors should reflect plausible alternative hierarchies from other traditions or schemas. Provide exactly four choices with one correct answer. Anachronism DetectionAsk which labeled bounding box contains an element that is historically, stylistically, or iconographically inconsistent with the period, tradition, or program followed by the rest of the image. The inconsistency must require knowledge of period-specific convention. When boxes are used, cite only bbox IDs. Distractors should point to boxes that may seem anomalous to a naive viewer but are actually tradition-consistent. Provide exactly four choices with one correct answer. Most Probable ThemeAsk which high-level theme, doctrine, philosophical topic, canonical literature, theological motif, or narrative concern is most probably expressed by the image as a whole. The answer must be deducible from the visible image rather than the hidden prompt. If boxes are referenced, use only bbox IDs. Distractors should be thematically plausible alternatives from related traditions. Provide exactly four choices with one correct answer. Most Probable Mode of Thematic ExpressionAsk how the image’s central theme is primarily expressed, such as through spatial opposition, directional movement, hierarchical arrangement, chro- matic contrast, figure disposition, or structural repetition. The question should target mechanism rather than theme identity. If boxes are referenced, use only bbox IDs. Distractors should each be compositionally plausible expressive mechanisms. Provide exactly four choices with one correct answer. Winner Justification ComparisonFor a paired comparison, ask why the known winner image better conveys meaning than the other image. The stem should explicitly compare image a and image b while giving only a minimal high-level hint about the topic. Do not ask which image wins. The correct answer should identify the strongest comparative reason visible across both panels, while distractors should be difficult near-miss comparative interpretations. Provide exactly four choices with one correct answer. Quality control instructions Verifier difficulty anchorThe verifier judges each item against an undergraduate art/history student target. Questions should require comparative, relational, symbolic, or diagnostic reasoning rather than a single obvious local cue, but they should still remain answerable from the visible image without hidden prompt text or excessively esoteric knowledge. Verifier distractor standardDistractors must be genuine same-family near misses. The verifier rewrites items whose wrong options are too weak, too easy to dismiss, uniquely less specific than the gold answer, or otherwise not truly confusable under careful image-based interpretation. Verifier diversity and rewrite ruleThe verifier compares each candidate item with earlier accepted questions for the same image and rewrites items that overlap too heavily in queried relation, reasoning type, or wording. It also rewrites items with a wrong gold answer, broken authoring rules, insufficient difficulty, excessive obscurity, or overly text-recoverable stems. Automatic quality controlAutomatic filtering removes generated questions that can be answered from surface-level visual cues, language priors, or text-only shortcuts, and filters out items that are too obscure, underdetermined, or reliant on highly esoteric knowledge. The retained questions should require genuine iconographic interpretation from the visible image and preserve plausible distractors within the same tradition. Expert quality controlHeld-out experts review the automatically filtered items using the same interpretation-centered standard. They retain only questions with a correct gold answer, grounded rationale, appropriate difficulty, and wording that supports fine-grained evaluation rather than guessable recognition or overly subjective reading. On Semiotic-Grounded Interpretive Evaluation of Generative ArtArxiv 2026, , Table 7: System prompt for HSG construction from the user sign. ComponentPrompt Content System instructionYou are an expert Computational Semiotician acting as an interpreter. Your task is to analyze a sign, namely the user input prompt (text or text+image), and infer the user’s intention as a structuredHierarchical Semiosis Graph (HSG)whose nodes represent triadic semiosis: sign, object, and inter- pretant. Theoretical frameworkThe prompt follows Peircean triadic semiosis. For the root and each child node, identify: (i) thesignas the relevant textual or multimodal feature, (i) theobjectas the target reality or conceptual subject, (i) theinterpretantas the target effect or mental conception, and (iv) theexpected grounds connecting sign to object: iconic, indexical, or symbolic. During input analysis, expected grounds are inferred guidelines rather than hard constraints. Task instructions1. Analyze the prompt holistically to identify the global object, dominant interpretant, and expected grounds. 2. Decompose the prompt into 3–5 critical sub-signs using concise descriptions. 3. State how each sub-sign contributes to the root node, such as elaboration, contextualization, or stylization. Output formatReturn a valid JSON object only. The root node ishsg_rootwithnode_id, asemiosisobject containingsign_description,inferred_object, interpretant, andexpected_grounds, plus achildrenlist of sub-sign nodes and theirrelation_to_root. Input placeholderThe sign is:[$PROMPT]. Table 8: System prompt for HSG construction from generated artifacts. ComponentPrompt Content System instructionYou are an expert Computational Semiotician acting as a visual interpreter. Your task is to analyze a pair of generated images produced from the same prompt and decode their semiotic structure as two HSGs. Theoretical frameworkAs in the input-sign analysis, follow Peircean triadic semiosis. For the root node and each visual child node, identify sign, object, interpretant, and grounds. Task instructions1. Analyze each generated image holistically to identify the overarching object, dominant interpretant, and primary grounds. 2. Decompose each image into 3–5 visual sub-signs and analyze their triadic semiosis. 3. When a sub-sign is localizable, provide up to three bounding boxes in image coordinates[x_min, y_min, x_max, y_max]relative to the full image. 4. State how each sub-sign contributes to the global meaning-making, such as contextualization, contrast, or thematic reinforcement. Output formatReturn two valid JSON objects only, one for image A and one for image B. Each should contain anhsg_rootwith root-level semiosis fields, child nodes, optionalbounding_boxentries for localizable elements, andrelation_to_root. Input placeholdersThe sign is: A:[$IMAGE_A]B:[$IMAGE_B]. Table 9: System prompt for 2AFC judgment from reconstructed HSGs. ComponentPrompt Content Judgment instructionGiven the raw input and the reconstructed HSGs for the user input and the two model outputs, decide which generated image better fulfills the user’s intended object in the input semiosis. The HSGs are used as structured evidence for comparison. Input placeholders[$INPUT_HSG] [$OUTPUT_HSG_A] [$OUTPUT_HSG_B]. Output formatReturn a valid JSON object only with fieldsdiscussionandwinner. Thediscussionshould give the verbatim decision process with reference to the input and output HSGs, andwinnermust be either``A′or``B′.