Paper deep dive
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation
Dexiang Hong, Yijie Guo, Weidong Chen, Xinyan Liu, Zixuan Zou, Zhendong Mao, Yongdong Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/18/2026, 11:50:56 AM
Summary
The paper introduces EmoStyle, a framework for emotion-aware artistic image generation that bridges the gap between text prompts and fine-grained affective conditions. It utilizes an LLM reasoner to predict affective cues (valence-arousal, dominant emotion, therapeutic-effect) and aspect ratio, encoding them into an affective condition vector injected via AdaLN-style modulation into a Z-Image-based diffusion model. The system employs bucket-specific LoRA adapters for different artistic styles and uses a VLM-guided candidate selection process to ensure prompt alignment, style consistency, and emotional expression. The method achieved first place in the AffectiveArt Challenge 2026 Track 1.
Entities (10)
Relation Signals (9)
EmoStyle → achieves → first place
confidence 95% · our USTC_PI_LAB_TEAM submission achieved first place
EmoStyle → participatesin → AffectiveArt Challenge 2026
confidence 95% · In Track 1 of the AffectiveArt Challenge 2026
EmoStyle → uses → Z-Image
confidence 95% · EmoStyle, a Z-Image-based framework
EmoStyle → uses → LLM Reasoner
confidence 92% · An LLM reasoner first predicts affective cues
EmoStyle → uses → LoRA Adapter
confidence 90% · train a dedicated LoRA adapter for each artistic style bucket
EmoStyle → uses → AdaLN-style modulation
confidence 88% · inject it into the denoising blocks through AdaLN-style modulation
LLM Reasoner → predicts → Valence-Arousal
confidence 85% · predicts affective cues (valence-arousal...)
EmoStyle → uses → VLM-guided candidate selection
confidence 85% · VLM-guided candidate selection step ranks the generated images
EmoArt → provides → style buckets
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affective attributes available in the training data are not explicitly provided at test time. Without these attributes, the generator has to decide not only what to depict, but also how the target emotion should be expressed through color, lighting, brushwork, composition, line, and layout. This creates a control gap between the available test prompt and the fine-grained conditions needed for emotion-aware artistic generation. To bridge this gap, we propose EmoStyle, a Z-Image-based framework that converts the input prompt into a structured generation state. An LLM reasoner first predicts affective cues (valence-arousal, dominant emotion, and therapeutic-effect labels) and an aspect-ratio decision. Instead of using these predictions only as additional prompt text, we encode the affective fields into an affective condition vector and inject it into the denoising blocks through AdaLN-style modulation. This allows the inferred control variables to directly guide the generation of intermediate features. Since emotional expression is also style-dependent, we further train a dedicated LoRA adapter for each artistic style bucket and select the corresponding expert during inference, enabling the same affective cues to be rendered with bucket-specific priors for color, texture, brushwork, and composition. Finally, a lightweight VLM-guided candidate selection step ranks the generated images based on prompt alignment, style consistency, emotional expression, and visual quality. In Track 1 of the AffectiveArt Challenge 2026, our USTC\_PI\_LAB\_TEAM submission achieved first place.
Tags
Links
- Source: https://arxiv.org/abs/2607.10165v1
- Canonical: https://arxiv.org/abs/2607.10165v1
Trouble viewing inline? Open PDF directly →
Full Text
58,676 characters extracted from source content.
Expand or collapse full text
\@ACM@balancefalse EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation Dexiang Hong hongdexiang@mail.ustc.edu.cn University of Science and Technology of ChinaHeFei, China , Yijie Guo guoyijie@ustc.edu University of Science and Technology of ChinaHeFei, China , Weidong Chen chenweidong@ustc.edu.cn University of Science and Technology of ChinaHeFei, China , Xinyan Liu xinyliu@hit.edu.cn Harbin Institute of Technology, WeiHaiWeiHai, China , Zixuan Zou zouzixuan@hit.edu.cn Harbin Institute of Technology, WeiHaiWeiHai, China , Zhendong Mao zdmao@ustc.edu.cn University of Science and Technology of ChinaHeFei, China and Yongdong Zhang zhyd73@ustc.edu.cn University of Science and Technology of ChinaHeFei, China (2026) Abstract. Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the target emotion. In this challenge, the main difficulty is that the visual and affective attributes available in the training data are not explicitly provided at test time. Without these attributes, the generator has to decide not only what to depict, but also how the target emotion should be expressed through color, lighting, brushwork, composition, line, and layout. This creates a control gap between the available test prompt and the fine-grained conditions needed for emotion-aware artistic generation. To bridge this gap, we propose EmoStyle, a Z-Image-based framework that converts the input prompt into a structured generation state. An LLM reasoner first predicts affective cues (valence-arousal, dominant emotion, and therapeutic-effect labels) and an aspect-ratio decision. Instead of using these predictions only as additional prompt text, we encode the affective fields into an affective condition vector and inject it into the denoising blocks through AdaLN-style modulation. This allows the inferred control variables to directly guide the generation of intermediate features. Since emotional expression is also style-dependent, we further train a dedicated LoRA adapter for each artistic style bucket and select the corresponding expert during inference, enabling the same affective cues to be rendered with bucket-specific priors for color, texture, brushwork, and composition. Finally, a lightweight VLM-guided candidate selection step ranks the generated images based on prompt alignment, style consistency, emotional expression, and visual quality. In Track 1 of the AffectiveArt Challenge 2026, our USTC_PI_LAB_TEAM submission achieved first place. multimedia, affective computing, emotion image generation, diffusion models †copyright: acmlicensed†journalyear: 2026†conference: Proceedings of the 34th ACM International Conference on Multimedia; November 10–14, 2026; Rio de Janeiro, Brazil†booktitle: Proceedings of the 34th ACM International Conference on Multimedia (M ’26)†doi: 10.1145/X.X†isbn: 979-8-4007-X-X/2026/10†ccs: Computing methodologies Computer vision 1. Introduction Large-scale artificial intelligence models have rapidly advanced a wide range of multimedia tasks, including efficient video understanding, retrieval and recognition, vision-language reasoning, sentiment-aware language generation, and controllable creative design (Meng et al., 2026; Qin et al., 2025; Chen et al., 2021, 2022, 2023b; Hong et al., 2021b, 2022; Li et al., 2022c; Hong et al., 2021a; Li et al., 2022b, a; Gu et al., 2026; Liu et al., 2025; Wang et al., 2025a; Ye et al., 2025a; Li et al., 2025; Jin et al., 2024b, a; Liu et al., 2024; Wang et al., 2023a; Li et al., 2024b; Fu et al., 2024; Zhou et al., 2025; Tian et al., 2023; Han et al., 2023; Zhao et al., 2023; Lin et al., 2024; Wen et al., 2025; Wang et al., 2025b). Emotion-aware artistic generation combines text-to-image synthesis, computational aesthetics, and affective computing. Recent diffusion and flow-based models have greatly improved prompt-conditioned image fidelity and controllability (Ho et al., 2020; Dhariwal and Nichol, 2021; Rombach et al., 2022; Saharia et al., 2022; Esser et al., 2024; Z-Image Team et al., 2025). Yet visual emotion datasets and artistic benchmarks show that perceived emotion depends not only on object semantics, but also on artistic style, color, lighting, brushwork, line, and composition (Achlioptas et al., 2021; AffectiveArt Challenge Organizers, 2026; Yang et al., 2023). Emotion-aware artistic generation must therefore preserve content, follow style, and express the target emotion through appropriate visual form. In the AffectiveArt Challenge, the training data contains rich visual and affective annotations, including visual attributes, valence-arousal values, dominant emotions, and therapeutic-effect labels, but such structured annotations are absent at test time. The model receives only the input prompt and target style bucket, and must infer the missing cues that determine how the target emotion should be expressed under the required artistic style. To address this issue, we propose EmoStyle, a Z-Image-based framework for emotion-aware artistic generation. Given a prompt and its target style bucket, EmoStyle first uses an LLM reasoner to construct a structured generation plan, including affective cues (valence-arousal, dominant emotion, and therapeutic-effect labels) and an aspect-ratio decision, following the broader idea that prompt optimization and expansion can improve controllable generation (Hao et al., 2023; Wang et al., 2023b; Cao et al., 2023; Datta et al., 2023). The affective fields are encoded into an affective condition vector and injected into the denoising network through AdaLN-style shift-scale modulation (Peebles and Xie, 2023). This allows emotion-related cues to guide intermediate generation features rather than being used only as extra prompt text. To preserve style consistency, we train one LoRA adapter for each style bucket and select the corresponding expert during inference. The selected LoRA provides bucket-specific priors for color, texture, brushwork, and composition, while the affective condition controls how the target emotion is expressed within those priors. Since a single sample may still fail to satisfy content, style, affect, and quality constraints simultaneously, we further introduce a lightweight VLM-guided test-time refinement step, drawing on recent VLM reasoning and test-time scaling studies (Liu et al., 2023; Bai et al., 2023; Snell et al., 2024; Brown et al., 2024). It evaluates generated candidates according to prompt alignment, style consistency, affective consistency, and visual quality, and selects the best output without additional training. Table 1. Leaderboard of the ACM M’26 AffectiveArt Challenge (Track 1). Our submission (Rank 1) is highlighted with a gray background. Rank Participant Overall Score↑ FID↓ FID Score↑ AAS↑ Content↑ Style↑ Attribute↑ !20 1 USTC_PI_LAB_TEAM 0.80 66.12 0.60 0.99 0.99 0.99 1.00 2 EmoForge 0.78 77.47 0.56 1.00 1.00 1.00 1.00 3 VIRlab 0.78 75.56 0.57 0.99 0.98 0.99 0.99 4 edaich 0.77 78.59 0.56 0.99 0.98 0.99 0.99 5 Latent Feeling 0.76 87.51 0.53 1.00 0.99 1.00 1.00 This work makes three main contributions: • We introduce an LLM-based affective planner that infers missing affective and layout cues from the input prompt and target style bucket. • We propose a framework on top of Z-Image that separates style selection from affective modulation, where artistic styles are modeled by bucket-specific LoRA experts and affective cues are injected into denoising blocks through AdaLN-style modulation. • We design a VLM-guided test-time refinement strategy that evaluates generated candidates, identifies residual failures in content, style, affect, or quality, and selects or regenerates the final output. With this framework, our submission ranked first in Track 1 of the AffectiveArt Challenge 2026. 2. Related Work 2.1. Visual Emotion Datasets Visual emotion analysis relies on datasets that connect visual content with human affective responses. Early benchmarks such as FI provide large-scale discrete emotion labels for image emotion recognition (You et al., 2016), while EMOTIC annotates natural scenes with both categorical emotions and continuous valence-arousal-dominance dimensions (Kosti et al., 2020). For artworks, ArtEmis introduces emotion attributions and natural-language explanations, enabling the study of why a painting evokes a particular feeling (Achlioptas et al., 2021). Recent datasets further improve annotation richness. EmoSet provides interpretable emotion-related attributes (Yang et al., 2023), EmoVerse builds an MLLM-driven emotion representation dataset for interpretable visual emotion analysis (Guo et al., 2025), and EmoArt provides multidimensional annotations for emotion-aware artistic generation, including painting styles, visual attributes, and emotion categories (Zhang et al., 2025a). These resources have advanced emotion recognition, explanation, and benchmarking, and recent emotional generation methods for images and videos also show the importance of grounding abstract affect in concrete visual content (Yang et al., 2024; Yuan et al., 2025; Ye et al., 2024; Chen et al., 2026d). Beyond direct generation, a related line of work studies structured affective reasoning and explanation, including fine-grained emotion-cause extraction (Chen et al., 2026c; Ye et al., 2025b), consensus-prompted affective explanation captioning (Song et al., 2026; Zhang et al., 2026b), and multimodal empathetic response generation (Wang et al., 2026; Chen et al., 2026b). However, their annotations do not directly specify how a compact challenge caption should be mapped to controllable generation variables, such as affective state, visual attributes, style prior, and aspect ratio. EmoStyle uses structured reasoning and affective conditioning to bridge this gap. 2.2. LoRA Adaptation and Composition Parameter-efficient adaptation is widely used to specialize large text-to-image models without retraining all parameters. LoRA represents weight updates as low-rank matrices, reducing adaptation cost while preserving the base model prior (Hu et al., 2022). Related personalization methods, including Textual Inversion and DreamBooth, learn new visual concepts or subject identities from a small number of examples (Gal et al., 2023; Ruiz et al., 2023). Custom Diffusion and Mix-of-Show further extend this approach to multi-concept customization and shared-scene composition (Kumari et al., 2023; Gu et al., 2023). Style-oriented diffusion methods, such as InST, StyleAligned, and ArtAdapter, further show that compact conditions or adapters can transfer artistic appearance and preserve style consistency (Zhang et al., 2023; Hertz et al., 2024; Chen et al., 2023a). These works show that compact adaptation modules can efficiently encode new concepts and styles. Recent studies also examine how multiple adapters interact. Multi-LoRA Composition proposes training-free strategies such as LoRA Switch and LoRA Composite (Zhong et al., 2024), while LoRA-Composer shows that naive use of multi-LoRA can lead to concept confusion or concept vanishing (Yang et al., 2025). LoraHub further explores dynamic LoRA composition across tasks (Huang et al., 2024). In creative graphic design, multi-conditional diffusion, layout-to-image generation, editable multi-layer poster generation, image-conditioned poster design, and generative parsing methods also highlight the need for explicit control over layout, elements, and editable visual structure (Zhang et al., 2026a, 2025b, 2025c; Hu et al., 2025; Chen et al., 2026a). These methods mainly target subjects, identities, tasks, design elements, or general concepts, rather than affective artistic variables such as valence, arousal, dominant emotion, visual attributes, and style-sensitive resolution. In contrast, we train LoRA experts for style buckets and use affective conditioning to control emotional expression, keeping style selection and affective modulation separate. 2.3. Expert-Based Style Specialization Mixture-of-Experts (MoE) models offer a general mechanism for conditional specialization. Sparsely-gated MoE activates a small subset of experts for each input (Shazeer et al., 2017), while Switch Transformers simplify the design by selecting the top-1 expert (Fedus et al., 2022b). Beyond language models, V-MoE applies sparse expert routing to vision transformers (Riquelme et al., 2021), Expert Choice Routing improves load balancing by letting experts select tokens (Zhou et al., 2022), and Graph-MoE combines memory-augmented routers with sparse experts for multivariate time-series anomaly detection (Huang et al., 2025). Recent reviews summarize the broader design trade-offs of sparse experts, including capacity, routing stability, and transfer (Fedus et al., 2022a). These works support the idea that different inputs can benefit from specialized parameter subsets. However, learning a free-form router requires a clear routing objective and sufficient data for each expert. In our setting, EmoArt already provides style-bucket annotations, which serve as an explicit and reliable specialization signal. Recent adapter-based MoE methods, such as MoLE and MixLoRA, show that LoRA modules can be organized as reusable experts (Wu et al., 2024; Li et al., 2024a). Unlike these learned-routing methods, we train one style LoRA expert per bucket and use a deterministic bucket-to-expert mapping. This keeps style selection interpretable while leaving affective variation to the proposed conditioning module. 3. Method 3.1. Framework Overview We address Track 1 of the AffectiveArt Challenge 2026, where each input caption specifies semantic content, artistic movement, and the desired emotional state. EmoArt (Zhang et al., 2025a) provides the affective fields and style buckets used in our method. For the i-th sample, we denote the input caption as xix_i and its target style bucket as yiy_i. Our goal is to generate artwork that is semantically faithful, emotionally aligned, and consistent with the target artistic style. Since test captions are compact, we formulate the task as structured affective conditioning over bucket-specific style experts. As shown in Fig. 1, EmoStyle follows a four-stage pipeline. First, an LLM reasoner expands the input caption into an affective plan iP_i, including aspect ratio, valence-arousal cues, dominant emotion, and therapeutic-effect labels. Second, the affective fields are encoded as i,affc_i,aff, while a rule-based bucket mapping selects the style LoRA expert. Third, Z-Image generates candidates from the text embedding, timestep embedding, selected LoRA expert, and AdaLN-style affective modulation. Finally, a VLM judge scores prompt alignment, style consistency, affective consistency, and image quality, then selects the best candidate or triggers regeneration when necessary. Figure 1. Overview of EmoStyle. An LLM reasoner builds an affective plan, rule-based style selection and affective modulation condition the Z-Image generator, and a VLM judge refines the final candidate. Pipeline diagram with four stages: LLM reasoning, affective control construction and style-bucket LoRA selection, Z-Image generation, and VLM-guided candidate refinement. 3.2. Prompt-to-Affect Planning Although the challenge caption specifies semantic content, artistic movement, and the desired emotional state, the structured affective annotations available in EmoArt are not explicitly provided at test time. Instead of rewriting the caption or generating free-form visual attributes, we use an LLM reasoner only to infer compact affective and layout variables that can be directly consumed by the generator. For the i-th sample, the reasoner takes the input caption xix_i, the target style bucket yiy_i, and a fixed instruction template Π as input: (1) (Fi,ρi)=Rψ(xi,yi;Π),(F_i, _i)=R_ψ(x_i,y_i; ), where RψR_ψ denotes the LLM reasoner, FiF_i is the predicted affective summary, and ρi _i is the predicted aspect ratio. Together they constitute the affective plan i=(Fi,ρi)P_i=(F_i, _i). The affective summary contains continuous valence-arousal scores and discrete affective labels: (2) Fi=(νi,ηi,ei,Hi),F_i=( _i, _i,e_i,H_i), where νi,ηi∈[0,1] _i, _i∈[0,1] denote valence and arousal, eie_i is the dominant emotion, and HiH_i denotes the therapeutic-effect labels. The dominant emotion and therapeutic-effect labels are selected from the predefined EmoArt label vocabularies, while valence and arousal are normalized continuous values. The reasoner also predicts the aspect ratio ρi _i from a predefined candidate set. The predicted ratio determines the canvas orientation and is converted into the final generation resolution. For landscape layouts, we set the short edge based on the target style bucket and derive the long edge from ρi _i; portrait layouts are handled symmetrically. The resulting resolution is rounded to the nearest multiple of 16 before generation. The resulting plan provides two explicit control signals. The affective summary FiF_i is converted into the affective condition vector in Sec. 3.4 through valence-arousal coordinate embedding and label embeddings, while the aspect-ratio decision ρi _i determines the output layout. The original caption xix_i is directly encoded by the frozen Z-Image text encoder as the semantic condition. We do not use LLM-generated caption rewriting or free-form visual attributes, which keeps the control state compact and avoids relying on unstable natural-language fields at test time. 3.3. Bucket-Specific Style LoRA Experts EmoArt provides a target style bucket for each sample, which gives a reliable supervision signal for style specialization. Instead of learning an additional router, we train one LoRA expert for each style bucket on top of the frozen Z-Image backbone. In our implementation, the bucket set contains inkwash, ukiyoe, gongbi, renaissance, abstract, soviet realism, baroque, and impressionism. Let pip_i denote the text prompt of the i-th training sample, IiI_i denote the corresponding artwork image, and yi∈y_i denote its style bucket. For each bucket y, we define the bucket-specific training subset as (3) y=(pi,Ii)∣yi=y.D_y=\(p_i,I_i) y_i=y\. We train a LoRA expert EyE_y with parameters θy _y on yD_y, while keeping the Z-Image backbone ΘZ _Z frozen: (4) θy⋆=argminθy(pi,Ii)∼y,t∼(0,1)[ℒFM(Ii,pi,t;ΘZ,θy)], _y = _ _yE_(p_i,I_i) _y,\;t (0,1) [L_FM (I_i,p_i,t; _Z, _y ) ], where t is the flow-matching timestep and ℒFML_FM is the standard text-conditioned flow-matching loss used by Z-Image. This bucket-wise training produces a compact style prior for each artistic category while retaining the general image-generation ability of the base model. At inference time, style selection is deterministic: (5) Ei=Eyi.E_i=E_y_i. The selected expert is used throughout generation and remains fixed during the subsequent affective-modulation stage. This design keeps style control simple, avoids poorly supervised routing, and separates style specialization from emotion-specific conditioning. 3.4. Affective Condition Modulation The structured plan in Sec. 3.2 provides two types of affective cues: continuous valence-arousal scores and discrete emotion labels. We do not encode the natural-language visual attributes into the affective condition, because these attributes are less stable at test time and may overlap with the text prompt. Instead, we represent affect through a compact label-coordinate embedding and inject it into the denoising blocks. For sample i, let νi,ηi∈[0,1] _i, _i∈[0,1] denote valence and arousal, eie_i denote the dominant emotion, and HiH_i denote the set of therapeutic-effect labels. We view valence and arousal as a two-dimensional affective coordinate i=[νi,ηi]⊤a_i=[ _i, _i] . To better represent this continuous affective space, we use a sinusoidal coordinate embedding similar to positional encoding. Let (6) ℬ=(2k,0),(0,2k),(2k,2k),(2k,−2k)k=0K−1B=\(2^k,0),(0,2^k),(2^k,2^k),(2^k,-2^k)\_k=0^K-1 be a fixed set of frequency directions, where K denotes the number of frequency bands. The valence-arousal embedding is defined as (7) ΦVA(i)=[i;sin(π⊤i),cos(π⊤i)∈ℬ]. _VA(a_i)=[a_i;\ ( a_i), ( a_i)\_b ]. This embedding preserves the continuous structure of the affective space while allowing the model to distinguish different emotional regions and intensity levels. The dominant emotion is encoded with a learnable label embedding e,i=Ee(ei)z_e,i=E_e(e_i). For the therapeutic-effect labels, which can be multi-label, we pool a learnable per-label embedding EHE_H over the active labels: (8) H,i=1max(1,|Hi|)∑h∈HiEH(h),z_H,i= 1 (1,|H_i|) _h∈ H_iE_H(h), where an empty label set is represented by a learnable null embedding. The continuous affective coordinate is projected to the same dimension by a learnable matrix WVAW_VA, i.e. VA,i=WVAΦVA(i)z_VA,i=W_VA _VA(a_i). We then build the affective feature by combining the continuous coordinate, the discrete labels, and their interaction: (9) i,aff=[VA,i;e,i;H,i;VA,i⊙e,i].s_i,aff=[z_VA,i;z_e,i;z_H,i;z_VA,i _e,i]. The interaction term helps distinguish cases with similar valence-arousal values but different dominant emotions. A lightweight encoder maps this feature into the final affective condition: (10) i,aff=gϕ(i,aff).c_i,aff=g_φ(s_i,aff). For each style bucket y, we pair its LoRA expert EyE_y with a bucket-specific residual modulation module ℳy=Ml,ylM_y=\M_l,y\_l. The affective encoder gϕg_φ is shared across buckets, while the residual modulation heads are style-specific. This allows the model to learn a common affective representation, while each style bucket learns how to express the same affective state under its own visual prior. We implement affective control as residual offsets to the native scale-gate modulation of Z-Image, rather than inserting an additional AdaLN layer. Let hi,l(t)h_i,l^(t) be the hidden state at denoising step t and block l, and let te_t be the timestep embedding. The frozen Z-Image block first produces its original modulation parameters: (11) (l,msaZ,(t),l,msaZ,(t),l,mlpZ,(t),l,mlpZ,(t))=AlZ(t),(s_l,msa^Z,(t),g_l,msa^Z,(t),s_l,mlp^Z,(t),g_l,mlp^Z,(t))=A_l^Z(e_t), where AlZA_l^Z denotes the native modulation network of Z-Image. For the active style bucket yiy_i, our affective head predicts residual offsets: (12) (Δi,l,msa(t),Δi,l,msa(t),Δi,l,mlp(t),Δi,l,mlp(t))=Ml,yi(t,i,aff).( _i,l,msa^(t), _i,l,msa^(t), _i,l,mlp^(t), _i,l,mlp^(t))=M_l,y_i(e_t,c_i,aff). The final scale and gate parameters are obtained by residual merging, for ⋆∈msa,mlp ∈\msa,mlp\: (13) i,l,⋆(t)=1+l,⋆Z,(t)+Δi,l,⋆(t),i,l,⋆(t)=tanh(l,⋆Z,(t)+Δi,l,⋆(t)). α_i,l, ^(t)=1+s_l, ^Z,(t)+ _i,l, ^(t), τ_i,l, ^(t)= \! (g_l, ^Z,(t)+ _i,l, ^(t) ). The selected style LoRA expert is applied inside the attention and MLP layers of the block, where RMSNorm(⋅)RMSNorm(·) denotes root-mean-square normalization. The resulting block computation is (14) h^i,l(t)=hi,l(t)+i,l,msa(t)⊙RMSNorm(Attnl,Eyi(RMSNorm(hi,l(t))⊙i,l,msa(t))), h_i,l^(t)=h_i,l^(t)+ τ_i,l,msa^(t) \! (Attn_l,E_y_i\! (RMSNorm(h_i,l^(t)) α_i,l,msa^(t) ) ), (15) hi,l+1(t)=h^i,l(t)+i,l,mlp(t)⊙RMSNorm(MLPl,Eyi(RMSNorm(h^i,l(t))⊙i,l,mlp(t))).h_i,l+1^(t)= h_i,l^(t)+ τ_i,l,mlp^(t) \! (MLP_l,E_y_i\! (RMSNorm( h_i,l^(t)) α_i,l,mlp^(t) ) ). The final projection of each residual modulation head is zero-initialized, so the residual offsets are initially zero. Therefore, at the beginning of training, the block exactly reduces to the native Z-Image block equipped with the selected style LoRA expert. During training, we keep the Z-Image backbone and its native modulation network frozen, and jointly optimize the shared affective encoder, the bucket-specific LoRA experts, and their corresponding residual modulation modules using the standard Z-Image flow-matching loss. For samples from bucket y, only EyE_y and ℳyM_y are updated, while gϕg_φ receives gradients from all buckets. Figure 2. Qualitative comparison of optimized results across different generation backbones. Each row corresponds to one model after applying our affective planning and refinement strategy, and each column represents one test sample with the corresponding ground-truth artwork shown in the last row. Qualitative comparison grid showing generated artworks from different backbones and the corresponding ground-truth artwork. 3.5. VLM-Guided Candidate Refinement Although the planned prompt and affective conditions provide an initial generation state, a single stochastic sample may still fail in content, style, affective expression, or image quality. Therefore, during inference, we generate multiple candidates for each prompt with different random seeds and rank them using a VLM judge. Let iC_i denote the generated candidate set for sample i. Given the input caption, target style, and affective plan, the judge scores each candidate across four aspects: prompt consistency, style consistency, affective-plan consistency, and no-reference visual quality: (16) J(I)=λpJp(I)+λsJs(I)+λaJa(I)+λqJq(I).J(I)= _pJ_p(I)+ _sJ_s(I)+ _aJ_a(I)+ _qJ_q(I). Here, λp _p, λs _s, λa _a, and λq _q are nonnegative weights; JpJ_p measures consistency with the input caption, JsJ_s measures consistency with the target style, JaJ_a measures consistency with the affective plan, and JqJ_q measures no-reference visual quality. The candidate with the highest score is selected as the final output: (17) Ii∗=argmaxI∈iJ(I).I_i = _I _iJ(I). This refinement strategy improves the final output without additional training. It is especially useful for correcting occasional failures caused by weak style expression, insufficient affective-plan consistency, or poor visual quality. 4. Experiments 4.1. Challenge Setting and Evaluation Metrics We follow Track 1 of the AffectiveArt Challenge 2026, namely Emotion-Aware Artistic Image Generation. Given a short caption that describes the semantic content, artistic movement, and the desired emotional state, participants must generate a single artistic image that satisfies all three conditions. The track therefore evaluates whether a generative model can preserve the requested visual content while aligning with both the specified artistic style and the target emotional state. The official dataset for the challenge is EmoArt, a large-scale, emotion-aware artwork dataset containing 132,664 painting images from public-domain art sources. EmoArt covers 56 artistic styles and provides fine-grained annotations for both visual appearance and affective semantics, including content descriptions, five visual attributes, valence-arousal values, dominant emotion labels, and therapeutic-potential labels. The five visual attributes are brushwork, composition, color, line, and light, which are closely related to how emotion is expressed in artworks. For local model selection and ablation, we construct a validation set from the official training split, matching the per-style sample count of the test set so that the validation distribution is consistent with the final evaluation setting. For the Gongbi category, where the training split contains only 32 images but the test set contains 178, we use visually related China Image samples as substitutes in the train and val sets. The Challenge evaluation metrics include Frechet Inception Distance (FID) and Attribute Alignment Score (AAS). FID measures the distance between the feature distributions of generated and real images (Heusel et al., 2017), with lower values indicating better visual fidelity and distributional realism. AAS evaluates how well the generated artwork aligns with the target artistic and affective attributes. Following the challenge protocol, the official AAS is computed using a MiniCPM-V-2.6 evaluator (Yao et al., 2024) fine-tuned on EmoArt annotations. This evaluator predicts an attribute description for each generated image and compares it with the ground-truth attribute text using CLIP similarity (Radford et al., 2021). Since the official fine-tuned evaluator is not available during local development, we use the original MiniCPM-V-2.6 evaluator without additional fine-tuning for validation experiments, while the final leaderboard score is obtained from the official challenge evaluator. Higher AAS indicates better conditional alignment. Table 2. Main Results and Cross-Backbone Comparison on Local Validation Set (P2A+R: Prompt-to-Affect + Refinement). Model Setting FID↓ AAS↑ Content↑ Style↑ Attribute↑ Z-Image Base 75.8357 0.7309 0.6711 0.6972 0.8243 gray!45 black GPT Image 2 P2A + R 75.0873 0.7616 0.6815 0.7549 0.8484 Nano Banana 2 P2A + R 79.2556 0.7487 0.6450 0.7477 0.8535 gray!45 black Flux2-klein-base Full EmoStyle 69.4151 0.7435 0.6591 0.6814 0.8901 Z-Image Full EmoStyle 58.6373 0.7864 0.7153 0.7548 0.8892 4.2. Implementation Details We use Z-Image as the backbone. For each style bucket we train a separate LoRA expert (rank 128, 3 epochs, learning rate 1×10−41× 10^-4) with the backbone frozen; in the subsequent affective-control stage the backbone and style experts are kept fixed, and only the lightweight affective encoder and modulation heads are optimized to inject emotional cues into the denoising process. The LLM reasoner in Prompt-to-Affect Planning is instantiated with Gemini-3.5-Flash (Google, 2026), which converts short captions and target styles into structured affective plans, and the VLM judge in VLM-Guided Candidate Refinement with Qwen3.5-397B-A17B, which ranks the generated candidates. 4.3. Main Results and Cross-Backbone Comparison We conduct experiments on the validation set using Z-Image as the generation backbone and report FID, AAS, and the three AAS components. Table 4.1 reports the main results and the cross-backbone comparison. Compared with the Z-Image base model, full EmoStyle reduces FID from 75.8357 to 58.6373 and increases AAS from 0.7309 to 0.7864, indicating improved visual realism and conditional alignment. The gains in content, style, and attribute scores show that the proposed structured conditioning helps the model better preserve the prompt semantics, target style, and affective attributes. In addition, to further evaluate the effectiveness of our method, we compare it with several open-source and closed-source generation models, including Flux2-klein-base, GPT Image 2, and Nano Banana 2. Because closed-source API models do not expose internal parameters for bucket-specific LoRA adaptation and AdaLN-style modulation, we apply only the model-agnostic components of our framework to these models. Specifically, the structured affective plan is serialized as part of the generation instruction, and the candidates are ranked by the same VLM-guided refinement procedure, while the full EmoStyle architecture is implemented on open-source backbones. Overall, Z-Image achieves the best FID and AAS, Flux2-klein-base the highest attribute alignment, and GPT Image 2 competitive style alignment. Table 3. Ablation study on the local validation set. Method FID↓ AAS↑ Content↑ Style↑ Attribute↑ EmoStyle(Ours) 58.6373 0.7864 0.7153 0.7548 0.8892 - w/o aspect-ratio prediction 72.5912 0.7866 0.7052 0.7784 0.8761 - w/o prompt-to-affect planning 68.4145 0.7957 0.7535 0.7715 0.8621 - w/o style-specific LoRA 66.8951 0.7653 0.6851 0.7098 0.8741 - w/o candidate refinement 60.7312 0.7751 0.6987 0.7451 0.8815 Figure 2 presents qualitative results across different generation backbones. The examples are consistent with the quantitative results in Table 4.1: Z-Image produces more balanced results in terms of content preservation, style consistency, and visual quality, other backbones show varying degrees of degradation in style or detail. 4.4. Ablation Study Table 3 evaluates the contribution of each component. Removing aspect-ratio prediction causes the largest FID degradation, from 58.6373 to 72.5912, indicating that layout planning is important for matching the visual distribution of EmoArt. Removing prompt-to-affect planning weakens attribute alignment from 0.8892 to 0.8621, suggesting that structured affective planning helps the model realize fine-grained visual attributes. The higher content, style, and AAS scores in this variant indicate that these automatic metrics may favor direct prompt matching, while the full model achieves better visual realism and attribute realization. Removing the style-specific LoRA results in a clear drop in style alignment, from 0.7548 to 0.7098, confirming the importance of style-specialized adaptation. Removing candidate refinement also degrades both FID and AAS, showing that VLM-guided candidate selection improves final output quality. Overall, the full method achieves the best FID and attribute alignment while maintaining competitive content and style alignment, demonstrating a balanced improvement in visual quality and conditional control. 5. Conclusion We present EmoStyle, a Z-Image-based framework for emotion-aware artistic image generation. Rather than relying solely on the input prompt, EmoStyle constructs a structured affective generation state that integrates prompt-to-affect reasoning, bucket-specific LoRA selection, AdaLN-style affective modulation, and VLM-guided test-time refinement, letting the generator jointly consider semantic content, artistic style, visual attributes, emotion cues, and layout while keeping the base diffusion backbone frozen. Experiments on EmoArt show that the framework improves visual quality and conditional alignment across different backbones, and ablations confirm the contributions of layout planning, affective reasoning, style-specific adaptation, and candidate refinement. In Track 1 of the AffectiveArt Challenge 2026, our system achieved first place, demonstrating the effectiveness of structured affective conditioning for controllable artistic image generation. References P. Achlioptas, M. Ovsjanikov, K. Haydarov, M. Elhoseiny, and L. Guibas (2021) ArtEmis: affective language for visual art. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual Conference, p. 11569–11579. External Links: Document Cited by: §1, §2.1. AffectiveArt Challenge Organizers (2026) AffectiveArt Challenge 2026: emotion-aware artistic image generation. ACM Multimedia Grand Challenge. External Links: Link Cited by: §1. J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou (2023) Qwen-VL: a versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966. External Links: Link Cited by: §1. B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini (2024) Large language monkeys: scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787. External Links: Link Cited by: §1. T. Cao, C. Wang, B. Liu, Z. Wu, J. Zhu, and J. Huang (2023) BeautifulPrompt: towards automatic prompt engineering for text-to-image synthesis. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track, Singapore, p. 1–11. Cited by: §1. D. Chen, H. Tennent, and C. Hsu (2023a) ArtAdapter: text-to-image style transfer using multi-level style encoder and explicit adaptation. arXiv preprint arXiv:2312.02109. External Links: Link Cited by: §2.2. W. Chen, D. Hong, Z. Mao, Y. Cheng, X. Liu, L. Zhang, and Y. Zhang (2026a) CreatiParser: generative image parsing of raster graphic designs into editable layers. arXiv preprint arXiv:2604.19632. External Links: Link Cited by: §2.2. W. Chen, D. Hong, Y. Qi, Z. Han, S. Wang, L. Qing, Q. Huang, and G. Li (2022) Multi-attention network for compressed video referring object segmentation. In Proceedings of the 30th ACM International Conference on Multimedia, p. 4416–4425. External Links: Document, Link Cited by: §1. W. Chen, G. Li, X. Zhang, S. Wang, L. Li, and Q. Huang (2023b) Weakly supervised text-based actor-action video segmentation by clip-level multi-instance learning. ACM Transactions on Multimedia Computing, Communications, and Applications 19 (1), p. 12:1–12:22. External Links: Document, Link Cited by: §1. W. Chen, G. Li, X. Zhang, H. Yu, S. Wang, and Q. Huang (2021) Cascade cross-modal attention network for video actor and action segmentation from a sentence. In Proceedings of the 29th ACM International Conference on Multimedia, p. 4053–4062. External Links: Document, Link Cited by: §1. W. Chen, C. Ye, Z. Mao, P. Song, X. Liu, L. Zhang, X. Chang, and Y. Zhang (2026b) FACE-net: factual calibration and emotion augmentation for retrieval-enhanced emotional video captioning. arXiv preprint arXiv:2603.17455. External Links: Link Cited by: §2.1. W. Chen, C. Ye, Z. Mao, L. Wang, X. Liu, and Y. Zhang (2026c) Towards accurate emotion-attributed video captioning via fine-grained emotion-cause pair extraction. arXiv preprint arXiv:2606.08566. External Links: Link Cited by: §2.1. W. Chen, C. Ye, P. Song, L. Zhang, Y. Zhang, and Z. Mao (2026d) Subjective-objective emotion-correlated generation network for subjective video captioning. IEEE Transactions on Image Processing 35, p. 540–555. External Links: Document, Link Cited by: §2.1. S. Datta, A. Ku, D. Ramachandran, and P. Anderson (2023) Prompt expansion for adaptive text-to-image generation. arXiv preprint arXiv:2312.16720. External Links: Link Cited by: §1. P. Dhariwal and A. Nichol (2021) Diffusion models beat GANs on image synthesis. Advances in Neural Information Processing Systems 34, p. 8780–8794. Cited by: §1. P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y. Marek, and R. Rombach (2024) Scaling rectified flow transformers for high-resolution image synthesis. arXiv preprint arXiv:2403.03206. External Links: Link Cited by: §1. W. Fedus, J. Dean, and B. Zoph (2022a) A review of sparse expert models in deep learning. arXiv preprint arXiv:2209.01667. External Links: Link Cited by: §2.3. W. Fedus, B. Zoph, and N. Shazeer (2022b) Switch transformers: scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research 23 (120), p. 1–39. External Links: Link Cited by: §2.3. F. Fu, S. Fang, W. Chen, and Z. Mao (2024) Sentiment-oriented transformer-based variational autoencoder network for live video commenting. ACM Transactions on Multimedia Computing, Communications, and Applications 20 (4), p. 104:1–104:24. External Links: Document, Link Cited by: §1. R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or (2023) An image is worth one word: personalizing text-to-image generation using textual inversion. In International Conference on Learning Representations, Kigali, Rwanda. External Links: Link Cited by: §2.2. Google (2026) Gemini 3.5 Flash. Note: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flashGemini API documentation. Model ID: gemini-3.5-flash; last updated June 24, 2026; accessed June 25, 2026 Cited by: §4.2. X. Gu, C. Li, X. Wang, D. Hong, L. Zhang, T. Luo, L. Wen, and H. Fan (2026) Structured context learning for generic event boundary detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 4808–4817. Cited by: §1. Y. Gu, X. Wang, J. Z. Wu, Y. Shi, Y. Chen, Z. Fan, W. Xiao, R. Zhao, S. Chang, W. Wu, Y. Ge, Y. Shan, and M. Z. Shou (2023) Mix-of-show: decentralized low-rank adaptation for multi-concept customization of diffusion models. In Advances in Neural Information Processing Systems, New Orleans, LA, USA. Cited by: §2.2. Y. Guo, D. Hong, W. Chen, Z. She, C. Ye, X. Chang, and Z. Mao (2025) EmoVerse: a MLLMs-driven emotion representation dataset for interpretable visual emotion analysis. arXiv preprint arXiv:2511.12554. External Links: Link Cited by: §2.1. J. Han, Q. Wang, L. Zhang, W. Chen, Y. Song, and Z. Mao (2023) Text style transfer with contrastive transfer pattern mining. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada, p. 7914–7927. External Links: Document, Link Cited by: §1. Y. Hao, Z. Chi, L. Dong, and F. Wei (2023) Optimizing prompts for text-to-image generation. In Advances in Neural Information Processing Systems, New Orleans, LA, USA. Cited by: §1. A. Hertz, A. Voynov, S. Fruchter, and D. Cohen-Or (2024) Style aligned image generation via shared attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, p. 4775–4785. Cited by: §2.2. M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter (2017) GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, Vol. 30, Long Beach, California, USA, p. 6626–6637. Cited by: §4.1. J. Ho, A. Jain, and P. Abbeel (2020) Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, Virtual Conference. Cited by: §1. D. Hong, C. Li, L. Wen, X. Wang, and L. Zhang (2021a) Generic event boundary detection challenge at CVPR 2021 technical report: cascaded temporal attention network (CastaNet). arXiv preprint arXiv:2107.00239. Cited by: §1. D. Hong, G. Li, K. Xu, L. Su, and Q. Huang (2021b) Siamese dynamic mask estimation network for fast video object segmentation. In 2020 25th International Conference on Pattern Recognition (ICPR), p. 9476–9482. Cited by: §1. D. Hong, G. Li, B. Zhong, Z. Han, L. Su, and Q. Huang (2022) CRNet: collaborative refinement network for self-supervised video object segmentation. In 2022 IEEE 5th International Conference on Multimedia Information Processing and Retrieval (MIPR), p. 172–177. Cited by: §1. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, Virtual Conference. External Links: Link Cited by: §2.2. X. Hu, H. Chen, Z. Qi, H. Zhang, D. Hong, J. Shao, and X. Wu (2025) DreamPoster: a unified framework for image-conditioned generative poster design. arXiv preprint arXiv:2507.04218. Cited by: §2.2. C. Huang, Q. Liu, B. Y. Lin, T. Pang, C. Du, and M. Lin (2024) LoraHub: efficient cross-task generalization via dynamic LoRA composition. In First Conference on Language Modeling, External Links: Link Cited by: §2.2. X. Huang, W. Chen, B. Hu, and Z. Mao (2025) Graph mixture of experts and memory-augmented routers for multivariate time series anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 17476–17484. External Links: Document, Link Cited by: §2.3. Y. Jin, W. Chen, Y. Tian, Y. Song, C. Yan, and Z. Mao (2024a) Improving radiology report generation with D2D^2-Net: when diffusion meets discriminator. In ICASSP 2024 – 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 2215–2219. External Links: Link Cited by: §1. Y. Jin, W. Chen, Y. Tian, Y. Song, and C. Yan (2024b) Improving radiology report generation with multi-grained abnormality prediction. Neurocomputing 600, p. 128122. External Links: Document, Link Cited by: §1. R. Kosti, J. M. Alvarez, A. Recasens, and A. Lapedriza (2020) Context based emotion recognition using EMOTIC dataset. IEEE Transactions on Pattern Analysis and Machine Intelligence 42 (11), p. 2755–2766. External Links: Document Cited by: §2.1. N. Kumari, B. Zhang, R. Zhang, E. Shechtman, and J. Zhu (2023) Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, p. 1931–1941. Cited by: §2.2. C. Li, X. Wang, D. Hong, Y. Wang, L. Zhang, T. Luo, and L. Wen (2022a) Structured context transformer for generic event boundary detection. arXiv preprint arXiv:2206.02985. Cited by: §1. C. Li, X. Wang, L. Wen, D. Hong, T. Luo, and L. Zhang (2022b) End-to-end compressed video representation learning for generic event boundary detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13967–13976. Cited by: §1. D. Li, Y. Ma, N. Wang, Z. Ye, Z. Cheng, Y. Tang, Y. Zhang, L. Duan, J. Zuo, C. Yang, and M. Tang (2024a) MixLoRA: enhancing large language models fine-tuning with LoRA-based mixture of experts. arXiv preprint arXiv:2404.15159. External Links: Link Cited by: §2.3. G. Li, D. Hong, K. Xu, B. Zhong, L. Su, Z. Han, and Q. Huang (2022c) Self supervised progressive network for high performance video object segmentation. IEEE Transactions on Neural Networks and Learning Systems 35 (6), p. 7671–7684. Cited by: §1. J. Li, Z. Mao, H. Li, W. Chen, and Y. Zhang (2024b) Exploring visual relationships via transformer-based graphs for enhanced image captioning. ACM Transactions on Multimedia Computing, Communications, and Applications 20 (5), p. 133:1–133:23. External Links: Document, Link Cited by: §1. Z. Li, L. Zhang, K. Zhang, W. Chen, Y. Zhang, and Z. Mao (2025) Rethinking pseudo word learning in zero-shot composed image retrieval: from an object-aware perspective. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, p. 833–843. External Links: Document, Link Cited by: §1. Z. Lin, W. Chen, Y. Song, and Y. Zhang (2024) Prompting few-shot multi-hop question generation via comprehending type-aware semantics. In Findings of the Association for Computational Linguistics: NAACL 2024, p. 3730–3740. External Links: Document, Link Cited by: §1. C. Liu, Y. Tian, W. Chen, Y. Song, and Y. Zhang (2024) Bootstrapping large language models for radiology report generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 18635–18643. External Links: Document, Link Cited by: §1. H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. In Advances in Neural Information Processing Systems, New Orleans, LA, USA, p. 34892–34916. External Links: Document Cited by: §1. X. Liu, W. Chen, Z. Qi, B. Zhang, and W. Zhang (2025) Matching street view and satellite images via drone imagery and semantic descriptions. In UAVM 2025 – Proceedings of the 3rd International Workshop on UAVs in Multimedia: Capturing the World from a New Perspective, Co-located with M 2025, p. 4–9. External Links: Document, Link Cited by: §1. Z. Meng, D. Hong, W. Chen, Z. Zhou, B. Hu, and Z. Mao (2026) Audio-visual exchange-aware token pruning for efficient audio-visual captioning. arXiv preprint arXiv:2606.10533. External Links: Link Cited by: §1. W. Peebles and S. Xie (2023) Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, p. 4195–4205. Cited by: §1. X. Qin, D. Hong, W. Chen, C. Ye, X. Liu, P. Song, and L. Zhang (2025) Query-based collaborative multimodal token pruning for audio-visual question answering. In 2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics (AIHCIR), External Links: Document, Link Cited by: §1. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, Virtual Conference, p. 8748–8763. Cited by: §4.1. C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. S. Pinto, D. Keysers, and N. Houlsby (2021) Scaling vision with sparse mixture of experts. Advances in Neural Information Processing Systems 34, p. 8583–8595. Cited by: §2.3. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2022) High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, p. 10684–10695. External Links: Document Cited by: §1. N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman (2023) DreamBooth: fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, p. 22500–22510. Cited by: §2.2. C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. Gontijo-Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi (2022) Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems, New Orleans, LA, USA, p. 36479–36494. Cited by: §1. N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean (2017) Outrageously large neural networks: the sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations, External Links: Link Cited by: §2.3. C. Snell, J. Lee, K. Xu, and A. Kumar (2024) Scaling LLM test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314. External Links: Link Cited by: §1. P. Song, Z. Zhang, W. Chen, J. Hu, X. Yang, and X. Chang (2026) Bridging subjectivity in affective explanation captioning via consensus-prompted emotion reasoning. IEEE Transactions on Image Processing P. Note: Online ahead of print; IEEE Xplore document 11574588 External Links: Document, Link Cited by: §2.1. Y. Tian, W. Chen, B. Hu, Y. Song, and F. Xia (2023) End-to-end aspect-based sentiment analysis with combinatory categorial grammar. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, p. 13597–13609. External Links: Document, Link Cited by: §1. C. Wang, W. Chen, X. Cui, Y. Zhao, Z. Qi, P. Huang, X. Liu, and W. Zhang (2025a) Combatting data imbalance and noise in micro-action recognition. In M 2025 – Proceedings of the 33rd ACM International Conference on Multimedia, Co-Located with M 2025, p. 14229–14235. External Links: Document, Link Cited by: §1. L. Wang, C. Ye, W. Chen, P. Song, B. Hu, and Z. Mao (2026) A multi-agent framework with structured reasoning and reflective refinement for multimodal empathetic response generation. arXiv preprint arXiv:2604.18988. External Links: Link Cited by: §2.1. T. Wang, W. Chen, Y. Tian, Y. Song, and Z. Mao (2023a) Improving image captioning via predicting structured concepts. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 360–370. External Links: Document, Link Cited by: §1. X. Wang, L. Wen, C. Li, and D. Hong (2025b) Feature extraction method and apparatus for video, slicing method and apparatus for video, and electronic device and storage medium. Note: US Patent App. 18/837,577 Cited by: §1. Y. Wang, S. Shen, and B. Y. Lim (2023b) RePrompt: automatic prompt editing to refine ai-generative art towards precise expressions. In Proceedings of the CHI Conference on Human Factors in Computing Systems, Hamburg, Germany, p. 1–29. Cited by: §1. L. Wen, X. Wang, D. Hong, and C. Li (2025) Information segmentation methods, apparatuses, and electronic devices. Note: US Patent 12,481,631 Cited by: §1. X. Wu, S. Huang, and F. Wei (2024) Mixture of LoRA experts. arXiv preprint arXiv:2404.13628. External Links: Link Cited by: §2.3. J. Yang, J. Feng, and H. Huang (2024) EmoGen: emotional image content generation with text-to-image diffusion models. arXiv preprint arXiv:2401.04608. External Links: Link Cited by: §2.1. J. Yang, Q. Huang, T. Ding, D. Lischinski, D. Cohen-Or, and H. Huang (2023) EmoSet: a large-scale visual emotion dataset with rich attributes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 20326–20337. External Links: Document Cited by: §1, §2.1. Y. Yang, W. Wang, L. Peng, C. Song, Y. Chen, H. Li, X. Yang, Q. Lu, D. Cai, B. Wu, and W. Liu (2025) LoRA-composer: leveraging low-rank adaptation for multi-concept customization in training-free diffusion models. IEEE Transactions on Image Processing 34, p. 8145–8158. External Links: Document Cited by: §2.2. Y. Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, et al. (2024) MiniCPM-V: a GPT-4V level MLLM on your phone. arXiv preprint arXiv:2408.01800. External Links: Link Cited by: §4.1. C. Ye, W. Chen, B. Hu, L. Zhang, Y. Zhang, and Z. Mao (2025a) Improving video summarization by exploring the coherence between corresponding captions. IEEE Transactions on Image Processing 34, p. 5369–5384. External Links: Document, Link Cited by: §1. C. Ye, W. Chen, J. Li, L. Zhang, and Z. Mao (2024) Dual-path collaborative generation network for emotional video captioning. In Proceedings of the 32nd ACM International Conference on Multimedia, p. 496–505. External Links: Document, Link Cited by: §2.1. C. Ye, W. Chen, P. Song, X. Liu, L. Zhang, and Z. Mao (2025b) Multi-round mutual emotion-cause pair extraction for emotion-attributed video captioning. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 3320–3329. External Links: Document, Link Cited by: §2.1. Q. You, J. Luo, H. Jin, and J. Yang (2016) Building a large scale dataset for image emotion recognition: the fine print and the benchmark. Proceedings of the AAAI Conference on Artificial Intelligence 30 (1), p. 308–315. External Links: Document Cited by: §2.1. K. Yuan, Y. Zhang, S. Gao, Y. Zhu, W. Chen, and Y. Yue (2025) CoEmoGen: towards semantically-coherent and scalable emotional image content generation. arXiv preprint arXiv:2508.03535. External Links: Link Cited by: §2.1. Z-Image Team, H. Cai, S. Cao, R. Du, P. Gao, A. Hao, S. Hoi, Z. Hou, S. Huang, D. Jiang, Y. Jiang, X. Jin, L. Li, Z. Li, Z. Li, D. Liu, D. Liu, et al. (2025) Z-Image: an efficient image generation foundation model with single-stream diffusion transformer. arXiv preprint arXiv:2511.22699. External Links: Link Cited by: §1. C. Zhang, H. Xie, B. Wen, S. Zuo, R. Zhang, and W. Cheng (2025a) EmoArt: a multidimensional dataset for emotion-aware artistic generation. arXiv preprint arXiv:2506.03652. External Links: Link Cited by: §2.1, §3.1. H. Zhang, D. Hong, Y. Wang, J. Shao, X. Wu, Z. Wu, and Y. Jiang (2025b) CreatiLayout: siamese multimodal diffusion transformer for creative layout-to-image generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 18487–18497. Cited by: §2.2. H. Zhang, D. Hong, M. Yang, Y. Cheng, Z. Zhang, W. Chen, J. Shao, X. Wu, Z. Wu, and Y. Jiang (2026a) CreatiDesign: a unified multi-conditional diffusion transformer for creative graphic design. In International Conference on Learning Representations, Rio de Janeiro, Brazil. External Links: Link Cited by: §2.2. Y. Zhang, N. Huang, F. Tang, H. Huang, C. Ma, W. Dong, and C. Xu (2023) Inversion-based style transfer with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, p. 10146–10156. Cited by: §2.2. Z. Zhang, Y. Cheng, D. Hong, M. Yang, G. Shi, L. Ma, H. Zhang, J. Shao, and X. Wu (2025c) CreatiPoster: towards editable and controllable multi-layer graphic design generation. arXiv preprint arXiv:2506.10890. Cited by: §2.2. Z. Zhang, P. Song, J. Hu, W. Chen, L. Ni, and X. Yang (2026b) Stimuli-aware emotion adaptor for enhancing LLM in affective explanation captioning. In ICASSP 2026 – 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 10662–10666. External Links: Document, Link Cited by: §2.1. B. Zhao, W. Chen, B. Hu, H. Xie, and Z. Mao (2023) Difference-aware iterative reasoning network for key relation detection. In 2023 IEEE International Conference on Multimedia and Expo (ICME), p. 276–281. External Links: Document, Link Cited by: §1. M. Zhong, Y. Shen, S. Wang, Y. Lu, Y. Jiao, S. Ouyang, D. Yu, J. Han, and W. Chen (2024) Multi-LoRA composition for image generation. arXiv preprint arXiv:2402.16843. External Links: Link Cited by: §2.2. Q. Zhou, J. Yao, S. Tang, W. Chen, L. Cheng, and J. Tang (2025) Hierarchical knowledge distillation for cross-lingual stance detection. In 2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics (AIHCIR), p. 1–5. External Links: Document, Link Cited by: §1. Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. Dai, Z. Chen, Q. Le, and J. Laudon (2022) Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems 35, p. 7103–7114. Cited by: §2.3.