Paper deep dive
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
Zhenyu Hu, Qing Wang, Te Cao, Luo Liao, Longfei Lu, Liqun Liu, Shuang Li, Hang Chen, Mengge Xue, Yuan Chen, Chao Deng, Peng Shu, Huan Yu, Jie Jiang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/13/2026, 12:39:52 AM
Summary
DSH-Bench is a comprehensive benchmark for evaluating subject-driven text-to-image (T2I) generation models. It addresses limitations in existing benchmarks by introducing a hierarchical taxonomy of 58 categories, a dual-classification scheme for subject difficulty and prompt scenarios, and a novel Subject Identity Consistency Score (SICS) metric that correlates better with human evaluation. The benchmark evaluates 19 leading models, providing diagnostic insights into subject preservation, prompt following, and image quality.
Entities (5)
Relation Signals (3)
DSH-Bench → utilizes → Subject Identity Consistency Score
confidence 100% · we introduce Subject Identity Consistency Score (SICS), which innovatively focuses on subject-level consistency
DSH-Bench → evaluates → Subject-driven T2I models
confidence 95% · enables systematic multi-perspective analysis of subject-driven T2I models
Qwen2.5-VL-7B → powers → Subject Identity Consistency Score
confidence 90% · we fine-tune Qwen2.5-VL-7B on this dataset... Finally, we use Kendall’s τ value to quantify the alignment
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions. However, evaluating these models remains a significant challenge. Existing benchmarks exhibit critical limitations: 1) insufficient diversity and comprehensiveness in subject images, 2) inadequate granularity in assessing model performance across different subject difficulty levels and prompt scenarios, and 3) a profound lack of actionable insights and diagnostic guidance for subsequent model refinement. To address these limitations, we propose DSH-Bench, a comprehensive benchmark that enables systematic multi-perspective analysis of subject-driven T2I models through four principal innovations: 1) a hierarchical taxonomy sampling mechanism ensuring comprehensive subject representation across 58 fine-grained categories, 2) an innovative classification scheme categorizing both subject difficulty level and prompt scenario for granular capability assessment, 3) a novel Subject Identity Consistency Score (SICS) metric demonstrating a 9.4\% higher correlation with human evaluation compared to existing measures in quantifying subject preservation, and 4) a comprehensive set of diagnostic insights derived from the benchmark, offering critical guidance for optimizing future model training paradigms and data construction strategies. Through an extensive empirical evaluation of 19 leading models, DSH-Bench uncovers previously obscured limitations in current approaches, establishing concrete directions for future research and development.
Tags
Links
- Source: https://arxiv.org/abs/2603.08090v1
- Canonical: https://arxiv.org/abs/2603.08090v1
Trouble viewing inline? Open PDF directly →
Full Text
63,226 characters extracted from source content.
Expand or collapse full text
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation Zhenyu Hu 1⋆ , Qing Wang 1⋆ , Te Cao 1⋆ , Luo Liao 1⋆ , Longfei Lu 1 , Liqun Liu 1† , Shuang Li 1 , Hang Chen 1 , Mengge Xue 1‡ , Yuan Chen 1 , Chao Deng 1 , Peng Shu 1 , Huan Yu 1 , and Jie Jiang 1 Tencent mapleshu, rosaliecao, graceqwang, berryxue, liqunliu@tencent.com Hard Medium Please generate an image based on the reference image and the given prompt. Subject Difficulty Level Classification Prompt Scenario Classification Image Quality Please change the image style to oil painting. 31.2/100 (Good) HPSv2 21.2/100 (Poor) HPSv2 Evaluation Dimension Prompt Following The cat is on the table with a sofa and a pot of flowers in theroom. 33.2/100 (Good) CLIP-TScore 15.7/100 (Poor) CLIP-TScore Subject Preservation Theperfume is outdoors. It is dusk now and there is a beautiful sunset glow. 4/5 (Good) SICSScore 1/5 (Poor) SICSScore Variation in subject viewpoint or size A wide-angle shot of a cat basking in the sun, captured with a high- angle perspective, surrounded by scattered autumnleaves. Attribute change A cat with sleek black fur.The background is basically the same as the original picture. Style change A watercolor painting of a cat sleeping amidst soft pastel tones and diffuse edges that blend seamlessly. Imagination A cat floating weightlessly in space, wearing a tiny astronaut helmet and pawing at sparkling stars nearby. Interaction with other entities A cat playing with a curious puppy in a garden, their movements creating dynamic, playful action. Background change A cat lounging peacefully on a grassy meadow, surrounded by wildflowers under a blue sky. Fig. 1: Overview of DSH-Bench. We curate a diverse dataset of subject images and categorize them into three difficulty levels—easy, medium, and hard—based on the complexity of preserving subject details. Leveraging GPT-4o’s capabilities, we systematically generate contextually appropriate prompts for various scenarios. The generated images are then rigorously evaluated across three key dimensions: Subject Preservation, Prompt Following, and Image Quality. Abstract.Significant progress has been achieved in subject-driven text- to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions. However, evaluating these models remains a significant challenge. Existing benchmarks exhibit criti- cal limitations: 1) insufficient diversity and comprehensiveness in subject images, 2) inadequate granularity in assessing model performance across ⋆ Equal contribution. † Corresponding author. ‡ Project leader. arXiv:2603.08090v1 [cs.CV] 9 Mar 2026 2Z. Hu et al. different subject difficulty levels and prompt scenarios, and 3) a pro- found lack of actionable insights and diagnostic guidance for subsequent model refinement. To address these limitations, we propose DSH-Bench, a comprehensive benchmark that enables systematic multi-perspective analysis of subject-driven T2I models through four principal innovations: 1) a hierarchical taxonomy sampling mechanism ensuring comprehensive subject representation across 58 fine-grained categories, 2) an innovative classification scheme categorizing both subject difficulty level and prompt scenario for granular capability assessment, 3) a novel Subject Identity Consistency Score (SICS) metric demonstrating a 9.4% higher correlation with human evaluation compared to existing measures in quantifying subject preservation, and 4) a comprehensive set of diagnostic insights derived from the benchmark, offering critical guidance for optimizing fu- ture model training paradigms and data construction strategies. Through an extensive empirical evaluation of 19 leading models, DSH-Bench un- covers previously obscured limitations in current approaches, establishing concrete directions for future research and development. Keywords: Single Subject-driven· Text-to-image Generation· Dataset and Benchmark· Subject Identity Consistency 1 Introduction Subject-driven text-to-image (T2I) generation, which synthesizes novel scenes conditioned on specific reference images and textual prompts, has emerged as a pivotal research frontier, propelled by rapid advancements in large-scale T2I diffusion models [3,6,10,11,14,27,51,55]. In subject-driven T2I generation, aside from image quality, two other fundamental criteria must be satisfied: Subject Preservation and Prompt Following. Subject Preservation requires that the generated image maintain the details of the reference subject. Prompt Following demands that the generated image consistently reflects the content in the prompt. For example, a user might request an image of "his dog traveling around the world". In this scenario, the generated image must depict a dog identical to the reference image while illustrating the act of traveling as described. Although significant progress has been made in subject-driven T2I generation in recent years [15, 17, 23, 30, 32, 49, 54, 62, 66, 73], the community still lacks a standardized protocol to comprehensively and effectively evaluate their true capa- bilities. An optimal benchmark should not only ensure unbiased, human-aligned evaluation but also provide granular diagnostics to guide research. However, current benchmarks [7, 30, 45, 54, 63] are limited by insufficient diversity and comprehensiveness in subject image collection, which restricts the thoroughness of model evaluation. Crucially, they fail to disentangle the inherent difficulty of the reference subjects from the complexity of the prompt scenarios. As shown in Fig. 2, our empirical analysis reveals a significant performance variance contingent on these factors: models that effortlessly reconstruct simple geometries (e.g., a tennis ball) frequently fail to preserve the intricate structural details of complex artifacts (e.g., a camera). This observation highlights a critical blind spot in DSH-Bench: A comprehensive benchmark for Subject-Driven T2I3 current evaluations: the necessity of stratifying test cases by the subject difficulty and prompt scenario. Moreover, subject-driven T2I generation can be broadly categorized into single-subject and multi-subject paradigms depending on the number of reference subjects. While the single-subject task can be conceptually considered a special case of the multi-subject scenario, the two paradigms impose fundamentally different demands on model capabilities. Regarding the scope of evaluation, we posit that the single-subject paradigm is the cornerstone of subject-driven generation. While multi-subject generation introduces interactional complexities, empirical evidence suggests that models still struggle to achieve saturated performance even in single-subject tasks. Consequently, establishing a rigorous benchmark for the single-subject scenario is a prerequisite for any reliable multi-subject assessment. To address the aforementioned limitations, we introduce DSH-Bench, a novel and comprehensive benchmark designed to provide multi-dimensional evaluations. DSH-Bench offers three distinct advantages: 1.The diversity of subject images in DSH-Bench is substantially greater To effectively mitigate potential evaluation bias stemming from lim- ited subject diversity, we construct a rigorous hierarchical taxonomy for image collection. Specifically, we incorporate established ontologies from COCO [34] and ImageNet [9] into our taxonomy design. As illustrated in Fig. 3(a), the widely utilized DreamBench comprises merely 6 categories and 30 subjects. In sharp contrast, our benchmark significantly scales up the dataset to 58 distinct categories and 459 unique subjects—representing a substantial increase of 8×and 15×, respectively. While DreamBench++ [45] offers 150 subjects, its diversity remains constrained by a limited collection scope. Notably, 33% of our cate- gories are entirely absent from the DreamBench++ distribution. Consequently, benefiting from DSH-Bench’s superior subject diversity, we facilitate a far more comprehensive and robust evaluation of generative models. 2. An innovative classification scheme for subject difficulty and prompt scenarios As illustrated in Fig. 2, model performance exhibits sig- nificant variance across different samples, underscoring the critical necessity for a dual classification of both subject images and prompts. While DreamBench++ attempts to categorize prompts based on perceived difficulty, the underlying criteria remain ambiguous and lack a systematic definition. Furthermore, it neglects to analyze the inherent difficulty levels associated with distinct subjects. To address these limitations, we systematically stratify subjects into three diffi- culty tiers (easy, medium, hard) based on the complexity of preserving visual appearance, and classify prompts into six distinct scenarios (background change, variation in subject viewpoint or size, interaction with other entities, attribute change, style transfer, and imagination). Thus, our approach enables a more comprehensive and granular diagnosis of the challenges faced by current models. 3. A human-aligned and more efficient metric for subject preserva- tion DreamBench++ replaces CLIP [50] and DINO [5] with GPT-4o [41] for evaluation, resulting in improved alignment with human evaluation. However, our benchmark reveals that per-model evaluation under this paradigm requires 4Z. Hu et al. Style change Variation in subject viewpoint or size 0 50 100 150 200 250 EasyMediumHard Difficulty Distribution of Subject DreambenchCustomConcept101Dreambench++Ours 0 200 400 600 800 1000 Background change Variation in subject viewpoint or size Interaction with other entities Attribute changeStyle changeImagination Scenario Distribution of Prompt DreambenchCustomConcept101Dreambench++Ours Interaction with other entities Easy Medium Hard [Prompt] A on a table indoors with a sofa and a bouquet of flowers in the background Attribute change Imagination Background change Reference subject A cat lounging peacefully on a grassy meadow, surrounded by wildflowers under a clear blue sky. A wide-angle shot of a cat basking in the sun, captured with a high- angle perspective, surrounded by scattered autumn leaves. A cat playing with a curious puppy in a garden, their movements creating dynamic, playful action. A cat with sleek black fur.The background is basically the same as the original picture. A watercolor painting of a cat sleeping amidst soft pastel tones and diffuse edges that blend seamlessly. A cat floating weightlessly in space, wearing a tiny astronaut helmet and pawing at sparkling stars nearby. Fig. 2: Qualitative comparison under different difficulty levels and scenarios. approximately 20,000 API calls to GPT-4o, incurring prohibitive computational costs exceeding $400 for each evaluation. To address the limitation, we introduce Subject Identity Consistency Score (SICS), which innovatively focuses on subject-level consistency rather than merely relying on embedding comparisons. Firstly, five annotators label a training dataset containing 5,000 image-text pairs, focusing on subject preservation evaluation. We then fine-tune Qwen2.5-VL- 7B [2] on this dataset, which leads the model to focus on core visual attributes rather than high-level semantics. Finally, we use Kendall’sτvalue to quantify the alignment between model outputs and human evaluation. Experimental results demonstrate that SICS achieves a statistically significant improvement, outperforming Dreambench++ by 9.4% in human evaluation correlation metrics. Takeaways We present some insightful findings from evaluating fifteen meth- ods: i) Our evaluation reveals that no single method demonstrates consistently robust performance across all categories. Therefore, implementing hierarchical taxonomy sampling of subject images is critical for mitigating potential eval- uation biases. i) All methods exhibit degraded performance on hard subject images. It is crucial to enhance models’ ability to encode and reconstruct complex subject details more effectively in future research. i) The subject-driven T2I model’s capability for different prompt scenarios is not robust. Future research on subject-driven T2I generation should focus on optimizing for adaptation to a variety of prompt scenarios. In summary, our contributions are as follows: 1) We employ a hierarchical taxonomy in image collection to ensure both the diversity and comprehensiveness of subject images. 2) We propose an innovative classification scheme to categorize subject difficulty levels and prompt scenarios. This scheme enables us to obtain valuable insights. 3) We propose a human-aligned metric to evaluate subject preservation, which offers greater efficiency compared to DreamBench++. We are open-sourcing DSH-Bench, including subject images, prompts and code. DSH-Bench: A comprehensive benchmark for Subject-Driven T2I5 (b) t-SNE Visualization of DreamBench++ vs. DSH-Bench (a) Distribution of SubjectImages Across DifferentCategories (Under Photorealistic Category) 0 5 10 15 20 25 30 35 40 VehicleMusical Instrument Public Facility Food and Beverage Medical Supply BookFurnitureHome Appliance AmphibianBuildingDigital Product InsectStationeryDaily Necessity PlantJewelryBeauty and Skincare ArtworkClothingSports Equipment Shoe, Bag and Accessory ToyMammalReptileBirdFishHalf-body or Full-body Photo Facial Close-up Artistic Image and Celebrity Subject Category Distribution DreamBenchCustomConcept101DreamBench++DSH-Bench DreamBench++ DSH-Bench Fig. 3: Distribution of subject images. (a) Category-wise image distribution for our benchmark versus prior benchmarks. (b) t-SNE comparison of images between DSH-Bench and DreamBench++. 2 Related Work 2.1 Subject-Driven Text-to-Image Generation In recent years, subject-driven T2I generation has attracted significant re- search attention [15–17, 23, 30, 32, 49, 54, 62, 66]. Within the context of diffu- sion models, optimization-based model [20, 24, 36, 61] enables subject-driven generation by introducing lightweight parameters and performs parameter- efficient fine-tuning for each subject. In contrast, the encoder-based meth- ods [8, 21, 22, 25, 26, 31, 33, 37, 42, 53, 56, 67, 70, 74] leverage additional image encoders and network layers to encode the reference image of the subject. IP- Adapter [73] introduces cross-attention through an additional image encoder to incorporate control signals. Furthermore, SSR-Encoder [76] enhances identity preservation without necessitating further fine-tuning when introducing new concepts. The Diffusion Transformers [44,47,52] uses transformer as a denoising network to iteratively refine noisy image tokens. Based on these foundation models, approaches like OminiControl [59] and UNO [67] explore the inherent image reference capabilities of transformers. 2.2 Subject-Driven T2I Generation Benchmark Evaluation for subject-driven T2I generation involves a variety of metrics focusing on different aspects. For image quality, several notable studies [1,29,64,68,72] have conducted. For subject preservation evaluation, learning-based metrics [12,48,75] compute the distances between image features extracted by deep neural net- works. Image embeddings from large vision models like CLIP [50], DINO [5] and image-retrieval score [35] have been utilized. To better match human per- ception, DreamSim [13] focuses on foreground objects when evaluating image similarity. In terms of semantic consistency, the CLIP score is frequently used. Dreambench is the first benchmark established for subject-driven T2I tasks. However, Dreambench is limited in the diversity of subjects and prompt scenarios. DreamBench-v2 [7] extends the evaluation by introducing 220 novel prompts. CustomConcept101 [30] lacks of non-photorealistic subject images. Although 6Z. Hu et al. DreamBench++ increases to 150 subject images, it can not provide a system- atic categorization of subjects and prompts, which makes it difficult to derive meaningful insights from the evaluation results. To bridge this gap, we intro- duce DSH-Bench, which enables a more comprehensive evaluation of models for subject-driven T2I generation and provids deeper insights. 3 DSH-Bench This section provides an overview of the primary components of DSH-Bench. Sec. 3.1 outlines the data construction process. Sec. 3.2 introduces the definitions and evaluation methods for three evaluation dimensions. A detailed explanation is available in the supplementary materials. Candidate category labels mining Wikipedia ImageNet COCO car cat book tableprinter printer bowl vase oven cabinet guitar pillow sofa scissors ...... ...... Photorealistic Root Non-photorealistic Person Object Animal Animal ObjectPerson Toys ClothingFurniture Vehicle Establish Hierarchical Category GPT-4o associate basedon categories. Human propose based on categories. ukuleledouble bedsofascissors ...... cat cap dinning table pen Collect Keywords Collect Images Unsplash Pinterest ... ... Search keywords Step1. Subject Image CollectionStep2. Subject Image Processing Filter Images Multiple subjects Low image quality Inappropriate proportions Goodfor generation Center the subject by cropping. Step3. Prompt Generation ...... Easy MediumHard ...... ...... ...... Classify Subject Difficulty Please assign appropriate labels to the images based on the complexity of the details contained within the subject. Human Annotators GPT-4o [TaskDescription]Inthegivenimage,theprimarysubjectisacat. Pleasegeneratepromptsforthissubject,addressingvarious dimensionsasspecifiedbythefollowingrequirements.Two promptsaregeneratedforeachdimension. (1) Background change: Only change the background environment while keeping the subject's default attributes. (2)Variationinsubjectviewpointorsize:Thesubject'slocation, perspective,lightingconditionscouldbeadjustedproperly.The imagemaybeshotinwide-angle,telephoto,bird's-eyeview,low- angleshot,high-angleshot,eye-level,close-up,mediumshot, longshotandsoon. (3) Interaction withother entities: Create interactions between the subject and other entities, or add effects like obstruction, reflection. (4) Attribute change: Alter the attributes of the subject, such as color, shape, material, or appearance, and so on. (5) Style change: Modify the style of the scene, including different art movements or artistic forms. (6) Imagination: Imagine some imaginative and unrealistic scenario for this subject. (1)Backgroundchange:Acat loungingpeacefullyonagrassy meadow,surroundedby wildflowersunderabluesky. (3)Interactionwithotherentities: Acatplayingwithacuriouspuppy inagarden,theirmovements creatingdynamic,playfulaction. (5)Stylechange:Awatercolor paintingofacatsleepingamidst softpasteltonesanddiffuseedges thatblendseamlessly. (2)Variationinsubject viewpointorsize:Awide-angle shotofacatbaskinginthesun, withahigh-angleperspective, surroundedbyautumnleaves. (4)Attributechange:Acatwith sleekblackfur.Thebackgroundis basicallythesameastheoriginal picture (6)Imagination:Acatfloating weightlesslyinspace,wearinga tinyastronauthelmetandpawing atsparklingstarsnearby. GPT-4o Generation Human Inspection Goodfor generation Human assessment Automatic assessment Fig. 4: Dataset construction process of DSH-Bench. We construct a hierarchical taxonomy to obtain a comprehensive set of keywords. Then we collect web images using these keywords. After performing both manual review and automated filtering of the images, we classify the difficulty of subject images and use GPT-4o to generate prompts for each subject image. 3.1 Benchmark Dataset Construction Subject Image Collection 1) Hierarchical Taxonomy Establishment As illustrated in Fig. 4, we construct a rigorous three-tier taxonomy. The first level distinguishes between Photorealistic and Non-photorealistic domains, with strictly aligned subcategories to ensure cross-domain consistency. For the second level, we refine the coarse "Living" category from prior benchmarks [45,54] into distinct Humans and Animals, acknowledging the unique demand for facial fidelity in human subjects versus the high variance in animals. We explicitly exclude abstract "Style" categories [30] to focus strictly on tangible entity customization. DSH-Bench: A comprehensive benchmark for Subject-Driven T2I7 At the third level, to balance granularity with generality, we compiled candidate labels from COCO and ImageNet, utilizing GPT-4o to consolidate them into 58 fine-grained categories. Notably, we stratify the Human category into celebrities, facial close-ups, and full-body shots, allowing us to disentangle foundation model bias (e.g., celebrity overfitting) from structural reconstruction capabilities. 2) Keyword Collection & Internet Image Collection In DreamBench++, keywords collection relies on GPT-4o and human input. The approach does not adequately ensure the diversity of the obtained keywords, potentially introducing bias during the image collection process. In contrast, DSH-Bench derives keywords from a hierarchical taxonomy. For each third-level category, we use GPT-4o to generate associated keywords, which are further supplemented by humans. All keywords are then consolidated and deduplicated, resulting in a final set of 400 unique keywords—surpassing DreamBench++’s 300. The specific keywords are provided in the Supplementary material (Sec. B). Given a set of selected keywords, we retrieve images from Unsplash [60] and Pinterest [46]. Each image’s copyright status has been verified for academic suitability. Subject Image Processing 1) Image Filtering To filter unsuitable images, we use aesthetic score [71] and SAM [28] to filter images with low image quality and inappropriate proportions of subject regions. The curated images are subsequently cropped to centralize the reference subject. 2) Subject Difficulty Level Classification As illustrated in Fig. 2, the model’s performance varies considerably across different samples. To derive meaningful insights, we classify the subject images according to the difficulty level that the model experiences in preserving details of the reference subject. We define three subject difficulty levels, including (1) Easy: Subjects characterized by minimal surface complexity and homogeneous textural properties, exemplified by smooth-surfaced objects such as a ceramic mug with uniform coloration. These cases present negligible challenges for detail preservation due to their structural regularity. (2) Medium: Subjects containing discernible high-frequency features while maintaining global structural coherence, such as cylindrical containers with legible typographic elements. These cases require intermediate detail preservation capabilities. (3) Hard: Subjects exhibiting non-uniform texture distributions and multi-scale geometric details, typified by objects like book covers containing fine-grained calligraphic elements. Such cases expose model limitations in maintaining structural fidelity and textural granularity under complex topological constraints. We utilize GPT-4o to classify the subject images according to the aforementioned criteria. Subsequently, all images are reviewed by five human annotators to ensure accuracy and consistency. As these images are curated based on a meticulously constructed taxonomy, they possess the comprehensiveness and representativeness required for rigor- ous evaluation. Given the inference efficiency of current generative models, an overabundance of redundant images would substantially inflate the evaluation overhead. Ultimately, we obtain a total of 459 high-quality images. 8Z. Hu et al. Prompt Generation Although DreamBench++ categorizes prompts based on their perceived difficulty, it does not provide empirical evidence to substantiate the criterion. To address this limitation, we organize the prompts according to specific application scenarios, dividing them into six categories, including (1) Background change (BC): scenarios involving changes in background elements. (2) Variation in subject viewpoint or size (VS): scenarios that entail changes in camera angle, which may include variations in subject size, lighting, or shadows. (3) Interaction with other entities (IE): scenarios requiring complex interactions with additional entities, potentially resulting in occlusion and necessitating adherence to physical plausibility. (4) Attribute change (AC): scenarios involving modifications to certain attributes of the subject, such as color or shape. (5) Style change (SC): scenarios involving alterations in the artistic or visual style of the subject. (6) Imagination (IM): scenarios where the target image depicts an imagined or fictional scene. We generate two prompts for each scenario. The specific instructions are depicted in Fig. 4. All prompts are reviewed by five human annotators to ensure they are ethical and free from defects. For the specific verification procedure, please refer to the Supplementary material (Sec. E.3). Finally, we obtain a total of 5,508 prompts. Fig. 2 shows the distribution of subject difficulty levels and prompt scenarios. We visualize the t-SNE of images from our benchmark and DreamBench++ in Fig. 3, which indicate DSH-Bench achieves superior diversity. Human annotation Supervised Fine-Tuning Image Pairs Dataset EvaluationInstruction As an experienced evaluator, you are given two images. The first image is the reference image. Your tasks are as follows: 1.[Focus only on the main subject of the reference image]When comparing, consider only the main subject in the first (reference) image. In the second image, look for the corresponding main subject and compare it to the reference. Ignore the background and any other objects or elements, including those interacting with the main subject. 2. [Criteria for comparison]Compare the main subject in the second image to the main subject in the reference image based on the following criteria: -Shape and structure. -Color and texture-Size and proportion-Distinctive features or markings 3. [Similarity score]Based on your analysis, assign a similarity score to the main subject in the second image compared to the main subject in thereference image according to the following scale: -Completely dissimilar(0): The main subject in the second image is entirely different from the reference, with no obvious similarities. -Very low similarity(1):The main subject in the second image has some similar features to the reference, but overall they are still very different. -Low similarity(2): The main subject in the second image has several similar parts to the reference, but they are clearly not alike as a whole. -Moderate similarity(3):The main subject in the second image can be recognized as the same subject as the reference, but there are differences that are relatively easy to spot. -High similarity(4):The main subject in the second image is very similar to the reference overall, but there are only subtle differences that require close inspection to notice. -Identical(5): The main subject in the second image is essentially the same as the reference, with no noticeable differences. 4. [Explanation:]Provide a detailed explanation describing the differences between the subject in the second image and the subject in the reference image. Focus only on the subjects and not the background or any other objects. SICS Score:2;Explanation:The vehicles share the same classic van style but differ significantly in color, length, and additional features. The left image shows a yellow van with a shorter design, while the right image showcases an orange van with an extended roof and roof rack. The license plates are also different, indicating they are distinct vehicles. [Input]Reference image at left. Generated image at right. Provide scoring results regarding the above instructions. Evaluation examples SICS Train Process Fig. 5: The training process of SICS. We constructed and annotated a dataset specifically tailored for subject consistency determination, and subsequently trained models using this dataset. 3.2 Evaluation Dimension Previous notable works [15,30,54,62] evaluate the performance of subject-driven T2I models from two perspectives: Subject Preservation and Prompt Follow- DSH-Bench: A comprehensive benchmark for Subject-Driven T2I9 VE: Vehicle TO: Toy AR: Artwork PUF: Public Facility FB: Food and Beverage MI: Musical Instrument MS: Medical Supply BO: Book FUR: Furniture HA: Home Appliance SE: Sports Equipment AM: Amphibian BU: Building CL: Clothing DP: Digital Product IN: Insect ST: Stationery PL: Plant JE: Jewelry DN: Daily Necessity SBA: Shoe, Bag, and Accessory BS: Beauty and Skincare HFB: Half or Full Body FC: Facial Close-up AC: Artistic and Celebrity BI: Bird FI: Fish MA: Mammal RE: Reptile Animal Object Human Photorealistic (Non-photorealistic) Fig. 6: Category hierarchy of the dataset. The top-level categories Photorealistic and Non-photorealistic share an identical set of sub-categories. ing. RealCustom++ [39] also uses ImageReward [72] to evaluate image quality. Therefore, DSH-Bench evaluates from the three aforementioned dimensions. Subject Preservation DreamBench++ utilizes GPT-4o for evaluation to improve alignment with human assessments. However, the GPT-4o-based method is prohibitively expensive. To address this limitation, we propose a novel metric— Subject Identity Consistency Score (SICS). As shown in Fig. 5, we establish a scoring criterion for assessing subject preservation firstly. Five annotators label the collected image pairs according to the criterion. During the annotation process, each image pair is not only assigned a score but also accompanied by an explanation. Previous work [65] has indicated that labeled data with explanatory reasoning can help models better understand the underlying logic and reasoning behind the labels. We then perform meticulous fine-tuning of the model using this annotated dataset. During fine-tuning, SICS leverages prompts to explicitly prioritize subject consistency rather than global semantics, mitigating background and style artifacts that commonly bias CLIP-based approaches and yielding closer alignment with the goals of subject-consistency evaluation. Although GPT-4o demonstrates outstanding performance across a wide range of tasks, it has not been specifically optimized for subject preservation evaluation. More details of the SICS metric can be found in Supplementary material (Sec. E.2). Prompt Following Prompt following primarily evaluates whether a model can generate images that accurately correspond to textual prompts. Dream- Bench++ has demonstrated that the CLIP-T score is highly consistent with human annotations. Therefore, we also adopt CLIP-T score as the evaluation metric for prompt following. Image Quality HPSv2 [68] utilizes professionally annotated data to more accurately reflect human aesthetic preferences for generated images. Previous studies [58] demonstrate that models optimized with HPSv2 achieve superior per- formance in image quality assessment compared to existing approaches. Therefore, we adopt HPSv2 for image quality evaluation. 10Z. Hu et al. 4 Experiment 4.1 Experiment Setup Implementation Details We conducted experiments on the following 19 models in total: 1) Textual Inversion(TI) [16], 2) DreamBooth, 3) Custom Diffusion, 4) Hiper [19], 5) NeTI [1], 6) BLIP-Diffusion [32], 7) IP-Adapter [73], 8) MS-Diffusion [63], 9) Emu2 [57], 10) OminiControl [59], 11) SSR-Encoder [76], 12) RealCustom++ [39], 13) OmniGen [69], 14)λ-Eclipse [43], 15) UNO [67], 16) ACE++ [38], 17) DreamO [40], 18) FLUX.1 Kontext [dev] [4], 19) Nano- Banana [18]. Our experiments are conducted using the official implementations to guarantee reliability and fairness. To demonstrate that our benchmark remains challenging for advanced closed-source models, we conducted experiments using Nano-Banana. More details can be found in supplementary material (Sec. E). Dataset Our benchmark ultimately comprises 459 subject images and 5,508 prompts. These images are distributed across 58 distinct categories. Detailed information regarding category distribution is provided in Fig. 6. Human Annotation All annotation tasks, including labeling of the SICS training datasets, were conducted by the same five human annotators. We provide the annotators with detailed labeling guidelines and sufficient training to ensure they fully understand the subject-driven T2I generation task and could provide unbiased and discriminative scores. For additional details regarding the human annotation process, please see the supplementary material (Sec. E.4). Table 1: The human alignment degree among different metrics, measured by Kendall’s τvalue and Spearman correlation coefficient value. H: Human, G: GPT-4o, D: DINO, Dv2: DINOv2, CB: CLIP-B, CL: CLIP-L, S: SICS. Bold font is used to denote the maximum value in a row. Method Kendall↑Spearman↑ H-CB H-CL H-D H-Dv2 H-GH-SH-CB H-CL H-D H-Dv2 H-GH-S BLIP-Diffusion0.228 0.176 0.285 0.1670.354 0.5310.285 0.215 0.350 0.2060.383 0.554 IP-Adapter0.294 0.296 0.258 0.2900.419 0.6220.364 0.371 0.325 0.3640.459 0.657 MS-Diffusion0.158 0.090 0.116 0.122 0.119 0.1780.194 0.109 0.144 0.1560.1310.189 OminiControl0.375 0.371 0.337 0.3480.650 0.7130.490 0.486 0.441 0.4530.729 0.764 SSR-Encoder0.264 0.338 0.295 0.3480.504 0.6640.328 0.421 0.368 0.4340.549 0.697 ACE++ 0.341 0.325 0.312 0.3120.425 0.5240.298 0.352 0.3230.4150.313 0.492 DreamO0.243 0.311 0.219 0.2430.321 0.3610.240 0.291 0.291 0.3110.322 0.383 FLUX.1 Kontext [dev]0.346 0.313 0.269 0.3400.413 0.4650.371 0.3940.430 0.3610.297 0.496 UNO0.249 0.2180.299 0.240 0.236 0.3850.340 0.2970.390 0.3120.268 0.426 RealCustom++0.181 0.128 0.206 0.2410.291 0.4640.229 0.162 0.266 0.3030.325 0.511 OmniGen0.465 0.396 0.344 0.3490.617 0.6210.579 0.497 0.440 0.456 0.6970.667 λ-Eclipse0.143 0.233 0.084 0.1030.325 0.3750.176 0.287 0.103 0.1270.352 0.393 Custom Diffusion 0.316 0.336 0.382 0.4250.487 0.6420.388 0.409 0.4700.5190.512 0.654 DreamBooth0.639 0.591 0.537 0.429 0.647 0.6920.733 0.721 0.661 0.5370.705 0.740 Textual Inversion0.482 0.459 0.447 0.4380.541 0.5680.587 0.559 0.545 0.5340.582 0.590 HiPer0.338 0.387 0.351 0.4040.584 0.6250.417 0.469 0.430 0.4960.629 0.655 NeTI0.469 0.456 0.431 0.4170.617 0.7280.573 0.561 0.529 0.5120.682 0.778 ALL0.416 0.411 0.350 0.3760.619 0.6770.529 0.522 0.451 0.4830.697 0.734 DSH-Bench: A comprehensive benchmark for Subject-Driven T2I11 Table 2: Evaluation of Subject-driven T2I generation on open-source models. DB: DreamBench, DB++: DreamBench++, HB: DSH-Bench. All scores are normalized to 0-1. Bold indicates the minimum value in each row for a given evaluation dimension. Experiments show that DSH-Bench is more difficult. Method Subject PreservationPrompt FollowingImage Quality DBDB++HBDBDB++HBDBDB++HB BLIP-Diffusion0.2290.216 0.2040.2910.278 0.2770.2670.254 0.223 IP-Adapter0.2300.244 0.2290.3210.318 0.3150.2910.296 0.266 MS-Diffusion0.3160.3460.3520.3320.3390.3380.3110.314 0.294 OminiControl 0.2790.268 0.2580.3250.3370.3340.3120.308 0.290 SSR-Encoder0.231 0.202 0.2020.290 0.2870.2950.2730.270 0.247 ACE++ 0.3240.303 0.2920.3210.316 0.3040.2940.303 0.252 DreamO0.4120.396 0.3910.3240.339 0.3260.3140.308 0.283 FLUX.1 Kontext [dev]0.4450.432 0.4240.3210.324 0.3190.2730.270 0.288 UNO0.4090.410 0.4090.3170.3220.3230.3040.297 0.278 Emu20.3600.343 0.3410.2910.3090.3040.2720.278 0.260 RealCustom++0.3770.380 0.3750.3250.3290.3320.3160.314 0.298 Table 3: DSH-Bench leaderboard. The models are ranked by the final scoreS h . MethodT2I Model SubjectPromptImage S h ↑ PreservationFollowingQuality Nano-Banana-0.4390.3370.3020.272 FLUX.1 Kontext [dev]FLUX.1 Kontext0.4240.3190.2880.256 UNOFLUX.1-dev0.4090.3230.2780.252 DreamOFLUX.1-dev0.3910.3260.2830.251 RealCustom++SDXL0.3750.3320.2940.251 MS-DiffusionSDXL0.3520.3380.2940.248 Emu2SDXL0.3410.3040.2600.228 OminiControlFLUX.1-schnell0.2580.3340.2900.218 ACE++FLUX.1-dev0.2920.3040.2520.214 IP-AdapterSDXL0.2560.2920.2660.199 λ-EclipseSDXL0.2290.3150.2420.198 OmniGenSD v1.50.2020.2950.2650.183 SSR-EncoderSDXL0.1880.3220.2470.181 NeTISD v1.40.1920.3010.2340.176 BLIP-DiffusionSD v1.50.2040.2770.2230.174 DreamBoothSD v1.50.1580.3210.2450.164 HiPerSD v1.40.1350.3180.2470.151 Textual InversionSD v1.50.1090.2990.2250.129 Custom DiffusionSD v1.40.0620.3230.2400.091 4.2 Main Results SICS Results Tab. 1 presents a rigorous study of human alignment using Kendall’sτvalue (KDV) and Spearman correlation coefficient value (SCV) (met- ric selection rationale in supplementary material (Sec. E.2)). Our experimental results demonstrate that SICS achieves superior alignment with human evaluations compared to existing methods, showing consistently higher agreement across both correlation metrics in most experimental settings. Although SICS attains second-highest correlation scores in MS-Diffusion and OmniGen, it significantly outperforms GPT-4o (GPT-4o refers to the evaluation method used in DreamBench++) by 9.37% (KDV) and 5.31% (SCV). This performance gap strongly suggests SICS’s enhanced capability in modeling human evaluation. Notably, GPT-4o exhibits greater consistency with human evaluation than CLIP 12Z. Hu et al. and DINO, aligning with DreamBench++ findings. Importantly, our proposed SICS metric surpasses all existing metrics in human judgment consistency. Quantitative & Qualitative Results Tab. 2 shows overall evaluation results. The results show that: i) DSH-Bench poses more significant chal- lenges than existing benchmarks. For subject preservation and image quality, the majority of methods consistently yield lower scores on DSH-Bench. The result can be attributed to the hierarchical taxonomy sampling method employed, which allows our dataset to more accurately represent the true data distribution. Moreover, it highlights that benchmarks derived from true distributions present greater challenges. i) For prompt following, DreamBench yields slightly lower scores than DSH-Bench for certain methods. In DreamBench, prompts requiring attribute change constitute 22.7%, which is higher than the 16.7% observed in DSH-Bench. Fig. 8b indicates that all methods exhibit relatively poor average performance on prompts involving attribute change. i) Tab. 3 shows that there exists a trade-off between subject preservation and prompt following. We plot the Pareto frontier (see in supplementary material (Sec. D.1)) using the data presented in Tab. 3. The primary objective is to identify a Pareto optimal solution that effectively balances the two objectives. Additional results and discussions are in supplementary material (Sec. D.2). Leaderboard In order to assess a model’s overall capability, we define the final score as: S h = 3 λ SP + γ PF + μ IQ (1) SP, PF, and IQ represent the scores for Subject Preservation, Prompt Following, and Image Quality, respectively.λ,γ,μare the weights assigned to the importance of each corresponding dimension. In this study, we setλ= 1.5,γ= 1.5,μ= 1, as subject preservation and prompt following are of paramount importance in subject-driven T2I generation. The harmonic mean requires strong performance across all dimensions to yield a high overall score. We rank models byS h scores in Tab. 3. Nano-Banana exhibits relatively strong overall performance. Among open-source models, FLUX.1 Kontext [dev] performs best. 5 Analysis In this section, we conduct a detailed analysis of the performance of all methods based on the hierarchical category, subject difficulty level, and prompt scenario: A scientific and comprehensive subject image sampling method is necessary Fig. 8c present the performance of various methods in the third-level categories. The results reveal that model robustness varies considerably among categories. For example, performance in category "Book" (both photorealistic and non-photorealistic) is substantially lower. This disparity suggests that the absence of subject images from specific categories can lead to biased evaluation results, highlighting the importance of data diversity. Furthermore, Fig. 8c also demonstrates that none of the current models perform well across all categories. We hypothesize that this may be related to the varying complexity of the subjects DSH-Bench: A comprehensive benchmark for Subject-Driven T2I13 BLIP- Diffusion IP-AdapterMS-DiffusionOminiControlSSR-Encoder Nano- Banana Emu2RealCustom++OmniGen흀−EclipseHiPerNeTI Custom Diffusion DreamBooth Textual Inversion Easy Hard Thegreenvelvetsofa situatedoutdoorsina gardensetting surroundedbyblooming flowersandlush greenery,maintaining thesamesofa attributes Apillowrestingagainst awoodencabinwall, surroundedbywarm, earthytones. Aneye-levelshotofthe kittensittingatthe edgeofapond, surroundedbyautumn leavesandgently ripplingwater reflectingthevibrant orangesandyellowsof thetrees. Reference Image Aplainbeigebowl placedonawooden tablewithascenic countrysideviewinthe background. Ayellowalarmclock photographedfroma bird's-eyeview,placed onamessyworkdesk filledwithscattered papers,pens,anda coffeecup. Medium Abooktitled'ABOOK FULLOFHOPE'layingon asandybeachwithsoft wavesvisibleinthe background. Fig. 7: Examples generated by methods listed in the leaderboard. within different categories. A more detailed analysis of model performance in different categories can be found in supplementary material (Sec D.1). Current subject-driven T2I models exhibit performance degrada- tion on hard level subjects As illustrated in Fig. 8a, the model exhibits substantial variation in performance across different difficulty levels: 1) For sub- ject preservation, there is a pronounced decline in performance as the difficulty of the subject images increases. The model achieves significantly better results on images classified as simple compared to those categorized as hard. This ob- servation supports the validity of our image difficulty classification scheme. 2) As demonstrated in Table 3, current closed-source models have achieved re- markable advancements in subject-driven T2I tasks, yielding highly competitive scores in our benchmark evaluation. Nevertheless, as illustrated in Figure 8, the performance of Nano-Banana on challenging categories leaves ample room for im- provement. Thus, our benchmark retains its formidable challenge. 3) For prompt following, Fig. 8a shows that model capability is minimally influenced by the subject difficulty level. This could be explained by the fact that CLIP-T primarily emphasizes overall semantic information. Hence, as long as the generated image correctly represents the general category and overall shape, the evaluation score is unlikely to be substantially reduced, even if finer details are not perfectly captured. Given these findings, it is crucial to enhance models’ ability to encode and reconstruct complex subject details more effectively in future research. The subject-driven T2I capability for different prompt scenarios is not robust Fig. 8b shows the average performance of all models across six prompt scenarios. The results show that: 1) In BC, VS, and IE scenarios, the model’s performance consistently declines across all evaluation dimensions. This trend suggests that scenario difficulty increases progressively from BC to IE. Notably, the finding that the IE scenario is more challenging than the BC scenario aligns with intuitive expectations. 2) For subject preservation, the model’s average performance across the AC, SC, and IM scenarios remains relatively low. This could be because the generated subjects undergo partial modifications relative to the original subjects in these three scenarios. Given these findings, more emphasis 14Z. Hu et al. BLIP-DiffusionIP-AdapterMS_DiffusionOminiControlSSR-EncoderUNOEmu2 RealcustomOmnigenCustom DiffusionDreamboothTextual Inversionλ-EclipseHiPer NeTIDreamOFlux.1 kontextACE++Nano-Banana 0 0.1 0.2 0.3 0.4 0.5 EasyMediumHard 0.24 0.28 0.32 0.36 EasyMediumHard 0.15 0.25 0.35 EasyMediumHard Prompt Following Image Quality Subject Preservation (a) DSH-Bench scores across different subject difficulty level. 0.250 0.260 0.270 0.280 0.290 0.300 0.310 0.320 BCVSIEACSCIM 0.300 0.305 0.310 0.315 0.320 0.325 BCVSIEACSCIM 0.246 0.248 0.250 0.252 0.254 0.256 0.258 0.260 0.262 BCVSIEACSCIM Prompt Following Image Quality Subject Preservation (b) DSH-Bench scores across different prompt scenarios. BLIP-DiffusionIP-AdapterMS_DiffusionOminiControlSSR-EncoderUNOEmu2 RealcustomOmnigenCustom DiffusionDreamboothTextual Inversionλ-EclipseHiPer NeTIDreamOFlux.1 kontextACE++Nano-Banana 0.16 0.2 0.24 0.28 0.32 0.36 VE MI PUF FB MS BO FUR HA AM BU DP IN ST DN PLJE BS AR CL SE SBA TO MA RE BI FI HFB FC AC 0.12 0.16 0.2 0.24 0.28 0.32 0.36 VE MI PUF FB MS BO FUR HA AM BU DP IN ST DN PLJE BS AR CL SE SBA TO MA RE BI FI HFB FC AC -0.25 -0.05 0.15 0.35 0.55 0.75 VE MI PUF FB MS BO FUR HA AM BU DP IN ST DN PLJE BS AR CL SE SBA TO MA RE BI FI HFB FC AC 0.17 0.21 0.25 0.29 0.33 0.37 VE MI PUF FB MS BO FUR HA AM BU DP IN ST DN PLJE BS AR CL SE SBA TO MA RE BI FI HFB FC AC 0.15 0.19 0.23 0.27 0.31 0.35 VE MI PUF FB MS BO FUR HA AM BU DP IN ST DN PLJE BS AR CL SE SBA TO MA RE BI FI HFB FC AC -0.35 -0.15 0.05 0.25 0.45 0.65 VE MI PUF FB MS BO FUR HA AM BU DP IN ST DN PLJE BS AR CL SE SBA TO MA RE BI FI HFB FC AC Photorealistic Non - photorealistic Prompt Following Image Quality Subject Preservation (c) DSH-Bench scores across different categories. Fig. 8: Comparison for DSH-Bench scores across different evaluation di- mensions. Note that the categories are divided into two types: Photorealistic and Non-photorealistic. The specific metric values are provided in the supplementary mate- rial (Sec. D.2). Category definitions refer to supplementary material (Sec. E.3). Best viewed when zoomed in. DSH-Bench: A comprehensive benchmark for Subject-Driven T2I15 should be placed on enhancing methods for IE prompt scenario. For instance, increasing the volume of training data tailored to these specific contexts. 6 Conclusion This paper introduces a novel benchmark called DSH-Bench, designed specifically for subject-driven T2I generation. Key features include: 1) a hierarchical category system in image collection to ensure both the diversity and comprehensiveness of subject images; 2) an innovative classification scheme for categorizing subject difficulty levels and prompt scenarios to obtain valuable insights; and 3) a human- aligned and more efficient metric for subject preservation. The benchmark will be publicly available to support the advancement in future research. References 1.Alaluf, Y., Richardson, E., Metzer, G., Cohen-Or, D.: A neural space-time repre- sentation for text-to-image personalization. ACM Transactions on Graphics (TOG) 42(6), 1–10 (2023) 2.Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., Lin, J.: Qwen2.5-vl technical report (2025), https://arxiv.org/abs/2502.13923 3.Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Zhang, Q., Kreis, K., Aittala, M., Aila, T., Laine, S., et al.: ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324 (2022) 4.Batifol, S., Blattmann, A., Boesel, F., Consul, S., Diagne, C., Dockhorn, T., English, J., English, Z., Esser, P., Kulal, S., et al.: Flux. 1 kontext: Flow matching for in- context image generation and editing in latent space. arXiv e-prints p. arXiv–2506 (2025) 5.Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 9650–9660 (2021) 6.Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.H., Murphy, K.P., Freeman, W.T., Rubinstein, M., et al.: Muse: Text-to-image generation via masked generative transformers. In: International Conference on Machine Learning. p. 4055–4075. PMLR (2023) 7.Chen, W., Hu, H., Li, Y., Ruiz, N., Jia, X., Chang, M.W., Cohen, W.W.: Subject- driven text-to-image generation via apprenticeship learning. Advances in Neural Information Processing Systems 36, 30286–30305 (2023) 8. Chen, Z., Fang, S., Liu, W., He, Q., Huang, M., Zhang, Y., Mao, Z.: Dreamidentity: Improved editability for efficient face-identity preserved image generation (2023), https://arxiv.org/abs/2307.00300 9. Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. p. 248–255. Ieee (2009) 10.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al.: Cogview: Mastering text-to-image generation via transformers. Advances in neural information processing systems 34, 19822–19835 (2021) 16Z. Hu et al. 11.Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., et al.: Dreamllm: Synergistic multimodal comprehension and creation. In: ICLR (2024) 12.Dosovitskiy, A., Brox, T.: Generating images with perceptual similarity metrics based on deep networks. In: Lee, D., Sugiyama, M., Luxburg, U., Guyon, I., Garnett, R. (eds.) Advances in Neural Information Processing Systems. vol. 29. Curran Associates, Inc. (2016),https://proceedings.neurips.c/paper_files/ paper/2016/file/371bce7dc83817b7893bcdeed13799b5-Paper.pdf 13.Fu, S., Tamir, N., Sundaram, S., Chai, L., Zhang, R., Dekel, T., Isola, P.: Dreamsim: Learning new dimensions of human visual similarity using synthetic data. In: Advances in Neural Information Processing Systems. vol. 36, p. 50742–50768 (2023) 14. Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., Taigman, Y.: Make- a-scene: Scene-based text-to-image generation with human priors. In: European Conference on Computer Vision. p. 89–106. Springer (2022) 15.Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A.H., Chechik, G., Cohen-Or, D.: An image is worth one word: Personalizing text-to-image generation using textual inversion (2022).https://doi.org/10.48550/ARXIV.2208.01618, https://arxiv.org/abs/2208.01618 16.Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A.H., Chechik, G., Cohen-or, D.: An image is worth one word: Personalizing text-to-image generation using textual inversion. In: The Eleventh International Conference on Learning Representations (2023), https://openreview.net/forum?id=NAQvF08TcyG 17.Gal, R., Arar, M., Atzmon, Y., Bermano, A.H., Chechik, G., Cohen-Or, D.: Encoder- based domain tuning for fast personalization of text-to-image models. ACM Trans- actions on Graphics (TOG) 42(4), 1–13 (2023) 18.Google: Introducing gemini 2.5 flash image, our state-of-the-art image model (2025), https://developers.googleblog.com/introducing-gemini-2-5-flash-image/, accessed: 2025-12-15 19. Han, I., Yang, S., Kwon, T., Ye, J.C.: Highly personalized text embedding for image manipulation by stable diffusion. arXiv preprint arXiv:2303.08767 (2023) 20.Hao, S., Han, K., Zhao, S., Wong, K.Y.K.: Vico: Plug-and-play visual condition for personalized text-to-image generation. arXiv preprint arXiv:2306.00971 (2023) 21.He, J., Tuo, Y., Chen, B., Zhong, C., Geng, Y., Bo, L.: Anystory: Towards unified single and multiple subject personalization in text-to-image generation (2025), https://arxiv.org/abs/2501.09503 22. Hu, H., Chan, K.C.K., Su, Y.C., Chen, W., Li, Y., Sohn, K., Zhao, Y., Ben, X., Gong, B., Cohen, W., Chang, M.W., Jia, X.: Instruct-imagen: Image generation with multi-modal instruction (2024), https://arxiv.org/abs/2401.01952 23.Hu, H., Chan, K.C., Su, Y.C., Chen, W., Li, Y., Sohn, K., Zhao, Y., Ben, X., Gong, B., Cohen, W., et al.: Instruct-imagen: Image generation with multi-modal instruction. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 4754–4763 (2024) 24.Hua, M., Liu, J., Ding, F., Liu, W., Wu, J., He, Q.: Dreamtuner: Single image is enough for subject-driven generation (2023),https://arxiv.org/abs/2312.13691 25.Huang, L., Lin, H., Zhou, Y., Xiao, K.: Flexip: Dynamic control of preservation and personality for customized image generation (2025),https://arxiv.org/abs/ 2504.07405 26. Huang, Z., Zhuang, S., Fu, C., Yang, B., Zhang, Y., Sun, C., Zhang, Z., Wang, Y., Li, C., Zha, Z.J.: Wegen: A unified model for interactive multimodal generation as we chat (2025), https://arxiv.org/abs/2503.01115 DSH-Bench: A comprehensive benchmark for Subject-Driven T2I17 27.Kang, M., Zhu, J.Y., Zhang, R., Park, J., Shechtman, E., Paris, S., Park, T.: Scaling up gans for text-to-image synthesis. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 10124–10134 (2023) 28.Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., et al.: Segment anything. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 4015–4026 (2023) 29.Kirstain, Y., Polyak, A., Singer, U., Matiana, S., Penna, J., Levy, O.: Pick-a-pic: An open dataset of user preferences for text-to-image generation. Advances in Neural Information Processing Systems 36, 36652–36663 (2023) 30. Kumari, N., Zhang, B., Zhang, R., Shechtman, E., Zhu, J.Y.: Multi-concept cus- tomization of text-to-image diffusion. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 1931–1941 (2023) 31.Le, D.H., Pham, T., Lee, S., Clark, C., Kembhavi, A., Mandt, S., Krishna, R., Lu, J.: One diffusion to generate them all (2024),https://arxiv.org/abs/2411.16318 32. Li, D., Li, J., Hoi, S.: Blip-diffusion: Pre-trained subject representation for con- trollable text-to-image generation and editing. Advances in Neural Information Processing Systems 36, 30146–30166 (2023) 33. Li, Z., Cao, M., Wang, X., Qi, Z., Cheng, M.M., Shan, Y.: Photomaker: Customizing realistic human photos via stacked id embedding (2023),https://arxiv.org/abs/ 2312.04461 34. Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Computer vision– ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13. p. 740–755. Springer (2014) 35.Liu, Z., Rodriguez-Opazo, C., Teney, D., Gould, S.: Image retrieval on real-life images with pre-trained vision-and-language models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 2125–2134 (2021) 36. Liu, Z., Zhang, Y., Shen, Y., Zheng, K., Zhu, K., Feng, R., Liu, Y., Zhao, D., Zhou, J., Cao, Y.: Cones 2: Customizable image synthesis with multiple subjects (2023), https://arxiv.org/abs/2305.19327 37.Ma, J., Liang, J., Chen, C., Lu, H.: Subject-diffusion:open domain personalized text-to-image generation without test-time fine-tuning (2024),https://arxiv.org/ abs/2307.11410 38. Mao, C., Zhang, J., Pan, Y., Jiang, Z., Han, Z., Liu, Y., Zhou, J.: Ace++: Instruction- based image creation and editing via context-aware content filling. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 1958–1966 (2025) 39. Mao, Z., Huang, M., Ding, F., Liu, M., He, Q., Zhang, Y.: Realcustom++: Represent- ing images as real-word for real-time customization. arXiv e-prints p. arXiv–2408 (2024) 40. Mou, C., Wu, Y., Wu, W., Guo, Z., Zhang, P., Cheng, Y., Luo, Y., Ding, F., Zhang, S., Li, X., et al.: Dreamo: A unified framework for image customization. In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers. p. 1–12 (2025) 41.OpenAI: Introducing gpt-4o and more tools to chatgpt free users (2024),https:// openai.com/index/gpt-4o-and-more-tools-to-chatgpt-free/, accessed: 2024- 06-15 42. Patashnik, O., Gal, R., Ostashev, D., Tulyakov, S., Aberman, K., Cohen-Or, D.: Nested attention: Semantic-aware attention values for concept personalization (2025), https://arxiv.org/abs/2501.01407 18Z. Hu et al. 43.Patel, M., Jung, S., Baral, C., Yang, Y.:λ-eclipse: Multi-concept personalized text-to-image diffusion models by leveraging clip latent space. arXiv preprint arXiv:2402.05195 (2024) 44.Peebles, W., Xie, S.: Scalable diffusion models with transformers (2023),https: //arxiv.org/abs/2212.09748 45. Peng, Y., Cui, Y., Tang, H., Qi, Z., Dong, R., Bai, J., Han, C., Ge, Z., Zhang, X., Xia, S.T.: Dreambench++: A human-aligned benchmark for personalized image gen- eration. In: The Thirteenth International Conference on Learning Representations (2025), https://openreview.net/forum?id=4GSOESJrk6 46. pin: https://w.pinterest.com/. https://w.pinterest.com/ 47.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: Sdxl: Improving latent diffusion models for high-resolution image synthesis (2023), https://arxiv.org/abs/2307.01952 48. Prashnani, E., Cai, H., Mostofi, Y., Sen, P.: Pieapp: Perceptual image-error assess- ment through pairwise preference. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 1808–1817 (2018).https://doi.org/10.1109/ CVPR.2018.00194 49.Qiu, Z., Liu, W., Feng, H., Xue, Y., Feng, Y., Liu, Z., Zhang, D., Weller, A., Schölkopf, B.: Controlling text-to-image diffusion by orthogonal finetuning. Ad- vances in Neural Information Processing Systems 36, 79320–79362 (2023) 50. Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. p. 8748–8763. PmLR (2021) 51.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 10684–10695 (2022) 52.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models (2022),https://arxiv.org/abs/2112. 10752 53.Rowles, C., Vainer, S., Nigris, D.D., Elizarov, S., Kutsy, K., Donné, S.: Ipadapter- instruct: Resolving ambiguity in image-based conditioning using instruct prompts (2024), https://arxiv.org/abs/2408.03209 54.Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., Aberman, K.: Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 22500–22510 (2023) 55.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al.: Photorealistic text- to-image diffusion models with deep language understanding. Advances in neural information processing systems 35, 36479–36494 (2022) 56. Shi, J., Xiong, W., Lin, Z., Jung, H.J.: Instantbooth: Personalized text-to-image generation without test-time finetuning (2023),https://arxiv.org/abs/2304. 03411 57.Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Wang, Y., Rao, Y., Liu, J., Huang, T., Wang, X.: Generative multimodal models are in-context learners. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 14398–14409 (2024) 58.Sun, S., Qu, B., Liang, X., Fan, S., Gao, W.: Ie-bench: Advancing the measure- ment of text-driven image editing for human perception alignment. arXiv preprint arXiv:2501.09927 (2025) DSH-Bench: A comprehensive benchmark for Subject-Driven T2I19 59.Tan, Z., Liu, S., Yang, X., Xue, Q., Wang, X.: Ominicontrol: Minimal and universal control for diffusion transformer. arXiv preprint arXiv:2411.15098 (2024) 60. uns: https://unsplash.com/. https://unsplash.com/ 61. Voynov, A., Chu, Q., Cohen-Or, D., Aberman, K.: P+: Extended textual condi- tioning in text-to-image generation (2023), https://arxiv.org/abs/2303.09522 62.Wang, H., Spinelli, M., Wang, Q., Bai, X., Qin, Z., Chen, A.: Instantstyle: Free lunch towards style-preserving in text-to-image generation. arXiv preprint arXiv:2404.02733 (2024) 63.Wang, X., Fu, S., Huang, Q., He, W., Jiang, H.: Ms-diffusion: Multi-subject zero- shot image personalization with layout guidance. arXiv preprint arXiv:2406.07209 (2024) 64.Wang, Y., Zang, Y., Li, H., Jin, C., Wang, J.: Unified reward model for multimodal understanding and generation. arXiv preprint arXiv:2503.05236 (2025) 65. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E.H., Le, Q.V., Zhou, D.: Chain-of-thought prompting elicits reasoning in large language models. In: Advances in Neural Information Processing Systems (NeurIPS) (2022) 66.Wei, Y., Zhang, Y., Ji, Z., Bai, J., Zhang, L., Zuo, W.: Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 15943–15953 (2023) 67.Wu, S., Huang, M., Wu, W., Cheng, Y., Ding, F., He, Q.: Less-to-more general- ization: Unlocking more controllability by in-context generation. arXiv preprint arXiv:2504.02160 (2025) 68.Wu, X., Sun, K., Zhu, F., Zhao, R., Li, H.: Human preference score: Better aligning text-to-image models with human preference. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 2096–2105 (2023) 69.Xiao, S., Wang, Y., Zhou, J., Yuan, H., Xing, X., Yan, R., Li, C., Wang, S., Huang, T., Liu, Z.: Omnigen: Unified image generation. arXiv preprint arXiv:2409.11340 (2024) 70. Xiong, Z., Xiong, W., Shi, J., Zhang, H., Song, Y., Jacobs, N.: Groundingbooth: Grounding text-to-image customization (2025),https://arxiv.org/abs/2409. 08520 71.Xu, J., Huang, Y., Cheng, J., Yang, Y., Xu, J., Wang, Y., Duan, W., Yang, S., Jin, Q., Li, S., et al.: Visionreward: Fine-grained multi-dimensional human preference learning for image and video generation. arXiv preprint arXiv:2412.21059 (2024) 72.Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., Dong, Y.: Imagereward: Learning and evaluating human preferences for text-to-image generation. Advances in Neural Information Processing Systems 36, 15903–15935 (2023) 73.Ye, H., Zhang, J., Liu, S., Han, X., Yang, W.: Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721 (2023) 74.Zeng, Y., Patel, V.M., Wang, H., Huang, X., Wang, T.C., Liu, M.Y., Balaji, Y.: Jedi: Joint-image diffusion models for finetuning-free personalized text-to-image generation (2024), https://arxiv.org/abs/2407.06187 75.Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 586–595 (2018).https://doi. org/10.1109/CVPR.2018.00068 76. Zhang, Y., Song, Y., Liu, J., Wang, R., Yu, J., Tang, H., Li, H., Tang, X., Hu, Y., Pan, H., et al.: Ssr-encoder: Encoding selective subject representation for subject- 20Z. Hu et al. driven generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 8069–8078 (2024)