Paper deep dive
Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design
Shenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, Mingbo Zhao
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semantic information of diverse design elements exhibits intricate coupling relationships in 3D representations, posing substantial challenges for generating diverse 3D garments. In this work, to handle the above problem, We introduce Fashion-3DLR, a novel 3D garment generation framework that utilizes diverse design elements to create high-quality, versatile 3D garment assets. Specifically, to bridge the semantic gaps between different fashion elements, we propose a Garment Feature Fusion Diffusion Transformer (GFF-DiT) module to integrate 2D fashion design elements, e.g., sketch and texture, into latent space. Within the latent space, we then employ a rectified flow transformer to generate geometry latents, which can be decoded into various 3D garment representations, including 3D Gaussians and meshes. Furthermore, we integrate Fashion-3DLR into downstream tasks, achieving the 3D Gaussian Splatting (3DGS)-driven cloth physical simulation and mesh-based virtual try-on. Experimental results indicate that Fashion-3DLR surpass the previous state-of-the-art methods, which verify that the proposed work can generate well-structured, non-watertight garments capable of physical simulation and virtual try-on, underscoring its potential as a versatile 3D garment design tool.
Tags
Links
- Source: https://arxiv.org/abs/2607.23189v1
- Canonical: https://arxiv.org/abs/2607.23189v1
Trouble viewing inline? Open PDF directly โ
Full Text
48,722 characters extracted from source content.
Expand or collapse full text
Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design Shenghao Yang 2242164@mail.dhu.edu.cn Donghua University Shanghai, China Hongtao Zhang 2201957@mail.dhu.edu.cn Donghua University Shanghai, China Yuhan Yi yuhanyi@mail.dhu.edu.cn Donghua University Shanghai, China Zhihao Tang 220995117@mail.dhu.edu.cn Donghua University Shanghai, China Zihao Cui 20995127@mail.dhu.edu.cn Donghua University Shanghai, China Lian Wen 210910807@mail.dhu.edu.cn Donghua University Shanghai, China Han Yan yanhan@dhu.edu.cn Donghua University Shanghai, China Yuan Gao gaoyuan@pjlab.org.cn Shanghai Artificial Intelligence Laboratory Shanghai, China Mingbo Zhao โ mzhao4@dhu.edu.cn Donghua University Shanghai, China Abstract AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digi- tal fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semantic in- formation of diverse design elements exhibits intricate coupling relationships in 3D representations, posing substantial challenges for generating diverse 3D garments. In this work, to handle the above problem, We introduce Fashion-3DLR, a novel 3D garment generation framework that utilizes diverse design elements to cre- ate high-quality, versatile 3D garment assets. Specifically, to bridge the semantic gaps between different fashion elements, we propose a Garment Feature Fusion Diffusion Transformer (GFF-DiT) module to integrate 2D fashion design elements, e.g., sketch and texture, into latent space. Within the latent space, we then employ a rec- tified flow transformer to generate geometry latents, which can be decoded into various 3D garment representations, including 3D Gaussians and meshes. Furthermore, we integrate Fashion-3DLR into downstream tasks, achieving the 3D Gaussian Splatting (3DGS)- driven cloth physical simulation and mesh-based virtual try-on. Experimental results indicate that Fashion-3DLR surpass the previ- ous state-of-the-art methods, which verify that the proposed work can generate well-structured, non-watertight garments capable of physical simulation and virtual try-on, underscoring its potential as a versatile 3D garment design tool. โ Corresponding author Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conference acronym โX, Woodstock, NY ยฉ 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X CCS Concepts โข Applied computingโArts and humanities;โข Computing methodologiesโ Neural networks. Keywords Intelligent fashion design, 3D generative model, image-to-image translation, fashion synthesis ACM Reference Format: Shenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, and Mingbo Zhao. 2026. Fashion-3DLR: A Con- trollable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design . In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym โX). ACM, New York, NY, USA, 10 pages. https://doi.org/X.X 1 Introduction Ready-to-wear fashion design necessitates iterative deconstruction and recombination of multi-modal elements, including conceptual sketches, fabric samples, and brush paintings, to transform creative ideas into tangible realities. Fortunately, the advent of generative models, notably Generative Adversarial Networks (GANs) [8] and Stable Diffusion (SD) [30], has precipitated the maturation of neural network (N)-driven methodologies in digital fashion design. In- tegration of generative artificial intelligence (AI) frameworks into the design workflow empowers designers to dynamically modify stylistic and structural garment attributes, thereby lowering the industryโs entry barrier for newcomers [2]. Concurrently, founda- tional advancements in diffusion models have catalyzed progress toward 3D digital garment synthesis [11]. An effective digital gar- ment creation tool should allow users to customize their outfits in a manner that highlights the individuality and diversity of human beings, with various garment attributes. Recent breakthroughs in latent diffusion models have signif- icantly advanced the state-of-the-art (SOTA) in digital garment asset generation [21]. These frameworks enable novice users to arXiv:2607.23189v1 [cs.CV] 25 Jul 2026 Conference acronym โX, June 03โ05, 2018, Woodstock, NYShenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, and Mingbo Zhao iteratively refine fashion designs through multi-modal conditions, including text, sketches, textures, and colors. Especially, Fashion- Diff [42] leverages paired fashion-related elements (e.g., contour, textures, and brush area) to creatively generate editable fashion item images by learning from human designersโ styles. Compared to 2D garments, 3D garments exhibit superior geometric fidelity and richer structural information. Despite growing industry de- mand for 3D garments, the field of 3D garment editing remains in the nascent stages of development. Developing a controllable 3D fashion generation model capa- ble of simultaneously processing multi-modal conditions presents three fundamental challenges. (1) Cross-modal semantic dis- ambiguation. The model must effectively bridge the substan- tial semantic gap between sketch-based structural representations and appearance-based auxiliary features, such as textures and col- ors [11,17,21]. Furthermore, preserving the unique characteristics of both elements within a highly integrated 3D representation is hard-handling. (2) High-quality generation. Current 3D genera- tion architectures [22,24,25,36,38] often suffer from significant geometric detail loss during both the feature extraction and de- coding phases. To provide users with realistic fashion designs, the model should generation high-quality 3D models with precise tex- ture details. (3) Intuitive edit-ability. The traditional paradigms of 3D modeling, which rely on specialized expertise and intricate manual parameter tuning, is both time-consuming and technically prohibitive [1]. Thus, our solution requires a user-friendly, interac- tive editing workflow, as intuitive as chatting with an LLM-based AI agent. Current AIGC systems still face fundamental limitations in per- forming high-level creative tasks such as 3D garment synthesis. In- tegrating heterogeneous design elements, maintaining fine-grained details under sparse training data, and achieving physically plau- sible simulations remain major challenges. We introduce a novel 3D garment generation framework, i.e., Fashion-3DLR, that simul- taneously addresses these limitations through cross-modal fusion of multi-semantic inputs, detail enhancement via 2D priors, and physics-aware Gaussian kernel deformation. Fashion-3DLR estab- lishes the first end-to-end pipeline capable of generating high- fidelity 3D garments with dynamic simulation properties while maintaining precise artistic control, significantly reducing the man- ual refinement typically required in professional design workflows. Our key contributions are as follows: โขWe propose Fashion-3DLR, a unified framework that syn- thesizes high-quality 3D garments from paired design ele- ments by integrating sketch-driven structural modeling with texture-guided appearance generation. The framework uni- fies cross-modal semantic fusion, rectified-flow-based latent generation, and multi-format decoding, enabling high-fidelity outputs in both 3DGS and mesh. โขWe introduce GFF-DiT, a diffusion-based cross-modal fusion module that jointly encodes sketch geometry and texture semantics through bidirectional modulation and multi-level feature aggregation. This design preserves structural fidelity while enhancing fabric-level details, enabling diverse and high quality 3D garment synthesis under limited 3D super- vision. โขFashion-3DLR can be integrated with 3D Gaussian-based physical simulation, where Gaussians are modeled as de- formable particles to enable realistic cloth dynamics. We enable realistic, fabric-specific dynamics and support the simulation of diverse materials, including cotton, silk, wool, and nylon, while preserving high-frequency visual details. โข Fashion-3DLR supports mesh decoding for downstream vir- tual try-on, producing garments with accurate geometry, consistent textures, and strong alignment with input sketches and textures. Compared with existing sewing pattern-based methods, Fashion-3DLR achieves superior realism, structural fidelity, and user-preferred try-on quality. 2 Related Work Heterogeneous Semantic Fusion Generation. In the fields 2D and 3D generation, researches on fusing inputs that share the same data type but contain different semantic elements continues to advance. In the 2D domain, methods [9,10,14,26,44] address unpaired data translation through specialized loss functions or shared latent spaces. Specifically, FashionDiff [42] focuses on fash- ion image generation by creating paired data such as sketches and textures, utilizing the attention-based fusion process to elim- inate semantic differences and merge them into the latent space, combined with the ControlNet-style generator to achieve flexible control over multiple elements [30,43]. In the 3D domain, excel- lent 3D generation works such as SDFusion [7], Trellis [39], and CraftsMan3D [19] not only ensure high-quality generation results but also possess good text comprehension abilities, capable of effec- tively balancing different semantics within prompts. However, their exploration of heterogeneous semantic understanding for image inputs remains preliminary, only supporting multi-image inputs with strong consistency, such as multi-view images of the same object [22,24,25,36,38]. In this paper, Fashion-3DLR achieves 3D generation through heterogeneous semantic image fusion. It can simultaneously learn and preserve features from both sketch im- ages and design element images that have no consistency between them to generate 3D clothing assets. 2D Generative Models for 3D Creation. The exceptional gener- alization capabilities of 2D generative models [6,23,24,33] have sparked considerable interest in research exploring their application to 3D asset creation [22,24,25,36,38] in recent years. DreamFu- sion [29] established the foundation for this direction by distilling knowledge from pretrained image diffusion models to optimize 3D assets. Subsequent works further improved 3D generation quality through more advanced distillation techniques, such as improved score distillation sampling, introduction of multi-view consistency constraints. However, 2D generative models inherently lack the ability to model three-dimensional spatial consistency. This limita- tion leads to perspective contradictions in the generated multi-view images. As a result, 3D assets reconstructed using these methods generally exhibit poorer geometric accuracy and less detailed fea- tures compared to native 3D generative models [22,34] that learn directly from 3D datasets. Meanwhile, Trellis [39] introduced an ef- fective visual feature aggregation method that maps image features to the 3D shape latent space, enhancing both structural integrity and material detail in the generated results. Drawing inspiration Fashion-3DLR.Conference acronym โX, June 03โ05, 2018, Woodstock, NY Figure 1: Given paired design elements, e.g., sketch and texture, Fashion-3DLR generates a diverse set of high-quality, wearable 3D garments. Fashion-3DLR synthesizes richly detailed 3D models by preserving the structural content encoded in the sketch while integrating material characteristics conveyed by the texture or brush strokes, thereby producing garments with faithful and realistic fabric appearance. Sketch & Texture Sketch & BrushArea (a) Input 3DGS Mesh (c) Output DiTFusion Block GFF Block GFF Block ... ... ... Noise Latent Self Att. Cross Att . FFN N ร Mod KV GFF-DiT Garment Latent Sp. Conv. De. Meshes 3DGSs Physic 3DGS Output (b)Fashion-3DLR Pos. Emb. Forward Figure 2: Overall pipeline for Fashion-3DLR. Our approach processes paired design element images by fusing their features through GFF-DiT. These features then interact via multi-head attention mechanisms and noisy latents in our custom rectified flow transformers. This interaction produces shape latents with integrated semantic features, which a VAE decoder transforms into high-quality 3D clothing assets. from this approach, we designed rectified flow transformers that effectively incorporate GFF-DiT fused image features into the 3D generation process. 3D Gaussian-Based Cloth Simulation. Recent works have ex- plored incorporating 3D Gaussian Splatting [15] into garment simu- lation pipelines. However, most existing methods [18,31,32] remain fundamentally mesh-based. In these approaches, cloth dynamics are first simulated using conventional mesh-based physical models, after which 3D Gaussians are applied as a texture or appearance layer mapped onto the mesh surface. Such hybrid designs largely in- herit the limitations of traditional graphics pipelines and introduce additional artifacts during the mapping process, often leading to loss of fine-grained details and rendering quality that is inferior to native mesh-based materials. To move beyond mesh-dependent for- mulations, PhysGaussians [41] proposed a preliminary attempt at directly simulating 3D Gaussians by leveraging the Material Point Method (MPM) [12], where 3D Gaussian kernels are interpreted as particle primitives to enable physically plausible motion. MPM is a hybrid simulation framework that combines Lagrangian parti- cles with Eulerian grids, and has demonstrated strong capability in handling large deformations, topological changes, and complex frictional interactions. It has been widely applied to multi-physics scenarios involving elastic materials, fluids, and granular media such as sand and snow [13,16,35]. Building upon these insights, we further advance this direction by developing a fully 3D Gaussian- based cloth simulation framework. Unlike prior mesh-dependent Conference acronym โX, June 03โ05, 2018, Woodstock, NYShenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, and Mingbo Zhao Input Input Mean Std Conv. ํ ํ ! " ํ # ! " ํ # ! # Conv. ResNet ํ ํ ! % ํฌ ! ํฉ ! โฑ ! Conv. Conv. Mean Std Conv. ํ ํ) ! " ํ) ! # ํ) ! Conv. ResNet ํฉ !"# ํฌ !"# โฑ # ! โฑ ! ํฅ ! ํต Conv. Conv. Input Figure 3: Details of the proposed GFF process. approaches, our method eliminates the need for intermediate mesh representations and operates directly on 3D Gaussian primitives. This design preserves the rich visual fidelity of 3D Gaussians as- sets while enabling realistic cloth dynamics, thereby expanding the applicability of 3D Gaussians for garment modeling and simulation. 3 Methods Overview. The Fashion-3DLR framework introduces a novel ap- proach for controllable 3D garment generation by leveraging pair- wise fashion elements, such as sketches and textures, to create high-quality and versatile 3D garment assets. The framework ad- dresses three key challenges in 3D fashion generation: cross-modal semantic disambiguation, high-quality generation, and intuitive edit-ability. The pipeline begins with the Garment Feature Fusion Diffusion Transformer (GFF-DiT), which extracts and fuses seman- tic information from heterogeneous design elements (e.g., sketches and textures) into a unified latent space. This fusion is achieved through a bidirectional modulation process that ensures the preser- vation of distinctive features from each input modality. The fused features are then passed to a rectified flow transformer, which gener- ates garment latents with integrated 3D structural and appearance details. These latents are subsequently decoded into multiple out- put formats, including 3DGS representations and meshes, enabling flexible downstream applications. 3.1 Garment Feature Fusion-based DiT The Garment Feature Fusion (GFF) serves as the feature extract and fusion module within Fashion-3DLR framework. Fig. 3 presents the proposed GFF module, which begins with the extraction of semantic information from two conditions, i.e., sketch and tex- ture, and then these elements are seamlessly integrated into latent spaces [42]. Especially,ํandํare fed into convolutional layers for down-sampling to gain the intermediate variable ห ํ ํ and ห ํ ํ , re- spectively. Subsequently, ห ํ ํ is first fed into a residual block and then passed through two different convolutional layers to produce two learnable factorsฮ ํ andฮ ํ , denoted asG(ํ) โ ฮ ํ ,ฮ ํ . The extracted sketch feature ห ํ ํ is convolved to obtain the mean and variance, which is calculated as ห ํ ํ ํ = ํํํํ( ห ํ ํ ) and ห ํ ํ ํ = ํฃํํ( ห ํ ํ ) , respectively. To enable interactive communication between the sketch and texture, we modulate the sketch feature ห ํ ํ with two learnable factorsฮ ํ andฮ ํ determined by feature ห ํ ํ , denoted as: F ฮ ํ ,ฮ ํ (ํ)= ( ฮ ํ | G(ํ) ) โ " ห ํ ํ โ ห ํ ํ ํ ห ํ ํ ํ # โ ( ฮ ํ | G(ํ) ) (1) whereโdenotes multiplication andโindicates element-wise ad- dition. To further facilitate bi-directional interaction between the sketch and texture conditions, we performG((F ฮ ํ ,ฮ ํ (ํ))) โ ฮ ํ+1 ,ฮ ํ+1 , thus we have: F ฮ ํ+1 ,ฮ ํ+1 (ํ)=ฮ ํ+1 โ ห ํ ํ โ ห ํ ํ ํ ห ํ ํ ํ โฮ ํ+1 (2) Finally,F ฮ ํ+1 ,ฮ ํ+1 (ํ)is passed to a zero convolution layerZto obtain the fusion feature ํฅ ํ , denoted as: ํฅ ํ =ZโขF ฮ ํ+1 ,ฮ ํ+1 (ํ)(3) It is important to highlight that the symmetrical treatment ofํ andํ, which has the potential to eliminate disparities within the data, ultimately leading to a robust and accurate representation of visual semantics. In a similar manner, we proceed with the afore- mentioned steps to derive the values ofํฅ ํ+1 ,ํฅ ํ+2 ,ํฅ ํ+3 , denoted X ํ =ํฅ ํ ,ํฅ ํ+1 ,ํฅ ํ+2 ,ํฅ ํ+3 . In the subsequent DiT fusion block, the input conditionX ํ is first transformed through a linear layer to produce latent noise. This latent noise, along with the time step embedding, is then processed by the DiT module [28] to generate fused conditional featuresM. These features subsequently serve as the controlling condition for the 3D garment generation. 3.2 Garment Latents Generation We introduce a rectified flow transformer (RFT) for generating garment latentsS, as shown in Fig. 2. The transformer block is sequentially composed of a self-attention layer, a cross-attention layer, and a feed-forward network. An input dense noisy grid is serialized into the noise latent, combined with positional encodings, and fed into the transformer blocks for denoising. We divide 3D garment assets into voxel grids, where each grid contains activated voxel markers and visual features that capture the detailed structure and appearance of the local area. Specifically, we render images by randomly sampling camera perspectives on a sphere, extract feature maps using pretrained encoders [27], project each voxel onto these multi-view feature maps to retrieve features from corresponding positions, and then average these features to create the final visual feature [39]. We align the resolution of voxelized features with garment latents, enabling transformer blocks to effectively denoise by combining visual features with both the strong representational capabilities of relevant features and the structural framework in activated voxels. ConditionMis injected through cross-attention layers as keys and values, and the noise latent is injected through the self-attention layer as a query. After passing throughํtransformer blocks, we obtain a garment latentSwhich not only reflects the input texture ํand sketchํbut also possesses the correct 3D features. Differ- ent convolutional up-sampling blocks are appended at the end of the transformer blocks to convert the garment latentSinto high- quality 3D assets in various formats, including 3DGSs and meshes. Fashion-3DLR.Conference acronym โX, June 03โ05, 2018, Woodstock, NY 3.3 Optimization of Fashion-3DLR We employ a two-stage training approach to optimize our proposed Fashion-3DLR framework. For the GFF-DiT module, in addition to the initial imageํง 0 and the noise imageํง ํก generated at timeํก, we incorporate fashion-related inputs including fashion-specific conditions (e.g., sketchํand textureํ) and the current time stepํก. We train the GFF-DiT as follows: L ํบํน (ํ)=E ํง 0 ,ํก,ํ,ํ,ํ โฅ ํโํ ํ ( ํง ํก ,ํก, f(ํ,ํ) )โฅ 2 2 (4) wherefrepresents the proposed GFF process andํ ํ is the DiT net- work. Upon completion of training, we freeze the relevant network weights of GFF-DiT. For the rectified flow transformer module, we utilize a linear interpolation forward process. Specifically, given data samplesํ 0 and noisesํwith a timestepํก, the forward and back- ward process is denoted asํ(ํก)=(1โํก)ํ 0 +ํกํandํฃ(ํ,ํก)=โ ํก ํ, respectively. Consequently, the rectified flow transformerํฃ ํ can be approximated by minimizing the conditional flow matching (CFM) objective, denoted as: L ํถํนํ (ํ)=E ํก,ํ 0 ,ํ โฅ ํฃ ํ (ํ,ํก)โ ( ํโํ 0 )โฅ 2 2 (5) After training, the garment latentsScan be generated by the Fashion-3DLR and converted into high-quality 3D assets. 3.4 Physical Garment 3DGS In this section, we highlighted that Fashion-3DLR supports 3DGS- based garment asset generation. To estimate fabric-specific physical parameters for cotton, silk, wool, and nylon, we adopt a physics integrated 3D Gaussian representation inspired by PhysGaussian. A static cloth instance is first reconstructed as a set of unstructured 3D Gaussian kernelsํฅ ํ , ํ ํ , ํด ํ , ํถ ํ ํโP , whereํฅ ํ ,ํ ํ ,ํด ํ ,ํถ ํ denote the Gaussian centers, opacities, covariance matrices, and spherical harmonic coefficients. For each fabric type, we instantiate a corresponding material model from a fabric library and generate its ground-truth dynamic behavior by simulating wind-driven cloth motion in a commercial cloth simulator [3]. We then construct a 3D-Gaussian cloth with identical initial geometry and endow each Gaussian with continuum-mechanics-based kinematics using the MPM solver. During parameter identification, we iteratively adjust a compact set of physically meaningful parametersํ= ํ ํ , ํ ํ , ํ ํ โ , ํ, whereํ ํ refers to bending stiffness,ํ ํ refers to stretching stiffness,ํ ํ โ refers to shear stiffness, andํrefers to area density. These parameters govern the evolution of the Gaussian field through ํด ํ (ํก)= ํน ํ (ํก),ํด ํ ,ํน ํ (ํก) โค , ํฅ ํ (ํก)= ํ(ํ ํ ,ํก)(6) whereํน ํ (ํก)andํ(ยท)denote the deformation gradient and defor- mation map. We optimize by minimizing the discrepancy between the simulated Gaussian dynamics and the reference cloth video: L gauss (ํ)= arg min ํ ํน video โ ํน gauss (ํ) 2 (7) whereํน video contains motion features extracted from the baseline cloth video, e.g., local vibration frequencies, modal responses, and motion-energy spectra, andํน gauss (ํ)denotes the same features computed from the physics-driven Gaussian evolution. Through the above process, we enables physically consistent and fabric-specific parameter identification directly within the 3D Gaussian domain. 4 Experiments In this section, we conduct both quantitative and qualitative compar- ative experiments to evaluate artistic 3D generation by integrating design elements. We include general 3D generation models support- ing multi-image inputs and existing clothing generation methods based on sewing patterns. 4.1 Experiment Setup Dataset. For training, we fine-tuned the pre-trained RFT struc- ture on two datasets: the SewFactory dataset [21] containing 19.1k high-quality 3D garment assets, and the Polyvore fashion image collection [42] comprising 92,488 items across 6 categories. For Sew- Factory, we rendered 120 multi-view images per 3D asset and used GPT-4o [1] to generate descriptive captions for both rendered and Polyvore images. To construct multi-modal supervision, we applied Holistically-Nested Edge Detection (HED) [40] to obtain sketch im- ages, employed neural painting to synthesize oil-brush stylization, and adopted 32ร32 random foreground window sampling to extract texture patches. This pipeline yielded paired instanceโelement data, which were further augmented in text, image, and 3D modalities, with all images standardized to a 256ร256 resolution. Metrics. We introduced the following metrics to evaluate Fashion- 3DLR. FID CLIP measures the semantic consistency between the generated garments and the input texture or brush by computing the distance between their CLIP-based feature distributions, where a higher score indicates better alignment in color and details. LPIPS evaluates the contour similarity between generated garments and sketch images by comparing deep feature distances, with lower values indicating closer alignment in shape and structural layout. CLIP-I score quantifies the semantic similarity between the gen- erated garments and the paired fashion elements by computing their feature alignment in the CLIP imageโimage embedding space, where a higher score indicates stronger semantic correspondence. Implementation details. We first trained GFF-DiT on large-scale image data to strengthen its ability to fuse semantically diverse visual inputs. The training dataset contains approximately 15.2K samples, with a total training duration of about 30 hours and a single inference time of approximately 40 seconds. The trained model was then frozen and integrated into a RFT structure for subsequent training on paired 3D instanceโelement data. Detailed training configurations are provided in the appendix. Table 1: Quantitative comparisons with baseline methods for the task of 3D generation with fusion of sketch images and texture images. Bold indicates best results. MethodFID CLIP โLPIPSโCLIP-I score โ Hunyuan3D-3.0 [4] [2025]0.46480.70010.7946 Tripo [37] [2024]0.45240.68030.7961 Trellis [39] [CVPR2025]0.71420.78420.6615 Sparc3D [20] [NeurIPS2025]0.83080.78250.6638 ReconViaGen [5] [ICLR2026]0.86370.81130.6395 Fashion-3DLR [ours]0.27430.59460.9271 Conference acronym โX, June 03โ05, 2018, Woodstock, NYShenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, and Mingbo Zhao Texture Sketch Texture Sketch Texture Sketch Texture Sketch Brush Sketch Brush Sketch Brush Sketch Brush Sketch Figure 4: Qualitative results using sketch images with texture images, and sketch images with oil brush images. 4.2 Controllable 3D Garment Generation We evaluate the 3D garment generation quality of Fashion-3DLR through both qualitative analysis and comparison with baseline methods. We first present representative results produced by our Fashion-3DLR. Results Conditioning on Sketch, Texture and Brush Area. As illustrated in Fig. 4, Fashion-3DLR effectively fuses input sketches, textures, and artistic oil-brush features to synthesize 3D garments exhibiting diverse and visually coherent styles. In terms of feature fusion, Fashion-3DLR demonstrates strong capability in translat- ing 2D visual cues into consistent 3D representations in two pri- mary aspects. First, it adapts minimalist sketch inputs into detailed 3D assets enriched with high-fidelity color textures. Second, it in- corporates brush-stroke-based artistic cues to generate distinctive blending and textural effects across garment surfaces. The resulting structured silhouettes align precisely with texture pattern distri- butions and color gradients, while visible brush traces impart a handcrafted aesthetic that mitigates the uniformity of conventional digital garment. Details Conditioning on Sketch, Texture and Brush Area. Re- garding shape generation, Fashion-3DLR faithfully preserves both the stylistic intent and geometric outline of the sketch while in- troducing refined local details. It is capable of generating complex band-like structures with bodysuit-level articulation, and perspec- tive renderings reveal non-closed topologies that ensure garment wear-ability. These include realistic openings such as circular neck- lines, ring-shaped cuffs, and open hems, as well as internal cavities that conform to the human body. Collectively, the above results demonstrate the robustness and versatility of Fashion-3DLR in multi-feature fusion and the creation of high-quality, artistically expressive 3D garment assets. 4.3 Comparison with SOTA Methods To comprehensively evaluate the effectiveness of our proposed Fashion-3DLR framework, we compared Fashion-3DLR with sev- eral open-sourced state-of-the-art 3D generation methods. The qualitative results are presented in Fig. 5 and Fig. 6. Qualitative Comparison for Sketch & Texture. When condi- tioned jointly on a texture map and a sketch, Fashion-3DLR achieves high-quality fusion of these heterogeneous elements and recon- structs 3D garments with structurally suitable, non-watertight topology. Competing methods, however, exhibit notable short- comings. Hunyuan3D occasionally recovers valid non-watertight forms but frequently misinterprets rare or atypical sketch geome- tries, producing front-only shell artifacts. Its heavy reliance on sketch background statistics makes it particularly vulnerable to white-pixel contamination, which often washes out the final texture and destroys color consistency. Tripo shows more severe geometric instability, yielding implausible outputs such as collapsed pancake structures, hollow ring-like artifacts, or even watertight meshes that contradict the intended garment topology. Its texture fusion remains limited and is similarly degraded by white-background in- terference. Trellis can sometimes preserve the input sketch but still Fashion-3DLR.Conference acronym โX, June 03โ05, 2018, Woodstock, NY Hunyuan3DTripoTrellisOursInput Figure 5: Qualitative comparisons with baselines for 3D garment generation with fusion of sketch, texture and oil brush. Figure 6: 360ยฐ demonstration of qualitative results using sketch images with oil brush images. suffers from unstable reconstruction with frequent topology col- lapse. Although it partially inherits texture colors, most generated surfaces become desaturated or entirely white, with substantial loss of fine-scale detail. In contrast, Fashion-3DLR consistently recon- structs sketch-aligned, non-watertight garments while preserving high-fidelity texture appearance, demonstrating clear advantages in geometric reliability and texture-geometry consistency. Qualitative Comparison for Sketch & Oil Brush. Under the more challenging setting of cross-domain conditioning with oil- brush artistic inputs and sketch constraints, Fashion-3DLR again exhibits remarkable robustness. Fashion-3DLR not only recovers the garment 3D structure dictated by the sketch but also gener- alizes the brush-stroke color distribution and stylistic cues into the final 3D asset, achieving coherent cross-modal style transfer. Because the oil-brush areas and sketch contain inherent domain discrepancies, baseline models fail to satisfy either modality. Hun- yuan3D simply overlays oil-brush areas onto the sketch without structural reasoning, while Tripo and Trellis are unable to produce valid garment shapes altogether. Qualitative comparisons demon- strate that Fashion-3DLR uniquely maintains structural fidelity, style consistency, and robustness across heterogeneous design in- puts, significantly outperforming existing approaches on this com- plex multi-feature fusion task. Quantitative Comparison. We quantitatively evaluate the gener- ation quality based on the rendered results of the final 3D garments, as summarized in Tab. 1. Under the instanceโelement paired test set derived from SewFactory, Fashion-3DLR consistently surpasses all baseline methods, demonstrating superior performance in both fine- grained detail fidelity and conditional consistency for high-quality 3D garment generation. User Study. We conducted a user study, as shown in Tab. 4. Partic- ipants scored the 3D models generated by each method based on two criteria: overall generation quality and consistency with the input images. A total of 24 participants were invited to evaluate 200 sets of generated results. Conference acronym โX, June 03โ05, 2018, Woodstock, NYShenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, and Mingbo Zhao Physics-based DynamicsStaticSketchTexture Physics-based DynamicsStaticSketchOilbrush Figure 7: Physical simulation demonstration of generated 3DGS clothing assets. 1 Garment2 Garments3 Garments ChatGarmentDressCode Fashion-3DLR Input Fashion-3DLRformultiple garmentstry-on Figure 8: Qualitative comparison of existing SOTA clothing generation methods based on sewing patterns with garments draped on human models. Table 2: Quantitative comparisons with baseline methods for the task of clothing generation on sewing patterns with gar- ments draped on human models. Bold indicates best results. MethodFID CLIP โLPIPSโCLIP-I score โ DressCode [11] [SIGGRAPH2024]0.45240.68030.8311 ChatGarment [2] [CVPR2025]0.46480.70010.7946 Fashion-3DLR [ours]0.27430.59460.9271 Table 3: Ablation study on components of Fashion-3DLR. MethodFID CLIP โLPIPSโCLIP-I score โ Ours w/o GFF-DiT0.68740.74290.6983 Ours (full model)0.2743 0.59460.9271 Table 4: Comparison on virtual try-on of User-Pref. MethodDressCode ChatGarment Fashion-3DLR User-Pref.(%)โ64.4173.2583.82 4.4 Ablations and Applications 3DGS Physical Experiments. We developed our simulation frame- work upon the MPM to model the physical dynamics of multiple 3DGS-generated garments. Initially, the generated 3DGS models are scaled, normalized, and positioned within a cubic simulation do- main. A small subset of kernels located at the top is fixed in a static state to serve as anchor points, while an external impulse is briefly applied to the remaining kernels, enabling natural motion governed by fabric dynamics. We simulated a variety of material types, includ- ing Cotton, Silk, Wool, and Nylon, each represented as a distinct 3DGS physical material. The corresponding results are presented in Fig. 7, where the simulation accurately preserves fine-scale geo- metric details, particularly in the slender band-like structures of bodysuit garments. As shown in Fig. 7, dynamic demonstrations of 3DGS clothing generated from sketch and oil-brush inputs re- veal realistic wave-like deformations and motion trajectories that faithfully capture the intrinsic physical behavior of different fabrics. Virtual Try-on. We compare our method with ChatGarment and DressCode, two representative approaches that have shown strong performance in this domain, to evaluate the garment quality and try-on realism achieved by Fashion-3DLR. Since ChatGarment and DressCode cannot directly take sketch images and texture images as input, we use our proposed GFF-DiT module to process these inputs. As shown in Fig. 8, ChatGarment and DressCode generate textures with noticeable color deviations and introduce unrealistic details, such as leather-like dotted patterns. In addition, the gen- erated garments exhibit a V-shaped hem that is inconsistent with the flat hem specified in the sketch image, and short sleeves are incorrectly produced. In contrast, our Fashion-3DLR generates gar- ments whose texture is highly consistent with the reference texture image, showing a more natural horizontal yarn-like distribution. The generated torso and sleeves are nearly equal in length, the hem remains flat, and the overall result faithfully preserves the characteristics of both the sketch image and the texture image. The quantitative results and user study are presented in Tab. 2 and Tab. 4. Our Fashion-3DLR achieves superior performance, demonstrating its strong potential for virtual try-on applications. Ablation Study. To further assess our Fashion-3DLR framework, we conducted an ablation study focusing on the GFF-DiT module. Experimental results indicate that the GFF-DiT component plays a crucial role in understanding and preserving features from se- mantically diverse image inputs. The pre-trained weights obtained Fashion-3DLR.Conference acronym โX, June 03โ05, 2018, Woodstock, NY from large-scale 2D image data substantially enhance 3D gener- ation quality by mitigating feature inconsistency and preventing generation collapse. As shown in Tab. 3, removing GFF-DiT leads to a marked decline in performance across all evaluation metrics. The consistent improvements observed in both generation fidelity and condition alignment, validate the pivotal role of GFF-DiT in achieving robust and high-quality 3D garment synthesis. 5 Conclusion We presented Fashion-3DLR, a controllable 3D garment generation framework that fuses heterogeneous fashion elements into a uni- fied latent space for high-quality 3D synthesis. Through GFF-DiT, flow-based latent modeling, and dual 3D Gaussians or mesh decod- ing, Fashion-3DLR effectively integrates multi-modal design cues and produces controllable 3D garments. Moreover, Fashion-3DLR supports fully 3D Gaussian-based physical simulation, enabling material-dependent dynamic behaviors that remain faithful to the characteristics of diverse fabrics. Experiments demonstrate superior controllability and generation quality, establishing Fashion-3DLR as an effective tool for intelligent 3D fashion design. Although Fashion-3DLR advances 3D fashion generation, it still faces chal- lenges. 3D Gaussian garments cannot interact with mesh-based human bodies, and mesh garments lack sewing-pattern structures. Addressing cross-representation simulation and pattern-level mod- eling will be key directions for future work. References [1]Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023). [2] Siyuan Bian, Chenghao Xu, Yuliang Xiu, Artur Grigorev, Zhen Liu, Cewu Lu, Michael J Black, and Yao Feng. 2025. Chatgarment: Garment estimation, gen- eration and editing via large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2924โ2934. [3]Katherine L Bouman, Bei Xiao, Peter Battaglia, and William T Freeman. 2013. Estimating the material properties of fabric from video. In Proceedings of the IEEE international conference on computer vision. 1984โ1991. [4] Siyu Cao, Hangting Chen, Peng Chen, Yiji Cheng, Yutao Cui, Xinchi Deng, Ying Dong, Kipper Gong, Tianpeng Gu, Xiusen Gu, et al.2025. HunyuanImage 3.0 Technical Report. arXiv preprint arXiv:2509.23951 (2025). [5]Jiahao Chang, Chongjie Ye, Yushuang Wu, Yuantao Chen, Yidan Zhang, Zhongjin Luo, Chenghong Li, Yihao Zhi, and Xiaoguang Han. 2025. ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation. arXiv preprint arXiv:2510.23306 (2025). [6] Yabo Chen, Jiemin Fang, Yuyang Huang, Taoran Yi, Xiaopeng Zhang, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, and Qi Tian. 2024. Cascade- zero123: One image to highly consistent 3d with self-prompted nearby views. In European Conference on Computer Vision. Springer, 311โ330. [7] Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexander G Schwing, and Liang-Yan Gui. 2023. Sdfusion: Multimodal 3d shape completion, reconstruction, and generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4456โ4465. [8] Antonia Creswell, Tom White, Vincent Dumoulin, Kai Arulkumaran, Biswa Sen- gupta, and Anil A Bharath. 2018. Generative adversarial networks: An overview. IEEE signal processing magazine 35, 1 (2018), 53โ65. [9]Yi Rui Cui, Qi Liu, Cheng Ying Gao, and Zhongbo Su. 2018. FashionGAN: Display your fashion design using conditional generative adversarial nets. In Computer Graphics Forum, Vol. 37. Wiley Online Library, 109โ119. [10]Haoye Dong, Xiaodan Liang, Yixuan Zhang, Xujie Zhang, Xiaohui Shen, Zhenyu Xie, Bowen Wu, and Jian Yin. 2020. Fashion editing with adversarial parsing learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8120โ8128. [11] Kai He, Kaixin Yao, Qixuan Zhang, Jingyi Yu, Lingjie Liu, and Lan Xu. 2024. Dresscode: Autoregressively sewing and generating garments from text guidance. ACM Transactions on Graphics (TOG) 43, 4 (2024), 1โ13. [12] Yuanming Hu, Yu Fang, Ziheng Ge, Ziyin Qu, Yixin Zhu, Andre Pradhana, and Chenfanfu Jiang. 2018. A moving least squares material point method with displacement discontinuity and two-way rigid body coupling. ACM Transactions on Graphics (TOG) 37, 4 (2018), 1โ14. [13]Chenfanfu Jiang, Craig Schroeder, Andrew Selle, Joseph Teran, and Alexey Stom- akhin. 2015. The affine particle-in-cell method. ACM Transactions on Graphics (TOG) 34, 4 (2015), 1โ10. [14]Shuhui Jiang, Jun Li, and Yun Fu. 2021. Deep learning for fashion style generation. IEEE Transactions on Neural Networks and Learning Systems 33, 9 (2021), 4538โ 4550. [15] Bernhard Kerbl, Georgios Kopanas, Thomas Leimkรผhler, and George Drettakis. 2023. 3D Gaussian splatting for real-time radiance field rendering. ACM Trans. Graph. 42, 4 (2023), 139โ1. [16]Gergely Klรกr, Theodore Gast, Andre Pradhana, Chuyuan Fu, Craig Schroeder, Chenfanfu Jiang, and Joseph Teran. 2016. Drucker-prager elastoplasticity for sand animation. ACM Transactions on Graphics (TOG) 35, 4 (2016), 1โ12. [17]Maria Korosteleva and Sung-Hee Lee. 2022. Neuraltailor: Reconstructing sewing pattern structures from 3d point clouds of garments. ACM Transactions on Graphics (TOG) 41, 4 (2022), 1โ16. [18] Boqian Li, Xuan Li, Ying Jiang, Tianyi Xie, Feng Gao, Huamin Wang, Yin Yang, and Chenfanfu Jiang. 2025. GarmentDreamer: 3DGS Guided Garment Synthesis with Diverse Geometry and Texture Details. In International Conference on 3D Vision. [19] Weiyu Li, Jiarui Liu, Hongyu Yan, Rui Chen, Yixun Liang, Xuelin Chen, Ping Tan, and Xiaoxiao Long. 2024. Craftsman3d: High-fidelity mesh generation with 3d na- tive generation and interactive geometry refiner. arXiv preprint arXiv:2405.14979 (2024). [20] Zhihao Li, Yufei Wang, Heliang Zheng, Yihao Luo, and Bihan Wen. [n. d.]. Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. [21] Lijuan Liu, Xiangyu Xu, Zhijie Lin, Jiabin Liang, and Shuicheng Yan. 2023. To- wards garment sewing pattern reconstruction from a single image. ACM Trans- actions on Graphics (TOG) 42, 6 (2023), 1โ15. [22]Minghua Liu, Ruoxi Shi, Linghao Chen, Zhuoyang Zhang, Chao Xu, Xinyue Wei, Hansheng Chen, Chong Zeng, Jiayuan Gu, and Hao Su. 2024. One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10072โ10083. [23]Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF international conference on computer vision. 9298โ 9309. [24]Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. 2023. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453 (2023). [25]Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al.2024. Wonder3d: Single image to 3d using cross-domain diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9970โ9980. [26] Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741 (2021). [27]Maxime Oquab, Timothรฉe Darcet, Thรฉo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, et al.2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023). [28]William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. In Proceedings of the IEEE/CVF international conference on computer vision. 4195โ4205. [29]Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988 (2022). [30]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjรถrn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684โ10695. [31]Boxiang Rong, Artur Grigorev, Wenbo Wang, Michael J Black, Bernhard Thomaszewski, Christina Tsalicoglou, and Otmar Hilliges. 2024. Gaussian gar- ments: Reconstructing simulation-ready clothing with photorealistic appearance from multi-view video. arXiv preprint arXiv:2409.08189 (2024). [32] Nikolaos Sarafianos, Tuur Stuyck, Xiaoyu Xiang, Yilei Li, Jovan Popovic, and Rakesh Ranjan. 2024. Garment3dgen: 3d garment stylization and texture genera- tion. arXiv preprint arXiv:2403.18816 (2024). [33] Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. 2023. Zero123++: a single image to consistent multi-view diffusion base model. arXiv preprint arXiv:2310.15110 (2023). [34]Ivan Skorokhodov, Sergey Tulyakov, Yiqun Wang, and Peter Wonka. 2022. Epigraf: Rethinking training of 3d gans. Advances in Neural Information Processing Systems Conference acronym โX, June 03โ05, 2018, Woodstock, NYShenghao Yang, Hongtao Zhang, Yuhan Yi, Zhihao Tang, Zihao Cui, Lian Wen, Han Yan, Yuan Gao, and Mingbo Zhao 35 (2022), 24487โ24501. [35]Alexey Stomakhin, Craig Schroeder, Lawrence Chai, Joseph Teran, and Andrew Selle. 2013. A material point method for snow simulation. ACM Transactions on Graphics (TOG) 32, 4 (2013), 1โ10. [36]Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. 2023. Dream- gaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653 (2023). [37]Dmitry Tochilkin, David Pankratz, Zexiang Liu, Zixuan Huang, , Adam Letts, Yangguang Li, Ding Liang, Christian Laforte, Varun Jampani, and Yan-Pei Cao. 2024. TripoSR: Fast 3D Object Reconstruction from a Single Image. arXiv preprint arXiv:2403.02151 (2024). [38]Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. 2021. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689 (2021). [39]Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. 2025. Structured 3d latents for scalable and versatile 3d generation. In Proceedings of the Computer Vision and Pattern Recognition Conference. 21469โ21480. [40]Saining Xie and Zhuowen Tu. 2015. Holistically-nested edge detection. In Pro- ceedings of the IEEE international conference on computer vision. 1395โ1403. [41] Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chen- fanfu Jiang. 2024. Physgaussian: Physics-integrated 3d gaussians for generative dynamics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4389โ4398. [42]Han Yan, Haijun Zhang, Xiangyu Mu, Jicong Fan, and Zhao Zhang. 2023. Fashion- diff: a controllable diffusion model using pairwise fashion elements for intelligent design. In Proceedings of the 31st ACM International Conference on Multimedia. 1401โ1411. [43]Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional con- trol to text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision. 3836โ3847. [44]Shizhan Zhu, Raquel Urtasun, Sanja Fidler, Dahua Lin, and Chen Change Loy. 2017. Be your own prada: Fashion synthesis with structural coherence. In Proceedings of the IEEE international conference on computer vision. 1680โ1688.