Paper deep dive
PartMat: Material-Aware 3D Part Decomposition with a Single Global Latent
Guangming Fu, Jin Song, Yiyun Fei, Guoqiu Li, Ruigao Yang, Jianan Jiang
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Part-level 3D generation has recently attracted increasing attention for producing structured and editable 3D assets. However, existing methods typically decompose objects according to functional semantics rather than the editable material boundaries (e.g., fabric, wood, metal) required in practical 3D applications such as interior design. Additionally, current methods often generate parts independently, causing computational costs to scale linearly with the part count. To address these limitations, we present PartMat, an efficient material-aware 3D part decomposition pipeline that represents multi-part geometry with a single global latent. Given a reference image and a single whole-object geometry, PartMat decomposes the object into parts that follow material boundaries. First, we propose PartVAE to learn such a unified representation and decode all material parts in a single forward pass, thereby decoupling inference cost from the number of parts. Second, with this representation, a diffusion model is trained for part generation and refined via reinforcement learning for accurate material assignment and overlap suppression. Finally, to recover fine-grained geometric details, we introduce a sparse-voxel flow-matching model with part attention for geometry post-processing. Extensive experiments demonstrate that PartMat significantly outperforms existing baselines in material-aware decomposition accuracy and achieves comparable geometric quality, while maintaining efficient inference.
Tags
Links
- Source: https://arxiv.org/abs/2608.01825v1
- Canonical: https://arxiv.org/abs/2608.01825v1
Trouble viewing inline? Open PDF directly →
Full Text
44,924 characters extracted from source content.
Expand or collapse full text
PartMat: Material-Aware 3D Part Decomposition with a Single Global Latent Guangming Fu1,2 , Jin Song2 , Yiyun Fei2, Guoqiu Li2, Ruigao Yang2, Jianan Jiang2 This work is done by Guangming Fu as an intern at Alibaba Group, supervised by Jin Song. Abstract Part-level 3D generation has recently attracted increasing attention for producing structured and editable 3D assets. However, existing methods typically decompose objects according to functional semantics rather than the editable material boundaries (e.g., fabric, wood, metal) required in practical 3D applications such as interior design. Additionally, current methods often generate parts independently, causing computational costs to scale linearly with the part count. To address these limitations, we present PartMat, an efficient material-aware 3D part decomposition pipeline that represents multi-part geometry with a single global latent. Given a reference image and a single whole-object geometry, PartMat decomposes the object into parts that follow material boundaries. First, we propose PartVAE to learn such a unified representation and decode all material parts in a single forward pass, thereby decoupling inference cost from the number of parts. Second, with this representation, a diffusion model is trained for part generation and refined via reinforcement learning for accurate material assignment and overlap suppression. Finally, to recover fine-grained geometric details, we introduce a sparse-voxel flow-matching model with part attention for geometry post-processing. Extensive experiments demonstrate that PartMat significantly outperforms existing baselines in material-aware decomposition accuracy and achieves comparable geometric quality, while maintaining efficient inference. Introduction Editability has become an essential requirement in image-to-3D generation. Rather than treating a generated asset as a single, undivided mesh, downstream applications require it to be decomposed into components that can be selected, replaced, and assigned independent attributes. Existing methods decompose objects by geometric structure and functional semantics, such as chair backs, seats, and legs. However, this organization does not account for material boundaries that determine appearance and physical behavior. Separating fabric, wood, metal, and glass regions enables independent editing in interior design and the assignment of distinct properties, such as density, friction, and stiffness, for embodied simulation and manipulation planning. Motivated by these requirements, we study material-aware 3D part decomposition: given a reference image and a single whole-object geometry mesh without part annotations, the goal is to decompose the input geometry into parts whose boundaries follow editable material assignments. As illustrated in Figure 1, our decomposition follows material boundaries rather than conventional function-oriented semantics. Here, “material-aware” refers to the geometric decomposition; predicting PBR parameters or physical properties remains a downstream task. Figure 1: Functional vs. material-aware decomposition. Functional parts follow structural roles, whereas our components follow editable material assignments, such as marble, painted wood, and brass. A material component may contain multiple disconnected surface regions. Figure 2: Material-Aware Part Decomposition Results. PartMat decomposes input 3D shapes into explicitly separated material-editable parts, each rendered in a distinct color for downstream PBR assignment and replacement. Left: A complex bedroom scene comprising various objects with rich materials. Middle: Material-aware decomposition of selected objects from the scene, as directly produced by our model. Right: Downstream applications enabled by our approach, demonstrating the decomposed assets seamlessly restyled with diverse PBR materials. Existing part-level methods mainly follow either segmentation-based or part-native pipelines. Segmentation-based methods decompose an input mesh primarily according to geometric structure or functional semantics (Yang et al. 2025a; Yan et al. 2026). Because geometry alone provides insufficient evidence for material boundaries, adapting these methods to material-aware decomposition requires first generating textures or material attributes on the mesh and then performing material segmentation and component reconstruction. This cascaded process lengthens the pipeline and propagates errors from appearance generation to the final decomposition. Part-native methods instead learn multi-part representations directly, but often encode or generate each component separately (Lin et al. 2025; Ding et al. 2025). Despite avoiding explicit segmentation, these methods have two main limitations. First, their representation size and decoding cost grow with the number of components, leading to inefficient inference for objects with many material regions. Second, separate per-component representations can weaken global coordination across material regions. Moreover, latent flow matching supervises only the latent space and does not directly enforce consistency in the decoded geometry. This can cause small regions to be misassigned, disconnected regions of the same material to be split across different components, and generated parts to be duplicated or overlap with one another. To address these challenges, we propose PartMat, a three-stage framework for image-guided material-aware decomposition. First, a global PartVAE encodes all material components into a single fixed-length latent and jointly decodes their signed-distance fields through a shared implicit backbone. Consequently, neither the latent size nor the number of decoder passes grows with the number of parts. Second, a conditional PartDiT generates this latent from the reference image and whole-object geometry. We then post-train PartDiT with direct-gradient reinforcement learning using differentiable rewards for component matching and overlap suppression. Third, to improve the geometric fidelity of the compactly decoded components, we introduce a high-resolution geometry post-processing stage. Figure 2 demonstrates PartMat’s material-aware decompositions in an indoor scene. Experiments on approximately 300K material-annotated furniture and household objects show that PartMat achieves state-of-the-art material decomposition accuracy on our benchmark. Contributions. Our main contributions are as follows: • We introduce PartVAE, to the best of our knowledge, which is the first VAE to represent multi-component geometry with a single global latent while decoupling the decoding latency from the number of parts. • We propose PartMat, a three-stage framework that first employs direct-gradient reinforcement learning for precise material decomposition and overlap suppression, and applies a part-aware sparse flow refiner to restore high-resolution geometry with seamless component boundaries. • Extensive experiments validate our method’s superior performance in part-level decomposition and generation. Related Work 3D Shape Representation Single-shape VAEs primarily use vector-set latents for implicit geometry (Zhang et al. 2023; Chen et al. 2025c; Lai et al. 2026) or sparse-voxel structures for high-resolution reconstruction (Ren et al. 2024; Wu et al. 2025; Xiang et al. 2026). Multi-part compression remains less explored. UniPart augments one holistic geometry latent with segmentation labels (He et al. 2026), whereas PartPacker packs non-contacting parts into two shared volume latents without dedicated per-part channels (Tang et al. 2025). PartMat instead compresses multiple explicit part geometries into one global latent. Object-Level Shape Generation Image-conditioned shape generation has progressed from transformer-based feed-forward reconstruction (Hong et al. 2024) to vector-set latent flow models trained at scale (Li et al. 2025; Zhao et al. 2025). Recent systems adopt sparse-structured latents and spatial sparse attention for efficient high-resolution geometry (Xiang et al. 2025, 2026; Wu et al. 2025), while LATTICE combines voxel structure with set-based modeling (Lai et al. 2026). Complementary to these latent shape representations, recent native mesh generators directly model explicit polygonal topology: SpaceMesh learns continuous manifold connectivity, MeshFlow generates continuous vertex-edge latents in parallel, and Nexus and LATO.2 factorize vertex and topology prediction (Shen et al. 2024; Li et al. 2026; Wang et al. 2026; Long et al. 2026). Despite improving geometric fidelity and scalability, these methods synthesize holistic assets without independently editable material parts. In contrast, PartMat natively decomposes shapes into independently editable material components, preserving high-frequency geometric details through a dedicated geometry refiner. Part Segmentation Classical geometry techniques (Golovinskiy and Funkhouser 2008; Shapira et al. 2008) and supervised point networks (Qi et al. 2017a, b; Wang et al. 2019; Chen et al. 2019; Zhao et al. 2021) predict part labels but require hand-crafted cues or dense annotations. Recent methods lift 2D priors (Kirillov et al. 2023; Abdelreheem et al. 2023; Liu et al. 2023; Tang et al. 2024; Thai et al. 2024; Umam et al. 2024; Zhong et al. 2024), learn open-world or native 3D features (Yang et al. 2024; Liu et al. 2025; Ma et al. 2025), or support language-grounded localization (Ahmed et al. 2025). Nevertheless, lifting can be inconsistent across views and occlusions, and these methods segment observed surfaces rather than recover complete material parts. Part-Level Shape Generation Early structured models use primitives or shape grammars, with PartNet providing fine-grained annotations (Tulsiani et al. 2017; Li et al. 2017; Mo et al. 2019). Part123 and PartGen segment views before reconstruction (Liu et al. 2024; Chen et al. 2025a), while HoloPart and X-Part complete segmented 3D geometry (Yang et al. 2025a; Yan et al. 2026); both can propagate segmentation errors. PartCrafter and FullPart model per-part geometry with costs that scale with part count (Lin et al. 2025; Ding et al. 2025). BANG enables controlled decomposition through exploded dynamics (Zhang et al. 2025), whereas AutoPartGen generates variable-length part sequences autoregressively (Chen et al. 2025b). Most existing approaches organize parts according to geometric structure or functional semantics. Material boundaries need not follow these definitions: one functional part may contain multiple materials, and disconnected regions may share one editable material assignment. PartMat directly generates material-aware parts from a single global latent, with each part represented by an explicit SDF field. Method The architecture of PartMat is illustrated in Figure 3. Given a reference RGB image I and a single geometry mesh ℳgM^g, such as one generated by an image-to-3D model, PartMat produces a material-decomposable 3D asset :=ℳkk=1MO:=\M_k\_k=1^M. Each ℳk=k,kM_k=\V_k,F_k\ is a separate mesh corresponding to one editable material component. All component meshes lie in the same object coordinate system as ℳgM^g, so they can be directly assembled into the complete object without additional transformations. PartMat has three stages: PartVAE jointly encodes and decodes all material components; PartDiT predicts the global latent from the image and input geometry and is further optimized using differentiable SDF rewards; and a part-aware sparse flow model refines the components at high resolution. Figure 3: PartMat overview. Stage I encodes a part-aware point cloud into a global part latent and jointly decodes the component SDFs. Stage I uses separate cross-attention to the reference image and full-geometry conditions to generate this latent, and aligns the decoded fields with matching and overlap rewards. Stage I conditions on the coarse components, global geometry, and reference image to produce refined meshes through global, part, and cross-attention. Snowflakes and flames denote frozen and trainable modules, respectively. Stage I: PartVAE Architecture. PartVAE follows the variational encoder–decoder Transformer of Hunyuan3D 2.1 ShapeVAE (Hunyuan3D et al. 2025): surface points are encoded into a sequence of latent tokens, from which spatial queries are decoded into SDF values. We adapt this architecture to material decomposition by adding material-part conditioning to the encoder and replacing the single SDF output with multiple part channels decoded from one global latent. Part-aware point encoder. Given an object with M≤KM≤ K annotated material parts, where K is the maximum number of part channels, we sample N points from the union of its part surfaces, including uniform surface samples and samples around sharp edges. The encoder input is =(i,i,mi,si)i=1N,P=\(x_i,n_i,m_i,s_i)\_i=1^N, (1) where ix_i, in_i, mim_i, and sis_i denote the position, normal, material part ID, and sharp-edge indicator of point i, respectively. Let ℛ=ℓ=1LR=\r_ \_ =1^L be voxel queries sampled from the occupied voxels of P. Our encoder Fourier-encodes the point and query coordinates, injects learned embeddings of mim_i into the point features, and uses the voxel queries to aggregate P through cross-attention and Transformer layers: =Encoderϕ(,ℛ).H=Encoder_φ(P,R). (2) A linear pre-KL projection of H predicts the Gaussian posterior parameters μ and log2 σ^2. The global latent is then sampled as =+⊙ϵ,ϵ∼(,).z= μ+ σ ε, ε (0,I). (3) where ∈ℝL×dz ^L× d is the global latent shared by all parts. Multi-channel SDF decoder. The decoder first maps z back to the Transformer width with a post-KL projection. A latent Transformer then produces the decoded latent decz_dec. For each spatial query q, the point-query decoder cross-attends to decz_dec, and a K-channel linear head predicts all part SDFs: _q =Dcross(squeryγ(),dec), =D_cross\! (W_squeryγ(q),z_dec ), (4) ψ(;) _ψ(q;z) =[f1(),…,fK()]=sdf+sdf. =[f_1(q),…,f_K(q)]=W_sdfh_q+b_sdf. The final projection produces one SDF channel per part. Thus, for a fixed K, the decoding cost is independent of the actual part count M. We extract each valid part mesh from its zero level set using Marching Cubes (Lorensen and Cline 1987). Training objective. Let =jj=1NqQ=\q_j\_j=1^N_q denote the sampled SDF query points. For each existing part k, fk()f_k(q) and gk()g_k(q) are respectively the predicted and ground-truth SDFs at ∈q . We use ∈0,1Ka∈\0,1\^K to mark existing parts and define =k∣ak=1V=\k a_k=1\. Channels in V are supervised by the reconstruction loss, while padding channels k∉k are kept away from the zero level set by the suppression loss. The overall objective is ℒVAE=ℒrecon+ℒsuppress+λKLℒKL.L_VAE=L_recon+L_suppress+ _KLL_KL. (5) Let ek()=fk()−gk()e_k(q)=f_k(q)-g_k(q). The reconstruction term applies MSE and L1 losses to all valid parts: ℒrecon=∑k∈∑∈[ek()2+|ek()|].L_recon= _k _q [e_k(q)^2+|e_k(q)| ]. (6) For nonexistent parts, the suppression term constrains predictions to the negative interval [slower,supper][s_lower,s_upper]: ℒsuppress=∑k∉∑∈[ _suppress= _k _q [ ReLU(fk()−supper)2 (f_k(q)-s_upper)^2 (7) + + ReLU(slower−fk())2], (s_lower-f_k(q))^2 ], where slower=−1.0s_lower=-1.0 and supper=−0.1s_upper=-0.1. Finally, ℒKL=DKL(qϕ(∣)∥(,))L_KL=D_KL(q_φ(z )\|N(0,I)). Stage I: PartDiT with RL Alignment PartDiT architecture. We freeze PartVAE and train a conditional generator in its latent space. PartDiT follows the Hunyuan3D 2.1 DiT architecture (Hunyuan3D et al. 2025). DINOv2 encodes the reference image into tokens Ic_I (Oquab et al. 2024). To obtain the full-geometry tokens gz^g, we reuse the frozen PartVAE encoder and treat the complete input mesh ℳgM^g as a single part. In each PartDiT block, the noisy part latent attends separately to Ic_I and gz^g through two cross-attention layers. PartDiT is trained with flow matching in the frozen PartVAE latent space. Let 1z_1 be the target part latent, 0∼(,)z_0 (0,I), t∼(0,1)t (0,1), and t=(1−t)0+t1z_t=(1-t)z_0+tz_1. With =(I,g)c=(c_I,z^g), the objective is ℒFM=[‖vθ(t,t;)−(1−0)‖22].L_FM=E\! [ \|v_θ(z_t,t;c)-(z_1-z_0) \|_2^2 ]. Differentiable SDF reward. To improve the decomposition accuracy of the PartDiT, a natural choice is RL post-training. Applying Flow-GRPO (Liu et al. 2026) with a mesh-IoU reward, however, requires explicitly decoding every sampled latent into component meshes before evaluation, resulting in high training cost. We instead compute a differentiable reward directly in the implicit SDF space of the frozen PartVAE. For reward queries ∈Rq _R, the predicted and target fields are converted to soft occupancies Pi()=σ(fi()/τs)P_i(q)=σ(f_i(q)/ _s) and Tj()=σ(gj()/τs)T_j(q)=σ(g_j(q)/ _s). Pairwise occupancy similarities SijS_ij are computed between all predicted and target channels. Because the predicted and target part channels may follow different orders, we use a differentiable soft assignment πij _ij to match predicted channel i with target channel j. Weighting SijS_ij by these assignments yields the order-invariant reward RmatchR_match. Meanwhile, RoverlapR_overlap penalizes simultaneous occupancy by different predicted components: Rmatch R_match =∑i,jπijSij, = _i,j _ijS_ij, (8) Roverlap R_overlap =−1|R|(K2)∑∈R∑1≤i<j≤KPi()Pj(). =- 1|Q_R| K2 _q _R _1≤ i<j≤ KP_i(q)P_j(q). The combined SDF reward is Rsdf=λmatchRmatch+λoverlapRoverlap.R_sdf= _matchR_match+ _overlapR_overlap. (9) Direct-gradient RL alignment. Following LeapAlign (Liang et al. 2026), we directly backpropagate the differentiable SDF reward to PartDiT. Let z denote a generated part latent; the frozen PartVAE decoder converts it into part SDFs for reward evaluation. The alignment objective is ℒRL=[max(0,c−Rsdf(^))],L_RL=E\! [ (0,c-R_sdf( z) ) ], (10) where c is the reward margin. Stage I: Part-Aware Sparse Geometry Refinement PartVAE provides coherent component geometry, but its compact latent and coarse SDF extraction can smooth thin structures and material interfaces. We refine all parts jointly with a sparse latent flow model conditioned on the coarse parts, the whole geometry, and the reference image. Sparse voxel. Diffusion on a dense 5123512^3 voxel grid is prohibitively expensive, whereas lowering the resolution loses surface detail; we therefore encode the coarse parts with a shape VAE and refine them on a fixed sparse voxel support (Xiang et al. 2025). The global input geometry and coarse parts are projected into this aligned sparse space to form the condition U. Coordinate construction and feature lookup are detailed in the supplementary material. Part attention. The refinement transformer must preserve part identity without processing every part in isolation. Let kX_k denote tokens with component index k. We define within-part attention as PartAttn()=concat1:kSelfAttn(k).PartAttn(X)=concat_1:kSelfAttn(X_k). (11) Each refinement block interleaves full self-attention over X, two part-attention layers, and cross-attention to the image. Part attention prevents features from losing their slot identity; full attention communicates object-level context across adjacent material interfaces. Sparse latent flow refinement. Let 1X_1 be the target features on the sparse structure ~ C and 0∼(,)X_0 (0,I). Following the rectified flow matching paradigm, we construct a linear interpolation path t=(1−t)0+t1X_t=(1-t)X_0+tX_1 and train the part-aware sparse transformer with ℒHR=[‖Hω(t,t,I,)−(1−0)‖22].L_HR=E\! [ \|H_ω(X_t,t,c_I,U)-(X_1-X_0) \|_2^2 ]. (12) At inference, we integrate the sparse flow on the fixed support and decode the sampled features. The resulting meshes remain in the shared object coordinate system and inherit the material slots produced by PartDiT. Experiments Experimental Setup Training Data. We train PartMat on approximately 300K material-aware furniture and household-object shapes. Component labels are derived directly from the material slots assigned to mesh faces, with each slot defining one material component. Detailed preprocessing and sampling configurations are provided in the supplementary material. Baselines and Metrics. For VAE reconstruction, we compare against Hunyuan2.1 (Hunyuan3D et al. 2025), TRELLIS2 (Xiang et al. 2026), PartCrafter (Lin et al. 2025), and PartPacker (Tang et al. 2025). All methods are evaluated at 5123512^3 resolution. We report Chamfer Distance (CD), F1-Score at threshold 0.01 (F1@0.01), average latent token count, and VAE decode latency for 1, 16, and 32 components. For image-conditioned generation, we compare against X-Part (Yan et al. 2026), PartCrafter (Lin et al. 2025), PartPacker (Tang et al. 2025), HoloPart (Yang et al. 2025a), and OmniPart (Yang et al. 2025b). We report CD and F1@0.01 for whole-geometry fidelity and Sem-IoU after optimal component matching for material decomposition accuracy. Further evaluation details are deferred to the supplementary material. Figure 4: Qualitative comparison of material-aware 3D reconstruction. PartVAE recovers complete geometry while preserving editable material-component labels, shown with distinct colors. Method CD(×104×\!10^4) ↓ F1@0.01 ↑ Sem-IoU ↑ X-Part (Yan et al. 2026) 5.13 55.40 43.00 PartCrafter (Lin et al. 2025) 12.32 10.27 15.11 PartPacker (Tang et al. 2025) 6.06 38.77 30.62 HoloPart (Yang et al. 2025a) 5.78 55.31 32.47 OmniPart (Yang et al. 2025b) 5.60 30.38 29.75 PartMat 4.90 68.23 46.92 PartMat w/ RL 3.67 67.45 50.51 PartMat w/ geometry refine 2.83 69.59 49.19 Table 1: Image-to-3D material-aware component generation comparison. Our PartMat significantly outperforms other methods on all metrics. Best results are marked in bold font. Implementation Details. PartVAE follows Hunyuan3D 2.1-VAE (Hunyuan3D et al. 2025) and jointly decodes K=32K=32 material-component SDF channels. PartDiT uses the Hunyuan3D 2.1 DiT backbone, and the geometry refiner augments TRELLIS2 (Xiang et al. 2026) with part attention. Complete training schedules and hyperparameters are provided in the supplementary material. All models are trained on NVIDIA H20 GPUs, while decoding latency is measured on a single NVIDIA RTX 3090 GPU. Figure 5: Image-conditioned material-aware part generation comparison. Our method generates components with higher geometric quality, better material-region correspondence, and editable component IDs. Method CD (×104×\!10^4) ↓ F1@0.01 ↑ Avg. Tokens Enc-Dec Time (s) ↓ Material Support 1 comp. 16 comps. 32 comps. Hunyuan2.1 2.179 75.8 4,096 3.572 – – × TRELLIS2 (Xiang et al. 2026) 1.690 73.7 2,149 0.154 – – × PartCrafter (Lin et al. 2025) 2.340 74.4 49152 1.782 12.958 26.508 ✓ per-component PartPacker (Tang et al. 2025) 2.270 74.7 8,192 13.674 10.617 11.628 † spatial PartVAE (Ours) 2.140 75.2 4096 1.860 1.302 1.698 ✓ material Table 2: VAE reconstruction and decode latency at 5123512^3 resolution (10K sample points, F1 threshold 0.01). Decode time is measured for 1, 16, and 32 material components on the same GPU. Lower CD / decode time is better; higher F1 is better. Material support: × = none, †= spatial only, ✓ = material-editable components. Best results are marked in bold font. Figure 6: Qualitative ablation of RL alignment. We compare the supervised fine-tuning (SFT) against different RL configurations. The matching reward improves component correspondence, while the overlap penalty reduces spatial conflicts; combining both yields the decomposition closest to the ground truth. Results Material-Aware VAE Reconstruction Table 2 compares PartVAE against single-object and part-aware VAE baselines at 5123512^3 resolution. Although TRELLIS2 achieves the lowest CD, and Hunyuan2.1 obtains the highest F1@0.01, neither of them provide material-editable component channels. PartCrafter preserves component identity but decodes each component independently, causing latency to grow with component count. PartPacker reduces this scaling cost with packed geometry, but uses twice the token budget and does not maintain one editable material slot per decoded channel. In contrast, PartVAE maintains a compact token budget of 4,096 and exhibits nearly constant decode latency as the component count scales to 32, all while delivering editable material parts and achieving geometric reconstruction competitive with these baselines. Figure 4 visualizes reconstruction results. PartVAE directly decodes part-specific SDF fields with material-component IDs and obtains reconstruction quality comparable to Hunyuan2.1. Holistic VAEs do not expose editable regions, and per-component decoding baselines show higher latency cost as component count grows. Image-to-3D Material-Aware Generation Leveraging the compact latent space of PartVAE, we train an efficient downstream generative model that jointly synthesizes all parts within this unified representation, enabling the decoding of all components in a single forward pass. Qualitative results are shown in Figure 5. Compared to prior methods, our RL alignment enables PartMat to produce sharper material boundaries and more accurate structural correspondences between the reference image and the 3D output. Meanwhile, the geometry refinement network further enhances fine geometric details. Throughout this process, PartMat natively exposes editable component IDs for downstream material assignment. These visual advantages are quantitatively corroborated in Table 1. Configuration CD(×104×\!10^4) ↓ F1@0.01 ↑ Local-box 3.03 71.8 Packed-sem 2.87 72.4 Global multi-channel (Ours) 2.53 74.3 Table 3: VAE architecture ablation for improving geometry. All variants use one global latent and up to K=32K=32 material-component channels. Best results are marked in bold font. Ablation Studies To rigorously validate our architectural choices, we conduct comprehensive ablations, deferring further algorithmic and implementation details to the supplementary material. First, we evaluate our VAE representation paradigm against two alternatives, with quantitative and qualitative results detailed in Table 3 and Figure 7. The Local-box approach decodes component SDFs within normalized local spaces and explicitly predicts their global assembly poses. However, this coupled pose-geometry optimization introduces severe training difficulty, particularly struggling to reconstruct thin structures. Alternatively, the Packed-sem strategy compresses components into a few spatially non-overlapping groups and relies on an auxiliary semantic classifier to disentangle them. This secondary classification frequently produces jagged material boundaries and destroys the continuous latent space strictly required by the downstream PartDiT. In contrast, our global multi-channel strategy natively models seamless junctions and guarantees strict material disentanglement without secondary classification noise. Next, we further explore the specific contributions of our generative and refinement stages via the ablation experiments detailed in Table 1. As visualized in Figure 6, our direct-gradient RL alignment explicitly penalizes overlapping regions, effectively mitigating the blurred boundaries and spatial conflicts caused by purely supervised flow-matching. This facilitates more precise delineations between material regions. Finally, our sparse latent flow refinement restores high-frequency details lost in the compact global latent. This recovers sharp edges and micro structures crucial for photorealistic downstream applications, as visualized in Figure 5. Figure 7: Qualitative comparison of VAE representation paradigms. Unlike the missing thin structures in Local-box and jagged boundaries in Packed-sem, our Global multi-channel strategy achieves complete geometry and strict material disentanglement. Conclusion We proposed PartMat, an efficient framework for material-aware 3D part generation. Integrating a global PartVAE, an RL-aligned PartDiT for overlap suppression, and a sparse flow refiner, PartMat significantly outperforms existing baselines. It uniquely maintains constant decoding latency regardless of the part count while natively generating disentangled, high-fidelity parts for seamless downstream PBR material workflows. Limitations. PartMat assumes a fixed component capacity (K=32K=32) and relies on high-quality material annotations, where weakly supervised alternatives may introduce boundary noise. Reconstructing tiny details remains challenging, requiring stronger refinement in future work. References A. Abdelreheem, I. Skorokhodov, M. Ovsjanikov, and P. Wonka (2023) SATR: zero-shot semantic segmentation of 3d shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 15166–15179. Cited by: Part Segmentation. M. Ahmed, J. Fei, J. Ding, E. M. Bakr, and M. Elhoseiny (2025) Kestrel: 3d multimodal llm for part-aware grounded description. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 8973–8983. Cited by: Part Segmentation. M. Chen, R. Shapovalov, I. Laina, T. Monnier, J. Wang, D. Novotny, and A. Vedaldi (2025a) Partgen: part-level 3d generation and reconstruction with multi-view diffusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 5881–5892. Cited by: Part-Level Shape Generation. M. Chen, J. Wang, R. Shapovalov, T. Monnier, H. Jung, D. Wang, R. Ranjan, I. Laina, and A. Vedaldi (2025b) AutoPartGen: autoregressive 3d part generation and discovery. In Advances in Neural Information Processing Systems, Cited by: Part-Level Shape Generation. R. Chen, J. Zhang, Y. Liang, G. Luo, W. Li, J. Liu, X. Li, X. Long, J. Feng, and P. Tan (2025c) Dora: sampling and benchmarking for 3d shape variational auto-encoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 16251–16261. Cited by: 3D Shape Representation. Z. Chen, K. Yin, M. Fisher, S. Chaudhuri, and H. Zhang (2019) BAE-net: branched autoencoder for shape co-segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 8490–8499. Cited by: Part Segmentation. L. Ding, S. Dong, Y. Li, C. Gao, X. Chen, R. Han, Y. Kuang, H. Zhang, B. Huang, Z. Huang, et al. (2025) FullPart: generating each 3d part at full resolution. arXiv preprint arXiv:2510.26140. Cited by: Introduction, Part-Level Shape Generation. A. Golovinskiy and T. Funkhouser (2008) Randomized cuts for 3d mesh analysis. In ACM SIGGRAPH Asia 2008 papers, p. 1–12. Cited by: Part Segmentation. X. He, Y. Wu, X. Guo, C. Ye, J. Zhou, T. Hu, X. Han, and D. Du (2026) Unipart: part-level 3d generation with unified 3d geom-seg latents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 34227–34236. Cited by: 3D Shape Representation. Y. Hong, K. Zhang, J. Gu, S. Bi, Y. Zhou, D. Liu, F. Liu, K. Sunkavalli, T. Bui, and H. Tan (2024) Lrm: large reconstruction model for single image to 3d. In International Conference on Learning Representations, Vol. 2024, p. 50678–50702. Cited by: Object-Level Shape Generation. T. Hunyuan3D, S. Yang, M. Yang, Y. Feng, X. Huang, S. Zhang, Z. He, D. Luo, H. Liu, Y. Zhao, et al. (2025) Hunyuan3d 2.1: from images to high-fidelity 3d assets with production-ready pbr material. arXiv preprint arXiv:2506.15442. Cited by: Stage I: PartVAE, Stage I: PartDiT with RL Alignment, Experimental Setup, Experimental Setup. A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, et al. (2023) Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 4015–4026. Cited by: Part Segmentation. Z. Lai, Y. Zhao, Z. Zhao, H. Liu, Q. Lin, J. Huang, C. Guo, and X. Yue (2026) Lattice: democratize high-fidelity 3d generation at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19982–19992. Cited by: 3D Shape Representation, Object-Level Shape Generation. J. Li, K. Xu, S. Chaudhuri, E. Yumer, H. Zhang, and L. Guibas (2017) GRASS: generative recursive autoencoders for shape structures. ACM Transactions on Graphics 36 (4). Cited by: Part-Level Shape Generation. W. Li, A. Toisoul, T. Monnier, R. Shapovalov, R. Ranjan, P. Tan, and A. Vedaldi (2026) MeshFlow: efficient artistic mesh generation via MeshVAE and flow-based diffusion transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Cited by: Object-Level Shape Generation. Y. Li, Z. Zou, Z. Liu, D. Wang, Y. Liang, Z. Yu, X. Liu, Y. Guo, D. Liang, W. Ouyang, et al. (2025) Triposg: high-fidelity 3d shape synthesis using large-scale rectified flow models. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: Object-Level Shape Generation. Z. Liang, T. Yang, J. Wu, C. Feng, and L. Zheng (2026) LeapAlign: post-training flow matching models at any generation step by building two-step trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 23238–23248. Cited by: Stage I: PartDiT with RL Alignment. Y. Lin, C. Lin, P. Pan, H. Yan, F. Yiqiang, Y. Mu, and K. Fragkiadaki (2025) Partcrafter: structured 3d mesh generation via compositional latent diffusion transformers. Advances in neural information processing systems 38, p. 35387–35415. Cited by: Introduction, Part-Level Shape Generation, Experimental Setup, Table 1, Table 2. A. Liu, C. Lin, Y. Liu, X. Long, Z. Dou, H. Guo, P. Luo, and W. Wang (2024) Part123: part-aware 3d reconstruction from a single-view image. In ACM SIGGRAPH 2024 Conference Papers, p. 1–12. Cited by: Part-Level Shape Generation. J. Liu, G. Liu, J. Liang, Y. Li, J. Liu, X. Wang, P. Wan, D. Zhang, and W. Ouyang (2026) Flow-grpo: training flow matching models via online rl. Advances in neural information processing systems 38, p. 40783–40818. Cited by: Stage I: PartDiT with RL Alignment. M. Liu, M. A. Uy, D. Xiang, H. Su, S. Fidler, N. Sharp, and J. Gao (2025) Partfield: learning 3d feature fields for part segmentation and beyond. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 9704–9715. Cited by: Part Segmentation. M. Liu, Y. Zhu, H. Cai, S. Han, Z. Ling, F. Porikli, and H. Su (2023) PartSLIP: low-shot part segmentation for 3d point clouds via pretrained image-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 21736–21746. Cited by: Part Segmentation. H. Long, T. Zhao, J. Lin, Y. Zhang, H. Guo, R. Liang, J. Xu, J. Hladký, M. Nießner, Y. Hu, and W. Yang (2026) LATO.2: factorized 3d mesh generation with vertex and topology flow. arXiv preprint arXiv:2607.10623. Cited by: Object-Level Shape Generation. W. E. Lorensen and H. E. Cline (1987) Marching cubes: a high resolution 3d surface construction algorithm. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), p. 163–169. Cited by: Stage I: PartVAE. C. Ma, Y. Li, X. Yan, J. Xu, Y. Yang, C. Wang, Z. Zhao, Y. Guo, Z. Chen, and C. Guo (2025) P3-sam: native 3d part segmentation. arXiv preprint arXiv:2509.06784. Cited by: Part Segmentation. K. Mo, S. Zhu, A. X. Chang, L. Yi, S. Tripathi, L. J. Guibas, and H. Su (2019) Partnet: a large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 909–918. Cited by: Part-Level Shape Generation. M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. (2024) DINOv2: learning robust visual features without supervision. Transactions on Machine Learning Research. Cited by: Stage I: PartDiT with RL Alignment. C. R. Qi, H. Su, K. Mo, and L. J. Guibas (2017a) Pointnet: deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 652–660. Cited by: Part Segmentation. C. R. Qi, L. Yi, H. Su, and L. J. Guibas (2017b) Pointnet++: deep hierarchical feature learning on point sets in a metric space. In Advances in neural information processing systems, Vol. 30. Cited by: Part Segmentation. X. Ren, J. Huang, X. Zeng, K. Museth, S. Fidler, and F. Williams (2024) Xcube: large-scale 3d generative modeling using sparse voxel hierarchies. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 4209–4219. Cited by: 3D Shape Representation. L. Shapira, A. Shamir, and D. Cohen-Or (2008) Consistent mesh partitioning and skeletonisation using the shape diameter function. The Visual Computer 24, p. 249–259. Cited by: Part Segmentation. T. Shen, Z. Li, M. Law, M. Atzmon, S. Fidler, J. Lucas, J. Gao, and N. Sharp (2024) SpaceMesh: a continuous representation for learning manifold surface meshes. In SIGGRAPH Asia 2024 Conference Papers, p. 1–11. Cited by: Object-Level Shape Generation. G. Tang, W. Zhao, L. Ford, D. Benhaim, and P. Zhang (2024) Segment any mesh: zero-shot mesh part segmentation via lifting segment anything 2 to 3d. arXiv preprint arXiv:2408.13679. Cited by: Part Segmentation. J. Tang, R. Lu, M. Li, Z. Hao, X. Li, F. Wei, S. Song, G. Zeng, M. Liu, and T. Lin (2025) Efficient part-level 3d object generation via dual volume packing. Advances in Neural Information Processing Systems 38, p. 27115–27137. Cited by: 3D Shape Representation, Experimental Setup, Table 1, Table 2. A. Thai, W. Wang, H. Tang, S. Stojanov, J. M. Rehg, and M. Feiszli (2024) 3× 2: 3d object part segmentation by 2d semantic correspondences. In European Conference on Computer Vision, p. 149–166. Cited by: Part Segmentation. S. Tulsiani, H. Su, L. J. Guibas, A. A. Efros, and J. Malik (2017) Learning shape abstractions by assembling volumetric primitives. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 2635–2643. Cited by: Part-Level Shape Generation. A. Umam, C. Yang, M. Chen, J. Chuang, and Y. Lin (2024) PartDistill: 3d shape part segmentation by vision-language model distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 3470–3479. Cited by: Part Segmentation. H. Wang, Y. Liu, Y. Guo, Q. Feng, Z. Zou, D. Liang, B. Zhang, and Y. Cao (2026) Nexus: native mesh generation with diffusion. ACM Transactions on Graphics (TOG) 45 (4), p. 1–14. Cited by: Object-Level Shape Generation. Y. Wang, Y. Sun, Z. Liu, S. E. Sarma, M. M. Bronstein, and J. M. Solomon (2019) Dynamic graph cnn for learning on point clouds. ACM Transactions on Graphics (tog) 38 (5), p. 1–12. Cited by: Part Segmentation. S. Wu, Y. Lin, F. Zhang, Y. Zeng, Y. Yang, Y. Bao, J. Qian, S. Zhu, X. Cao, P. H. S. Torr, and Y. Yao (2025) Direct3d-s2: gigascale 3d generation made easy with spatial sparse attention. Advances in Neural Information Processing Systems 38, p. 170778–170804. Cited by: 3D Shape Representation, Object-Level Shape Generation. J. Xiang, X. Chen, S. Xu, R. Wang, Z. Lv, Y. Deng, H. Zhu, Y. Dong, H. Zhao, N. J. Yuan, et al. (2026) Native and compact structured latents for 3d generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14419–14429. Cited by: 3D Shape Representation, Object-Level Shape Generation, Experimental Setup, Experimental Setup, Table 2. J. Xiang, Z. Lv, S. Xu, Y. Deng, R. Wang, B. Zhang, D. Chen, X. Tong, and J. Yang (2025) Structured 3d latents for scalable and versatile 3d generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 21469–21480. Cited by: Object-Level Shape Generation, Stage I: Part-Aware Sparse Geometry Refinement. X. Yan, J. Xu, Y. Li, C. Ma, Y. Yang, C. Wang, Z. Zhao, Z. Lai, Y. Zhao, Z. Chen, et al. (2026) X-part: high fidelity and structure coherent shape decomposition and completion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 27062–27071. Cited by: Introduction, Part-Level Shape Generation, Experimental Setup, Table 1. Y. Yang, Y. Guo, Y. Huang, Z. Zou, Z. Yu, Y. Li, Y. Cao, and X. Liu (2025a) HoloPart: generative 3d part amodal segmentation. arXiv preprint arXiv:2504.07943. Cited by: Introduction, Part-Level Shape Generation, Experimental Setup, Table 1. Y. Yang, Y. Huang, Y. Guo, L. Lu, X. Wu, E. Y. Lam, Y. Cao, and X. Liu (2024) SAMPart3D: segment any part in 3d objects. arXiv preprint arXiv:2411.07184. Cited by: Part Segmentation. Y. Yang, Y. Zhou, Y. Guo, Z. Zou, Y. Huang, Y. Liu, H. Xu, D. Liang, Y. Cao, and X. Liu (2025b) Omnipart: part-aware 3d generation with semantic decoupling and structural cohesion. In Proceedings of the SIGGRAPH Asia 2025 Conference Papers, p. 1–12. Cited by: Experimental Setup, Table 1. B. Zhang, J. Tang, M. Niessner, and P. Wonka (2023) 3dshape2vecset: a 3d shape representation for neural fields and generative diffusion models. ACM Transactions On Graphics (TOG) 42 (4), p. 1–16. Cited by: 3D Shape Representation. L. Zhang, Q. Zhang, H. Jiang, Y. Bai, W. Yang, L. Xu, and J. Yu (2025) BANG: dividing 3d assets via generative exploded dynamics. ACM Transactions on Graphics (TOG) 44 (4), p. 1–21. Cited by: Part-Level Shape Generation. H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun (2021) Point transformer. In Proceedings of the IEEE/CVF international conference on computer vision, p. 16259–16268. Cited by: Part Segmentation. Z. Zhao, Z. Lai, Q. Lin, Y. Zhao, H. Liu, S. Yang, Y. Feng, M. Yang, S. Zhang, X. Yang, et al. (2025) Hunyuan3d 2.0: scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202. Cited by: Object-Level Shape Generation. Z. Zhong, Y. Xu, J. Li, J. Xu, Z. Li, C. Yu, and S. Gao (2024) MeshSegmenter: zero-shot mesh semantic segmentation via texture synthesis. In European Conference on Computer Vision, p. 182–199. Cited by: Part Segmentation.