Paper deep dive
Interact3D: Compositional 3D Generation of Interactive Objects
Hui Shan, Keyang Luo, Ming Li, Sizhe Zheng, Yanwei Fu, Zhen Chen, Xiangru Huang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/22/2026, 5:29:52 AM
Summary
Interact3D is a novel framework for generating physically plausible, interactive 3D compositional objects from single images and text prompts. It employs a two-stage pipeline: global-to-local geometric registration for anchoring primary objects, and SDF-based optimization for collision-aware integration of subsequent components. The system further incorporates a VLM-based agentic refinement loop to iteratively correct geometric mismatches, producing high-fidelity, collision-aware 3D scenes.
Entities (5)
Relation Signals (3)
Interact3D â employs â PartField
confidence 95% ¡ To recover the spatial relationship between components, we apply PartField to segment M scene into two parts
Interact3D â integrates â VLM
confidence 95% ¡ we integrate an agentic, VLM-driven refinement loop to iteratively edit the scene
Interact3D â utilizes â TRELLIS2
confidence 95% ¡ Subsequently, we use TRELLIS2 to generate M scene and M comp meshes separately based on images I scene and I comp.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional objects from single images--particularly under occlusions--remains challenging. Existing methods often degrade geometric details in hidden regions and fail to preserve the underlying object-object spatial relationships (OOR). We present a novel framework Interact3D designed to generate physically plausible interacting 3D compositional objects. Our approach first leverages advanced generative priors to curate high-quality individual assets with a unified 3D guidance scene. To physically compose these assets, we then introduce a robust two-stage composition pipeline. Based on the 3D guidance scene, the primary object is anchored through precise global-to-local geometric alignment (registration), while subsequent geometries are integrated using a differentiable Signed Distance Field (SDF)-based optimization that explicitly penalizes geometry intersections. To reduce challenging collisions, we further deploy a closed-loop, agentic refinement strategy. A Vision-Language Model (VLM) autonomously analyzes multi-view renderings of the composed scene, formulates targeted corrective prompts, and guides an image editing module to iteratively self-correct the generation pipeline. Extensive experiments demonstrate that Interact3D successfully produces promising collsion-aware compositions with improved geometric fidelity and consistent spatial relationships.
Tags
Links
- Source: https://arxiv.org/abs/2603.16085v1
- Canonical: https://arxiv.org/abs/2603.16085v1
Trouble viewing inline? Open PDF directly â
Full Text
54,565 characters extracted from source content.
Expand or collapse full text
Interact3D: Compositional 3D Generation of Interactive Objects Hui Shan 1,2,3 , Keyang Luo, Ming Li 1,2,3 , Sizhe Zheng 2,3 , Yanwei Fu 2,5 , Zhen Chen 4 , and Xiangru Huang 3â 1 Zhejiang University 2 Shanghai Innovation Institute 3 Westlake University 4 Adobe 5 Fudan University huangxiangru@westlake.edu.cn âPlace this basketball flat on a blue, hollow triangular stand.â User-given Mesh input Complementary Mesh 3D Interactive Scene output Interact3DInteract3D Dataset Fig. 1. Given a text prompt and a user-given mesh, Interact3D synthesizes a high- quality, geometrically-compatible complementary mesh. It then seamlessly composes these two assets into an interactive 3D scene. Unlike existing approaches, our framework automatically generates collision-aware and physically-sound 3D environments. Abstract. Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D com- positional objects from single imagesâparticularly under occlusionsâremains challenging. Existing methods often degrade geometric details in hidden regions and fail to preserve the underlying object-object spatial rela- tionships (OOR). We present a novel framework Interact3D designed â Corresponding author. arXiv:2603.16085v1 [cs.CV] 17 Mar 2026 2H. Shan et al. to generate physically plausible interacting 3D compositional objects. Our approach first leverages advanced generative priors to curate high- quality individual assets with a unified 3D guidance scene. To physically compose these assets, we then introduce a robust two-stage composi- tion pipeline. Based on the 3D guidance scene, the primary object is anchored through precise global-to-local geometric alignment (registra- tion), while subsequent geometries are integrated using a differentiable Signed Distance Field (SDF)-based optimization that explicitly penal- izes geometry intersections. To reduce challenging collisions, we further deploy a closed-loop, agentic refinement strategy. A Vision-Language Model (VLM) autonomously analyzes multi-view renderings of the com- posed scene, formulates targeted corrective prompts, and guides an im- age editing module to iteratively self-correct the generation pipeline. Extensive experiments demonstrate that Interact3D successfully pro- duces promising collsion-aware compositions with improved geometric fidelity and consistent spatial relationships. The code will be released on https://github.com/SII-Hui/Interact3D. Keywords: 3D Generation¡ Interaction¡ Optimization 1 Introduction Robotic manipulation shows immense potential in domestic assistance, indus- trial automation, and disaster response. In recent years, learning-based manip- ulation paradigms, particularly Reinforcement Learning (RL) [18,39,30], have achieved remarkable progress. However, deploying these algorithms directly in the real world is expensive and unsafe. Therefore, training agents in physically simulated virtual interactive environments (Sim2Real) has become a dominant paradigm [19,20,35]. To generalize to diverse real-world scenarios, these simula- tion platforms require a massive and diverse collection of interactive 3D assets with realistic geometry, plausible physical properties, and correct objectâobject relationships (OOR). However, the availability of high-quality interactive 3D data remains severely limited. Curating these 3D datasets heavily relies on tedious manual modeling and annotation, making it difficult to scale. To bypass the scarcity of interac- tive 3D assets, recent works attempt to leverage large-scale human videos as an alternative supervision signal for robot learning [11,9,41]. While video-driven approaches provide rich semantic priors and action trajectories, 2D videos in- herently lack precise 3D geometric grounding and physical collision constraints. This fundamental limitation hinders the robotâs ability to comprehend complex spatial relationships and perform fine-grained interactions in 3D space. Concurrently, 3D generation models have achieved impressive results in syn- thesizing high-fidelity 3D assets from text or image prompts [33,44,32]. How- ever, these models typically output monolithic and fused geometries which lack OOR. The generated meshes are âbakedâ into a single geometry, severely lacking independent geometry and physically meaningful spatial relationship. Naively Interact3D: Compositional 3D Generation of Interactive Objects3 Generated Scene (3D Guidance) Individual Meshes byPartField Individual Meshes & 3D Compositional Scene byInteract3D 2D Compositional Image Prompt output output output output input input Fig. 2. Given a generated scene, PartField directly segments it to extract individual 3D assets, whereas Interact3D leverages it as 3D spatial guidance to compose high- quality independent geometries. (Though generating scenes and meshes via separate TRELLIS2 inferences yields minor texture variations, their spatial guidance remains robust to these discrepancies.) segmenting such geometry, for instance, directly using PartField [14], often in- troduces difficult-to-repair holes and inaccurate 3D assets (see Figure 2). Hun- yuan3D [27] and Rodin [4] websites can segment scenes and inpaint geometries, but suffer from the loss of the original texture in their remeshed geometry. At- tempting to manually segment and optimize these monolithic assets is extremely labor-intensive and fails to generalize across diverse topologies. Therefore, the key challenge lies in how to reliably generate geometrically-compatible 3D com- positional assets guided by user-given text or image prompts. An intuitive solu- tion is to train a feed-forward network using large amounts of 3D compositional data. However, the current amount of data does not support such an approach, and the training cost is expensive. To address these challenges, we propose Interact3D, a novel pipeline for scalable and automated synthesis of interactive 3D scenes, see Figure 3. Our key insight is to leverage the spatial priors implicitly encoded in image-guided 3D generation as guidance for geometric composition. By fully exploiting the priors of advanced 2D and 3D generative models, we reformulate compositional gener- ationâtraditionally requiring complex geometric reasoningâas a structured 3D registration problem. However, unlike traditional 3D registration, compositional generation re- quires not only accurate alignment but also physically plausible interactions between objects. To this end, we first anchor the primary object using a global- to-local registration strategy, and then integrate additional components through an Signed Distance Field (SDF)-based optimization that explicitly penalizes in- terpenetrations, resulting in collsion-aware compositions. Moreover, inherent 2D occlusions often cause generated assets with complex joints to suffer from severe geometric mismatches, which cannot be fixed by simple rigid transformations. To address this, we propose a VLM-based agentic refinement strategy that formulates corrective text prompts to guide 2D gen- erative image editing, thereby iteratively fixing the geometric flaws. Extensive 4H. Shan et al. experiments demonstrate the effectiveness of our methods in multi-object com- position. Our contributions are summarized in the following: 1. We propose Interact3D, a generate-then-compose framework for training- free compositional 3D generation, featuring a two-stage, extensible compo- sition pipeline, that enables the creation of collsion-aware interactive 3D scenes. 2. We introduce an interactive 3D dataset, which contains more than 8,000 interactive pairs and will be publicly released to facilitate future research. 3. We develop a VLM-based agentic refinement strategy that automatically resolves severe geometric mismatches and improves compositional coherence. 2 Related Work 3D Generation. Recent advances in 3D generation have increasingly shifted away from slow, optimization-based paradigms [15,43,13,1,24] to large-scale foun- dational models [42,44,33,26,32,12], drastically improving both synthesis speed and geometric fidelity. CLAY [42] has advanced native 3D geometry generation, using 3D datasets from Objaverse [5] to produce highly controllable, realistic, and watertight assets. Hunyuan3D [44] achieves industrial-grade asset quality by ef- fectively unifying text-to-3D and image-to-3D pipelines within robust large-scale flow-based diffusion transformers. TRELLIS [33] introduces a unified structured latent representation (SLAT) and rectified flow transformers that significantly mitigate multi-view inconsistencies and artifacts in complex geometries. Build- ing upon this foundation, its recent successor, TRELLIS2 [32] further scales the transformer architecture within the 3D latent space, achieving state-of-the-art in high-resolution 3D generation. 3D Part Segmentation and Generation. The segmentation and generation of 3D parts, which are essential tasks in 3D generation, require a compositional understanding of the structures of 3D objects. SAMPart3D [36], PartField, and P 3 -SAM [17] capture 3D features through feed-forward training and perform segmentation. Although it can segment parts generated by advanced 3D gen- eration models, the resulting components often contain difficult-to-repair holes. In contrast, OmniPart [37] and SAM3D [2] are capable of directly generating complete 3D parts from masked images through expensive data curation and feed-forward training. Although their meshes generated are watertight, their ge- ometric accuracy, PBR materials, and alignment with image prompts still fall short when compared to 3D foundation models like TRELLIS2. 3D Shape Composition. 3D shape composition [34,40] typically restores complete objects or scenes from multiple segmented components. Jigsaw [16], NSM [3] and RGL-NET [21] focus on fragment composition without relying on predefined semantic information, as fragments often contain complex and precise geometric details. Based on these, Puzzlefusion [10], Diffassemble [25], and Puzzlefusion++ [29] make a trial to use diffusion model to refine the poses Interact3D: Compositional 3D Generation of Interactive Objects5 of fragments. 2BY2 [22] is the first to propose daily pairwise object composi- tion task and contribute the corresponding dataset for the task. Compared to fragment composition, daily pairwise components always lack clear geometric relationships. Therefore, COPY-TRANSFORM-PASTE [7] utilizes text-guided optimization with vision-language models to supervise composition and achieves promising results. Although 2BY2 introduces a daily composition dataset, it con- sists of only 517 pairs, most of which are selected from existing 3D datasets [5,31]. In contrast, we propose a fully automated pipeline for generating large-scale brand-new 3D composition datasets. 3 Preliminary Our method is designed for physically-sound 3D compositional generation. In this section, we first formally define the task of 3D compositional generation (section 3.1), then briefly formulate 3D point clouds registration task and intro- duce one classical solution, Iterative Closest Point (ICP) (section 3.2). 3.1 Problem Setup Given a 3D mesh M = (V, F), where V âR NĂ3 denotes the vertex positions and F âN N f Ă3 denotes the face indices, together with a text prompt that specifies compositional semantics, our goal is to generate a complementary 3D component M comp that forms a coherent compositional scene with M. In addition to synthesizing M comp , we optimize the transformation param- eters θ = (Ď, R, s) between both meshes to achieve geometrically compatible composition, where Ď âR 3 denotes translation, R â SO(3) denotes rotation, and sâR + denotes uniform scaling. 3.2 3D Point Clouds Registration Given source and target point clouds P = p i N i=1 and Q = q j M j=1 , we adopt scale-aware ICP [38] for similarity registration. The objective jointly estimates a uniform scale sâR + , rotation RâSO(3), and translation Ď âR 3 : min s,R,Ď X (p i ,q j )âC âĽs¡ R¡ p i +Ď â q j ⼠2 2 ,(1) where C denotes point correspondences. At each iteration, correspondences are updated using the scheme mentioned in [38], and the optimal similarity trans- formation (s, R,Ď ) is computed in closed form using Umeyamaâs method [28]. 4 Method We present a comprehensive framework for generating physically-sound 3D shape components. After curating high-quality individual assets and a 3D guidance 6H. Shan et al. scene, we process them through a novel two-stage composition pipeline (Sec- tion 4.1). A global-to-local geometric alignment (Section 4.2) is used to accu- rately anchor the reference object in stage one. A SDF-based physics-aware op- timization (Section 4.3) is introduced to align subsequent objects while avoiding spatial intersections in stage two. For cases of severe and unavoidable collisions, we integrate an agentic, VLM-driven refinement loop (Section 4.4) to iteratively edit the scene until a low-collsion configuration is achieved. The main notations are included in Table 1. Table 1. Summary of main notations. SymbolDescription M,I renderd Input 3D mesh, and corresponding rendered image I scene ,I comp Generated compositional and complementary image M scene , M comp Scene and complementary mesh reconstructed from I scene and I comp M anchor , M remain Anchor mesh and remaining mesh selected from M, M comp M Ⲡ, M Ⲡcomp Segmented meshes extracted from M scene M Ⲡanchor , M Ⲡremain Segmented counterparts from M scene θ(p) = s¡ Rp +Ď Transformation function in terms of translationĎ , rotation R, scale s) ÎŚ M (p)Signed distance from point p to mesh M 4.1 Data Curation and Workflow We begin by rendering the input mesh M from a canonical frontal viewpoint, obtaining an image I rendered . Conditioning on this rendered image and a user- provided textual prompt, we use Nano Banana Pro [8] to generate a composi- tional scene image I scene . To obtain the complementary component, we further prompt the model to remove the original object from I scene , producing an im- age I comp that contains only the newly introduced geometry. Intuitively, I comp captures the complementary part needed to complete the composition. Subsequently, we use TRELLIS2 to generate M scene and M comp meshes sep- arately based on images I scene and I comp . To recover the spatial relationship between components, we apply PartField to segment M scene into two parts: M Ⲡand M Ⲡcomp . Although these segmented meshes often contain holes and geomet- ric artifacts that make them unsuitable as final assets, they still provide reliable spatial cues. In particular, the transformation θ extracted from M Ⲡand M Ⲡcomp serve as geometric guidance for composing the original mesh M and the high- quality complementary mesh M comp through 3D point cloud registration (see Figure 3). Two-stage Composition Pipeline. We decompose composition into two se- quential stages with distinct roles. Interact3D: Compositional 3D Generation of Interactive Objects7 Image I rendered rendered from M Image I scene Image I comp Nano Banana Pro Nano Banana Pro âPlace this pink back pillow on a wooden chair. The picture should show the complete object geometry and realistic proportions. White background.â âRemove the pink back pillow in the picture and complete the occupied part of the content. Other parts remain unchanged. White background.â TRELLIS2 TRELLIS2 Mesh M scene Render PartField Mesh M comp User-given mesh M Mesh Mâ comp Mesh Mâ 3D Spatial Guidance Stage 1: Global-to-local Accurate Geometric Alignment Stage 2: SDF-based Collision-aware Optimization 3D Compositional Scene Agentic Optimization Individual 3D Asset 1 2 2 3 3 4 5 6 7 Fig. 3. Overview of Interact3D. Given a user-provided mesh M, we render it and use Nano Banana Pro to synthesize a guided scene image I scene and a comple- mentary image I comp . TRELLIS2 then reconstructs these into 3D meshes (M scene and M comp ) to provide spatial guidance. Relying on coarse parts segmented by PartField, we execute a two-stage composition. Stage 1 performs a global-to-local registration on the âlargestâ mesh to establish the anchor pose, while Stage 2 applies an SDF-based collision-aware optimization on the remaining mesh to resolve spatial intersections. (In this case, M comp is considered as M anchor , M is considered as M remain .) Finally, a VLM-based agentic refinement handles unavoidable collisions, yielding a physically- sound interactive 3D scene. Stage 1: Anchor Alignment. We first select an anchor object M anchor from the candidate meshes (M or M comp ). Empirically, the object with the larger pro- jected area (measured by 2D bounding box size) in I scene provides a more stable reference. We apply the global-to-local registration strategy (Section 4.2) exclu- sively to this anchor object to recover an accurate initial pose with respect to the guidance mesh M Ⲡanchor . Stage 2: Collision-aware Composition. Once the anchor object is fixed, we op- timize the pose of the remaining component(s) M remain using the SDF-based collision-aware formulation (Section 4.3) with the guidance M Ⲡremain . This stage explicitly balances geometric alignment with collision avoidance, ensuring phys- ically plausible composition. 4.2 Global-to-local Accurate Geometric Alignment This stage is applied only to the selected anchor object mentioned in Section 4.1. Its objective is to recover a reliable initial pose under partial overlop, occulsion, and reconstruction noise. 8H. Shan et al. In practice, the generated compositional image I scene inevitably introduces occlusions, and the individual meshes are not perfectly consistent with M scene . Moreover, the subsequent segmentation process in PartField often produces structural holes and incomplete surfaces (see Figure 2). As a result, the resulting point cloud pairs (M anchor and M Ⲡanchor ) typically exhibit low-overlap ratios and contain a significant number of outlier points. Under such conditions, traditional local registration methods, such as ICP, are highly sensitive to initialization and may converge to undesirable local minima. To effectively address the aforementioned challenges, we adopt a robust and accurate global-to-local registration paradigm. We begin by resolving the scale discrepancies using an Oriented Bounding Box (OBB) approach to estimate the geometric extents and compute the initial scale factor s. Following this scale alignment, we employ GeoTransformer [23], a transformer-based 3D registration network, which is exceptionally robust in handling geometries with low-ratio overlaps. This step provides reliable global estimates of the initial global trans- lation Ď and rotation R. Finally, with these global estimates of (s, R,Ď ) as initialization, a scale-aware ICP algorithm is deployed to achieve precise align- ment. In summary, this global initialization provides a warm start that prevents the algorithm from falling into local minima, allowing the subsequent local opti- mization to focus on minimizing residual geometric discrepancies for an accurate final alignment. 4.3 SDF-Based Collision-Aware Composition After global-to-local registration, the primary object is fixed as the spatial an- chor. The remaining components M remain are then placed relative to this anchor. While we can use the similar global-to-local alignment method for the rest com- ponents to provide a reasonable pose estimate, it does not guarantee physically valid placement. Small geometric discrepancies may lead to interpenetration or unnatural spacing between objects. To refine the placement, we introduce an SDF-guided optimization stage. Given the anchor mesh M anchor , we precompute its Signed Distance Field (SDF) ÎŚ anchor (p), which encodes the signed distance from a query point p to the surface of M anchor and negative values indicate penetration into the anchor object. To start with, followed by our global-to-local approach, we first get the initial transformation θ between M remain and M Ⲡremain using OBB and GeoTransformer as a starting point. For any point p in M remain , we define the collision loss L col using the ReLU operator [¡] + after transformation θ: L col (θ) = X pâM remain   ďŁ [âÎŚ anchor (θ(p))] 2 + | z Hard Penalty +Ν¡ [Îľâ ÎŚ anchor ((θ(p))] + |z Soft Repulsion    .(2) Here, Îľ denotes the safety margin and Îť is set to a minimal value to ensure a smooth transition of the penalty field near the boundary of the object. The final Interact3D: Compositional 3D Generation of Interactive Objects9 placement is obtained by minimizing: min θ X pâM remain âĽÎ¸(p)â p Ⲡ⼠2 + β (k) L col (θ),(3) where the first term preserves the alignment between M remain and M Ⲡremain , and β (k) balances geometric consistency and collision avoidance. Specifically, β (k) is initialized as 0, so that it is equivalent to the scale-aware ICP solver at the beginning and increases linearly to β max throughout the optimization process. At the k-th optimization stage, p Ⲡis first chosen from M Ⲡremain for each θ(p) following [38], and then optimize Equation (3) with p Ⲡfixed. This schedule allows the network to prioritize global geometric alignment in the early stages and progressively enforce physical plausibility as the pose converges. By treating the anchor as a fixed spatial reference and refining the remaining components through SDF-guided optimization, our method generates geometri- cally compatible and physically plausible composition. 4.4 Agentic Optimization Although SDF-based optimization effectively resolves moderate geometric dis- crepancies, it cannot correct fundamentally incompatible geometries (See Fig- ure 9). In practice, the complementary mesh generated from 2D guidance may deviate significantly from the intended spatial configuration, leading to persis- tent collisions or implausible object relationships that cannot be resolved through transformation alone. To address this limitation, we introduce an agentic refinement mechanism that operates at the semantic level. Instead of further adjusting pose parame- ters, we revise the complementary geometry itself. Specifically, we render the current compositional scene from multiple viewpoints and feed the rendered images, along with the original compositional prompt, into a Vision-Language Model (VLM, e.g., Gemini 3 Pro [8]). The VLM analyzes spatial inconsistencies and generates corrective feedback in the form of targeted editing instructions. These instructions are then used to guide the 2D image editing model to refine the complementary component. The updated 2D image is subsequently recon- structed into a new 3D mesh and re-integrated into the composition pipeline. This closed-loop process enables semantic-level correction beyond purely geo- metric optimization. By iteratively combining geometric alignment and language- guided refinement, our framework can resolve severe mismatches that are oth- erwise difficult to handle with traditional registration or SDF-based methods alone. 5 Experiments 5.1 Experimental Setup We use Nano Banana Pro to generate 4K resolution images and adopt TRELLIS2 (and partially Hunyuan3D) for image-conditioned 3D Generation. Our proposed 10H. Shan et al. âPlace flowers into this classically carved vase. The picture should show the complete object geometry and realistic proportions. White background.â Mesh M scene (3D Guidance) User-given mesh M âRemove the carved vase in the picture and complete the occupied part of the content. Other parts remain unchanged. White background.â Mesh M comp Spatial logic anomaly: The image contains obvious errors in composition or rendering. The rose is not inserted inside the vase, but the entire bouquet is suspended to the left rear of the vase opening. âTrim the foliage of the flowers, and cut the stem short. The rest should remain unchanged. White background.â Mesh M comp The lower stems and leaves exceed the vase's diameter, so the modification aims to bring them closer to the central axis. âShorten the flower branches, and fold the flower branches inward a little. White background.â Mesh M comp Agentic Optimization Image I scene Image I comp Image I comp Image I comp 3D Compositional Scene 3D Compositional Scene input 3D Compositional Scene output output output output output output 1 2 3 4 Fig. 4. Agentic Refinement. When generating flowers inside a vase, severe occlusion in image I scene cause the lost of object-object spatial relationships during TRELLIS2 reconstruction. As a result, the flower stem in the complementary image I comp (gener- ated by Nano Banana Pro) may not align correctly with the vase mesh M. This leads to geometric intersections after composition (see top-right image). To resolve this, we ren- der muli-view images, including internal cross-sections. These renderings are analyzed by a VLM, which generates a corrective text prompt to update the complementary image via Nano Banana Pro. This process continues until no more geometric intersec- tions are found or the maximum iteration limit is reached. framework is entirely training-free. Specifically, we set Îť = 0.003 to smooth the penalty field in Equation (2), set k max = 100,β max = 3.0 in Equation (3) to ensure physical plausibility progressively. Furthermore, we set max iteration in agentic optimization to 5. Baselines. We evaluate the compositional performance of Interact3D against three categories of baselines. 1. Geometry-guided composition baselines. We compare with Jigsaw [16] (for fractured object assembly), and 2BY2 [22], (for daily pairwise object composition). Both methods take the user-provided mesh M and the TREL- LIS2 generated complementary mesh M comp as input, without additional image or text conditioning. For Jigsaw, we use the official pretrained check- points. As 2BY2 lacks released weights, we train it from scratch following their official protocol. Notably, neither baseline optimizes object scale s. Interact3D: Compositional 3D Generation of Interactive Objects11 âPlease place several pairs of shoes in the shoe cabinet. In the upper compartment, place two pairs: leather shoes on the left and high-heeled shoes on the right. In the lower compartment, put one pair of sneakers on the left side. The picture should show the complete object geometry and realistic proportions. White background.â User-given mesh M Mesh M scene Mesh M scene (3D Guidance) Mesh M scene (3D Guidance) Mesh M scene (3D Guidance) Image I scene âKeep the leather shoes on the left side of the upper compartment in shoe cabinet, and remove the other shoes and the shoe cabinet. White background.â Image I comp Mesh M comp âKeep the high-heeled shoes on the right side of the upper compartment in shoe cabinet, and remove the other shoes and the shoe cabinet. White background.â Image I comp PartField Mesh M comp âKeep the sneakers on the left side of the lower compartment in shoe cabinet, and remove the other shoes and the shoe cabinet. White background.â Image I comp Mesh M comp Interact3DInteract3DInteract3D 3D Compositional SceneFinal 3D Compositional Scene3D Compositional Scene input output output output output 1 2 3 3 4 4 2 2 5 5 Fig. 5. More than two parts composition results. Given a cabinet mesh, we render an image from it, add three pairs of shoes to the image using Nano Banana Pro and generate it with TRELLIS2 to obtain the object-object spatial relationships (OOR). Next, we use Nano Banana Pro again to extract each individual object in the shoe cabinet and generate them with TRELLIS2. Finally, using the OOR information, we sequentially add them to the cabinet mesh to form an interactive 3D scene . 2. Direct segmentation baseline. To evaluate the impact of our generate- then-compose design, we construct a baseline that directly segments the TRELLIS2-generated scene mesh M scene using PartField. This approach re- quires no extra pose estimation and purely relies on the scene-level segmenta- tion. However, the resulting meshes typically contain geometric artifacts and structural holes, which are hard to repair, and modify the user-provided as- sets. Comparing against this baseline highlights the advantage of our method in preserving high-fidelity individual geometries. 3. Classical registration baseline. To further demonstrate the effectiveness of our proposed registration approach (Section 4.2 and 4.3), we employ the RANSAC [6] algorithm to independently register the two inputs, M and M comp , based on M Ⲡand M Ⲡcomp , yielding the final composed result (Part- Field + RANSAC). This baseline isolates the contribution of our robust registration and collision-aware optimization design. 12H. Shan et al. 5.2 Qualitative Evaluation Two-part Composition. In Figure 6, we show the comparisons for two-part composition using different methods. Both Jigsaw and 2BY2 rely heavily on their training data distributions, which limits their generalization capabilities. As a result, their composed outputs often exhibit misalignment or implausible spatial configurations. Directly segmenting TRELLIS2-generated scene meshes using PartField leads to visible geometric artifacts and structural holes, since the generated geometries are typically fused. Additionally, the RANSAC algorithm struggles to achieve precise geometric alignment, often resulting in inverted poses (e.g., the pen is inverted in the pen holder, as observed in line 1,2). Finally, registering the two objects independently fails to account for collisions, leading to severe geometry intersections in the final composition (in line 3,4,5,6,9,10). More parts Composition. Our framework naturally extends to compose an arbitrary number of objects, enabling the generation of rich interactive scenes. Similarly to our standard pipeline, after the user provides an initial 3D mesh, we first render it into an image and use Nano Banana Pro to edit the rendered image and introduce additional objects in the image space. We then use TRELLIS2 to generate a corresponding 3D scene. This provides spatial relationship priors that guide the placement of all components. Our multi-part composition is achieved through sequential two-object composition. As illustrated in Figure 5, the user- provided mesh M is first treated as the anchor object M anchor . A complementary object (e.g., leather shoes) is generated and treated as the remaining component M remain , which is aligned to the anchor using our composition pipeline. The resulting composed scene is then treated as a new anchor, and another object (e.g., high-heeled shoes) is introduced as the remaining component. This process is repeated until all objects are incorporated into the final scene. Due to space limitation, more cases are provided in the Appendix. 5.3 Quantitative Evaluation We quantitatively evaluate composition results across 10 test cases based on two key aspects: the semantic fidelity, and the physical validity (see Table 2). Our test set consists of the 5 cases in Figure 6 and 5 additional cases in Figure 7. Due to space limitation, we show the generated guidance image I scene and our com- positional results of additional cases. All other qualitative results are provided in the Appendix. Compositional Semantic Fidelity. As PartField does not perform composi- tion, our quantitative comparison focuses on the other three baselines. We utilize CLIP to measure the semantic alignment of 2D renderings of the final compo- sitional 3D scene from pre-defined viewpoints, against both the user-given text prompt and the 2D image generated by Nano Banana Pro. Higher CLIP scores denote stronger semantic consistency (line 1,2 in Table 2). Interact3D: Compositional 3D Generation of Interactive Objects13 Jigsaw User-given Mesh / Generated Mesh 2BY2 PartField PartField + RANSAC Ours Generated Scene Fig. 6. Results of two-part composition. Data-driven baselines (Jigsaw, 2BY2) struggle to generalize to novel inputs. 3D segmentation of generated scenes (TREL- LIS2 + PartField) inherently produces severe geometric holes (circled regions). Further- more, âPartField + RANSACâ suffers from orientation inversion (1st and 2nd rows) and severe geometry intersections (circled region in the corresponding result column). Conversely, our framework consistently synthesizes physically plausible, and collision- aware components. 14H. Shan et al. Generated Image Ours Fig. 7. Additional cases in quantitative evaluation. We show the generated images by Nano Banana Pro and the compositional 3D scenes by Interact3D. Table 2. Quantitative comparison of 3D compositional generation. We eval- uate semantic alignment (text CLIP and image CLIP), and geometric intersection rate (surface and volume). â/â denotes metrics not applicable to the baseline. The best re- sults are highlighted in bold. Jigsaw2BY2PartFieldPartField+RANSACOurs Text CLIP (Avg) â0.30250.2780/0.31290.3307 Image CLIP (Avg) â0.74070.6905/0.80820.8248 R surface (Ă10 â3 ) â2.72781.5302/2.15230.6766 R volume (Ă10 â3 ) â14.5717.6195/6.77443.2467 Geometric Intersection Rate. To evaluate collision avoidance, we measure interpenetration from both surface and volumetric perspectives, where lower values indicate better physical compatibility. Surface Intersection Ratio. We compute the total surface area involved in triangleâtriangle intersections between meshes A and B. Let A int denote the area of intersected surface regions, and A A , A B the total surface areas of A and B. The surface intersection rate is defined as: R surface = A int A A +A B . (4) Volume Intersection Ratio. Since the compositional scene consists of solid objects represented by surface meshes, we estimate volumetric interpenetration via Monte Carlo sampling. Specifically, we uniformly sample N (= 10 6 ) points within a bounding box enclosing both meshes. The volumetric intersection ratio is defined as the fraction of sampled points that lie inside both objects among those that lie inside at least one object: R volume = # points inside both A and B # points inside A or B .(5) 6 Conclusion We present Interact3D, a training-free generate-then-compose framework for interactive 3D scene synthesis. By leveraging spatial guidance from generative Interact3D: Compositional 3D Generation of Interactive Objects15 models, our method reformulates compositional generation as a collision-aware geometric alignment problem. Combining global-to-local registration, SDF re- finement, and VLM-driven correction, our pipeline produces geometrically com- patible and physically plausible scenes. We also release a curated compositional 3D dataset to facilitate future research in structured scene generation and robotic simulation. Future Work. Although Interact3D generalizes well to daily objects, it strug- gles with fine-grained components (e.g., screws or tightly coupled joints). In these cases, severe occlusions in the 2D guidance images cause 3D geometric ambiguities, highlighting the inherent limitations of relying on 2D spatial priors for complex components. Future work will explore native 3D compositional gen- eration to reason directly in 3D space, bypassing 2D dependencies and robustly handling intricate structures. References 1. Chen, R., Chen, Y., Jiao, N., Jia, K.: Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 22246â22256 (2023) 4 2. Chen, X., Chu, F.J., Gleize, P., Liang, K.J., Sax, A., Tang, H., Wang, W., Guo, M., Hardin, T., Li, X., et al.: Sam 3d: 3dfy anything in images. arXiv preprint arXiv:2511.16624 (2025) 4 3. Chen, Y.C., Li, H., Turpin, D., Jacobson, A., Garg, A.: Neural shape mating: Self- supervised object assembly with adversarial shape priors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 12724â 12733 (2022) 4 4. Deemos Technology: Hyper3d.ai. https://hyper3d.ai/ (2026), accessed: 2026-02 3 5. Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., Farhadi, A.: Objaverse: A universe of annotated 3d objects. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 13142â13153 (2023) 4, 5 6. Fischler, M.A., Bolles, R.C.: Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communi- cations of the ACM 24(6), 381â395 (1981) 11 7. Gatenyo, R., Fried, O.: Copy-trasform-paste: Zero-shot object-object align- ment guided by vision-language and geometric constraints. arXiv preprint arXiv:2601.14207 (2026) 5 8. Google: Gemini. https://gemini.google.com/ (2026), accessed: 2026-02 6, 9 9. Hoque, R., Huang, P., Yoon, D.J., Sivapurapu, M., Zhang, J.: Egodex: Learn- ing dexterous manipulation from large-scale egocentric video. arXiv preprint arXiv:2505.11709 (2025) 2 10. Hossieni, S.S., Shabani, M.A., Irandoust, S., Furukawa, Y.: Puzzlefusion: Unleash- ing the power of diffusion models for spatial puzzle solving. Advances in Neural Information Processing Systems 36, 9574â9597 (2023) 4 16H. Shan et al. 11. Jain, V., Attarian, M., Joshi, N.J., Wahid, A., Driess, D., Vuong, Q., Sanketi, P.R., Sermanet, P., Welker, S., Chan, C., et al.: Vid2robot: End-to-end video-conditioned policy learning with cross-attention transformers. arXiv preprint arXiv:2403.12943 (2024) 2 12. Li, W., Zhang, X., Sun, Z., Qi, D., Li, H., Cheng, W., Cai, W., Wu, S., Liu, J., Wang, Z., et al.: Step1x-3d: Towards high-fidelity and controllable generation of textured 3d assets. arXiv preprint arXiv:2505.07747 (2025) 4 13. Lin, C.H., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M.Y., Lin, T.Y.: Magic3d: High-resolution text-to-3d content creation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition. p. 300â309 (2023) 4 14. Liu, M., Uy, M.A., Xiang, D., Su, H., Fidler, S., Sharp, N., Gao, J.: Partfield: Learning 3d feature fields for part segmentation and beyond. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 9704â9715 (2025) 3 15. Liu, Z., Li, Y., Lin, Y., Yu, X., Peng, S., Cao, Y.P., Qi, X., Huang, X., Liang, D., Ouyang, W.: Unidream: Unifying diffusion priors for relightable text-to-3d generation. In: European Conference on Computer Vision. p. 74â91. Springer (2024) 4 16. Lu, J., Sun, Y., Huang, Q.: Jigsaw: Learning to assemble multiple fractured objects. Advances in Neural Information Processing Systems 36, 14969â14986 (2023) 4, 10 17. Ma, C., Li, Y., Yan, X., Xu, J., Yang, Y., Wang, C., Zhao, Z., Guo, Y., Chen, Z., Guo, C.: P3-sam: Native 3d part segmentation. arXiv preprint arXiv:2509.06784 (2025) 4 18. Ma, Y.J., Liang, W., Wang, G., Huang, D.A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., Anandkumar, A.: Eureka: Human-level reward design via coding large language models. arXiv preprint arXiv:2310.12931 (2023) 2 19. Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., et al.: Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv preprint arXiv:2108.10470 (2021) 2 20. Narang, Y., Storey, K., Akinola, I., Macklin, M., Reist, P., Wawrzyniak, L., Guo, Y., Moravanszky, A., State, G., Lu, M., et al.: Factory: Fast contact for robotic assembly. arXiv preprint arXiv:2205.03532 (2022) 2 21. Narayan, A., Nagar, R., Raman, S.: Rgl-net: A recurrent graph learning framework for progressive part assembly. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. p. 78â87 (2022) 4 22. Qi, Y., Ju, Y., Wei, T., Chu, C., Wong, L.L., Xu, H.: Two by two: Learning multi- task pairwise objects assembly for generalizable robot manipulation. In: Proceed- ings of the Computer Vision and Pattern Recognition Conference. p. 17383â17393 (2025) 5, 10 23. Qin, Z., Yu, H., Wang, C., Guo, Y., Peng, Y., Ilic, S., Hu, D., Xu, K.: Geo- transformer: Fast and robust point cloud registration with geometric transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(8), 9806â9821 (2023) 8 24. Qiu, L., Chen, G., Gu, X., Zuo, Q., Xu, M., Wu, Y., Yuan, W., Dong, Z., Bo, L., Han, X.: Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to-3d. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 9914â9925 (2024) 4 Interact3D: Compositional 3D Generation of Interactive Objects17 25. Scarpellini, G., Fiorini, S., Giuliari, F., Moreiro, P., Del Bue, A.: Diffassemble: A unified graph-diffusion model for 2d and 3d reassembly. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 28098â 28108 (2024) 4 26. Shan, H., Li, M., Yang, H., Zheng, K., Zheng, S., Fu, Y., Huang, X.: Ni- tex: Non-isometric image-based garment texture generation. arXiv preprint arXiv:2511.18765 (2025) 4 27. Tencent: Tencent hunyuan3d. https://3d.hunyuan.tencent.com/ (2026), ac- cessed: 2026-02 3 28. Umeyama, S.: Least-squares estimation of transformation parameters between two point patterns. IEEE Transactions on pattern analysis and machine intelligence 13(4), 376â380 (2002) 5 29. Wang, Z., Chen, J., Furukawa, Y.: Puzzlefusion++: Auto-agglomerative 3d fracture assembly by denoise and verify. arXiv preprint arXiv:2406.00259 (2024) 4 30. Wu, P., Escontrela, A., Hafner, D., Abbeel, P., Goldberg, K.: Daydreamer: World models for physical robot learning. In: Conference on robot learning. p. 2226â 2240. PMLR (2023) 2 31. Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al.: Sapien: A simulated part-based interactive environment. In: Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 11097â11107 (2020) 5 32. Xiang, J., Chen, X., Xu, S., Wang, R., Lv, Z., Deng, Y., Zhu, H., Dong, Y., Zhao, H., Yuan, N.J., et al.: Native and compact structured latents for 3d generation. arXiv preprint arXiv:2512.14692 (2025) 2, 4 33. Xiang, J., Lv, Z., Xu, S., Deng, Y., Wang, R., Zhang, B., Chen, D., Tong, X., Yang, J.: Structured 3d latents for scalable and versatile 3d generation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 21469â21480 (2025) 2, 4 34. Xiong, Y., Ma, W.C., Wang, J., Urtasun, R.: Learning compact representations for lidar completion and generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 1074â1083 (2023) 4 35. Xu, Y., Wan, W., Zhang, J., Liu, H., Shan, Z., Shen, H., Wang, R., Geng, H., Weng, Y., Chen, J., et al.: Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal generation and goal-conditioned policy. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 4737â4746 (2023) 2 36. Yang, Y., Huang, Y., Guo, Y.C., Lu, L., Wu, X., Lam, E.Y., Cao, Y.P., Liu, X.: Sampart3d: Segment any part in 3d objects. arXiv preprint arXiv:2411.07184 (2024) 4 37. Yang, Y., Zhou, Y., Guo, Y.C., Zou, Z.X., Huang, Y., Liu, Y.T., Xu, H., Liang, D., Cao, Y.P., Liu, X.: Omnipart: Part-aware 3d generation with semantic decoupling and structural cohesion. In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers. p. 1â12 (2025) 4 38. Ying, S., Peng, J., Du, S., Qiao, H.: A scale stretch method based on icp for 3d data registration. IEEE Transactions on automation science and engineering 6(3), 559â565 (2009) 5, 9 39. Yuan, Y., Cui, H., Huang, Y., Chen, Y., Ni, F., Dong, Z., Li, P., Zheng, Y., Hao, J.: Embodied-r1: Reinforced embodied reasoning for general robotic manipulation. arXiv preprint arXiv:2508.13998 (2025) 2 18H. Shan et al. 40. Zakka, K., Zeng, A., Lee, J., Song, S.: Form2fit: Learning shape priors for gen- eralizable assembly from disassembly. In: 2020 IEEE International Conference on Robotics and Automation (ICRA). p. 9404â9410. IEEE (2020) 4 41. Zhang, C., Wang, J., Gao, Z., Su, Y., Dai, T., Zhou, C., Lu, J., Tang, Y.: Clap: Contrastive latent action pretraining for learning vision-language-action models from human videos. arXiv preprint arXiv:2601.04061 (2026) 2 42. Zhang, L., Wang, Z., Zhang, Q., Qiu, Q., Pang, A., Jiang, H., Yang, W., Xu, L., Yu, J.: Clay: A controllable large-scale generative model for creating high-quality 3d assets. ACM Transactions on Graphics (TOG) 43(4), 1â20 (2024) 4 43. Zhang, Y., Liu, Y., Xie, Z., Yang, L., Liu, Z., Yang, M., Zhang, R., Kou, Q., Lin, C., Wang, W., et al.: Dreammat: High-quality pbr material generation with geometry- and light-aware diffusion models. ACM Transactions on Graphics (TOG) 43(4), 1â18 (2024) 4 44. Zhao, Z., Lai, Z., Lin, Q., Zhao, Y., Liu, H., Yang, S., Feng, Y., Yang, M., Zhang, S., Yang, X., et al.: Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202 (2025) 2, 4 Interact3D: Compositional 3D Generation of Interactive Objects19 A Experiments A.1 Dataset Description Our dataset comprises approximately 8,300 3D compositional scenes, predomi- nantly representing common everyday scenarios. Specifically, about 7,700 scenes consist of two interactive objects, while about 600 scenes are composed from more than two (typically 3 to 5) objects. Each scene includes the individual 3D assets along with their corresponding compositional poses θ. Notably, the user- provided reference meshes are highly flexible: they can be either high-fidelity assets obtained through manual 3D modeling or outputs synthesized by arbi- trary 3D generative models (e.g., Hunyuan, TRELLIS2, Rodin). A.2 More Results Effectiveness of Agentic Optimization. Severe 2D occlusions inevitably lead to collisions in the final 3D scene. As demonstrated in the additional case (Figure 9), our agentic optimization effectively resolves these conflicts by gen- erating an updated complementary mesh, ensuring a physically plausible final composition. Two-part Composition Qualitative Evaluation. Figure 10 presents a qual- itative comparison of the five cases introduced in Figure 7 across all baselines. Compared to Jigsaw and 2BY2, Interact3D demonstrates superior scale aware- ness. Furthermore, our generated 3D assets exhibit significantly higher geometric fidelity than those produced by direct segmentation in PartField. When evalu- ated against âPartField+RANSACâ, Interact3D showcases more robust physical plausibility and spatial alignment capabilities. Notably, in the 3rd and 4th rows, the baselines struggle with severe artifacts, resulting in unrealistic upside-down inversions of the computer. More (than two) Parts Composition. To further validate the scalability and robustness of Interact3D, we evaluate its performance on more complex multi-part composition tasks (more than two parts), as visualized in Figures 11 and 12. Although compositional scenes with more parts inherently introduce compounded spatial conflicts and severe occlusions, our framework consistently yields highly coherent compositions. Even in these highly constrained scenarios, Interact3D successfully maintains precise spatial alignment and strict physical plausibility across all constituent assets. Multi-view Rendering Results. Since two-viewpoint evaluations may ob- scure geometric flaws, Figures 13, 14, and 15 showcase our compositional scenes from multiple viewing angles. This exhaustive multi-view display validates In- teract3Dâs capability. It maintains the high quality of the constituent 3D assets and also enforces collision-aware constraints, resulting in physically plausible composition from any viewing angle. 20H. Shan et al. Generated ImageGenerated Scene 3D Compositional Scene Generated ImageGenerated Scene3D Compositional Scene Fig. 8. Failure cases. (left) Strong geometric symmetry in books can lead to upside- down orientations despite accurate spatial localization. (right) Heavy 2D occlusions cause severe mesh interpenetrations in the TRELLIS2-generated prior, propagating these physical collisions directly to the final compositional scene. A.3 Failure Cases Since our registration in composition tasks relies on geometry-aware GeoTrans- former without texture awareness, objects with strong geometric symmetry (such as cuboid books) can cause orientation ambiguity, occasionally leading to upside- down inversions (see Figure 8 (left)). Additionally, since our composition pipeline is explicitly supervised by the spatial priors of M scene , initial spatial misalign- ments generated in extremely complex scenarios can propagate downstream, ultimately compromising the composition quality (see Figure 8 (right)). âPut a teddy bear in a brown and white canvas bag. The picture should show the complete object geometry and realistic proportions. White background.â Mesh M scene (3D Guidance) User-given mesh M âRemove the canvas bag in the picture and complete the occupied part of the content. Other parts remain unchanged. White background.â Mesh M comp Leg penetration: As can be clearly seen in figures, the legs and feet of the teddy bear directly penetrated the front body of the handbag. In the real physical world, two solid surfaces cannot occupy the same space. âFold the legs completely and tuck them deeply underneath the body so that only the bottom edges of the soles are visible, rather than sticking out.â Mesh M comp Front-to-back space conflict inside the bag: On the right side of the handbag, the teddy bear's furry part is clipping directly through the side panel. âChange the teddy bear's pose from sitting to standing.â Mesh M comp Agentic Optimization Image I scene Image I comp Image I comp Image I comp 3D Compositional Scene 3D Compositional Scene input 3D Compositional Scene output output output output output output 1 2 3 4 Fig. 9. Visualizations of Agentic Optimization. Progressive geometric refinement of a teddy bear to achieve a collision-free fit inside a canvas bag. Interact3D: Compositional 3D Generation of Interactive Objects21 Jigsaw User-given Mesh / Generated Mesh 2BY2 PartField PartField + RANSAC Ours Generated Scene Fig. 10. Additional qualitative comparison. Interact3D outperforms Jigsaw and 2BY2 in scale awareness, and PartField in asset fidelity. Unlike âPartField+RANSACâ, our method ensures physical plausibility and spatial alignment capabilities. 22H. Shan et al. âPlace a watermelon, an apple, and a peach in a wooden basket. The picture should show the complete object geometry and realistic proportions. White background.â User-given mesh M Mesh M scene Mesh M scene (3D Guidance) Mesh M scene (3D Guidance) Mesh M scene (3D Guidance) Image I scene âKeep the watermelon in the figure and remove the other fruit and basket. White background.â Image I comp Mesh M comp âKeep the apple in the figure and remove the other fruit and basket. White background.â Image I comp PartField Mesh M comp âKeep the peach in the figure and remove the other fruit and basket. White background.â Image I comp Mesh M comp Interact3D Interact3D Interact3D 3D Compositional Scene Final 3D Compositional Scene 3D Compositional Scene input output output output output 1 2 3 3 4 4 2 2 5 5 Fig. 11. Multi-part composition. A user-provided basket is rendered and popu- lated with fruits via Nano Banana Pro, then reconstructed by TRELLIS2 to extract object-object spatial relationships (OOR). Individual fruits are similarly synthesized and sequentially registered into the basket using the OOR guidance, forming a highly coherent, interactive 3D scene. âPlace a black over-ear headphone, a cute little yellow duck figurine, and a pink square eyeglass case in an open, shallow white storage box. The picture should show the complete object geometry and realistic proportions. White background.â User-given mesh M Mesh M scene Mesh M scene (3D Guidance) Mesh M scene (3D Guidance) Mesh M scene (3D Guidance) Image I scene âKeep the black over-ear headphone in the figure and remove the other things and box. White background.â Image I comp Mesh M comp âKeep the cute little yellow duck figurine in the figure and remove the other things and box. White background.â Image I comp PartField Mesh M comp âKeep the pink square eyeglass case in the figure and remove the other things and box. White background.â Image I comp Mesh M comp Interact3D Interact3D Interact3D 3D Compositional Scene Final 3D Compositional Scene 3D Compositional Scene input output output output output 1 2 3 3 4 4 2 2 5 5 Fig. 12. Multi-part composition. A rendered box is populated with a headphone, duck, and eyeglass case via Nano Banana Pro, then reconstructed by TRELLIS2 to extract object-object relationships (OOR). Individual items are then independently synthesized and sequentially registered into the box using this OOR guidance. Interact3D: Compositional 3D Generation of Interactive Objects23 Text Prompt / User-given Mesh Generated Comple- mentary Mesh Front/Back Views of Generated Scene Other Views of Generated Scene âHang the over-ear headphones on the black standing headphone hook.â âAdd a wooden tray for this teapot.â âAdd a grey mouse pad for this mouse.â âPlace a pair of shoes in this cabinet.â Fig. 13. Multi-view rendering results. Given a user-provided mesh and a textual prompt, Interact3D synthesizes a complementary mesh and compose them into a co- herent, interactive 3D scene. We visualize the compositional scenes rendered from six distinct viewpoints. 24H. Shan et al. Text Prompt / User-given Mesh Generated Comple- mentary Mesh Front/Back Views of Generated Scene Other Views of Generated Scene âplace this yellow coffee mug on the coffee machine to prepare for the coffee.â âPlace this coin into a golden cartoon piggy bank.â âPlace this cute cartoon bear figurine upright on the yellow round display tray.â âAdd a red cap to this yellow cartoon-shaped beverage bottle.â Fig. 14. Multi-view rendering results. We visualize the compositional scenes ren- dered from six distinct viewpoints. Interact3D: Compositional 3D Generation of Interactive Objects25 Text Prompt / User-given Mesh Generated Comple- mentary Mesh Front/Back Views of Generated Scene Other Views of Generated Scene âPut this wooden stick into an empty jar.â âPlace these sunglasses in an open pink eyeglass case.â âPlace these blue sneakers in the shoebox.â âPlace this basketball flat on a blue, hollow triangular stand.â Fig. 15. Multi-view rendering results. We visualize the compositional scenes ren- dered from six distinct viewpoints.