Paper deep dive
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield, Ankit Goyal, Stan Birchfield, Fabio Ramos, Jonathan Tremblay
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/14/2026, 1:56:27 AM
Summary
RoboLab is a high-fidelity simulation benchmarking framework designed to evaluate the generalization capabilities of task-generalist robotic policies. It addresses the limitations of existing benchmarks by providing a scalable, LLM-enabled workflow for generating diverse scenes and tasks across three competency axes: visual, procedural, and relational. The framework includes the RoboLab-120 benchmark and a suite of metrics, including sensitivity analysis, to quantify policy performance and robustness against environmental perturbations.
Entities (5)
Relation Signals (3)
RoboLab → includes → RoboLab-120
confidence 100% · We introduce RoboLab... With this, we propose the RoboLab-120 benchmark
RoboLab → utilizes → IsaacLab
confidence 95% · using human-readable USD and Python interfaces in IsaacLab
RoboLab-120 → evaluates → π 0.5
confidence 90% · the state-of-the-art π 0.5 achieves only a ~30% success rate on RoboLab-120
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing. Existing benchmarks often exhibit significant domain overlap between training and evaluation, trivializing success rates and obscuring insights into robustness. We introduce RoboLab, a simulation benchmarking framework designed to address these challenges. Concretely, our framework is designed to answer two questions: (1) to what extent can we understand the performance of a real-world policy by analyzing its behavior in simulation, and (2) which external factors most strongly affect that behavior under controlled perturbations. First, RoboLab enables human-authored and LLM-enabled generation of scenes and tasks in a robot- and policy-agnostic manner within a physically realistic and photorealistic simulation. With this, we propose the RoboLab-120 benchmark, consisting of 120 tasks categorized into three competency axes: visual, procedural, relational competency, across three difficulty levels. Second, we introduce a systematic analysis of real-world policies that quantify both their performance and the sensitivity of their behavior to controlled perturbations, indicating that high-fidelity simulation can serve as a proxy for analyzing performance and its dependence on external factors. Evaluation with RoboLab exposes significant performance gap in current state-of-the-art models. By providing granular metrics and a scalable toolset, RoboLab offers a scalable framework for evaluating the true generalization capabilities of task-generalist robotic policies.
Tags
Links
- Source: https://arxiv.org/abs/2604.09860v1
- Canonical: https://arxiv.org/abs/2604.09860v1
Trouble viewing inline? Open PDF directly →
Full Text
124,059 characters extracted from source content.
Expand or collapse full text
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies Xuning Yang 1 , Rishit Dagli 2,4 , Alex Zook 1 , Hugo Hadfield 1 , Ankit Goyal 1 , Stan Birchfield 1 , Fabio Ramos 1,3 , and Jonathan Tremblay 1 1 NVIDIA, 2 University of Toronto, 3 The University of Sydney, 4 Work done during internship at NVIDIA Fig. 1: Overview of RoboLab. RoboLab addresses the simulation-to-real gap by evaluating robotics policies on entirely held-out domains. By featuring a streamlined generation pipeline for new scenes and tasks (top row), RoboLab enables rapid extensibility for testing generalization capabilities. Our accompanying benchmark introduces visual, relational, and procedural testing axes, paired with robust metrics designed to reveal how modern models perform when faced with novel, out-of-distribution challenges (bottom row). Abstract—The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing. Existing benchmarks often exhibit significant domain overlap between training and evaluation, trivializing success rates and obscuring insights into robustness. We introduce RoboLab, a simulation benchmarking framework designed to address these challenges. Concretely, our framework is designed to answer two questions: (1) to what extent can we understand the performance of a real-world policy by analyzing its behavior in simulation, and (2) which external factors most strongly affect that behavior under controlled perturbations. First, RoboLab enables human-authored and LLM-enabled generation of scenes and tasks in a robot- and policy-agnostic manner within a physically realistic and photorealistic simulation. With this, we propose the RoboLab-120 benchmark, consisting of 120 tasks categorized into three competency axes: visual, procedural, relational competency, across three difficulty levels. Second, we introduce a systematic analysis of real-world policies that quantify both their performance and the sensitivity of their behavior to controlled perturbations, indicating that high-fidelity simulation can serve as a proxy for analyzing performance and its dependence on external factors. Evaluation with RoboLab exposes significant performance gap in current state-of-the-art models. By providing granular metrics and a scalable toolset, RoboLab offers a scalable framework for evaluating the true generalization capabilities of task-generalist robotic policies. I. INTRODUCTION The pursuit of generality has been a longstanding challenge in modern robotics. Recent advances have produced impressive generalist robot policies that demonstrate success in challenging and novel tasks in the real-world. Despite this progress, benchmarks for evaluating whether these policies are truly task- general has been slow. Evaluating models in the real world remains prohibitively expensive and logistically intractable, motivating the rise of simulation-based benchmarks as an appealing alternative. Current robotics benchmarks [19,35,11,16] face several crit- ical limitations: (1) a lack of high-fidelity simulation capable of supporting real-world policies; (2) rapid performance saturation on static task sets; and (3) a lack of granular analysis regarding policy failure modes. For instance, popular benchmarks like arXiv:2604.09860v1 [cs.RO] 10 Apr 2026 Fig. 2: Three approaches for robotic benchmarks. LEFT: To date, pure simulation based benchmarks have exhibited low visual quality, creating a large sim2real transfer gap. MIDDLE: Real2sim benchmarks address this issue by using techniques to bring real-world visual texture into simulation. However, these environments are extremely costly with reported per-scene generation time of∼1hr [10]. RIGHT: Our approach achieves a high degree of realism with low overhead. LIBERO [19] often utilize nearly identical environments for both training and evaluation. When policies are fine-tuned on these simulation-specific demonstrations, the lack of a meaningful domain gap trivializes the evaluation process and obscures the model’s true generalization capabilities. Many existing platforms have limited realism or are difficult to extend due to using rigid architectures that make it cumbersome to introduce new objects, tasks, or robots (Fig. 2). To address these limitations, we present RoboLab (Fig. 1), a simulation platform and benchmarking suite designed for rigorous robotics evaluation. Unlike prior benchmarks that rely on PDDL or rigid scene-graph definitions [19], RoboLab introduces an easy-to-use interface that enables human-authored and LLM-scaled scene and task generation. RoboLab enables generation and validation of new scenes and tasks from natural language prompts. This system enables the creation of over 800 diverse scenarios (see supplemental), providing a scalable framework that mitigates benchmark saturation and ensures long-term value. RoboLab introduces novel task axes and robust metrics to provide deeper diagnostic insights, paired with a streamlined toolset for generation of scenes and tasks (see Fig. 1).. To pro- vide a granular assessment of policy behavior in our proposed benchmark, we evaluate three axes: Visual (perceptual attributes like color and size), Procedural (action-oriented logic such as stacking and reorientation), Relational (spatial and linguistic logic like “and/or”) spanning across three difficulty levels depending on task length and language nuance. Policy execution on these tasks is then evaluated with metrics on graded task completion, failure and error occurrences, and trajectory quality. Finally, we highlight novel metrics for evaluation; including sensitivity analysis to identify environmental factors that most strongly influence policy performance, e.g., camera placement. We introduce the RoboLab-120 benchmark, comprising 120 tasks generated via our automated workflow and verified by humans. These tasks span varying difficulties (35 simple, 28 moderate, 16 complex) and multiple competency axes (33 relational, 54 visual, and 9 procedural). To prevent overfitting to the simulation domain, we evaluate policies trained exclusively on the real-world DROID [13] dataset. This creates an environment that reflects “in the wild” conditions; for instance, the state-of-the-artπ 0.5 [9] achieves only a∼30% success rate on RoboLab-120, highlighting the benchmark’s difficulty. In summary, our contributions are: 1)RoboLab: A novel simulation platform designed for evaluating modern robotics policies with a scalable, LLM- based workflow capable of procedurally generating over 800 unique scenes and tasks using human-readable USD and Python interfaces in IsaacLab[21]. 2)RoboLab-120 Benchmark: Comprising 120 tasks eval- uated across three distinct competency axes (visual, procedural, relational) and supported by four new robust- ness metrics. We also present five policies evaluated on RoboLab-120. 3) Policy Analysis: We introduce a suite of analysis tools that gives insight into the model performance beyond binary success rates and broader understanding of policy performance. I. RELATED WORK Simulation-Based Benchmarks.Simulation provides a scalable and reproducible environment for evaluating robot manipulation policies. Widely used benchmarks such as RLBench [11], MetaWorld [33], and robosuite [36], Man- iSkill2 [7], CALVIN [20], LIBERO [19], and BEHAVIOR- 1K [15], offer standardized task suites for learning and evalua- tion in simulation across pre-defined task families and object configurations. However, in these settings, policies are typically trained and evaluated in the same simulated environments, which encourages overfitting to simulator-specific quirks, leads to rapid benchmark saturation, and makes real-world general- ization hard to assess. [35]. In our setting, policies are instead trained on large-scale real-world data (e.g., DROID [13]), while high-fidelity simulation is used only as a controlled evaluation environment, so training and evaluation domains are decoupled and measured performance more closely reflects robustness in the real world. Real-to-sim Evaluation.Recent work have focused on leveraging 3D reconstruction to build photorealistic simulation scenes from real-world videos in order to achieve closer visual alignment between simulation and real world photorealism [16,12,10,37]. These works typically use Gaussian splatting, 3D segmentation, and multi-view inpainting, often operated at a per-scene level, which entails costly optimization and Fig. 3: Task progression of a few tasks, illustrating errors encountered during policy rollout. Top row: Although the task is successfully completed, errors were encountered during execution: 1) The robot drops the milk jug too early, missing the bin. 2) the robot grasps an orange (wrong object) and puts it in the bin. Mid row: An extraneous object was reoriented before the actual intended object. Final row: Intended objects were attempted unsuccessfully, and the policy tended to two wrong other objects. makes it slow to scale beyond a small number of environments [10, 34, 28]. In contrast, our framework produces large-scale, photorealistic scenes and tasks within minutes rather than hours, while preserving sufficient geometric and visual fidelity for policy evaluation, thereby making real-to-sim benchmarking practical at the scale needed for modern generalist robot policies. I. ROBOLAB Evaluating real-world, generalist robotics policies in simu- lation remains a significant challenge. RoboLab is a bench- marking framework that introduces three novel task axes and three original metrics tailored for modern robotics systems. RoboLab enables a multifaceted analysis of Vision-Language- Action (VLA) models, providing deeper insights into their scalability and task generalization. A. RoboLab-120 Inspired by the Large Language Model (LLM) community’s use of Visual Question and Answering (VQA) benchmarks, we introduce RoboLab-120 Benchmark that focus on evaluating specific competency axes spanning three difficulty levels. This Fig. 4: Example of language instructions in RoboLab-120. taxonomic decomposition enables fine-grained analysis of policy capabilities by systematically assessing performance. Figure 4 shows examples of these questions accompanied by scene examples. Visual Competency: Assesses recognition of color, semantics, and size, capturing the policy’s capability to link perceptual attributes with higher-level reasoning. Procedural Competency: Evaluates the ability to perform tasks that involve action-oriented reasoning, including affordances, reorientation, or stacking. Relational Competency: Tests understanding of language con- junctions (e.g., ‘and’, ‘or’), counting, and spatial relationships, measuring how effectively the policy interprets multi-object instructions and scene structure. Tasks from these competencies can span one of the following difficulty levels: simple, medium, complex. These are deter- mined as a function of two aspects: whether if the language was straightforward in describing the task, as well as the number of required reasoning steps for the task. B. Metrics for Evaluation We establish a comprehensive suite of evaluation metrics that captures the full spectrum of policy performance characteristics. While task success rate remains a fundamental metric, prior work [14] has demonstrated that they fail to reveal nuanced aspects of policy behavior and failure modes. Unlike approaches relying on human judgment [12], we define a set of discrete and continuous metrics to characterize policy performance. We regroup these novel metrics as follow, failure cases scoring, trajectory metrics, and sensitivity analysis. Failure cases.In addition to success rate, we compute Fig. 5: Comparison of policy performance for bowl-in-bin manipulation. Rows represent distinct policies shown in chronological order (left to right). Successful execution involves grasping the central red bowl and depositing it into the gray bin on the right. Unsuccessful attempts are characterized by aimless arm trajectories and a lack of object interaction. a normalized graded scoreSc(T ) = 1 |T| P τ∈T Sc(τ ) . For example, for the instruction “pick the lemon and the lime,” the subtasks τ “pick lemon” and “pick lime” includes steps such as “grasp” and “drop”. The final task score is the normalized subtask scores. Our benchmark automatically records instances of events; including wrong object grasped, object dropped, and gripper collisions. Fig. 3 demonstrates a successful episode; however, the policy incorrectly grasped an extraneous object. Such errors highlight potential biases in the policy not captured by other metrics. Trajectory Metrics. Trajectory quality metrics capture char- acteristics of motion efficiency and optimality. We compute the following: Spectral arc-length (SPARC), which evaluates motion smoothness [2] via the arc length of the normalized Fourier magnitude spectrum of the velocity profile. Given a speed profilev(t)of the end effector over time interval[0,T ] SPARC =− Z ω c 0 v u u t 1 ω c 2 + d ˆ V (ω) dω ! 2 dω(1) where ˆ V (ω) = V (ω)/V (0)represents the normalized Fourier magnitude spectrum. Smoother motions yield values closer to zero, while jerkier trajectories produce more negative values. We employ an adaptive cutoff frequencyω c = min(10 Hz,ω α ), whereω α = max k∈K ω k andK = k | ˆ V (ω k ) ≥ αdenotes the set of frequency bins exceeding thresholdα = 0.05. This adaptive strategy ensures that the smoothness evaluation focuses on relevant frequency components. Lastly, trajectory optimality is assessed through end effector speedv(t), and path length l = P N−1 k=0 ∥p k+1 − p k ∥, wherep k denotes the end-effector position at timestepk. Shorter path lengths indicate more direct trajectories and generally reflect superior motion quality. C. Sensitivity Analysis We present a Bayesian framework for evaluating policy robustness across diverse environmental conditions using Simulation-Based Inference (SBI). This analysis provides insight into which scene parameters are most strongly linked to success and failure outcomes by learning an approximate posterior distribution over them given evaluation data. Let θ = (θ cont ,θ disc )denote the environment parameters comprising of continuous variables (e.g., object distance, camera displace- ment) and/or discrete variables. After evaluating policyπunder varied conditions, we generate episodesD = (θ i ,x i ) N i=1 with observed outcomesx i (e.g., task success). The posterior distributionp(θ | x) ∝ p(x | θ)p(θ)is approximated using Mixed Neural Posterior Estimation (MNPE), which trains a neural density estimatorq φ (θ | x)to directly learn the mapping from observations to parameter distributions. The resulting posteriorq φ (θ | x)characterizes which scene variables are most associated with a target observationx. Our approach provides systematic assessment of which variables most strongly influence performance outcomes. Further details are in Appendix B. D. Robolab scene and task generation RoboLab offers a user-friendly workflow that mirrors the process of preparing a real-world robot evaluation (Fig. 1): 1) create a scene by positioning and orienting objects in a workspace; 2) define a task as language instructions for a goal state in the scene; 3) instantiate an environment by selecting a robot, policy, and variations of scene features including camera, lighting, and backgrounds for a task. We make this process reproducible and scalable by decoupling task definitions from environments, allowing reuse over new embodiments and policies. Our approach automates the process of environment assembly, reducing manual labor when evaluating a new robot or policy. In addition, we developed an automated workflow Fig. 6: Example scene variations, lighting variations, and camera pose variations in RoboLab. to generate new scenes and tasks to facilitate extension of our evaluations and to mitigate benchmark saturation in the future. Formally, define a sceneS = (b i , p i , q i ) N i=1 , whereb i represents an object instance selected from the available catalog of objectsBandp i ∈ R 3 , q i ∈ SO(3)denote its position and orientation. Define a taskT = S,l, wherelis the language instruction to complete in the scene. Define a policy π : O → Awhere the action spaceA ∈ A joint ,A E ,... and observation spaceO = (O proprio ,O rgb ,O depth · )is policy dependent. An environmentE = (T ,R,O,A,ξ)consists of a task, robot embodimentR, policy parameters (A,O), and scene variationsξ = (ξ camera ,ξ light ,ξ background ,ξ pose ). More details on the specific objects, scenes and tasks in RoboLab can be found in Appendix A. 1) Scaling Scene Generation: We enable scaling scene generation through an automated pipeline that: 1) prompts an LLM to generate a structured scene plan for asset placement; Fig. 7: Examples of language ablation experiments. Top: Same scene and goal, but the instruction wording ranges from precise to increasingly vague. Middle: Same scene, but the instruction specifies different tasks to perform. Bottom: Same instruction, but the scene becomes progressively more complex. 2) uses a geometric solver and physics simulation to check asset placement validity; and 3) refines the scene if it is not valid. First, the LLM is prompted with a theme (e.g., “messy counter”) to generate a structured scene plan consisting of a subset of objectsB ⊂Band spatial predicatesPgoverning the layout. The LLM is provided with the full catalog of objects B containing names and bounding box dimensions d i ∈ R 3 . Second, a spatial solver converts the relational predicatesP into valid pose configurations(p, q)). Objects are processed in dependency order, with support surfaces placed before objects on those surfaces (Algorithm 1). To check physical stability, the scene is then forward simulated in Isaac Sim [23] for300 steps under gravity. An objectb i is flagged as unstable if it’s maximum Euclidean displacement is larger than a threshold (typically0.02m). Third, If any object is unstable, we generate a text error describing the failure (e.g., “Object ‘apple’ fell off ‘plate’ with displacement 0.15m”). This feedback is provided to the LLM to refine the scene plan and repeat the process. Further details, including on the spatial and physical solvers, are provided in Appendix C. 2) Scaling Task Generation: We enable scaling task genera- tion through an automated pipeline that: 1) generates task code from information including the scene and competency axes; 2) validates code syntax; 3) validates asset selections in the scene; and 4) refines the task if it is not valid. First, we prompt an LLM with detailed task information: 1) the scene object catalogB S and metadata (including bounding boxes and semantics) with dimensions; 2) task examples demonstrating the task structure; 3) the complete predicate library defining sub-task success and termination; 4) Competency-axes language templates with placeholders for objects, spatial verbs, and attributes; and 5) constraints including difficulty levels and physical feasibility requirements (e.g., containment size constraints, stacking stability). The prompt forbids referencing objects not present inB S and includes previously generated tasks to prevent duplicates (see Appendix C for details). Second, tasks TABLE I: Overall performance of VLAs on RoboLab. While recent VLAs exhibit emerging capabilities across diverse task dimensions, overall success rates and consistency remain limited. † Overall MetricsDifficulty (succ%)Procedural (succ%)Relational (succ%)Visual (succ%) ModelSucc% (↑)Score (↑)SPARC (↑)Speed (↑)simplemoderatecomplexaffordancereorientationstackingconjunctioncountingspatialcolorsemanticssize π 0.5 [9]31.90.11 −9.10 ±4.55.7 ±1.940.628.219.420.053.316.076.072.023.930.021.535.0 π 0 -FAST [27]18.20.05 −9.14 ±4.74.7 ±1.926.017.13.10.020.00.042.056.020.08.312.135.0 GR00T N1.6 [24]8.00.06 −6.44 ±2.04.2 ±1.312.37.10.00.00.00.04.038.07.05.65.90.0 π 0 [5] 5.90.04 −9.45 ±3.44.4 ±1.411.72.10.00.00.00.026.010.03.01.71.80.0 PaliGemma [4]1.90.01 −17.6 ±11.81.4 ±1.21.13.90.00.00.00.00.022.01.30.02.15.0 † Current paper results were from an older version of the benchmark which contained 80 tasks; updated results on the full benchmark (120 tasks) will be available soon. are check for syntax validity as code. Third, asset validation checks that all objects are not in the forbidden set and, for containment tasks (e.g., “placeb i insideb j ”), that inner objects fit inside containers with some clearance. Fourth, if validation fails, feedback is gathered into a fix promptQ fix that includes the original promptQ, the invalid output, and an error message Edescribing syntax errors or invalid asset references. The fix prompt is provided to the LLM to refine the task and repeat the process. We evaluated our task generation approach using an LLM- as-judge framework. We generated812tasks across59scenes evenly across the competency axes with o1 [26]. We then extracted each natural-language instruction and its program- matic termination conditions from the generated code, and prompted a second o1 judge to score instruction–criterion alignment across relation, target, object, and quantifier match, plus instruction clarity and physical feasibility (each on a 0– 1 scale), and to assign an aligned/partial/misaligned verdict. Overall, tasks achieved0.91alignment,0.96clarity,0.92 feasibility, and0.95semantic match, with76%judged fully aligned (misaligned≈ 1%) and covered88%of objects. These results show our approach can scale to generate diverse tasks that are semantically aligned to their language instructions (see Appendix D). IV. EXPERIMENTS We evaluate several off-the-shelf VLA policies on RoboLab- 120, controlled ablations, and environmental perturbations to identify which competencies generalize and where failures concentrate. The experiments are designed to address the following questions: Q1: How well does a real-world policy perform in our simulated benchmark? Q2: How well does a policy generalize with language variations? Q3: When and why does a policy fail? A. Experiment Setup We evaluate 80 tasks † of varying difficulty levels (35 simple, 28 moderate, 16 complex) and spanning competency axes (33 relational, 54 visual, and 9 procedural). Each task was assigned to one or more competency axes. Our experiments used the DROID robot [13], which is commonly used to benchmark VLAs [10,1]. DROID has a 7-DOF Franka Panda robot arm with a Robotiq-2F-85 gripper, an externally mounted ZED 2i camera withf =2.1m, and a ZED mini as the wrist camera. We evaluated VLAs with off-the-shelf checkpoints fine-tuned on the DROID dataset [13]:π 0.5 [9],π 0 -FAST[27],π 0 [5], PaliGemma [4], and GR00T N1.6 [24]. The action space is 7-DOF Franka joint positions and a 1-DOF binary gripper command. The environments were composed of a default office- like background and natural lighting to mimic typical setups in the DROID dataset [13], with wrist and external camera poses designed to match the real-world DROID robot. Each task was repeated 10 times with a fixed seed to address uncontrolled stochasticity in the physics simulation and robot policy. B. Task Results Table I shows the overall results on our benchmark. Overall success rates were low, with the best policy (π 0.5 ) reaching 31.9% success. This matches prior observations on out-of- domain generalization for VLAs [35]. Below, we discussπ 0.5 to illustrate how RoboLab supports targeted diagnosis of policy capabilities and suggests concrete directions for improvement. Competency axes also highlighted asymmetric generalization in relational reasoning tasks:π 0.5 handled conjunctions (76.0% success) and counting (60.0%) better than spatial relations (23.9%). In visual grounding, performance remained low across attribute types (35.0% for size, 30.0% for color, and 21.5% for semantics), indicating brittle language-to-object binding beyond a narrow set of familiar object descriptions. Procedural understanding proved most challenging:π 0.5 achieved modest success on reorientation (53.3%) but struggled with affordances (20.0%) and stacking (16.0%). Together, these results show how RoboLab isolates where generalization fails, supporting diagnosis that can inform data collection and training priorities. C. Ablation Experiments To further probe robustness and language grounding, we performed controlled ablations that varied the instruction, scene, or task in isolation (Fig. 7). Varying instruction specificity in a fixed scene. Table IIa shows how VLAs respond to varying levels of instruction specificity. Results reveal that VLAs lack grounding for abstract or implied goals. Interestingly, we observe that the for the vague command “Empty the grey bin”,π 0.5 tries to grab the bin instead of clearing its content. These results indicate that current VLAs may rely on keyword matching within instructions rather than demonstrating the linguistic inference necessary to identify implicit task goals. As shown in Table IIa, π 0.5 exhibits higher performance on specific instructions than on vague ones, indicating a sensitivity to unde-rspecified task goals. TABLE I: Language understanding ablations. (a) VLA perfor- mance degrades with abstract or vague language instructions. (b) Performance drops as scene complexity increases. (c) VLAs show brittle language grounding, with consistent object confusion patterns across different instructions in the same scene. (a) Effect of language specificity on task performance. π 0.5 π 0 -FAST TaskSucc %ScoreSucc %Score Bananas Out Of Bin Task “Take all the bananas out of the grey bin and put it on the table.” 500.13300.05 “Take the bananas out”400.22100.15 “Empty the grey bin”100.07700.11 White Mugs In Bin Task “Put the white mugs in the grey bin”800.50200.22 “Put the mugs in the bin”900.50100.11 “Put away mugs”00.0000.00 Remove Measuring Spoons from the Plate Task “Put the orange measuring cup and the blue measuring cup outside of the plate” 200.4700.31 “Clear the plate”00.0800.10 (b) Effect of scene complexity on task performance. Sceneπ 0.5 π 0 -FASTπ 0 GR00T N1.6 Task: “Pack boxed foods into the bin” 1 Box/Can10000 2 Boxes/Cans0000 3 Boxes/Cans0000 Task: “Pack canned foods into the bin” 1 Box/Can703000 2 Boxes/Cans301000 3 Boxes/Cans20000 (c) Instruction sensitivity within fixed scenes. Task / Promptπ 0.5 π 0 -FASTπ 0 GR00T N1.6 Fruit Plate Scene “Move an orange or a lime to the wood bowl” 50000 “Move an orange to the white bowl”0000 “Put the onion in the wood bowl”70102020 “Put the onion on the plate”0000 Tools Cleanup Scene “Put hammers in the right bin”20000 “Put hammers in the left bin”10000 Tools Selection Scene “Select the cordless drill and put it on the table” 70502030 “Select the blue hammer and put it on the table” 00100 Varying scene complexity with a fixed instruction. Table IIb isolates the effect of scene complexity by increasing the number of objects to manipulate. Success rates drop as object count increases: from 70% with a single target object to 20% with three objects. More revealing is the type of failure: VLAs exhibit systematic geometric biases, frequently grasping cylindrical objects (cans) when instructed to manipulate boxes. This suggests that training data distributions create strong shape priors that override language-specified targets, a critical limitation for real-world deployment where objects vary widely TABLE I: Robustness to controlled environmental variations over two simple tasks (BananaInBowl, BananaAndCubeInBowl). PaliGemma is excluded as it fails to achieve meaningful results. π 0.5 π 0 -FASTπ 0 VariationSucc.%Time (s)Succ.%Time (s)Succ.%Time (s) Lighting Color96.714.5 ±7.993.317.9 ±10.76.731.1 ±4.3 Shadows100.016.0 ±6.090.012.4 ±3.50.0- Dim90.09.1 ±2.170.013.1 ±2.870.035.5 ±9.7 Overexposed100.013.9 ±4.4100.09.6 ±1.70.0- Visual Variations Background85.014.4 ±8.770.021.3 ±10.725.031.6 ±11.8 Table texture87.519.0 ±13.860.019.0 ±12.922.528.1 ±6.9 Object Pose 10cm95.016.2 ±9.255.026.5 ±13.822.534.7 ±13.3 20cm95.019.7 ±10.240.021.8 ±9.520.037.4 ±11.2 30cm62.518.9 ±8.735.024.3 ±11.917.524.3 ±11.9 Camera Pose external85.017.4 ±11.345.027.7 ±10.650.027.4 ±16.0 wrist60.021.9 ±13.425.020.1 ±9.910.035.3 ±10.4 in geometry. Varying tasks in a fixed scene. Table IIc shows whether VLAs can flexibly respond to different instructions while holding the scene fixed. Results expose brittle language grounding: π 0.5 was highly sensitive to object choice; 70% success for “Select the cordless drill and put it on the table” but 0% when replacing “cordless drill” with “blue hammer”. Error analysis reveals consistent object confusion patterns: visually similar distractors (pumpkin vs. orange, drill vs. hammer) frequently override language-specified targets (see Table. VI in Appendix for detailed analysis). These findings indicate that VLA language grounding is highly sensitive to the specific object-instruction pairings seen during training, rather than reflecting generalizable language-to-object binding. D. Sensitivity and robustness We perform a set of variations given two basic tasks and observe the outcome, for example, we only change the target object to pick or the target for place, this is akin to domain randomization [29] as illustrated in Fig. 6. The following variations are considered, variations 1) in wrist and external camera poses; 2) object poses; 3) in visual features, including background and table textures; and 4) in lighting, including saturation and hue. Table I illustrates the results for all experiments. Visual and Lighting variations.We vary the lighting conditions via color temperature shifts, lighting exposure and strong directional light that generates shadows as the robot is moving. Lighting: VLAs were robust to changes in lighting conditions, with 90–100% success across shadow variations, color temperature shifts, and 500×intensity changes. Visual appearance: Variations over 10 background textures and 4 table textures had minimal impact (<5% degradation), suggesting generalization to scene appearance changes. Camera variation sensitivity analysis. We infer posteriors over camera displacement conditioned on task success (Fig. 8). Fig. 8: Results of the sensitivity analysis using MNPE. Policies were highly sensitive to wrist-camera displacement from the nominal pose, indicating strong dependence on wrist-mounted camera calibration. Success also peaked for objects placed at approximately 0.5m from the robot, likely due to robot reachability. TABLE IV: Overall success rates across real and simulation environments across 6 selected simple tasks. Environmentπ 0.5 π 0 -FASTπ 0 PaliGemma Real79.534.163.20.0 Sim74.042.018.04.0 Camera poses were randomized in both orientation and position for 10 episodes each. Displacement is calculated with respect to the nominal position of the cameras. Across all policies, the wrist-camera posterior is sharply concentrated near zero, indicating that successful execution often required the wrist camera to remain close to its nominal pose, while performance is more tolerant to external camera position changes. This indicates performance is critically dependent on wrist camera than external camera. Object pose variation sensitivity analysis. We randomize initial object poses via a uniform distribution of 10cm, 20cm, and 30cm within its nominal placement (usually in front of the robot) for 10 episodes each. We then infer posteriors over initial object poses conditioned on task success (see Fig. 8), relative to the robot pose. We observe a strong peak over 0.5m from the robot’s origin, suggesting that objects placed at this distance has the highest probability of success, likely due to reachability. E. Real Robot Verification We evaluated the same policies on a small set of six simple real-robot tasks and compared success rates to matched simulation evaluations (Table IV).π 0.5 achieved 79.5% success in the real world, close to its 74.0% success in simulation, suggesting that RoboLab can provide a reasonable proxy for this policy on these task types.π 0 -FASTachieved 34.1% success in the real world and 42.0% in simulation, showing a similar trend. π 0 was a notable outlier, reaching 63.2% success on the real robot but only 18.0% in simulation; qualitatively, this policy appeared tuned to reliably grasp single objects, which matched the selected real-robot tasks. We leave deeper investigation of policy-specific sim-to-real deviations to future work. V. LIMITATIONS While RoboLab provides a flexible and scalable framework for evaluating language-conditioned manipulation, it currently focuses on rigid-body tabletop scenes and does not fully capture the challenges of deformable object manipulation (e.g., cloth, cables, bags). Moreover, many contact-rich skills that require precise force control, compliant interaction, or complex frictional dynamics are underrepresented and dependent on the physics simulation fidelity, limiting Robolab’s coverage of fine- grained, low-level control tasks. Finally, although evaluation in high-fidelity simulation is a strong proxy for real-world performance, a residual visual distribution shift remains. This gap needs to be characterized further both by analyzing the behavior and robustness of the visual perception stack and through extensive validation on real-world deployments. VI. CONCLUSION Recent benchmarking efforts have made significant strides in scalable robot evaluation, but they primarily assess robustness to perturbations of training environments rather than true task generalization to novel scenarios. RoboLab addresses this gap by evaluating real world policies in a high-fidelity simulation, structured evaluation vectors that decompose policy competence into visual, procedural, and relational dimensions, and a set of sensitivity analysis set of novel analysis that provides insight into policy behavior for robotics. Our benchmarking framework enables the community to critically answer the question of generalization and performance. At the same time, the framework is designed to be pragmatically usable: new tasks can be authored in minutes by arranging objects on a tabletop and attaching language instructions, and a generative scene–task–environment workflow that supports continuous benchmark evolution. REFERENCES [1]Pranav Atreya, Karl Pertsch, Tony Lee, Moo Jin Kim, Arhan Jain, Artur Kuramshin, Clemens Eppner, Cyrus Neary, Edward Hu, Fabio Ramos, et al. Roboarena: Distributed real-world evaluation of generalist robot policies. In Proceedings of the Conference on Robot Learning (CoRL 2025), 2025. [2]Sivakumar Balasubramanian, Alejandro Melendez- Calderon, and Etienne Burdet. A robust and sensitive metric for quantifying movement smoothness. IEEE TransactionsonBiomedicalEngineering,59(8): 2126–2136, 2012. doi: 10.1109/TBME.2011.2179545. [3]Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, Linguang Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Julian Engel, and Tomas Hodan. HOT3D: Hand and object tracking in 3D from egocentric multi-view videos. CVPR, 2025. [4]Lucas Beyer, Andreas Steiner, Andr ́ e Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Key- sers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bo ˇ snjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, and Xiaohua Zhai. Paligemma: A versatile 3b vlm for transfer, 2024. URL https://arxiv.org/ abs/2407.07726. [5]Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xi- aoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky.π 0 : A vision-language- action flow model for general robot control, 2026. URL https://arxiv.org/abs/2410.24164. [6]Rishit Dagli, Donglai Xiang, Vismay Modi, Charles Loop, Clement Fuji Tsang, Anka He Chen, Anita Hu, Gavriel State, David I.W. Levin, and Maria Shugrina. Vomp: Predicting volumetric mechanical property fields. arXiv preprint, 2025. [7]Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, et al. Maniskill2: A unified benchmark for generalizable manipulation skills. arXiv preprint arXiv:2302.04659, 2023. [8]Andrew Guo, Bowen Wen, Jianhe Yuan, Jonathan Trem- blay, Stephen Tyree, Jeffrey Smith, and Stan Birchfield. HANDAL: A dataset of real-world manipulable object categories with pose annotations, affordances, and recon- structions. In IROS, 2023. [9] Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Homer Walke, Anna Walling, Haohuan Wang, Lili Yu, and Ury Zhilinsky.π 0.5 : a vision-language- action model with open-world generalization, 2025. URL https://arxiv.org/abs/2504.16054. [10]Arhan Jain, Mingtong Zhang, Kanav Arora, William Chen, Marcel Torne, Muhammad Zubair Irshad, Sergey Zakharov, Yue Wang, Sergey Levine, Chelsea Finn, Wei- Chiu Ma, Dhruv Shah, Abhishek Gupta, and Karl Pertsch. Polaris: Scalable real-to-sim evaluations for generalist robot policies, 2025.URL https://arxiv.org/abs/2512. 16881. [11] Stephen James, Zicong Ma, David R. Arrojo, and Andrew J. Davison. RLBench: The Robot Learning Benchmark & Learning Environment. RAL, 2020. [12] Yash Jangir, Yidi Zhang, Kashu Yamazaki, Chenyu Zhang, Kuan-Hsun Tu, Tsung-Wei Ke, Lei Ke, Yonatan Bisk, and Katerina Fragkiadaki. Robotarena∞: Scalable robot benchmarking via real-to-sim translation, 2025. URL https://arxiv.org/abs/2510.23571. [13]Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ash- win Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yun- liang Chen, Kirsty Ellis, Peter David Fagan, Joey Hejna, Masha Itkina, Marion Lepert, Yecheng Jason Ma, Patrick Tree Miller, Jimmy Wu, Suneel Belkhale, Shivin Dass, Huy Ha, Arhan Jain, Abraham Lee, Youngwoon Lee, Marius Memmel, Sungjae Park, Ilija Radosavovic, Kaiyuan Wang, Albert Zhan, Kevin Black, Cheng Chi, Kyle Beltran Hatch, Shan Lin, Jingpei Lu, Jean Mercat, Abdul Rehman, Pannag R Sanketi, Archit Sharma, Cody Simpson, Quan Vuong, Homer Rich Walke, Blake Wulfe, Ted Xiao, Jonathan Heewon Yang, Arefeh Yavary, Tony Z. Zhao, Christopher Agia, Rohan Baijal, Mateo Guaman Castro, Daphne Chen, Qiuyu Chen, Trinity Chung, Jaimyn Drake, Ethan Paul Foster, Jensen Gao, Vitor Guizilini, David Antonio Herrera, Minho Heo, Kyle Hsu, Jiaheng Hu, Muhammad Zubair Irshad, Donovon Jackson, Char- lotte Le, Yunshuang Li, Kevin Lin, Roy Lin, Zehan Ma, Abhiram Maddukuri, Suvir Mirchandani, Daniel Morton, Tony Nguyen, Abigail O’Neill, Rosario Scalise, Derick Seale, Victor Son, Stephen Tian, Emi Tran, Andrew E. Wang, Yilin Wu, Annie Xie, Jingyun Yang, Patrick Yin, Yunchu Zhang, Osbert Bastani, Glen Berseth, Jeannette Bohg, Ken Goldberg, Abhinav Gupta, Abhishek Gupta, Dinesh Jayaraman, Joseph J Lim, Jitendra Malik, Roberto Mart ́ ın-Mart ́ ın, Subramanian Ramamoorthy, Dorsa Sadigh, Shuran Song, Jiajun Wu, Michael C. Yip, Yuke Zhu, Thomas Kollar, Sergey Levine, and Chelsea Finn. Droid: A large-scale in-the-wild robot manipulation dataset. 2024. [14]Hadas Kress-Gazit, Kunimatsu Hashimoto, Naveen Kup- puswamy, Paarth Shah, Phoebe Horgan, Gordon Richard- son, Siyuan Feng, and Benjamin Burchfiel.Robot learning as an empirical science: Best practices for policy evaluation, 2024. URL https://arxiv.org/abs/2409.09491. [15]Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Mart ́ ın-Mart ́ ın, Chen Wang, Gabrael Levine, Wensi Ai, Benjamin Martinez, et al. Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation. arXiv preprint arXiv:2403.09227, 2024. [16] Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941, 2024. [17] Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela, and Stan Birchfield. Multi-view fusion for multi- level robotic scene understanding. In IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 6817–6824, 2021. doi: 10.1109/IROS51168. 2021.9635994. [18] Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. Evaluating text-to-visual generation with image-to-text generation, 2024. URL https://arxiv.org/abs/2404.01291. [19] Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36:44776– 44791, 2023. [20]Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard. CALVIN: A Benchmark for Language- Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks. IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022. [21]Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano-Mu ̃ noz, Xinjie Yao, Ren ́ e Zurbr ̈ ugg, Nikita Rudin, Lukasz Wawrzyniak, Milad Rakhsha, Alain Denzler, Eric Heiden, Ales Borovicka, Ossama Ahmed, Iretiayo Akinola, Abrar Anwar, Mark T. Carlson, Ji Yuan Feng, Animesh Garg, Renato Gasoto, Lionel Gulich, Yijie Guo, M. Gussert, Alex Hansen, Mihir Kulkarni, Chenran Li, Wei Liu, Viktor Makoviychuk, Grzegorz Malczyk, Hammad Mazhar, Ma- soud Moghani, Adithyavairavan Murali, Michael Nosewor- thy, Alexander Poddubny, Nathan Ratliff, Welf Rehberg, Clemens Schwarke, Ritvik Singh, James Latham Smith, Bingjie Tang, Ruchik Thaker, Matthew Trepte, Karl Van Wyk, Fangzhou Yu, Alex Millane, Vikram Ramasamy, Remo Steiner, Sangeeta Subramanian, Clemens Volk, CY Chen, Neel Jawale, Ashwin Varghese Kuruttukulam, Michael A. Lin, Ajay Mandlekar, Karsten Patzwaldt, John Welsh, Huihua Zhao, Fatima Anes, Jean-Francois Lafleche, Nicolas Mo ̈ enne-Loccoz, Soowan Park, Rob Stepinski, Dirk Van Gelder, Chris Amevor, Jan Carius, Jumyung Chang, Anka He Chen, Pablo de Heras Ciechom- ski, Gilles Daviet, Mohammad Mohajerani, Julia von Muralt, Viktor Reutskyy, Michael Sauter, Simon Schirm, Eric L. Shi, Pierre Terdiman, Kenny Vilella, Tobias Widmer, Gordon Yeoman, Tiffany Chen, Sergey Grizan, Cathy Li, Lotus Li, Connor Smith, Rafael Wiltz, Kostas Alexis, Yan Chang, David Chu, Linxi ”Jim” Fan, Farbod Farshidian, Ankur Handa, Spencer Huang, Marco Hutter, Yashraj Narang, Soha Pouya, Shiwei Sheng, Yuke Zhu, Miles Macklin, Adam Moravanszky, Philipp Reist, Yun- rong Guo, David Hoeller, and Gavriel State. Isaac Lab: A GPU-accelerated simulation framework for multi-modal robot learning. arXiv preprint arXiv:2511.04831, 2025. URL https://arxiv.org/abs/2511.04831. [22]Nicolas Moenne-Loccoz, Ashkan Mirzaei, Or Perel, Riccardo de Lutio, Janick Martinez Esturo, Gavriel State, Sanja Fidler, Nicholas Sharp, and Zan Gojcic.3d gaussian ray tracing: Fast tracing of particle scenes. ACM Transactions on Graphics and SIGGRAPH Asia, 2024. [23] NVIDIA. Isaac Sim. URL https://github.com/isaac-sim/ IsaacSim. [24]NVIDIA, Johan Bjorck, Nikita Cherniadev Fer- nando Casta ̃ neda, Xingye Da, Runyu Ding, Linxi ”Jim” Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, Guilin Liu, Edith Llontop, Loic Magne, Ajay Mandlekar, Avnish Narayan, Soroush Nasiriany, Scott Reed, You Liang Tan, Guanzhi Wang, Zu Wang, Jing Wang, Qi Wang, Jiannan Xiang, Yuqi Xie, Yinzhen Xu, Zhenjia Xu, Seonghyeon Ye, Zhiding Yu, Ao Zhang, Hao Zhang, Yizhou Zhao, Ruijie Zheng, and Yuke Zhu. GR00T N1: An open foundation model for generalist humanoid robots. In ArXiv Preprint, March 2025. [25]OpenAI, Josh Achiam, Steven Adler, Sandhini Agar- wal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Ale- man, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna- Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Sim ́ on Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan,Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo,Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David M ́ ely, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeon- woo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cer ́ on Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774. [26]OpenAI, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexander Prokofiev, Alexander Wei, Allison Tam, Ally Bennett, Ananya Kumar, Andre Saraiva, Andrea Vallone, Andrew Duberstein, Andrew Kondrich, Andrey Mishchenko, Andy Applebaum, Angela Jiang, Ashvin Nair, Barret Zoph, Behrooz Ghorbani, Ben Rossen, Benjamin Sokolowsky, Boaz Barak, Bob McGrew, Borys Minaiev, Botao Hao, Bowen Baker, Brandon Houghton, Brandon McKinzie, Brydon Eastman, Camillo Lugaresi, Cary Bassin, Cary Hudson, Chak Ming Li, Charles de Bourcy, Chelsea Voss, Chen Shen, Chong Zhang, Chris Koch, Chris Orsinger, Christopher Hesse, Clau- dia Fischer, Clive Chan, Dan Roberts, Daniel Kappler, Daniel Levy, Daniel Selsam, David Dohan, David Farhi, David Mely, David Robinson, Dimitris Tsipras, Doug Li, Dragos Oprica, Eben Freeman, Eddie Zhang, Edmund Wong, Elizabeth Proehl, Enoch Cheung, Eric Mitchell, Eric Wallace, Erik Ritter, Evan Mays, Fan Wang, Fe- lipe Petroski Such, Filippo Raso, Florencia Leoni, Foivos Tsimpourlas, Francis Song, Fred von Lohmann, Freddie Sulit, Geoff Salmon, Giambattista Parascandolo, Gildas Chabot, Grace Zhao, Greg Brockman, Guillaume Leclerc, Hadi Salman, Haiming Bao, Hao Sheng, Hart Andrin, Hessam Bagherinezhad, Hongyu Ren, Hunter Lightman, Hyung Won Chung, Ian Kivlichan, Ian O’Connell, Ian Osband, Ignasi Clavera Gilaberte, Ilge Akkaya, Ilya Kostrikov, Ilya Sutskever, Irina Kofman, Jakub Pachocki, James Lennon, Jason Wei, Jean Harb, Jerry Twore, Jiacheng Feng, Jiahui Yu, Jiayi Weng, Jie Tang, Jieqi Yu, Joaquin Qui ̃ nonero Candela, Joe Palermo, Joel Parish, Johannes Heidecke, John Hallman, John Rizzo, Jonathan Gordon, Jonathan Uesato, Jonathan Ward, Joost Huizinga, Julie Wang, Kai Chen, Kai Xiao, Karan Singhal, Karina Nguyen, Karl Cobbe, Katy Shi, Kayla Wood, Kendra Rimbach, Keren Gu-Lemberg, Kevin Liu, Kevin Lu, Kevin Stone, Kevin Yu, Lama Ahmad, Lauren Yang, Leo Liu, Leon Maksin, Leyton Ho, Liam Fedus, Lilian Weng, Linden Li, Lindsay McCallum, Lindsey Held, Lorenz Kuhn, Lukas Kondraciuk, Lukasz Kaiser, Luke Metz, Madelaine Boyd, Maja Trebacz, Manas Joglekar, Mark Chen, Marko Tintor, Mason Meyer, Matt Jones, Matt Kaufer, Max Schwarzer, Meghan Shah, Mehmet Yatbaz, Melody Y. Guan, Mengyuan Xu, Mengyuan Yan, Mia Glaese, Mianna Chen, Michael Lampe, Michael Malek, Michele Wang, Michelle Fradin, Mike McClay, Mikhail Pavlov, Miles Wang, Mingxuan Wang, Mira Murati, Mo Bavarian, Mostafa Rohaninejad, Nat McAleese, Neil Chowdhury, Neil Chowdhury, Nick Ryder, Nikolas Tezak, Noam Brown, Ofir Nachum, Oleg Boiko, Oleg Murk, Olivia Watkins, Patrick Chao, Paul Ashbourne, Pavel Izmailov, Peter Zhokhov, Rachel Dias, Rahul Arora, Randall Lin, Rapha Gontijo Lopes, Raz Gaon, Reah Miyara, Reimar Leike, Renny Hwang, Rhythm Garg, Robin Brown, Roshan James, Rui Shu, Ryan Cheu, Ryan Greene, Saachi Jain, Sam Altman, Sam Toizer, Sam Toyer, Samuel Miserendino, Sandhini Agarwal, Santiago Hernandez, Sasha Baker, Scott McKinney, Scottie Yan, Shengjia Zhao, Shengli Hu, Shibani Santurkar, Shra- man Ray Chaudhuri, Shuyuan Zhang, Siyuan Fu, Spencer Papay, Steph Lin, Suchir Balaji, Suvansh Sanjeev, Szymon Sidor, Tal Broda, Aidan Clark, Tao Wang, Taylor Gordon, Ted Sanders, Tejal Patwardhan, Thibault Sottiaux, Thomas Degry, Thomas Dimson, Tianhao Zheng, Timur Garipov, Tom Stasi, Trapit Bansal, Trevor Creech, Troy Peterson, Tyna Eloundou, Valerie Qi, Vineet Kosaraju, Vinnie Monaco, Vitchyr Pong, Vlad Fomenko, Weiyi Zheng, Wenda Zhou, Wes McCabe, Wojciech Zaremba, Yann Dubois, Yinghai Lu, Yining Chen, Young Cha, Yu Bai, Yuchen He, Yuchen Zhang, Yunyun Wang, Zheng Shao, and Zhuohan Li. Openai o1 system card, 2024. URL https://arxiv.org/abs/2412.16720. [27]Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. Fast: Efficient action tokenization for vision-language-action models, 2025. URL https://arxiv. org/abs/2501.09747. [28]Mohammad Nomaan Qureshi, Sparsh Garg, Francisco Yandun, David Held, George Kantor, and Abhishesh Silwal. Splatsim: Zero-shot sim2real transfer of rgb manipulation policies using gaussian splatting, 2024. URL https://arxiv.org/abs/2409.10161. [29]Jonathan Tremblay, Aayush Prakash, David Acuna, Mark Brophy, Varun Jampani, Cem Anil, Thang To, Eric Camer- acci, Shaad Boochoon, and Stan Birchfield. Training deep networks with synthetic data: Bridging the reality gap by domain randomization. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 969–977, 2018. [30]Qi Wu, Janick Martinez Esturo, Ashkan Mirzaei, Nicolas Moenne-Loccoz, and Zan Gojcic.3dgut: Enabling distorted cameras and secondary rays in gaussian splatting. Conference on Computer Vision and Pattern Recognition (CVPR), 2025. [31]Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. 2018. [32]Yandan Yang, Baoxiong Jia, Shujie Zhang, and Siyuan Huang. Sceneweaver: All-in-one 3d scene synthesis with an extensible and self-reflective agent, 2025. URL https: //arxiv.org/abs/2509.20414. [33]Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Avnish Narayan, Hayden Shively, Adithya Bellathur, Karol Hausman, Chelsea Finn, and Sergey Levine. Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning, 2021. URL https://arxiv.org/ abs/1910.10897. [34]Kaifeng Zhang, Shuo Sha, Hanxiao Jiang, Matthew Loper, Hyunjong Song, Guangyan Cai, Zhuo Xu, Xiaochen Hu, Changxi Zheng, and Yunzhu Li. Real-to-sim robot policy evaluation with gaussian splatting simulation of soft-body interactions. arXiv preprint arXiv:2511.04665, 2025. [35]Xueyang Zhou, Yangming Xu, Guiyao Tie, Yongchao Chen, Guowen Zhang, Duanfeng Chu, Pan Zhou, and Lichao Sun. Libero-pro: Towards robust and fair evalua- tion of vision-language-action models beyond memoriza- tion. [arXiv preprint arXiv:2510.03827], 2025. [36]Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Mart ́ ın- Mart ́ ın, Abhishek Joshi, Kevin Lin, Soroush Nasiriany, and Yifeng Zhu.robosuite: A modular simulation framework and benchmark for robot learning. In arXiv preprint arXiv:2009.12293, 2020. [37]Alex Zook, Fan-Yun Sun, Josef Spjut, Valts Blukis, Stan Birchfield, and Jonathan Tremblay.Grs: Generating robotic simulation tasks from real-world images. In Pro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 594–603, 2025. APPENDIX A DETAILS ON THE ROBOLAB BENCHMARK In this section we provide detail on the benchmark. RoboLab provides a set of∼300 object assets from well- known 3D pose estimation benchmarks; including YCB [31], HOT3D [3], HOPE [17], HANDAL [8], and VoMP [6]. Each asset contains a visual and collision mesh, with mass and friction properties added. Each object has a language description and an object label attached to it. This forms the catalog of objects used in the scenes, in either manual or LLM-scaled scene environments. RoboLab the RoboLab-120 benchmark contains 120 manu- ally generated tasks. These tasks span one or more categories. More details on the benchmark implementation will be provided in the code repository, which will be open sourced. We provide additional details on the results reported in the paper. Please refer to Table V for expanded results on overall performance on our benchmark; Table VI for more details on the language ablation; Table VII for details on the real- sim verification experiments; and Tables VIII-XII for detailed per-task results for each policy. APPENDIX B DETAILS OF MNPE SENSITIVITY ANALYSIS MNPE allows us to analyze the relationship between scene parameters and policy outcomes in a likelihood-free Bayesian inference setting. Variable Definitions. Letθ ∈ Θdenote the vector of varia- tion parameters. In our camera pose sensitivity experiments, θ = (d ext ,d wrist ) ∈ R 2 whered ext andd wrist represent the displacement of the external and wrist cameras from their reference configurations, respectively, in SE(3). Letx∈0, 1denote the binary task success indicator, and let π denote the robot policy being evaluated. Handling Mixed Parameters. For experiments involving both continuous parameters (e.g., pose distances) and discrete parameters (e.g., lighting levels, table materials), MNPE handles mixed continuous-discrete parameters through fac- torization:q φ (θ | x) = q φ (θ cont | θ disc ,x) · q φ (θ disc | x), where discrete components use softmax distributions and continuous components use normalizing flows. In our camera pose experiments, all parameters are continuous. Pose Distance Metric.Poses are represented as 7-DoF transformationsT = (p, q)comprising positionp∈ R 3 and unit quaternion orientationq ∈ H. We compute a weighted distance from the reference configuration: d(T, T ref ) =∥p− p ref ∥ 2 + β· d SO(3) (q, q ref ),(2) where the geodesic distance on SO(3) is: d SO(3) (q 1 , q 2 ) = 2 arccos (min(1,|q 1 · q 2 |)).(3) The weighting factorβ = 1.0balances translational (meters) and rotational (radians) components. For camera displacement, reference poses correspond to nominal camera mounting positions. For object pose, reference pose is the origin of the robot base. Prior Specification. We adopt non-informative uniform priors to avoid biasing the inference toward any particular parameter region. For continuous parameters normalized to the unit interval: p(θ) = m Y j=1 Uniform(0, 1) = 1, θ j ∈ [0, 1].(4) Training Objective. The neural network parametersφare optimized by minimizing the negative log-likelihood over the training dataset, D =(θ i ,x i ) N i=1 : L(φ) =− 1 N N X i=1 logq φ (θ i | x i ).(5) We train for 50 epochs using the Adam optimizer on data collected from the camera pose variation and initial pose variation experiments (see Table I). Importance Sampling Correction. Since the experimental data may sample parameters non-uniformly, we apply impor- tance sampling to recover the posterior under a uniform prior: p(θ | x)≈ p(θ) ̃p(θ) q φ (θ | x), where ̃p(θ)is the empirical proposal distribution. We correct posterior samples using importance weights: w i = p(θ i ) ̃p(θ i ) ,(6) where ̃p(θ)is estimated via Gaussian kernel density estimation on the training data. The effective sample sizeESS = 1/ P i ̄w 2 i quantifies the efficiency of this correction. Posterior Inference. Given a query observationx o (e.g., x o = 1for successful task completion), we drawN s = 5000 samples from the learned posterior: n θ (i) o N s i=1 ∼ q φ (θ | x o ).(7) Posterior Statistics.For each continuous parameter, we compute the posterior mean and 95% credible interval: ˆμ j = 1 N s N s X i=1 θ (i) j ,(8) CI (j) 95% = h Q 0.025 θ (i) j , Q 0.975 θ (i) j i ,(9) where Q α (·) denotes the α-quantile. This analysis reveals which variation parameters are most strongly associated with successful task outcomes: a posterior distribution tightly concentrated near zero indicates high sensitivity to that parameter (the policy requires it to remain near the reference value), while a broad posterior indicates robustness to variation. APPENDIX C DETAILS ON SCALING SCENE GENERATION We present additional implementation details on scene generation (Section I-D1). TABLE V: Expanded results of Table I: Overall performance of VLAs on RoboLab (Best viewed with zoom). Task Categories # π 0.5 π 0 -FASTGR00T N1.6π 0 Succ %ScoreSPARCSpeed (cm/s)Succ %ScoreSPARCSpeed (cm/s)Succ %ScoreSPARCSpeed (cm/s)Succ %ScoreSPARCSpeed (cm/s) Total12025.0%0.00−10.11 (±7.05)5.7 (±1.8)15.4%0.00−9.51 (±6.88)4.6 (±1.8)8.0%0.09−6.44 (±2.04)4.2 (±1.3)5.3%0.00−9.48 (±4.07)4.4 (±1.4) simple6426.7%0.00 −8.46 (±5.73)5.8 (±1.7)21.1%0.00 −8.00 (±6.63)4.6 (±1.7)12.3%0.12 −5.79 (±1.96)4.3 (±1.5)7.7%0.00 −8.52 (±4.31)4.5 (±1.4) moderate3927.7%0.01 −10.87 (±8.35)5.9 (±1.9)11.3%0.01 −10.17 (±5.89)4.7 (±2.0)7.1%0.08 −6.70 (±1.91)4.1 (±1.2)3.7%0.00 −10.36 (±3.41)4.3 (±1.4) complex1712.4%0.01 −14.55 (±6.06)5.0 (±1.6)3.5%0.00 −13.67 (±7.93)4.1 (±1.3)0.0%0.07 −7.36 (±1.99)4.3 (±1.1)0.0%0.00 −11.67 (±3.94)4.1 (±1.2) Procedural3417.9%0.01−12.85 (±8.28)5.1 (±1.7)3.8%0.00−11.04 (±6.31)4.5 (±1.5)0.0%0.06−6.35 (±1.47)3.9 (±1.2)0.3%0.00−10.77 (±3.43)4.2 (±1.2) affordance1215.0%0.00 −13.02 (±10.43)4.8 (±1.7)1.7%0.00 −10.73 (±4.68)4.1 (±1.5)0.0%0.00 −5.39 (±1.59)4.6 (±1.1)0.0%0.00 −10.59 (±3.01)4.5 (±1.5) reorientation615.0%0.07 −14.94 (±9.48)4.7 (±1.5)8.3%0.03 −11.35 (±6.43)4.7 (±1.5)0.0%0.03 −6.89 (±1.44)3.2 (±0.6)1.7%0.01 −12.06 (±2.30)4.0 (±0.7) stacking626.7%0.00 −10.04 (±4.68)6.2 (±1.5)1.7%0.00 −8.82 (±2.46)5.1 (±1.5)0.0%0.09 −6.23 (±1.35)4.2 (±1.3)0.0%0.00 −9.52 (±3.35)3.8 (±0.9) Relational4232.6%0.00−9.09 (±5.84)6.1 (±1.8)24.8%0.00−8.73 (±6.71)5.1 (±2.0)11.2%0.07−6.30 (±2.22)4.2 (±1.4)7.4%0.00−9.26 (±3.46)4.7 (±1.5) conjunction848.8%0.00 −8.23 (±7.00)6.3 (±1.5)30.0%0.00 −7.57 (±4.15)5.4 (±1.7)4.0%0.13 −6.54 (±2.13)3.7 (±1.1)17.5%0.00 −7.90 (±2.63)4.8 (±1.2) counting757.1%0.00 −9.54 (±9.70)7.4 (±2.4)42.9%0.00 −10.04 (±6.83)5.5 (±2.3)38.0%0.14 −6.75 (±2.59)4.3 (±1.5)10.0%0.00 −11.60 (±4.90)4.2 (±1.3) spatial2920.0%0.00 −9.21 (±3.97)5.7 (±1.5)17.2%0.00 −8.74 (±7.19)4.9 (±2.0)7.0%0.04 −6.15 (±2.15)4.3 (±1.4)3.4%0.00 −9.07 (±2.95)4.8 (±1.7) Visual8418.2%0.00−10.63 (±6.35)5.5 (±1.8)8.9%0.00−10.43 (±7.81)4.2 (±1.5)5.8%0.10−6.73 (±2.05)4.2 (±1.2)2.3%0.00−9.63 (±3.73)4.2 (±1.4) color2614.6%0.01 −10.34 (±6.02)5.5 (±1.6)7.3%0.00 −9.15 (±6.72)4.6 (±1.5)5.6%0.13 −6.05 (±2.00)4.2 (±1.3)2.7%0.00 −9.28 (±3.05)4.2 (±1.4) semantics6019.3%0.00 −10.87 (±6.51)5.6 (±1.8)11.3%0.00 −11.16 (±7.72)4.1 (±1.6)5.9%0.09 −7.10 (±2.02)4.3 (±1.2)3.1%0.00 −9.95 (±4.03)4.3 (±1.4) size613.3%0.00 −9.43 (±6.08)4.7 (±2.1)8.3%0.00 −8.66 (±11.42)4.2 (±1.3)0.0%0.00 −6.51 (±1.34)2.5 (±0.3)0.0%0.00 −8.01 (±2.59)3.7 (±0.9) Fig. 9: (Left) We show one of our Gaussian Splat + Mesh scenes in RoboLab. This scene has a Gaussian splat background with a collision mesh for the splat estimated with 3DGRUT [22,30], and a mesh foreground. All objects in the scene have spatially varying density, and thus mass is estimated with VoMP [6]. (Right) We show a VLA running a task in this scene. A. Stage I: Predicates for Semantic Planning The following predicates are used: • PlaceIn(x,y): Objectxmust be contained withiny(e.g., fruit in a bowl). • PlaceOn(x,y): Objectxis supported byy(e.g., mug on a coaster). • ClusterAround(x,y i ) : Objectxacts as an anchor for a group B i . • PlaceAnywhere(x): Objectxis placed freely on the global support surface (table). If a predicate refers to a non-existent anchor, it is downgraded to aPlaceAnywhereconstraint to preserve the object in the scene. B. Stage I: Geometric Constraint Solving For Global Placement (PlaceAnywhere), we utilize rejec- tion sampling on the global table surface bounds, checking collision against all currently placed objects using SAT on OBBs. To handle high-density scenes, we employ two strategies: (1) an adaptive relaxation loop that progressively increases collision margins if a valid layout is not found, and (2) a stochastic perturbation step that randomly jitters all object positions when the solver converges to a local minimum (Algorithm 1). For Stacking (PlaceOn(b i ,b support )), we sample positions on the top surface ofb support using rejection sampling (up to K = 20attempts) to find a positionp xy such thatOBB(b i )∩ OBB(b existing ) = ∅for all previously placed objects on the same support (Algorithm 2). For Containment (PlaceIn(b i ,b container )), we compute the available interior volume ofb container using its bounding box dimensionsd b . We employ a packing heuristic that discretizes the container’s floor into a grid with resolution s = max(d x i , d y i ) + ε margin . A cell(u,v)is considered valid if it is unoccupied and within the container’s bounds scaled by factorγ = 0.7to avoid edge collisions. We assigno i to the first valid cell, setting its height z i = z container + h container /2. C. Baseline Method To validate the efficacy of our hierarchical approach, we implement a robust baseline inspired by standard domain randomization techniques. The baseline operates in a single pass without iterative feedback. The process begins with the LLM selecting a list of objectsOand suggesting a grid layout (rows R×columnsC) for the table surface. The table surface is then divided intoR× Crectangular cells, and objects are assigned to cells sequentially. Within each cellk, the object’s position is jittered uniformly:p xy ∼U (center k −w/4, center k +w/4). This ensures basic separation but precludes complex stacking or containment, as objects are simply placed at a safe height z = z table +h obj /2. Finally, we run the same physics simulation pass as in our method to allow objects to settle under gravity, resolving minor inter-penetrations but without the capability to correct semantic or structural failures. D. Experiments We compare our scene generation with the baseline method using popular scene generation metrics, VQA score [18], GPT preference where it is shown two images each from the baseline or our method and is asked to pick one, and following Yang et al.[32], we report the visual realism (Real.), functionality (Func.), layout correctness (Lay.), Quality (Qual.) and scene completeness (Comp.) scores. We use GPT-4o [25] to generate scenes from our method and baseline; and use GPT-4.1 [25] for the evaluations. These metrics are computed on rendered RGB images from two viewpoints: a frontal view aligned with the table axis (camera at(1.0, 0.0, 0.7)looking at table center) and an angled perspective view (camera at (−0.3, 0.3, 0.7)). TABLE VI: Expanded results of Table I: Language understanding ablations.(a) VLA performance degrades with abstract or vague language instructions.π 0 , PaliGemma and GR00T N1.6 excluded as no tasks were able to be completed. (b) Performance drops as scene complexity increases. (c) VLAs show brittle language grounding, with consistent object confusion patterns across different instructions in the same scene. (a) Effect of language specificity on task performance. π 0.5 π 0 -FAST TaskSucc %ScoreTime (s)Wrong ObjSPARCl (m)v (cm/s)Succ %ScoreTime (s)Wrong ObjSPARCl (m)v (cm/s) Bananas Out Of Bin Task “Take all the bananas out of the grey bin and put it on the table.”500.1326.9618-10.542.265.7300.0534.294-6.462.715.3 “Take the bananas out”400.2238.7714-9.902.635.2100.1524.533-9.462.895.0 “Empty the grey bin”100.0779.47120-13.024.064.4700.1151.2514-7.463.796.2 White Mugs In Bin Task “Put the white mugs in the grey bin”800.5044.970-6.503.397.1200.2254.007-7.393.926.3 “Put the mugs in the bin”900.5035.131-5.793.228.4100.1151.8022-6.744.287.0 “Put away mugs”00.00-0-7.514.647.300.00-53-7.204.056.3 Remove Measuring Spoons from the Plate Task “Put the orange measuring cup and the blue measuring cup outside of the plate”200.4795.1725-17.914.973.000.31-5-12.7510.205.4 “Clear the plate”00.08-15-13.439.765.100.10-0-23.864.932.5 (b) Effect of scene complexity on task performance. π 0.5 π 0 -FASTπ 0 GR00T N1.6 SceneSucc %Wrong object grabbedSucc %Wrong object grabbedSucc %Wrong object grabbedSucc %Wrong object grabbed Task: “Pack boxed foods into the bin” 1 Box/Can10-0soup can0soup can0soup can, mustard 2 Boxes/Cans0-0soup can0soup can0soup can, bin 3 Boxes/Cans0spam can, soup can0soup can0soup can, spam can0spam can, soup can Task: “Pack canned foods into the bin” 1 Box/Can70bin30-0-0- 2 Boxes/Cans30bin10-0-0- 3 Boxes/Cans20bin, pudding box0-0pudding box0- (c) Instruction sensitivity within fixed scenes. π 0.5 π 0 -FASTπ 0 GR00T N1.6 TaskSucc %Wrong object grabbedSucc %Wrong object grabbedSucc %Wrong object grabbedSucc %Wrong object grabbed Fruit Plate Scene “Move an orange or a lime to the wood bowl”50pumpkin, redonion0-0-0redonion, woodenbowl “Move an orange to the white bowl”0pumpkin0-0pumpkin, woodenbowl0redonion, woodenbowl, pomegranate “Put the onion in the wood bowl”70pumpkin, woodenbowl10-20-20woodenbowl, pumpkin, storagebox “Put the onion on the plate”0pumpkin0-0woodenbowl, lime0woodenbowl, pumpkin Tools Cleanup Scene “Put hammers in the right bin and ignore everything else”20rightbin, drill0drill0drill0- “Put hammers in the left bin”10leftbin0drill0drill0springclamp Tools Selection Scene “Take out all the hammers and put it on the table”0clamp, leftbin0-0drill, centerbin, rightbin0leftbin, centerbin “Select the cordless drill and put it on the table”70redhammer, woodhammer, bluehammer50clamp20centerbin, rightbin30bluehammer, redhammer, clamp “Select the blue hammer and put it on the table”0leftbin0redhammer10centerbin, rightbin, leftbin0leftbin, centerbin We show quantitative comparisons across 100 generated scenes for our method compared to the baselines in Table XIII. We show the quantitative comparisons split by the number of objects in the scene ([1, 5]objects,[6, 15]objects, and[16, 20] objects) in Table XIV. We show the quantitative comparisons split across the 10 scene themes we use in Table XV. We find our method consistently outperforms the baseline across all metrics, with particularly large gains in visual realism and semantic functionality. APPENDIX D DETAILS ON TASK GENERATION EVALUATION We evaluated the quality of our task generation method using an LLM-as-judge framework. Tasks were generated by prompting an LLM (o1 [26]) with scene descriptions and our Category templates 1 . For each generated task, we extracted the natural language instruction and the corresponding termination conditions (success criteria implemented as predicate functions) through static analysis of the generated Python code. We then prompted an LLM (also o1 [26]) to assess alignment between the instruction and the programmatic success conditions across six dimensions: relation match (whether the spatial/logical relationship is preserved), target match (correctness of goal state), object match (whether referenced objects are correct), quantifier match (handling of “all,” “any,” or specific counts), instruction clarity (unambiguous and well-formed language), and physical feasibility (whether the task is achievable given typical robot capabilities). Each dimension was scored on a 0–1 scale, and we computed an aggregate alignment score as the 1 For 2 simple scenes, we generated 1 task for each of the 7 categories and for the remaining 57 scenes, we generated 2 tasks. This produced57∗ 7∗ 2 + 2∗ 7∗ 1 = 812 tasks TABLE VII: Expanded results of Table IV: Comparison of success rates (%) between real and simulation environments per task. Success % TaskEnvπ 0.5 π 0 -FASTπ 0 PaliGemma BananaInBowl Real90.980.080.0- Sim100.070.020.00.0 BananaAndCubeInBowl Real100.030.090.00.0 Sim90.060.050.00.0 BananasOutOfBin Real83.360.0100.0- Sim50.030.00.00.0 FoodPacking2Cans Real40.00.040.0- Sim---- PickOranges Real33.30.040.0- Sim60.00.00.00.0 ToolsPickingDrill Real100.00.00.00.0 Sim70.050.020.020.0 Total Real79.534.163.20.0 Sim74.042.018.04.0 weighted mean. The model additionally provided a categorical verdict—aligned, partially aligned, or misaligned—based on whether the termination conditions would correctly evaluate task success as described in the instruction. Table XVI shows that our method can successfully generate a variety of types of tasks appropriate to the assets in the scene. The Alignment score represents the overall instruction-code alignment, aggregating the six sub-dimensions. Clarity mea- sures whether instructions are unambiguous and grammatically well-formed. Feasibility assesses physical realizability of the task. Match combines the four semantic dimensions (relation, target, object, quantifier) into a single score reflecting how accurately the code captures the instruction’s intent. Verdict reports the percentage of tasks judged as fully aligned versus partially aligned (misaligned tasks, comprising approximately 1% of the dataset, are omitted for brevity). We additionally compute scene coverage metrics: object coverage measures the fraction of manipulable objects in each scene that appear in at least one generated task, while predicate coverage measures the fraction of available termination predicates used across tasks for that scene. All evaluations use temperature 0 for reproducibility, with automatic retry logic to handle rate limits. The evaluation reveals strong overall task generation quality, with0.91mean alignment and76%of tasks receiving full align- ment verdicts. Performance varies by category: conjunction and recognition tasks achieve the highest alignment (0.97and0.96), likely because their success conditions map straightforwardly to compositional predicates, while color-based tasks show lower alignment (0.81), reflecting the challenge of grounding color references to specific object instances. High clarity (0.96) and semantic match (0.95) scores indicate that the generated instructions are well-formed and the termination conditions capture the intended semantics, though feasibility scores are slightly lower for spatial tasks (0.89) where precise placement requirements may exceed typical manipulation tolerances. The 88%object coverage demonstrates good utilization of scene Algorithm 1 Spatial Constraint Solver Input: Objects B, Predicates P , Table Bounds L max Output:2Dcoordinates(x,y,θ)forallbase objects 1: Margins M ← [μ, 1.25μ, 1.5μ, 2.0μ] 2: for all margin∈ M do 3: Phase 1: Initialization 4:Randomize (x,y) for all loose objects inside L max 5:for all p∈ P do 6:if p.type == place-on-base then 7:p.object.(x,y,θ)← (p.x,p.y,p.yaw) 8:else if p.type == cluster-around then 9:PolarPlace(p.targets, p.anchor, p.radius) 10:end if 11:end for 12: Phase 2: Relative Constraints 13:while constraints not satisfied do 14:ApplyRelativeConstraints(P ) 15:end while 16:ApplyOrientations(P ) 17: Phase 3: Collision Resolution 18:for k = 1 to K max do 19: C ← FindCollisions(B, margin) 20:if C =∅ then 21:return Success 22:end if 23:if |C| not decreasing for 10 steps then 24:PerturbPositions(B) 25:end if 26:for all (o i ,o j )∈ C do 27:ResolveOverlap(b i ,b j , margin) 28:ClampToBounds(b i ,b j ,L max ) 29:end for 30:end for 31: end for 32: return Failure assets, while the lower predicate coverage (29%) suggests the generator favors a subset of reliable predicates rather than exploring the full space of available success conditions— a conservative strategy that likely contributes to the high alignment scores. Overall, these results demonstrate our method can successfully generate a variety of types of tasks appropriate to the assets in the scene. Algorithm 2 Physical Placement Solver Input: Objects B, Predicates P , Solved Base Poses Output:3Dcoordinates(x,y,z)forallob- jects 1: Solve Stacking: 2: for all p∈ P where p.type == place-on do 3: s← p.support 4: B peers ←b ′ | b ′ is already on s 5:(x,y)← FindSpot(s,p.object,B peers ) 6: p.object.z ← s.z + s.height + p.object.height/2 7: p.object.(x,y)← (x,y) 8: end for 9: Solve Containment: 10: for all p∈ P where p.type == place-in do 11: c← p.container 12:if TotalArea(p.objects) > 0.8× Area(c) then 13: p.objects← SortAndFilter(p.objects,c.capacity) 14:end if 15:(R,C)← ComputeGridDimensions(c.dims,|p.objects|) 16:for i = 0 to |p.objects|− 1 do 17:(r,c)← (i//C,i%C) 18:(x loc ,y loc )← GridCellCenter(r,c,c.dims) 19:Jitter(x loc ,y loc ) 20: p.objects[i].(x,y)← c.(x,y) + (x loc ,y loc ) 21: p.objects[i].z ← c.z + c.height/2 + buffer 22:end for 23: end for 24: return Success TABLE VIII: Detailed results forπ 0.5 . Task NameSucc%ScoreTime(s)SPARCPathLen(m)Speed(cm/s)WrongObjNames TOTAL (120 tasks)25.00.00327.55 ± 28.49−10.11 ± 7.055.58 ± 10.135.7 ± 1.8 AnimalsInBinTask0.00.000-−9.62 ± 2.854.58 ± 1.255.1 ± 1.4- AppleAndYogurtInBowlTask80.00.00048.43 ± 23.22−20.74 ± 33.6221.15 ± 49.406.2 ± 1.3- BBQSauceInBinTask 0.00.000-−8.98 ± 1.764.08 ± 1.194.4 ± 1.2- BagelsOnPlateTask 0.00.000-−8.30 ± 2.312.07 ± 0.434.2 ± 0.5- BananaInBowlTask 100.0-14.95 ± 8.96−4.50 ± 0.840.97 ± 0.376.6 ± 1.0- BananaOnPlateTask100.0-10.43 ± 1.99−4.10 ± 1.010.74 ± 0.066.9 ± 1.2- BananaThenRubiksCubeTask 70.00.00027.54 ± 8.97−9.44 ± 10.737.43 ± 16.686.2 ± 0.5- BananasInBinOneMoreTask 100.0-10.45 ± 3.24−4.26 ± 1.190.98 ± 0.229.3 ± 1.4- BananasInBinThreeTotalTask100.0-12.35 ± 5.17−4.14 ± 0.851.11 ± 0.408.9 ± 1.1- BananasInCrateTask 90.00.00011.03 ± 4.75−9.87 ± 18.529.59 ± 26.929.4 ± 2.1- BananasOutOfBinTask90.00.00041.70 ± 9.24−9.07 ± 2.612.67 ± 0.845.6 ± 0.8- BigPumpkinInBinTask 0.00.000-−10.15 ± 4.222.21 ± 0.973.5 ± 1.5- BlackItemsInBinTask 0.00.000-−12.21 ± 2.524.55 ± 0.723.8 ± 0.5- BlockStackingOrderAgnosticTask30.00.00062.84 ± 15.40−11.33 ± 4.829.40 ± 14.176.0 ± 1.5- BlockStackingSpecifiedOrderTask 0.00.000-−12.01 ± 2.774.80 ± 0.995.0 ± 1.1- BlocksInBinTask 0.00.000-−10.66 ± 3.209.55 ± 1.246.1 ± 0.8- BowlInBinTask60.00.00028.43 ± 13.32−12.21 ± 11.784.06 ± 7.574.6 ± 1.4- BowlStackingLeftOnRightTask 10.00.00015.33−6.34 ± 1.311.25 ± 0.316.0 ± 1.3- BowlStackingRightOnLeftTask30.00.00014.00 ± 2.43−6.73 ± 1.961.06 ± 0.215.6 ± 1.4- ButterAboveRaisinTask0.00.000-−9.53 ± 2.021.92 ± 0.344.7 ± 0.8- CannedFoodInBinTask40.00.00030.80 ± 13.63−12.25 ± 12.394.65 ± 6.915.0 ± 0.9- ClampInRightBinTask0.00.000-−7.70 ± 1.483.28 ± 0.755.2 ± 1.2- CleanUpToysTask0.00.000-−14.94 ± 1.4619.68 ± 1.316.3 ± 0.4- ClearOrganicObjectsTask0.00.000-−13.93 ± 1.7918.60 ± 2.277.4 ± 0.8- ClutterPlasticTask10.00.000178.73−13.05 ± 3.2311.30 ± 2.276.1 ± 1.2- ClutterPumpkinTask0.00.000-−10.78 ± 2.543.73 ± 1.474.3 ± 1.3- CoffeePotInBinTask0.00.000-−9.58 ± 2.061.79 ± 0.543.0 ± 0.8- CondimentsInBinTask0.00.000-−17.09 ± 6.065.38 ± 2.102.9 ± 1.1- CookingClearPlateTask0.00.000-−18.86 ± 6.105.20 ± 2.172.8 ± 1.2- CookingPickPastaToolTask0.00.000-−8.90 ± 1.204.18 ± 0.716.5 ± 1.1- CubesAndBlocksInBinTask10.00.000127.27−12.86 ± 2.8518.51 ± 13.275.9 ± 1.0- DishesInBinTask20.00.000132.00 ± 9.99−13.96 ± 1.9113.36 ± 4.966.7 ± 0.6- ElectronicsInBinTask10.00.00064.60−14.90 ± 5.2911.49 ± 14.914.3 ± 1.2- FoodPacking1BoxesTask 10.00.00027.47−7.40 ± 1.175.18 ± 4.656.5 ± 1.2- FoodPacking1CansTask 60.00.00028.13 ± 14.07−6.49 ± 1.502.89 ± 1.047.3 ± 1.4- FoodPacking2BoxesTask0.00.000-−13.39 ± 2.319.95 ± 1.955.3 ± 1.0- FoodPacking2CansTask30.00.000139.09 ± 38.78−14.15 ± 4.1112.40 ± 8.324.9 ± 0.7- FoodPacking3BoxesTask0.00.000-−14.86 ± 2.9810.43 ± 3.224.4 ± 1.1- FoodPacking3CansTask10.00.000117.60−14.06 ± 3.2522.17 ± 27.305.6 ± 1.6- FoodPackingByColorTask 0.00.000-−10.39 ± 2.387.24 ± 1.765.8 ± 1.3- FruitsGreenLimesOnPlateTask 30.00.00073.69 ± 13.51−10.90 ± 2.365.02 ± 3.924.4 ± 0.7- FruitsMovingOrangeOrLimeTask 70.00.00018.92 ± 10.44−5.40 ± 2.391.55 ± 0.676.3 ± 1.8- FruitsMovingTask10.00.00049.67−8.89 ± 3.053.57 ± 1.235.8 ± 1.9- FruitsOnPlate3Task30.00.00081.71 ± 88.70−18.49 ± 11.7418.86 ± 24.734.3 ± 1.6- FruitsOnPlateTask0.00.000-−18.01 ± 4.1012.82 ± 2.904.2 ± 1.0- FruitsOnionTask60.00.00016.87 ± 6.94−11.00 ± 17.167.03 ± 16.656.4 ± 1.6- FruitsOnionToPlateTask 60.00.00029.06 ± 12.55−11.54 ± 13.225.88 ± 11.345.6 ± 1.3- FruitsOrangesOnPlateTask 80.00.00010.58 ± 8.37−4.36 ± 2.131.69 ± 1.807.6 ± 1.5- GrabABagelTask 0.00.000-−6.11 ± 0.412.07 ± 0.326.5 ± 1.0- GrabAFruitTask0.00.000-−7.02 ± 0.911.72 ± 0.115.4 ± 0.3- GreenSpoonsInPotTask 0.00.000-−22.14 ± 4.244.79 ± 1.972.8 ± 1.0- HammersInLeftBinTask0.00.000-−11.68 ± 2.348.06 ± 2.034.7 ± 1.0- JugsOnShelfTask 0.00.000-−10.23 ± 2.429.49 ± 2.327.5 ± 1.8- KeyboardOutOfBinTask0.00.000-−8.38 ± 1.452.40 ± 0.833.8 ± 1.3- LargerObjectRaisinBoxInBinTask0.00.000-−8.26 ± 3.201.15 ± 0.493.9 ± 1.4- MarkerInMugTask 0.00.000-−10.64 ± 1.991.13 ± 0.272.7 ± 0.7- MouseOnKeyboardTask50.00.00032.03 ± 19.35−14.21 ± 13.275.22 ± 11.174.4 ± 1.1- MoveBananaToBagelPlateTask0.00.000-−10.75 ± 1.445.83 ± 0.666.0 ± 0.7- MustardAboveRaisinTask 100.0-9.42 ± 1.24−3.38 ± 0.480.57 ± 0.045.7 ± 0.4- MustardInLeftBinTask50.00.00015.65 ± 8.47−4.58 ± 0.811.46 ± 0.556.6 ± 1.6- MustardInRightBinTask100.0-7.80 ± 1.58−4.09 ± 1.070.70 ± 0.088.6 ± 1.3- NonHammerToolsInRightBinTask 0.00.000-−16.02 ± 4.187.27 ± 1.314.1 ± 0.6- OneBottleInSquarePailTask 100.0-11.49 ± 3.62−4.43 ± 1.170.91 ± 0.207.6 ± 0.8- OneBottleOnShelfTask0.00.000-−6.79 ± 1.225.63 ± 0.698.9 ± 1.1- PhoneOrRemoteInBinTask 20.00.0004.10 ± 0.24−13.12 ± 14.333.09 ± 4.884.8 ± 2.1- PickDrillTask0.00.000-−8.68 ± 1.242.45 ± 0.215.8 ± 0.5- PickGlassesTask0.00.000-−5.74 ± 0.812.70 ± 0.358.5 ± 1.2- PickOrangeObjectTask 10.00.00014.40−9.28 ± 2.615.13 ± 4.476.4 ± 0.6- PickUpBluePitcherTask0.00.000-−8.39 ± 1.531.54 ± 0.174.9 ± 0.5- PickUpGreenObjectTask 0.00.000-−8.54 ± 1.481.58 ± 0.234.9 ± 0.7- PinkSpoonInPotTask50.00.00030.36 ± 11.56−9.15 ± 5.005.12 ± 9.535.4 ± 1.6- PlasticBottlesInSquarePailTask30.00.00052.44 ± 13.22−14.20 ± 12.5226.68 ± 56.466.2 ± 2.0- PutBowlOnShelfTopTask0.00.000-−9.14 ± 1.323.64 ± 0.765.7 ± 1.1- PutMugsOnShelfTask 0.00.000-−11.73 ± 2.7111.27 ± 2.666.1 ± 1.4- PutTwoMugsOnShelfTask0.00.000-−13.98 ± 4.1410.44 ± 3.305.8 ± 1.7- RecycleCartonTask 30.00.00046.44 ± 15.70−10.82 ± 8.607.58 ± 9.876.0 ± 1.2- RecycleCartonsOnBoxTask 0.00.000-−8.75 ± 0.706.45 ± 0.866.8 ± 0.9- RecycleCartonsVerticalCrateTask0.00.000-−9.28 ± 1.495.96 ± 0.806.3 ± 0.8- RedDishesInBinTask30.00.00040.64 ± 3.97−6.60 ± 1.093.47 ± 0.496.2 ± 0.9- RedItemsInBinTask30.00.00046.09 ± 6.57−6.73 ± 1.503.99 ± 0.836.9 ± 1.3- ReorientAllMugsTask0.00.100-−14.21 ± 2.334.81 ± 0.625.1 ± 0.6- ReorientJugTask0.00.000-−12.81 ± 1.593.10 ± 0.544.7 ± 0.8- ReorientRedMugTask30.00.00024.42 ± 26.24−11.97 ± 6.314.43 ± 6.215.9 ± 1.8- ReorientWhiteMugsTask60.00.75010.38 ± 4.92−14.03 ± 21.016.86 ± 17.105.5 ± 0.8- RubiksCubeAndBananaTask100.0-26.65 ± 13.07−5.76 ± 1.771.75 ± 0.616.7 ± 1.0- RubiksCubeBehindBowlTask100.0-10.59 ± 4.21−4.16 ± 1.200.84 ± 0.257.7 ± 1.1- RubiksCubeInFrontOfBowlTask0.00.000-−7.28 ± 0.422.00 ± 0.196.4 ± 0.6- RubiksCubeLeftOfBowlTask10.00.0006.13−6.42 ± 1.342.22 ± 0.777.7 ± 1.5- RubiksCubeOrBananaTask 100.0-12.61 ± 5.23−4.75 ± 1.070.88 ± 0.276.9 ± 1.1- RubiksCubeRightOfBowlTask0.00.000-−7.11 ± 0.632.05 ± 0.376.4 ± 1.1- RubiksCubeTask100.0-10.15 ± 3.47−4.15 ± 0.970.70 ± 0.146.8 ± 1.3- RubiksCubeThenBananaTask20.00.00042.67 ± 8.77−7.03 ± 1.283.65 ± 0.816.3 ± 1.0- RubiksCubesInBinTask 100.0-73.98 ± 24.10−7.61 ± 1.374.77 ± 1.246.3 ± 0.9- SauceBottlesCrateTask20.00.00016.43 ± 11.55−6.28 ± 1.152.11 ± 0.675.9 ± 1.0- SmallPumpkinInBinTask0.00.000-−8.89 ± 3.572.31 ± 0.883.7 ± 1.4- SmallerObjectButterInBinTask 40.00.00015.97 ± 9.90−7.37 ± 2.510.90 ± 0.154.7 ± 2.8- SmartphoneInBinTask10.00.00028.67−7.98 ± 1.674.20 ± 4.654.9 ± 1.2- SpoonInMugTask50.00.00026.23 ± 6.50−8.69 ± 3.402.20 ± 0.855.3 ± 1.7- SpoonsInPotTask 0.00.000-−14.50 ± 2.787.85 ± 2.174.2 ± 1.1- Stack3RubiksCubeTask20.00.00034.73 ± 4.71−9.22 ± 1.433.96 ± 2.815.5 ± 0.8- StackWhiteMugsTask60.00.00036.32 ± 17.10−7.59 ± 4.616.54 ± 9.797.8 ± 1.1- StackYellowOnRedTask 40.00.00029.25 ± 22.79−9.31 ± 8.634.71 ± 6.676.0 ± 1.8- TakeMeasuringSpoonOutTask10.00.00029.60−10.29 ± 2.092.19 ± 2.193.8 ± 1.1- TakeMugsOffOfShelfTask0.00.000-−14.95 ± 1.7210.84 ± 0.795.7 ± 0.4- TakeSpatulaOffShelfTask 0.00.000-−10.16 ± 1.503.14 ± 0.665.0 ± 1.0- ThrowAwayAppleTask0.00.000-−7.66 ± 1.314.14 ± 0.276.4 ± 0.4- ThrowAwaySnacksTask0.00.000-−9.17 ± 2.357.03 ± 2.335.5 ± 1.8- ToolOrganizationBothTask0.00.000-−13.35 ± 2.228.80 ± 1.755.1 ± 0.7- ToolOrganizationTask0.00.000-−13.44 ± 2.408.66 ± 1.064.8 ± 0.5- ToolsPickingAllHammersTask0.00.000-−18.32 ± 1.279.80 ± 1.504.0 ± 0.6- ToolsPickingDrillTask0.00.000-−9.24 ± 1.123.21 ± 0.745.1 ± 1.1- ToolsPickingHammerTask0.00.000-−11.27 ± 1.213.32 ± 0.305.3 ± 0.4- ToyInBinTask60.00.00017.87 ± 15.55−12.25 ± 17.984.95 ± 10.625.8 ± 2.1- UnstackRubiksCubeTask 10.00.00062.67−10.78 ± 0.668.88 ± 7.937.0 ± 0.4- UtensilsInMugTask 0.00.000-−20.54 ± 3.161.54 ± 0.241.8 ± 0.3- WhiteMugInCenterOfTableTask40.00.00020.88 ± 6.82−6.75 ± 0.471.66 ± 0.466.0 ± 0.8- WhiteMugsInBinTask0.00.000-−9.06 ± 1.734.94 ± 0.547.7 ± 0.8- WoodSpatulaToBowlTask10.00.00031.20−7.92 ± 2.394.25 ± 2.935.8 ± 1.0- YellowAndWhiteObjectsInBinTask0.00.000-−7.16 ± 1.224.38 ± 1.037.1 ± 1.6- YogurtInBowlTask10.00.00015.33−7.72 ± 1.753.09 ± 2.466.0 ± 1.4- TABLE IX: Detailed results forπ 0 -FAST. Task NameSucc%ScoreTime(s)SPARCPathLen(m)Speed(cm/s)WrongObjNames TOTAL (120 tasks)15.40.00221.05 ± 15.90−9.51 ± 6.884.18 ± 7.164.6 ± 1.8 AnimalsInBinTask0.00.000-−12.28 ± 3.073.68 ± 0.773.9 ± 0.8- AppleAndYogurtInBowlTask10.00.00026.20−12.43 ± 5.457.29 ± 8.804.0 ± 1.3- BBQSauceInBinTask0.00.000-−8.34 ± 2.453.37 ± 0.723.6 ± 0.8- BagelsOnPlateTask0.00.000-−9.85 ± 2.881.46 ± 0.442.9 ± 0.6- BananaInBowlTask60.00.00039.19 ± 8.84−8.13 ± 1.612.77 ± 4.283.4 ± 0.9- BananaOnPlateTask80.00.00016.14 ± 8.03−5.27 ± 1.971.00 ± 0.325.2 ± 1.5- BananaThenRubiksCubeTask50.00.00030.84 ± 12.06−8.05 ± 8.366.76 ± 12.776.0 ± 1.0- BananasInBinOneMoreTask80.00.00028.66 ± 13.89−9.78 ± 12.178.27 ± 19.676.6 ± 1.6- BananasInBinThreeTotalTask100.0-23.27 ± 17.44−5.04 ± 1.571.60 ± 1.007.2 ± 1.4- BananasInCrateTask100.0-12.50 ± 3.73−3.64 ± 0.691.13 ± 0.248.7 ± 1.0- BananasOutOfBinTask10.00.11137.93−11.91 ± 3.324.42 ± 3.633.8 ± 1.0- BigPumpkinInBinTask0.00.000-−7.84 ± 1.402.75 ± 0.784.4 ± 1.3- BlackItemsInBinTask0.00.000-−11.04 ± 2.653.76 ± 0.823.3 ± 0.5- BlockStackingOrderAgnosticTask10.00.00079.67−8.64 ± 2.323.77 ± 1.024.0 ± 1.0- BlockStackingSpecifiedOrderTask0.00.000-−8.54 ± 1.464.73 ± 0.915.0 ± 1.0- BlocksInBinTask0.00.000-−11.30 ± 1.578.17 ± 1.345.2 ± 0.8- BowlInBinTask90.00.00026.01 ± 12.94−5.11 ± 1.031.42 ± 0.704.8 ± 0.8- BowlStackingLeftOnRightTask80.00.0008.51 ± 4.80−3.20 ± 1.140.78 ± 0.247.9 ± 2.2- BowlStackingRightOnLeftTask100.0-9.49 ± 0.85−2.49 ± 0.160.71 ± 0.047.1 ± 0.4- ButterAboveRaisinTask 0.00.000-−7.05 ± 1.361.57 ± 0.334.0 ± 0.6- CannedFoodInBinTask 20.00.00040.83 ± 10.23−6.26 ± 1.932.73 ± 0.734.8 ± 1.4- ClampInRightBinTask0.00.000-−7.35 ± 2.062.12 ± 0.533.5 ± 0.8- CleanUpToysTask 0.00.000-−14.09 ± 1.1817.95 ± 2.275.7 ± 0.7- ClearOrganicObjectsTask 0.00.000-−17.33 ± 4.918.19 ± 1.943.3 ± 0.8- ClutterPlasticTask0.00.000-−21.50 ± 15.595.05 ± 1.512.7 ± 0.8- ClutterPumpkinTask 0.00.000-−10.11 ± 2.483.22 ± 0.693.4 ± 0.7- CoffeePotInBinTask 0.00.000-−8.23 ± 1.222.41 ± 0.403.8 ± 0.7- CondimentsInBinTask0.00.000-−12.66 ± 1.076.45 ± 1.073.5 ± 0.6- CookingClearPlateTask0.00.000-−12.47 ± 4.098.78 ± 1.914.6 ± 1.0- CookingPickPastaToolTask0.00.000-−8.87 ± 1.803.47 ± 0.625.3 ± 0.9- CubesAndBlocksInBinTask0.00.000-−12.50 ± 1.8512.15 ± 3.404.9 ± 1.4- DishesInBinTask0.00.000-−12.58 ± 2.5410.25 ± 1.705.4 ± 0.9- ElectronicsInBinTask0.00.000-−13.79 ± 1.876.35 ± 1.503.4 ± 0.8- FoodPacking1BoxesTask 0.00.000-−8.44 ± 1.822.34 ± 0.423.7 ± 0.6- FoodPacking1CansTask 10.00.00012.40−8.72 ± 3.392.78 ± 2.193.9 ± 1.2- FoodPacking2BoxesTask0.00.000-−13.34 ± 2.214.96 ± 1.322.7 ± 0.7- FoodPacking2CansTask0.00.000-−14.42 ± 3.394.93 ± 1.282.6 ± 0.7- FoodPacking3BoxesTask0.00.000-−14.20 ± 7.487.72 ± 4.193.2 ± 1.7- FoodPacking3CansTask0.00.000-−17.09 ± 2.657.40 ± 2.033.0 ± 0.8- FoodPackingByColorTask0.00.000-−10.25 ± 2.425.13 ± 1.254.1 ± 1.0- FruitsGreenLimesOnPlateTask0.00.000-−9.44 ± 1.913.14 ± 0.393.4 ± 0.4- FruitsMovingOrangeOrLimeTask0.00.000-−10.06 ± 1.992.06 ± 0.153.3 ± 0.2- FruitsMovingTask0.00.000-−7.87 ± 1.252.27 ± 0.523.6 ± 0.8- FruitsOnPlate3Task0.00.000-−15.06 ± 1.876.34 ± 0.923.0 ± 0.5- FruitsOnPlateTask0.00.000-−18.83 ± 4.888.34 ± 1.762.6 ± 0.6- FruitsOnionTask40.00.00023.93 ± 6.30−5.18 ± 0.841.72 ± 0.474.1 ± 1.5- FruitsOnionToPlateTask20.00.00013.23 ± 0.42−6.81 ± 2.341.72 ± 0.533.8 ± 1.6- FruitsOrangesOnPlateTask 20.00.00024.13 ± 3.77−13.53 ± 6.627.66 ± 15.223.7 ± 1.1- GrabABagelTask0.00.000-−6.38 ± 1.411.06 ± 0.303.4 ± 0.9- GrabAFruitTask 0.00.000-−6.29 ± 1.401.42 ± 0.154.4 ± 0.5- GreenSpoonsInPotTask0.00.000-−21.52 ± 7.965.11 ± 2.522.9 ± 1.4- HammersInLeftBinTask0.00.000-−12.74 ± 1.826.26 ± 0.743.4 ± 0.4- JugsOnShelfTask 0.00.000-−11.30 ± 3.216.60 ± 1.825.3 ± 1.5- KeyboardOutOfBinTask0.00.000-−8.43 ± 1.723.15 ± 0.324.9 ± 0.5- LargerObjectRaisinBoxInBinTask0.00.000-−6.09 ± 0.780.89 ± 0.133.6 ± 0.4- MarkerInMugTask 0.00.000-−8.48 ± 2.421.38 ± 0.453.4 ± 1.1- MouseOnKeyboardTask0.00.000-−8.61 ± 2.133.55 ± 0.615.5 ± 0.9- MoveBananaToBagelPlateTask0.00.000-−14.34 ± 3.443.04 ± 0.613.1 ± 0.6- MustardAboveRaisinTask 90.00.00011.07 ± 4.30−13.02 ± 30.558.49 ± 24.665.8 ± 0.8- MustardInLeftBinTask70.00.00010.21 ± 3.33−4.13 ± 1.380.97 ± 0.396.3 ± 1.6- MustardInRightBinTask 90.00.00012.87 ± 5.87−3.44 ± 0.480.95 ± 0.366.5 ± 1.1- NonHammerToolsInRightBinTask0.00.000-−14.97 ± 6.925.34 ± 1.443.0 ± 0.8- OneBottleInSquarePailTask 90.00.00013.84 ± 10.54−10.72 ± 22.949.02 ± 25.317.5 ± 2.2- OneBottleOnShelfTask0.00.000-−7.24 ± 1.542.95 ± 0.874.7 ± 1.4- PhoneOrRemoteInBinTask30.00.00011.24 ± 7.98−7.18 ± 3.681.50 ± 0.804.2 ± 2.1- PickDrillTask0.00.000-−7.01 ± 1.431.98 ± 0.374.7 ± 0.9- PickGlassesTask 0.00.000-−6.17 ± 1.331.81 ± 0.325.8 ± 0.9- PickOrangeObjectTask0.00.000-−9.44 ± 1.792.82 ± 0.704.4 ± 1.0- PickUpBluePitcherTask0.00.000-−6.17 ± 1.221.75 ± 0.405.3 ± 1.2- PickUpGreenObjectTask0.00.000-−5.76 ± 1.181.44 ± 0.274.5 ± 0.7- PinkSpoonInPotTask0.00.000-−7.09 ± 1.142.89 ± 0.554.6 ± 0.8- PlasticBottlesInSquarePailTask50.00.00071.17 ± 22.29−18.54 ± 26.7622.00 ± 51.285.0 ± 1.9- PutBowlOnShelfTopTask0.00.000-−7.38 ± 2.361.93 ± 0.773.5 ± 1.1- PutMugsOnShelfTask 0.00.000-−12.80 ± 3.387.91 ± 2.254.3 ± 1.2- PutTwoMugsOnShelfTask 0.00.000-−10.75 ± 4.968.68 ± 4.034.8 ± 2.1- RecycleCartonTask0.00.000-−11.09 ± 2.682.30 ± 0.462.4 ± 0.5- RecycleCartonsOnBoxTask0.00.000-−11.61 ± 2.193.18 ± 0.513.3 ± 0.5- RecycleCartonsVerticalCrateTask0.00.000-−14.98 ± 4.081.76 ± 0.581.9 ± 0.6- RedDishesInBinTask50.00.00033.85 ± 5.58−8.09 ± 7.185.29 ± 8.575.9 ± 1.3- RedItemsInBinTask10.00.00042.67−5.46 ± 0.933.07 ± 0.325.0 ± 0.7- ReorientAllMugsTask0.00.067-−8.76 ± 1.784.91 ± 0.845.0 ± 0.9- ReorientJugTask10.00.00036.33−8.09 ± 1.512.94 ± 0.454.7 ± 0.7- ReorientRedMugTask0.00.000-−8.64 ± 2.552.70 ± 0.404.2 ± 0.6- ReorientWhiteMugsTask40.00.00038.60 ± 19.35−7.57 ± 2.604.62 ± 2.527.1 ± 0.7- RubiksCubeAndBananaTask40.00.00024.48 ± 8.53−5.91 ± 1.542.69 ± 1.125.8 ± 1.1- RubiksCubeBehindBowlTask40.00.0007.62 ± 1.15−4.26 ± 1.201.81 ± 1.017.9 ± 0.8- RubiksCubeInFrontOfBowlTask 0.00.000-−4.75 ± 0.372.37 ± 0.177.5 ± 0.5- RubiksCubeLeftOfBowlTask0.00.000-−4.61 ± 0.532.71 ± 0.778.4 ± 2.3- RubiksCubeOrBananaTask60.00.00023.62 ± 4.25−4.45 ± 0.682.04 ± 0.417.6 ± 1.4- RubiksCubeRightOfBowlTask 0.00.000-−5.27 ± 0.771.77 ± 0.185.5 ± 0.6- RubiksCubeTask100.0-10.34 ± 2.07−3.30 ± 0.450.65 ± 0.066.0 ± 0.6- RubiksCubeThenBananaTask60.00.00025.38 ± 9.72−5.42 ± 1.632.41 ± 0.966.1 ± 0.8- RubiksCubesInBinTask 0.00.000-−10.48 ± 1.945.93 ± 1.164.8 ± 0.9- SauceBottlesCrateTask70.00.00011.59 ± 8.97−13.36 ± 27.346.83 ± 18.836.3 ± 2.4- SmallPumpkinInBinTask0.00.000-−7.43 ± 1.092.43 ± 0.893.9 ± 1.5- SmallerObjectButterInBinTask 0.00.000-−5.95 ± 1.260.82 ± 0.163.4 ± 0.6- SmartphoneInBinTask10.00.0008.47−7.87 ± 3.603.57 ± 2.475.0 ± 1.1- SpoonInMugTask0.00.000-−8.42 ± 1.283.33 ± 0.505.2 ± 0.8- SpoonsInPotTask 0.00.000-−13.49 ± 5.068.17 ± 2.204.4 ± 1.1- Stack3RubiksCubeTask0.00.000-−7.37 ± 1.133.15 ± 0.654.9 ± 1.0- StackWhiteMugsTask 0.00.000-−7.85 ± 2.212.86 ± 0.594.5 ± 0.9- StackYellowOnRedTask0.00.000-−8.39 ± 2.872.87 ± 0.734.5 ± 1.1- TakeMeasuringSpoonOutTask10.00.00025.60−7.00 ± 1.013.23 ± 4.374.9 ± 1.3- TakeMugsOffOfShelfTask 0.00.000-−19.66 ± 2.965.14 ± 1.222.6 ± 0.6- TakeSpatulaOffShelfTask0.00.000-−11.08 ± 3.001.21 ± 0.721.9 ± 1.1- ThrowAwayAppleTask 0.00.000-−9.01 ± 1.312.15 ± 0.403.4 ± 0.7- ThrowAwaySnacksTask0.00.000-−13.99 ± 3.263.46 ± 1.012.7 ± 0.8- ToolOrganizationBothTask 0.00.000-−12.72 ± 2.907.44 ± 1.894.0 ± 0.9- ToolOrganizationTask 0.00.000-−13.79 ± 4.304.72 ± 1.672.7 ± 0.8- ToolsPickingAllHammersTask0.00.000-−15.80 ± 5.6510.16 ± 2.454.1 ± 0.9- ToolsPickingDrillTask 0.00.000-−7.08 ± 2.952.82 ± 0.804.8 ± 1.5- ToolsPickingHammerTask 0.00.000-−6.53 ± 1.613.39 ± 0.315.3 ± 0.5- ToyInBinTask30.00.00031.00 ± 5.13−8.74 ± 4.815.20 ± 8.065.3 ± 1.6- UnstackRubiksCubeTask 0.00.000-−12.11 ± 1.437.28 ± 1.477.5 ± 1.5- UtensilsInMugTask0.00.000-−10.29 ± 2.273.42 ± 0.643.7 ± 0.6- WhiteMugInCenterOfTableTask20.00.00022.53 ± 1.32−5.45 ± 0.941.37 ± 0.194.5 ± 0.5- WhiteMugsInBinTask0.00.000-−8.56 ± 1.274.56 ± 0.337.0 ± 0.5- WoodSpatulaToBowlTask0.00.000-−8.31 ± 1.922.50 ± 0.703.9 ± 1.1- YellowAndWhiteObjectsInBinTask0.00.000-−6.96 ± 1.183.15 ± 0.845.1 ± 1.3- YogurtInBowlTask0.00.000-−6.12 ± 0.592.10 ± 0.344.9 ± 0.8- TABLE X: Detailed results forπ 0 . Task NameSucc%ScoreTime(s)SPARCPathLen(m)Speed(cm/s)WrongObjNames TOTAL (120 tasks)5.30.00123.09 ± 16.87−9.48 ± 4.074.24 ± 3.924.4 ± 1.4 AnimalsInBinTask0.00.000-−7.29 ± 1.115.20 ± 0.765.5 ± 0.8- AppleAndYogurtInBowlTask0.00.000-−8.17 ± 2.476.74 ± 1.715.3 ± 1.4- BBQSauceInBinTask0.00.000-−7.25 ± 1.054.18 ± 0.734.5 ± 0.8- BagelsOnPlateTask0.00.000-−6.30 ± 1.002.96 ± 0.654.5 ± 1.0- BananaInBowlTask0.00.000-−6.81 ± 1.002.30 ± 0.264.5 ± 0.5- BananaOnPlateTask70.00.00025.59 ± 9.52−12.71 ± 17.946.12 ± 15.214.7 ± 1.4- BananaThenRubiksCubeTask10.00.00047.47−9.06 ± 2.733.10 ± 0.635.1 ± 1.1- BananasInBinOneMoreTask0.00.000-−9.46 ± 1.412.58 ± 0.424.1 ± 0.7- BananasInBinThreeTotalTask20.00.00047.73 ± 4.15−9.94 ± 1.892.36 ± 0.184.0 ± 0.7- BananasInCrateTask50.00.00033.97 ± 21.15−9.00 ± 8.633.78 ± 4.625.5 ± 1.7- BananasOutOfBinTask0.00.083-−11.48 ± 1.994.95 ± 1.395.3 ± 1.5- BigPumpkinInBinTask0.00.000-−7.56 ± 0.982.27 ± 0.403.6 ± 0.6- BlackItemsInBinTask0.00.000-−7.96 ± 1.945.96 ± 1.734.7 ± 1.3- BlockStackingOrderAgnosticTask0.00.000-−7.63 ± 0.963.01 ± 0.653.2 ± 0.7- BlockStackingSpecifiedOrderTask0.00.000-−7.90 ± 0.842.74 ± 0.272.9 ± 0.3- BlocksInBinTask0.00.000-−8.98 ± 1.567.51 ± 0.834.8 ± 0.5- BowlInBinTask0.00.000-−5.68 ± 0.603.96 ± 0.146.2 ± 0.2- BowlStackingLeftOnRightTask0.00.000-−7.13 ± 1.171.11 ± 0.275.2 ± 1.2- BowlStackingRightOnLeftTask0.00.000-−5.43 ± 1.221.39 ± 0.176.6 ± 0.8- ButterAboveRaisinTask0.00.000-−6.36 ± 1.062.58 ± 0.316.1 ± 0.7- CannedFoodInBinTask0.00.000-−5.94 ± 0.762.64 ± 0.404.3 ± 0.6- ClampInRightBinTask0.00.000-−6.98 ± 1.712.15 ± 0.713.5 ± 1.1- CleanUpToysTask0.00.000-−18.35 ± 1.8615.55 ± 2.014.8 ± 0.6- ClearOrganicObjectsTask0.00.000-−14.38 ± 2.248.88 ± 0.863.5 ± 0.3- ClutterPlasticTask0.00.000-−11.23 ± 2.086.76 ± 0.603.7 ± 0.3- ClutterPumpkinTask0.00.000-−8.04 ± 0.993.69 ± 0.664.0 ± 0.7- CoffeePotInBinTask0.00.000-−5.61 ± 0.952.84 ± 0.714.6 ± 1.2- CondimentsInBinTask0.00.000-−9.71 ± 2.219.36 ± 1.735.0 ± 0.9- CookingClearPlateTask0.00.000-−15.81 ± 3.497.91 ± 0.744.2 ± 0.3- CookingPickPastaToolTask0.00.000-−10.52 ± 1.013.07 ± 0.524.7 ± 0.8- CubesAndBlocksInBinTask0.00.000-−10.66 ± 3.8711.11 ± 1.724.5 ± 0.7- DishesInBinTask0.00.000-−12.25 ± 1.918.24 ± 1.784.4 ± 1.0- ElectronicsInBinTask0.00.000-−10.63 ± 2.837.26 ± 1.883.8 ± 1.0- FoodPacking1BoxesTask0.00.000-−7.07 ± 1.152.96 ± 0.434.6 ± 0.7- FoodPacking1CansTask0.00.000-−7.83 ± 0.891.89 ± 0.253.0 ± 0.4- FoodPacking2BoxesTask0.00.000-−11.44 ± 1.507.17 ± 1.003.8 ± 0.5- FoodPacking2CansTask0.00.000-−12.78 ± 2.425.59 ± 0.703.0 ± 0.4- FoodPacking3BoxesTask0.00.000-−10.85 ± 2.5610.02 ± 1.514.0 ± 0.6- FoodPacking3CansTask0.00.000-−11.95 ± 2.608.75 ± 1.083.5 ± 0.4- FoodPackingByColorTask0.00.000-−9.99 ± 1.814.31 ± 0.493.4 ± 0.4- FruitsGreenLimesOnPlateTask0.00.000-−8.75 ± 1.152.99 ± 0.233.1 ± 0.2- FruitsMovingOrangeOrLimeTask0.00.000-−7.08 ± 0.712.85 ± 0.244.6 ± 0.4- FruitsMovingTask0.00.000-−7.07 ± 0.463.39 ± 0.305.5 ± 0.5- FruitsOnPlate3Task0.00.000-−11.04 ± 1.747.30 ± 0.593.5 ± 0.3- FruitsOnPlateTask0.00.000-−13.38 ± 3.0410.44 ± 0.603.3 ± 0.2- FruitsOnionTask20.00.00042.47 ± 11.22−7.79 ± 1.592.97 ± 0.365.1 ± 0.6- FruitsOnionToPlateTask0.00.000-−7.75 ± 1.192.65 ± 0.224.2 ± 0.3- FruitsOrangesOnPlateTask0.0--−19.87 ± 0.7925.16 ± 3.693.1 ± 0.5- GrabABagelTask0.00.000-−8.29 ± 1.011.37 ± 0.124.2 ± 0.4- GrabAFruitTask0.00.000-−7.79 ± 0.821.42 ± 0.184.4 ± 0.5- GreenSpoonsInPotTask 0.00.000-−14.29 ± 3.255.73 ± 1.103.0 ± 0.6- HammersInLeftBinTask0.00.000-−10.88 ± 2.485.71 ± 1.033.2 ± 0.5- JugsOnShelfTask0.00.000-−8.75 ± 1.137.21 ± 1.275.8 ± 1.0- KeyboardOutOfBinTask 0.00.000-−12.61 ± 1.471.91 ± 0.372.9 ± 0.6- LargerObjectRaisinBoxInBinTask0.00.000-−5.97 ± 1.011.11 ± 0.263.5 ± 0.8- MarkerInMugTask 0.00.000-−9.16 ± 2.361.30 ± 0.293.1 ± 0.7- MouseOnKeyboardTask0.00.000-−9.17 ± 3.482.42 ± 0.913.8 ± 1.4- MoveBananaToBagelPlateTask0.00.000-−15.15 ± 3.533.72 ± 1.063.8 ± 1.1- MustardAboveRaisinTask 0.00.000-−9.34 ± 0.852.13 ± 0.154.9 ± 0.3- MustardInLeftBinTask40.00.00011.53 ± 3.82−4.98 ± 0.871.48 ± 0.596.4 ± 0.9- MustardInRightBinTask0.00.000-−5.23 ± 0.782.04 ± 0.386.5 ± 1.2- NonHammerToolsInRightBinTask0.00.000-−12.45 ± 4.255.01 ± 2.002.8 ± 1.0- OneBottleInSquarePailTask40.00.00012.22 ± 6.19−5.89 ± 2.902.16 ± 1.545.5 ± 2.3- OneBottleOnShelfTask0.00.000-−7.60 ± 1.573.30 ± 0.945.2 ± 1.5- PhoneOrRemoteInBinTask 20.00.00042.17 ± 3.72−6.79 ± 1.141.65 ± 0.372.7 ± 0.5- PickDrillTask0.00.000-−8.79 ± 0.891.84 ± 0.164.3 ± 0.4- PickGlassesTask0.00.000-−10.01 ± 1.601.16 ± 0.183.6 ± 0.6- PickOrangeObjectTask0.00.000-−12.00 ± 1.852.15 ± 0.373.3 ± 0.6- PickUpBluePitcherTask 0.00.000-−8.66 ± 1.360.80 ± 0.102.5 ± 0.3- PickUpGreenObjectTask0.00.000-−8.03 ± 2.340.67 ± 0.132.1 ± 0.4- PinkSpoonInPotTask0.00.000-−9.33 ± 1.983.00 ± 0.544.6 ± 0.8- PlasticBottlesInSquarePailTask0.00.000-−12.01 ± 2.367.95 ± 2.644.2 ± 1.4- PutBowlOnShelfTopTask0.00.000-−10.34 ± 2.443.32 ± 0.845.3 ± 1.3- PutMugsOnShelfTask 0.00.000-−10.53 ± 2.0210.45 ± 2.605.5 ± 1.4- PutTwoMugsOnShelfTask 0.00.000-−11.34 ± 1.636.99 ± 1.384.0 ± 0.6- RecycleCartonTask0.00.000-−8.11 ± 1.893.37 ± 0.533.6 ± 0.6- RecycleCartonsOnBoxTask0.00.000-−16.26 ± 3.092.43 ± 0.672.5 ± 0.7- RecycleCartonsVerticalCrateTask0.00.000-−15.44 ± 1.952.18 ± 0.492.3 ± 0.5- RedDishesInBinTask 0.00.000-−6.11 ± 0.653.56 ± 0.185.6 ± 0.3- RedItemsInBinTask0.00.000-−6.25 ± 1.423.07 ± 0.524.9 ± 0.8- ReorientAllMugsTask0.00.033-−11.55 ± 2.254.04 ± 0.634.3 ± 0.7- ReorientJugTask0.00.000-−11.08 ± 1.472.66 ± 0.384.1 ± 0.6- ReorientRedMugTask10.00.00058.33−11.02 ± 1.632.53 ± 0.344.0 ± 0.5- ReorientWhiteMugsTask0.00.000-−11.65 ± 1.172.84 ± 0.304.5 ± 0.5- RubiksCubeAndBananaTask20.00.00055.90 ± 5.33−8.75 ± 2.653.32 ± 0.585.4 ± 0.9- RubiksCubeBehindBowlTask 40.00.00012.63 ± 4.35−6.06 ± 1.521.64 ± 0.607.1 ± 1.1- RubiksCubeInFrontOfBowlTask10.00.00018.33−7.74 ± 1.561.95 ± 0.396.4 ± 1.0- RubiksCubeLeftOfBowlTask10.00.0007.27−8.30 ± 1.901.92 ± 0.536.6 ± 0.8- RubiksCubeOrBananaTask 50.00.0009.61 ± 3.96−5.79 ± 1.671.02 ± 0.515.4 ± 1.4- RubiksCubeRightOfBowlTask0.00.000-−7.74 ± 1.770.93 ± 0.153.0 ± 0.5- RubiksCubeTask90.00.00017.40 ± 5.78−5.44 ± 1.221.02 ± 0.325.3 ± 1.2- RubiksCubeThenBananaTask40.00.00031.00 ± 18.08−7.30 ± 2.192.66 ± 0.865.5 ± 1.0- RubiksCubesInBinTask 0.00.000-−8.38 ± 1.295.16 ± 0.504.1 ± 0.4- SauceBottlesCrateTask60.00.0005.24 ± 1.63−4.32 ± 1.661.03 ± 0.677.7 ± 3.1- SmallPumpkinInBinTask0.00.000-−7.30 ± 1.452.63 ± 0.494.1 ± 0.8- SmallerObjectButterInBinTask 0.00.000-−5.89 ± 1.170.90 ± 0.102.8 ± 0.3- SmartphoneInBinTask20.00.00018.17 ± 7.31−11.81 ± 15.854.05 ± 6.254.0 ± 0.7- SpoonInMugTask0.00.000-−9.65 ± 1.362.30 ± 0.503.6 ± 0.8- SpoonsInPotTask 0.00.000-−12.77 ± 1.987.51 ± 1.024.0 ± 0.5- Stack3RubiksCubeTask0.00.000-−8.61 ± 1.582.80 ± 0.364.4 ± 0.6- StackWhiteMugsTask0.00.000-−8.26 ± 2.133.13 ± 0.385.0 ± 0.5- StackYellowOnRedTask 0.00.000-−8.76 ± 1.362.20 ± 0.343.4 ± 0.5- TakeMeasuringSpoonOutTask0.00.000-−10.29 ± 1.382.04 ± 0.224.7 ± 0.5- TakeMugsOffOfShelfTask0.00.000-−17.62 ± 1.1511.34 ± 0.896.0 ± 0.5- TakeSpatulaOffShelfTask0.00.000-−9.70 ± 2.162.63 ± 0.584.4 ± 0.7- ThrowAwayAppleTask0.00.000-−12.57 ± 1.562.29 ± 0.413.5 ± 0.6- ThrowAwaySnacksTask0.00.000-−9.87 ± 1.955.63 ± 0.914.4 ± 0.7- ToolOrganizationBothTask0.00.000-−10.09 ± 1.936.78 ± 1.633.7 ± 0.9- ToolOrganizationTask0.00.000-−9.70 ± 1.396.31 ± 0.943.4 ± 0.5- ToolsPickingAllHammersTask0.00.000-−10.46 ± 1.2916.62 ± 2.686.6 ± 1.0- ToolsPickingDrillTask0.00.000-−6.49 ± 1.044.07 ± 0.456.4 ± 0.7- ToolsPickingHammerTask0.00.000-−7.62 ± 0.903.26 ± 0.405.2 ± 0.6- ToyInBinTask20.00.00039.47 ± 28.94−9.12 ± 4.182.87 ± 2.183.7 ± 1.0- UnstackRubiksCubeTask0.00.000-−15.98 ± 2.543.59 ± 0.703.8 ± 0.7- UtensilsInMugTask0.00.000-−9.07 ± 1.186.35 ± 1.826.6 ± 1.9- WhiteMugInCenterOfTableTask0.00.000-−8.39 ± 1.071.25 ± 0.203.8 ± 0.6- WhiteMugsInBinTask0.00.000-−12.55 ± 2.242.66 ± 0.394.2 ± 0.6- WoodSpatulaToBowlTask0.00.000-−8.06 ± 1.422.70 ± 0.284.3 ± 0.4- YellowAndWhiteObjectsInBinTask0.00.000-−7.74 ± 1.913.19 ± 0.685.2 ± 1.0- YogurtInBowlTask0.00.000-−9.29 ± 1.561.62 ± 0.313.7 ± 0.7- TABLE XI: Detailed results for PaliGemma. Task NameSucc%ScoreTime(s)SPARCPathLen(m)Speed(cm/s)WrongObjNames TOTAL (79 tasks)1.90.01233.44 ± 17.42−17.60 ± 11.831.18 ± 1.371.4 ± 1.2 BagelsOnPlateTask0.00.000-−25.17 ± 8.080.21 ± 0.070.3 ± 0.1- BananaInBowlTableTask0.00.000-−6.30 ± 1.690.42 ± 0.270.8 ± 0.5- BananaOnPlateTableTask0.00.050-−9.23 ± 2.460.67 ± 0.251.6 ± 0.6- BananasInBinOneMoreTask70.00.00041.35 ± 14.72−10.10 ± 3.431.29 ± 0.273.2 ± 2.2greybin(1) BananasInBinThreeTotalTask0.00.000-−13.10 ± 3.821.15 ± 0.231.9 ± 0.3- BananasInCrateTask40.00.08336.43 ± 7.62−8.38 ± 4.742.12 ± 0.934.2 ± 1.8- BananasOutOfBinSlightlyVagueTask0.00.033-−14.15 ± 4.801.45 ± 0.682.3 ± 1.1greybin(3) BananasOutOfBinTask0.00.033-−21.73 ± 7.620.89 ± 0.581.4 ± 0.9greybin(1) BananasOutOfBinVagueTask0.00.017-−18.31 ± 5.101.87 ± 1.042.0 ± 1.1banana(1), greybin(1) BlockStackingOrderAgnosticTask0.00.000-−12.69 ± 3.861.65 ± 0.261.9 ± 0.3- BlockStackingSpecifiedOrderTask0.00.000-−19.45 ± 12.841.40 ± 0.851.5 ± 0.9- BowlInBinTask0.00.000-−8.67 ± 2.881.35 ± 0.722.1 ± 1.1- BowlStackingLeftOnRightTask0.00.000-−6.76 ± 0.220.04 ± 0.000.2 ± 0.0- BowlStackingRightOnLeftTask0.00.000-−6.68 ± 1.060.05 ± 0.010.2 ± 0.0- ButterAboveRaisinTask0.00.000-−11.91 ± 3.460.96 ± 0.332.4 ± 0.8- ClearOrganicObjectsTask0.00.000-−35.14 ± 10.241.98 ± 0.800.8 ± 0.3- ClutterPlasticTask0.00.000-−19.77 ± 2.381.04 ± 0.130.6 ± 0.1- ClutterPumpkinTask0.00.000-−12.49 ± 2.240.45 ± 0.170.5 ± 0.2- CookingClearPlateSpecificTask0.00.375-−14.17 ± 4.758.05 ± 3.094.2 ± 1.5- CookingClearPlateVagueTask0.00.125-−17.41 ± 7.222.79 ± 1.821.6 ± 1.1- CookingPickPastaToolTask0.00.000-−27.11 ± 6.220.38 ± 0.100.6 ± 0.2- DishesInBinTask0.00.000-−14.83 ± 4.732.63 ± 1.091.5 ± 0.5- FoodPacking1BoxesTask0.00.000-−12.81 ± 5.170.58 ± 0.520.7 ± 0.7- FoodPacking1CansTask0.00.000-−14.32 ± 4.880.88 ± 0.170.9 ± 0.2- FoodPacking2BoxesTask0.00.000-−38.48 ± 15.021.17 ± 0.240.6 ± 0.1- FoodPacking2CansTask0.00.000-−18.34 ± 8.702.39 ± 0.651.3 ± 0.3- FoodPacking3BoxesTask0.00.000-−26.20 ± 13.962.45 ± 0.631.0 ± 0.3- FoodPacking3CansTask0.00.000-−23.29 ± 8.812.96 ± 0.921.2 ± 0.3- FoodPackingByColorTask0.00.000-−34.80 ± 15.020.84 ± 0.450.7 ± 0.4- FruitsGreenLimesOnPlateTask0.00.000-−16.59 ± 3.001.17 ± 0.171.2 ± 0.2- FruitsLimesOnPlateTask0.00.000-−12.82 ± 6.331.39 ± 0.461.6 ± 0.5- FruitsMovingOrangeOrLimeTask0.00.000-−26.58 ± 10.560.93 ± 0.601.0 ± 0.6- FruitsMovingTask0.00.000-−24.39 ± 6.820.47 ± 0.220.7 ± 0.3- FruitsOnPlate3Task0.00.000-−16.62 ± 2.253.82 ± 1.461.9 ± 0.7- FruitsOnPlateTask0.00.000-−23.04 ± 5.174.83 ± 1.831.5 ± 0.6- FruitsOnionTask0.00.000-−30.73 ± 4.410.47 ± 0.130.7 ± 0.2- FruitsOnionToPlateTask0.00.000-−42.76 ± 9.360.67 ± 0.120.7 ± 0.1- FruitsOrangesOnPlateTask0.00.000-−10.46 ± 3.001.35 ± 0.301.5 ± 0.3- LargerObjectRaisinBoxInBinTask0.00.000-−9.01 ± 3.110.80 ± 0.222.9 ± 0.8- MustardAboveRaisinTask0.00.000-−22.18 ± 6.580.35 ± 0.140.8 ± 0.3- PickDrillTask0.00.000-−9.25 ± 0.850.10 ± 0.010.2 ± 0.0- PickOrangeObjectTask0.00.000-−10.79 ± 3.353.15 ± 0.514.9 ± 0.8- PickUpBluePitcherTask0.00.000-−27.40 ± 11.710.23 ± 0.040.7 ± 0.1- PickUpGreenObjectTask0.00.000-−6.05 ± 3.220.68 ± 0.152.1 ± 0.5- PutBowlOnShelfTopTask0.00.000-−11.55 ± 4.802.12 ± 0.812.2 ± 0.9- RecycleCartonTask0.00.000-−25.68 ± 15.830.34 ± 0.130.4 ± 0.2- RecycleCartonsOnBoxTask0.00.000-−22.53 ± 18.841.66 ± 1.311.7 ± 1.4- RecycleCartonsVerticalCrateTask0.00.000-−13.83 ± 6.240.56 ± 0.350.6 ± 0.4- RedDishesInBinTask0.00.050-−6.10 ± 0.961.47 ± 0.482.5 ± 0.8- RedItemsInBinTask0.00.000-−24.04 ± 6.850.17 ± 0.020.3 ± 0.0- ReorientAllMugsTask0.00.000-−12.20 ± 1.611.04 ± 0.191.1 ± 0.2- ReorientRedMugTask0.00.000-−19.78 ± 3.300.21 ± 0.010.3 ± 0.0- ReorientWhiteMugsTask 0.00.000-−17.99 ± 8.430.40 ± 0.150.6 ± 0.2- RubiksCubeAndBananaTask0.00.100-−12.45 ± 3.610.78 ± 0.671.2 ± 1.0- RubiksCubeBehindBowlTask0.00.000-−12.54 ± 7.830.23 ± 0.310.7 ± 0.9rubikscube(1) RubiksCubeInFrontOfBowlTask 0.00.000-−8.72 ± 2.540.09 ± 0.030.3 ± 0.1- RubiksCubeLeftOfBowlTask0.00.000-−9.96 ± 3.200.58 ± 0.761.8 ± 2.4rubikscube(4), banana(3) RubiksCubeOrBananaTask0.00.000-−7.55 ± 3.140.88 ± 0.672.7 ± 2.1- RubiksCubeRightOfBowlTask 0.00.000-−10.72 ± 2.310.84 ± 0.162.5 ± 0.5- RubiksCubeTask10.00.00039.67−7.49 ± 3.280.76 ± 0.461.8 ± 1.1- RubiksCubeThenBananaTask0.00.100-−14.03 ± 8.571.33 ± 0.862.1 ± 1.3- SauceBottlesCrateTask 0.00.000-−8.03 ± 3.301.06 ± 0.392.5 ± 0.9- SmallerObjectButterInBinTask10.00.00026.47−7.79 ± 1.490.70 ± 0.082.6 ± 0.5- Stack3RubiksCubeTask0.00.000-−22.98 ± 6.240.27 ± 0.020.4 ± 0.0- StackWhiteMugsTask 0.00.000-−17.33 ± 7.930.94 ± 0.401.6 ± 0.7- TakeMeasuringSpoonOutTask 0.00.000-−8.89 ± 1.240.14 ± 0.020.3 ± 0.0- ToolOrganizationBothTask0.00.000-−22.23 ± 13.222.60 ± 1.111.4 ± 0.5- ToolOrganizationLeftTask0.00.000-−15.34 ± 5.371.82 ± 0.941.1 ± 0.5- ToolOrganizationSpecificTask0.00.000-−25.98 ± 9.191.91 ± 0.531.2 ± 0.4- ToolsPickingAllHammersTask 0.00.000-−57.19 ± 12.121.49 ± 0.480.6 ± 0.2- ToolsPickingDrillTask20.00.0000.17 ± 0.05−24.12 ± 12.760.30 ± 0.171.0 ± 0.8- ToolsPickingHammerTask0.00.000-−24.27 ± 10.950.44 ± 0.090.7 ± 0.1- UnstackRubiksCubeTask0.00.000-−13.78 ± 0.690.22 ± 0.010.2 ± 0.0- WhiteMugInCenterOfTableTask0.00.000-−16.69 ± 2.210.26 ± 0.010.8 ± 0.0- WhiteMugsInBinMoreVagueTask0.00.000-−12.05 ± 3.451.37 ± 0.872.1 ± 1.3bananafar(1) WhiteMugsInBinTask0.00.000-−38.60 ± 16.950.40 ± 0.150.6 ± 0.2- WhiteMugsInBinVagueTask0.00.000-−7.83 ± 1.070.43 ± 0.540.7 ± 0.9- WoodSpatulaToBowlTask 0.00.000-−27.02 ± 5.460.51 ± 0.150.8 ± 0.2- YellowAndWhiteObjectsInBinTask 0.00.000-−12.78 ± 7.680.75 ± 0.521.2 ± 0.9- TABLE XII: Detailed results for GR00T N1.6. Task NameSucc%ScoreTime(s)SPARCPathLen(m)Speed(cm/s)WrongObjNames TOTAL (79 tasks)8.00.05619.29 ± 14.27−6.44 ± 2.043.77 ± 3.194.2 ± 1.3 BagelsOnPlateTask0.00.000-−5.98 ± 0.682.46 ± 0.233.9 ± 0.4banana(2) BananaInBowlTableTask10.00.00014.13−6.25 ± 1.191.87 ± 0.463.9 ± 0.7- BananaOnPlateTableTask0.00.050-−6.85 ± 1.311.15 ± 0.292.8 ± 0.7- BananasInBinOneMoreTask 50.00.00026.21 ± 15.72−8.99 ± 3.221.03 ± 0.233.2 ± 1.3- BananasInBinThreeTotalTask 100.0-19.84 ± 13.56−3.72 ± 0.540.99 ± 0.315.6 ± 1.4- BananasInCrateTask40.00.00040.43 ± 17.45−5.65 ± 1.672.35 ± 0.604.6 ± 1.4purplecrate(2) BananasOutOfBinSlightlyVagueTask 0.00.000-−6.97 ± 2.112.21 ± 0.533.6 ± 0.7greybin(24), banana03(16) BananasOutOfBinTask 0.00.033-−6.34 ± 1.312.55 ± 1.354.1 ± 2.1greybin(4) BananasOutOfBinVagueTask 0.00.000-−6.29 ± 1.184.35 ± 0.754.6 ± 0.8greybin(21), banana03(2) BlockStackingOrderAgnosticTask0.00.000-−6.31 ± 0.974.32 ± 1.124.5 ± 1.1yellowblock(28), blueblock(3) BlockStackingSpecifiedOrderTask0.00.000-−6.11 ± 1.035.64 ± 0.926.0 ± 0.9blueblock(13) BowlInBinTask10.00.00020.00−4.35 ± 0.852.63 ± 0.914.7 ± 1.4mustard(6), mug(2) BowlStackingLeftOnRightTask0.00.000-−3.87 ± 0.430.94 ± 0.094.5 ± 0.4bowl1(30) BowlStackingRightOnLeftTask 0.00.000-−3.77 ± 0.480.96 ± 0.044.6 ± 0.3bowl1(36) ButterAboveRaisinTask 0.00.000-−7.49 ± 1.121.39 ± 0.263.4 ± 0.7raisinbox(8) ClearOrganicObjectsTask0.00.009-−7.41 ± 1.5412.86 ± 1.605.2 ± 0.6rightbin(1) ClutterPlasticTask0.00.000-−8.75 ± 0.9011.75 ± 1.536.3 ± 0.8pomegranate01(27), orange01(1) ClutterPumpkinTask0.00.000-−8.21 ± 1.613.55 ± 0.693.9 ± 0.7pomegranate01(17), orange01(12), lime01(5), rightbin(1) CookingClearPlateSpecificTask0.00.125-−8.16 ± 1.326.09 ± 1.013.3 ± 0.5storagebox01(9), redonion(8), woodenbowl(6), servingspoon(6) CookingClearPlateVagueTask0.00.300-−6.50 ± 1.187.91 ± 0.694.2 ± 0.3redonion(8), storagebox01(4), woodenbowl(3), potatomasher(1) CookingPickPastaToolTask0.00.000-−5.99 ± 1.772.15 ± 0.423.6 ± 0.7storagebox01(17), servingspoon(15), redonion(2), servingbowl(1) DishesInBinTask0.00.100-−9.59 ± 4.047.50 ± 2.834.2 ± 1.3greybin(1) FoodPacking1BoxesTask0.00.000-−6.82 ± 1.364.26 ± 0.684.5 ± 0.7tomatosoupcan(13), mustard(1) FoodPacking1CansTask0.00.000-−6.51 ± 0.523.73 ± 0.834.0 ± 0.8- FoodPacking2BoxesTask0.00.000-−6.60 ± 1.489.15 ± 0.694.9 ± 0.3tomatosoupcan(19), bina06(1) FoodPacking2CansTask0.00.000-−7.24 ± 1.458.17 ± 1.824.3 ± 1.0- FoodPacking3BoxesTask0.00.000-−8.03 ± 1.0911.95 ± 1.874.8 ± 0.6spamcan(8), tomatosoupcan(3) FoodPacking3CansTask0.00.167-−7.29 ± 0.7610.41 ± 1.084.2 ± 0.4- FoodPackingByColorTask0.00.000-−5.48 ± 0.916.89 ± 0.735.5 ± 0.6mustard(7), tomatosoupcan(3), sugarbox(1) FruitsGreenLimesOnPlateTask0.00.000-−9.89 ± 1.662.44 ± 0.722.6 ± 0.7pomegranate01(2) FruitsLimesOnPlateTask0.00.000-−7.65 ± 2.153.73 ± 1.034.0 ± 1.1pomegranate01(10), lemon01(2), redonion(1) FruitsMovingOrangeOrLimeTask 0.00.015-−5.46 ± 2.192.99 ± 1.534.3 ± 1.1redonion(12), woodenbowl(3) FruitsMovingTask 0.00.000-−5.85 ± 0.853.16 ± 0.445.0 ± 0.7redonion(20), woodenbowl(2), pomegranate01(1) FruitsOnPlate3Task0.00.000-−8.37 ± 1.327.99 ± 1.293.9 ± 0.6pumpkinlarge(7), redonion(2) FruitsOnPlateTask0.00.000-−9.25 ± 2.0411.96 ± 3.333.9 ± 1.1redonion(4), pumpkinlarge(2) FruitsOnionTask20.00.00035.77 ± 22.96−5.16 ± 1.232.90 ± 0.795.1 ± 1.2woodenbowl(21), pumpkinlarge(7), pumpkinsmall(6), storagebox(4), woodenspoons(3) FruitsOnionToPlateTask0.00.050-−6.27 ± 1.704.55 ± 0.884.8 ± 0.9woodenbowl(7), pumpkinsmall(2), pumpkinlarge(1) FruitsOrangesOnPlateTask 0.00.000-−7.04 ± 1.153.88 ± 1.604.2 ± 1.7pomegranate01(13), redonion(2), lemon01(1), pumpkinlarge(1) LargerObjectRaisinBoxInBinTask 0.00.000-−6.57 ± 1.730.70 ± 0.122.4 ± 0.4butter(1) MustardAboveRaisinTask0.00.100-−5.45 ± 1.281.92 ± 0.474.6 ± 1.1- PickDrillTask0.00.000-−6.54 ± 1.451.43 ± 0.203.4 ± 0.5cordlessdrill(54) PickOrangeObjectTask 0.00.000-−4.98 ± 0.802.94 ± 0.594.6 ± 0.9redonion(8), servingspoon(8), woodenbowl(5), storagebox01(4), servingbowl(1) PickUpBluePitcherTask0.00.000-−5.39 ± 1.591.46 ± 0.364.6 ± 1.1pitcher(9) PickUpGreenObjectTask 0.00.000-−4.00 ± 1.101.05 ± 0.293.4 ± 0.9utilityjuga02(6) PutBowlOnShelfTopTask0.00.000-−8.97 ± 2.322.46 ± 0.392.7 ± 0.4rackl04(2) RecycleCartonTask 0.00.000-−5.65 ± 0.874.09 ± 0.914.4 ± 1.0ketchupbottle(24), containera01(24) RecycleCartonsOnBoxTask0.00.000-−6.30 ± 1.354.13 ± 0.614.5 ± 0.7cubeboxa02(7), ketchupbottle(3) RecycleCartonsVerticalCrateTask 0.00.000-−5.72 ± 1.583.68 ± 0.704.0 ± 0.7ketchupbottle(37), containera01(21) RedDishesInBinTask 0.00.450-−5.43 ± 0.993.11 ± 0.464.9 ± 0.7- RedItemsInBinTask10.00.25056.53−5.88 ± 1.352.92 ± 0.444.7 ± 0.7greybin(4) ReorientAllMugsTask0.00.100-−7.80 ± 1.073.01 ± 0.383.3 ± 0.4ceramicmug(4), redmug(2), cordlessdrill(2) ReorientRedMugTask0.00.000-−7.06 ± 1.321.63 ± 0.152.6 ± 0.2sidewayswhitemug(3) ReorientWhiteMugsTask0.00.000-−5.81 ± 1.262.35 ± 0.323.7 ± 0.5uprightwhitemug(8), cordlessdrill(4), redmug(2) RubiksCubeAndBananaTask0.00.125-−6.85 ± 0.811.98 ± 0.253.2 ± 0.4- RubiksCubeBehindBowlTask20.00.00012.00 ± 3.58−3.87 ± 0.671.27 ± 0.284.9 ± 1.5rubikscube(19) RubiksCubeInFrontOfBowlTask10.00.00022.33−5.19 ± 0.991.36 ± 0.214.5 ± 0.9rubikscube(72), banana(14), bowl(10) RubiksCubeLeftOfBowlTask50.00.00013.61 ± 5.11−4.67 ± 0.551.21 ± 0.515.3 ± 0.8rubikscube(16), banana(3) RubiksCubeOrBananaTask10.00.02812.07−5.52 ± 1.260.97 ± 0.173.4 ± 0.8- RubiksCubeRightOfBowlTask40.00.00013.92 ± 2.89−4.96 ± 1.391.54 ± 0.596.2 ± 1.1rubikscube(22), bowl(7), banana(1) RubiksCubeTask 100.0-18.23 ± 4.91−4.69 ± 0.840.95 ± 0.245.0 ± 0.6- RubiksCubeThenBananaTask10.00.13957.13−9.12 ± 1.971.60 ± 0.402.6 ± 0.7- SauceBottlesCrateTask80.00.5008.12 ± 1.11−3.72 ± 1.130.81 ± 0.436.5 ± 1.5- SmallerObjectButterInBinTask0.00.000-−6.45 ± 0.900.73 ± 0.042.5 ± 0.2- Stack3RubiksCubeTask0.00.000-−6.03 ± 0.872.24 ± 0.323.5 ± 0.5rubikscube(10), rubikscube2(7), rubikscube1(3) StackWhiteMugsTask0.00.000-−6.51 ± 2.571.76 ± 0.333.0 ± 0.5redmug(9) TakeMeasuringSpoonOutTask30.00.14315.24 ± 5.36−6.17 ± 1.961.36 ± 0.714.2 ± 1.6bowl(9), measuringcup(8), cordlessdrill(7) ToolOrganizationBothTask0.00.000-−8.45 ± 1.7910.53 ± 1.085.0 ± 0.5- ToolOrganizationLeftTask0.00.000-−7.43 ± 1.427.04 ± 1.383.8 ± 0.7springclamp(2) ToolOrganizationSpecificTask0.00.000-−8.38 ± 1.457.26 ± 0.963.8 ± 0.5- ToolsPickingAllHammersTask0.00.438-−7.06 ± 1.039.43 ± 1.263.9 ± 0.5leftbin(6), centerbin(1) ToolsPickingDrillTask30.00.0000.13 ± 0.12−7.04 ± 5.292.00 ± 1.415.4 ± 4.2bluehammer(11), redhammer(3), clamp(1), huskyhammer(1) ToolsPickingHammerTask0.00.450-−7.09 ± 0.792.35 ± 0.313.9 ± 0.5leftbin(8), centerbin(1) UnstackRubiksCubeTask0.00.425-−6.18 ± 0.683.57 ± 0.393.8 ± 0.4rubikscubebottom(5) WhiteMugInCenterOfTableTask10.00.00029.40−4.52 ± 0.891.34 ± 0.194.3 ± 0.6mug(13) WhiteMugsInBinMoreVagueTask0.00.000-−6.98 ± 1.052.08 ± 0.623.3 ± 1.0ketchupbottle(4), rubikscubebottom(2) WhiteMugsInBinTask0.00.212-−5.50 ± 0.942.93 ± 0.644.8 ± 1.1rubikscubemiddle(3), ketchupbottle(2), greybin(1) WhiteMugsInBinVagueTask0.00.050-−6.92 ± 1.262.46 ± 0.464.0 ± 0.7ketchupbottle(2), greybin(2), rubikscubebottom(1), rubikscubemiddle(1) WoodSpatulaToBowlTask0.00.025-−5.45 ± 0.693.85 ± 0.536.2 ± 0.8pumpkinlarge(8) YellowAndWhiteObjectsInBinTask0.00.250-−5.75 ± 1.733.05 ± 0.394.9 ± 0.6greybin(6) TABLE XIII: Quantitative comparison for Scene Generation. We evaluate our against the baseline across diverse metrics measuring the visual realism (Real.), functionality (Func.), layout correctness (Lay.), Quality (Qual.), VQA score [18], and GPT Preference. MethodVQA (↑)Real. (↑)Func. (↑)Lay. (↑)Compl. (↑)Qual. (↑)# Obj (↑)GPT Pref. (↑) Baseline0.398(±0.04)6.889(±0.36)6.221(±0.58)6.166(±0.42)4.687(±0.57)5.991(±0.24)13.750(±10.52)18.000 Ours0.554 (±0.03)8.755 (±0.27)8.951 (±0.33)7.919 (±0.28)8.207 (±0.40)8.458 (±0.25)26.870 (±24.90)82.000 You are a scene generation expert creating REALISTIC robot manipulation scenarios. REAL-WORLD SCENE PRINCIPLES: 1. Objects form CLUSTERS - not evenly spaced grids 2. Containers (bowls, bins) have objects INSIDE them 3. Supports (plates, trays) have objects ON TOP 4. Objects scatter naturally AROUND containers 5. Orientations VARY - not all aligned to 0 ◦ /90 ◦ COORDINATE SYSTEM: - Table bounds: X=[0.25 to 0.85], Y=[-0.40 to 0.40] (meters) - Table center: (0.55, 0.0) - Front=+X, Back=-X, Left=+Y, Right=-Y PLACEMENT TYPES: 1. place-on-base: Object directly on table “type”: “place-on-base”, “object”: “bowl 0”, “x”: 0.4, “y”: 0.1, “yaw”: 23 VARY yaw angles (15, 47, 123, not just 0/90/180). Position matters for anchors, less for loose objects. 2. place-in: Objects INSIDE a container “type”: “place-in”, “objects”: [“apple 01”, “orange01”], “container”: “bowl0” Container MUST be placed first with place-on-base. Great for fruits in bowls, items in bins. 3. place-on: Object ON TOP of support “type”: “place-on”, “object”: “banana”, “support”: “platelarge”, “position”: “center” Support MUST be placed first. position: “center”, “edge”, or “random” Great for food on plates, items on trays. 4. cluster-around: Objects scattered NEAR an anchor “type”: “cluster-around”, “objects”: [“mug”, “spoon”], “anchor”: “bowl 0”, “radius”: 0.15 Creates natural groupings. radius: how far from anchor (0.10–0.20m typical) SCENE STRUCTURE (follow this pattern): 1. Place 1-2 ANCHOR objects (containers/supports) on table 2. Put objects INSIDE containers (place-in) 3. Put objects ON supports (place-on) 4. Cluster objects AROUND anchors (cluster-around) 5. Add a few LOOSE objects to fill space REALISTIC SPACING: - Anchors: 0.25-0.35m apart - Clustered objects: 0.08-0.15m from anchor - Loose objects: fill remaining space naturally Fig. 10: System prompt for Stage I (Semantic Planning). This prompt instructs the LLM to generate physically plausible scene layouts using structured predicates rather than raw coordinates. OUTPUT FORMAT (JSON only, no markdown): “objects”: [ “name”: “bowl 0”, “name”: “platelarge”, “name”: “apple 01”, “name”: “orange01”, “name”: “banana”, “name”: “mug”, “name”: “spoon” ], “predicates”: [ “type”: “place-on-base”, “object”: “bowl 0”, “x”: 0.40, “y”: 0.15, “yaw”: 23, “type”: “place-on-base”, “object”: “platelarge”, “x”: 0.65, “y”: -0.10, “yaw”: 156, “type”: “place-in”, “objects”: [“apple01”, “orange01”], “container”: “bowl0”, “type”: “place-on”, “object”: “banana”, “support”: “platelarge”, “position”: “center”, “type”: “cluster-around”, “objects”: [“mug”, “spoon”], “anchor”: “bowl0”, “radius”: 0.12 ] CRITICAL RULES: 1. Object names MUST match EXACTLY from catalog 2. Containers/supports MUST be placed before objects go in/on them 3. Create INTERESTING scenes with containment, stacking, AND clustering 4. VARY yaw angles - real scenes aren’t grid-aligned 5. Return ONLY valid JSON, no markdown Fig. 11: Continued System prompt for Stage I (Semantic Planning). This prompt instructs the LLM to generate physically plausible scene layouts using structured predicates rather than raw coordinates. TABLE XIV: Quantitative comparison across Difficulty Splits. We evaluate our method against the baseline on Easy, Medium, and Hard splits. Our method consistently outperforms the baseline across all difficulty levels. MethodVQA (↑)Real. (↑)Func. (↑)Lay. (↑)Compl. (↑)Qual. (↑)# Obj (↑)GPT Pref. (↑) Easy ([0, 5] objects) Baseline0.458(±0.02)6.767(±0.36)6.269(±0.57)6.079(±0.48)4.737(±0.61)5.963(±0.26)6.467(±0.52)33.333 Ours0.525 (±0.05)8.331 (±0.24)8.423 (±0.27)7.636 (±0.21)7.626 (±0.25)8.004 (±0.09)12.800 (±11.54)66.667 Medium ([6, 15] objects) Baseline0.401(±0.02)6.933(±0.37)6.223(±0.59)6.199(±0.41)4.634(±0.56)5.997(±0.22)10.957(±2.11)17.143 Ours0.561 (±0.02)8.779 (±0.18)8.996 (±0.23)7.932 (±0.24)8.235 (±0.29)8.485 (±0.13)22.414 (±19.15)82.857 Hard ([16, 20] objects) Baseline0.326(±0.02)6.808(±0.30)6.162(±0.59)6.099(±0.42)4.883(±0.58)5.988(±0.30)34.067(±14.89)6.667 Ours0.553 (±0.03)9.067 (±0.13)9.271 (±0.17)8.142 (±0.27)8.659 (±0.25)8.785 (±0.07)61.733 (±28.80)93.333 SCENE REQUEST:theme from dataset TARGET: target object count objects TABLE SIZE: 0.7m × 1.0m = 0.70m 2 (objects must fit with spacing!) SIZE LIMITS (max 1-2 large objects, prefer smaller for 8+ items): Large (> 0.08m 2 ):computed from catalog footprint Avoid picking multiple large objects - they won’t all fit! AVAILABLE OBJECTS: CONTAINERS (can hold objects inside): filled from catalog SUPPORTS (can stack objects on):filled from catalog FOOD:filled from catalog TOOLS:filled from catalog OTHER:filled from catalog MEDIUM SCENE STRATEGY (10-14 objects): - Use 1-2 containers/supports as ANCHORS - Put 2-4 objects IN containers (place-in) - Stack 1-2 items ON supports (place-on) - Cluster 2-3 objects near anchors (cluster-around) - VARY yaw angles - no grid alignment! SUGGESTED for diversity (use only if they fit the theme): preselected objects Fig. 12: User prompt template for Stage I (medium target count). The highlighted fields are populated at runtime (theme, target count, size warnings, catalog subsets, and diversity suggestions). Analogous strategy blocks are used for sparse (fewer than 10) and dense (15+) targets. PREVIOUS ATTEMPT FAILED: feedback string produced by spatial/physical solver or grammar checks Please fix the issues. Common fixes: - Use MORE containment (place-in) to reduce table crowding - Use MORE stacking (place-on) to utilize vertical space - Use clustering (cluster-around) instead of individual placements - Select SMALLER objects if collisions persist Fig. 13: Feedback block appended to the user prompt when spatial solving, physical placement, grammar checks, or intersection validation fails. The highlighted region is the dynamic diagnostic message. TABLE XV: Per-Scene Quantitative Analysis. We report the performance breakdown across 10 distinct scene themes. Our method demonstrates robust generalization, outperforming the baseline in nearly all metrics across diverse environments. ThemeMethodVQA (↑)Real. (↑)Func. (↑)Lay. (↑)Compl. (↑)Qual. (↑)# Obj (↑)GPT Pref. (↑) Bathroom Counter Baseline0.405(±0.02)6.901(±0.37)6.196(±0.73)6.260(±0.26)4.805(±0.56)6.041(±0.19)15.00(±0.00)10.00 Ours0.564 (±0.02)8.913 (±0.16)9.159 (±0.22)7.967 (±0.28)8.276 (±0.33)8.579 (±0.15)28.50 (±18.47)90.00 Classroom Supplies Baseline0.401(±0.02)6.881(±0.31)6.458(±0.50)5.931(±0.41)5.018(±0.55)6.072(±0.17)10.40(±1.26)50.00 Ours0.562 (±0.02)8.790 (±0.18)8.990 (±0.18)7.842 (±0.28)8.406 (±0.26)8.507 (±0.12)32.90 (±23.70)50.00 Craft Station Baseline0.399 (±0.02)7.062(±0.49)6.385(±0.56)6.171(±0.39)4.590(±0.60)6.052(±0.24)9.70(±0.48)10.00 Ours0.560 (±0.03)8.761 (±0.19)9.045 (±0.28)7.938 (±0.22)8.179 (±0.33)8.481 (±0.06)17.70 (±13.61)90.00 Garage Workstation Baseline0.410 (±0.01)7.032(±0.34)6.308(±0.63)6.246(±0.47)4.432(±0.62)6.005(±0.26)9.00(±0.47)0.00 Ours0.566 (±0.02)8.796 (±0.21)9.004 (±0.20)7.881 (±0.28)8.207 (±0.29)8.472 (±0.13)13.00 (±11.68)100.00 Garden Tools Baseline0.400(±0.02)7.018(±0.45)6.155(±0.53)6.315(±0.47)4.489(±0.53)5.994(±0.21)10.30(±0.95)30.00 Ours0.561 (±0.02)8.778 (±0.16)8.949 (±0.15)7.950 (±0.20)8.175 (±0.27)8.463 (±0.12)13.80 (±11.36)70.00 Kitchen Cabinet Baseline0.327(±0.02)6.862(±0.31)6.281(±0.61)6.008(±0.39)5.069(±0.48)6.055(±0.26)37.38(±19.66)12.50 Ours0.554 (±0.03)9.052 (±0.13)9.313 (±0.17)8.070 (±0.25)8.710 (±0.25)8.786 (±0.06)48.88 (±25.63)87.50 Laundry Sorting Baseline0.396(±0.01)6.795(±0.24)6.327(±0.63)6.269(±0.34)4.536(±0.47)5.982(±0.21)12.30(±1.70)0.00 Ours0.558 (±0.03)8.742 (±0.17)8.975 (±0.26)8.021 (±0.16)8.202 (±0.28)8.485 (±0.12)35.30 (±27.34)100.00 Office Desk Baseline0.457(±0.02)6.747(±0.36)6.183(±0.25)6.101(±0.52)4.697(±0.59)5.932(±0.25)7.00(±0.00)57.14 Ours0.499 (±0.05)8.296 (±0.22)8.535 (±0.16)7.524 (±0.16)7.609 (±0.32)7.991 (±0.08)17.57 (±16.03)42.86 Storage Room Baseline0.324 (±0.02)6.747(±0.29)6.027(±0.58)6.203(±0.45)4.670(±0.64)5.912(±0.35)30.29(±5.91)0.00 Ours0.552 (±0.02)9.085 (±0.14)9.224 (±0.17)8.224 (±0.28)8.601 (±0.24)8.783 (±0.08)76.43 (±26.40)100.00 Tea Time Baseline0.458(±0.02)6.784(±0.39)6.345(±0.77)6.061(±0.48)4.773(±0.67)5.991(±0.29)6.00(±0.00)12.50 Ours0.547 (±0.03)8.361 (±0.26)8.325 (±0.31)7.735 (±0.21)7.642 (±0.18)8.016 (±0.10)8.63 (±1.85)87.50 Workshop Bench Baseline0.394 (±0.02)6.838(±0.35)5.733(±0.38)6.199(±0.47)4.564(±0.52)5.833(±0.22)10.00(±0.47)20.00 Ours0.556 (±0.03)8.675 (±0.10)8.846 (±0.22)7.928 (±0.26)8.200 (±0.30)8.412 (±0.16)15.70 (±10.36)80.00 TABLE XVI: LLM-judged quality metrics for812automatically generated manipulation tasks across59scenes and7competency axes. LLM JudgeCoverage CategoryNAlignmentClarityFeasibilityMatchAligned%Partial%ObjectPredicate color1160.810.940.800.9057400.880.29 conjunction1160.970.981.000.989190.880.29 counting1160.870.970.900.9260380.880.29 recognition1160.960.970.960.9785150.880.29 semantics1160.890.950.940.9472270.880.29 sorting1160.940.950.970.9686140.880.29 spatial1160.920.980.890.9580170.880.29 Overall8120.910.960.920.9576230.880.29