Paper deep dive
ModuLoop : Low-Level Code Generation using Modular Synthesizer and Closed-Loop Debugger for Robotic Control
Gina Yoon, Sumin Lee, Joo Yong Sim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/9/2026, 12:49:04 AM
Summary
ModuLoop is a robotic control framework that leverages pre-trained Large Language Models to generate and iteratively refine modular Python code from natural language instructions. It integrates a Modular Code Synthesizer for task decomposition and a Closed-Loop Debugger that uses execution feedback and simulation validation to autonomously correct errors and improve accuracy. Validated on hand-eye calibration and pick-and-place tasks, the framework demonstrates high execution accuracy, autonomy, and scalability without requiring task-specific fine-tuning.
Entities (9)
Relation Signals (8)
ModuLoop â appliesto â Hand-eye Calibration
confidence 96% ¡ We apply the proposed framework to the calibration of an RGB-D camera and a robotic arm
ModuLoop â uses â Modular Code Synthesizer
confidence 95% ¡ ModuLoop leverages a Modular Code Synthesizer to decompose tasks and generate initial code
ModuLoop â uses â Closed-Loop Debugger
confidence 95% ¡ ModuLoop employs a Closed-Loop Debugger to iteratively refine code based on execution results
ModuLoop â appliesto â Pick-and-place Task
confidence 94% ¡ through a subsequent pick-and-place task, we demonstrate not only the accuracy of the calibration but also the potential extensibility of the framework
Modular Code Synthesizer â generates â Python Code
confidence 94% ¡ transforms it into automatically executable code... Python script
Closed-Loop Debugger â refines â Python Code
confidence 94% ¡ iteratively refines the code based on execution results and accuracy metrics
ModuLoop â validateswith â Isaac Sim
confidence 93% ¡ employ a simulation-based validation environment in Isaac Sim to check workspace reachability
ModuLoop â controls â UR3 Robotic Manipulator
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) have demonstrated impressive performance across various domains, including code generation and problem solving. However, their application in robotic control, particularly in low-level tasks that require precise manipulation, real-time feedback, and environment-dependent execution, remains limited. To address this challenge, we propose the Closed-Loop Modular Code Synthesizer framework. This framework leverages a pre-trained LLM without any task-specific fine-tuning to perform modular code planning and generation, and iteratively executes the generated code while inserting debugging probes to observe its behavior. This closed-loop structure facilitates systematic debugging and refinement, ultimately producing executable control programs. We apply the proposed framework to the calibration of an RGB-D camera and a robotic arm, validating its effectiveness in real-world settings. Furthermore, through a subsequent pick-and-place task, we demonstrate not only the accuracy of the calibration but also the potential extensibility of the framework. Across both tasks, the framework achieved high execution accuracy and autonomy, illustrating the practicality and scalability of LLM-based robotic control using our framework.
Tags
Links
- Source: https://arxiv.org/abs/2606.03047v1
- Canonical: https://arxiv.org/abs/2606.03047v1
Trouble viewing inline? Open PDF directly â
Full Text
38,921 characters extracted from source content.
Expand or collapse full text
ModuLoop : Low-Level Code Generation using Modular Synthesizer and Closed-Loop Debugger for Robotic Control Gina Yoon Sumin Lee Joo Yong Sim Manuscript received: May, 24, 2025; Revised: August, 29, 2025; Accepted; October, 3, 2025.This paper was recommended for publication by Editor Chao-Bo Yan upon evaluation of the Associate Editor and Reviewersâ comments.This work was supported by National Research Foundation of Korea(NRF) grant (No. RS-2025-02216282, RS-2025-16070288), Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2022-I220025) and by Ministry of Trade, Industry and Energy (MOTIE RS 2023 00258591). (Gina Yoon and Sumin Lee are co-first authors.) (Corresponding author: Joo Yong Sim.)Gina Yoon, Sumin Lee, and Joo Yong Sim are with the Department of Mechanical Systems Engineering, Sookmyung Womenâs University, Cheongpa-ro 47-gil 100, Yongsan-gu, Seoul, 04310, South Korea (e-mail: ppippi272@sookmyung.ac.kr, sstone@sookmyung.ac.kr, jysim@sookmyung.ac.kr).Digital Object Identifier (DOI): see top of this page. Abstract Large Language Models (LLMs) have demonstrated impressive performance across various domains, including code generation and problem solving. However, their application in robotic controlâparticularly in low-level tasks that require precise manipulation, real-time feedback, and environment-dependent executionâremains limited. To address this challenge, we propose the Closed-Loop Modular Code Synthesizer framework. This framework leverages a pre-trained LLM without any task-specific fine-tuning to perform modular code planning and generation, and iteratively executes the generated code while inserting debugging probes to observe its behavior. This closed-loop structure facilitates systematic debugging and refinement, ultimately producing executable control programs. We apply the proposed framework to the calibration of an RGB-D camera and a robotic arm, validating its effectiveness in real-world settings. Furthermore, through a subsequent pick-and-place task, we demonstrate not only the accuracy of the calibration but also the potential extensibility of the framework. Across both tasks, the framework achieved high execution accuracy and autonomy, illustrating the practicality and scalability of LLM-based robotic control using our framework. I Introduction Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of domains, including natural language understanding, code generation, and symbolic reasoning. Advanced models such as the GPT series [18] have shown the ability to generate executable code from natural language instructions, highlighting their growing utility in industrial and applied settings [11]. In the field of robotics, recent approaches have sought to enhance LLMsâ task interpretation and planning capabilities through large scale dataset fine-tuning or few-shot learning[3, 23, 16]. Against this backdrop, recent LLM-based robotic control studies have generally progressed in two directions: (i) translating natural language commands into high-level plans while relying on external modules for execution [2, 23], and (i) directly generating executable low-level control or perception code through user-defined APIs from natural language instructions [16, 20, 10]. While the former limits LLM involvement in execution, the latter often suffers from prompt dependency and poor generalizationâissues particularly critical in low-level control where precise motion and real-time error handling are required. To address these challenges and enable LLMs to participate actively in the control loop, we propose ModuLoop, a framework that enables LLMs to participate in the full control loop of low-level robotic tasks. Fig. 1 illustrates the overall architecture of the ModuLoop based robotic control system. The pipeline begins with a high-level task command expressed in natural language. ModuLoop decomposes the given instruction into fine-grained subtasks and converts each subtask into executable Python code. Based on execution feedbackâincluding runtime errors and performance metricsâthe framework dynamically refines and corrects the generated code, forming a closed-loop structure in which the robot autonomously learns and improves through iterative execution. We validate ModuLoop in two representative tasks: (1) handâeye calibration between a camera and robotic manipulator, and (2) a pick-and-place manipulation task requiring perception and motion planning. Although originally designed for calibration without any user-defined APIs, the framework demonstrated extensibility to manipulation tasks with minimal perception guidance. Our main contributions are threefold: 1. A simulation-based LLM feedback loop that secures collision-free executable coordinates for handâeye calibration, thereby enhancing data efficiency and accuracy. 2. A modular code synthesizer framework that decomposes high-level natural language commands into executable low-level Python modules for robot control. 3. A closed-loop debugging mechanism that uses execution feedbackâsuch as runtime errors and accuracy metricsâto autonomously revise and improve generated code. Together, these contributions position ModuLoop as a prac- tical and generalizable framework that elevates LLMs from passive planners to autonomous agents capable of generating, executing, and adapting robot control code in real-world environments. Figure 1: The architecture of an LLM-based control framework that enables a robotic arm to autonomously execute tasks based on natural language commands. ModuLoop serves as a code synthesis framework that generates executable control code for both low-level handâeye calibration and high-level pick-and-place manipulation. Before code generation, the system collects task-specific environmental information â simulation-based reachable coordinate verification for calibration, and SAM- and LLM-based object detection for pick-and-place. Using these inputs, ModuLoop performs modular code synthesis and closed-loop debugging, producing verified control code that enables accurate calibration and autonomous manipulation. I Related Work I-A LLM-Driven Robotic Control Recent research on integrating Large Language Models (LLMs) into robotic systems has generally followed two directions. The first focuses on high-level action planning, as seen in systems such as SayCan [2] and PaLM-E, RT-1, RT-2 [6, 1, 23], which combine LLMs with affordance models or multimodal inputs to translate language commands into abstract action sequences. However, these approaches do not produce executable low-level control code and instead rely on downstream modules. The second direction explores directly generating robot control code from natural language. Code-as-Policies [16], for example, uses few-shot prompting with API specifications and annotated examples, while ProgPrompt [20] employs structured prompts with action primitives, and RobotGPT [12] generates and simulates code via ChatGPT with reinforcement learning. Despite producing executable code, most approaches remain limited by their reliance on structured inputs and lack of feedback-based refinement, reducing adaptability across environments. RobotScript [4] also demonstrates LLM-based code generation. It not only integrates perception tools like AnyGrasp [8] for grip pose prediction but also exposes them at the API level for direct invocation within the generated script. I-B Closed-Loop Refinement for Code Generation using LLM Existing LLM-based code generation systems (e.g., Codex [5]) operate in an open-loop manner, producing static code without execution feedback. AutoGen [22] introduces a multi-agent framework, while Reflexion[19] allows a single LLM to refine reasoning through self-reflection, both enabling iterative improvement. However, these methods remain confined to abstract programming and lack applicability to robotic control, where continuous sensing and physical feedback are essential. MCCoder [15] extends this line by incorporating sensor feedback for microcontroller-based robots, but its reliance on hardware-specific prompts and rigid templates limits generalizability. In contrast, ModuLoop autonomously synthesizes and improves executable control code by directly leveraging real-world feedback. I-C Hand-eye Calibration Hand-eye calibration [9] estimates the spatial transformation between a robotâs end-effector and a mounted camera, enabling vision-based manipulation. It aligns the coordinate frames of the robot and the camera so that perceived object positions can be accurately used for motion control. Recent learning-based methods automate this process, either by regressing extrinsics from RGB or point clouds without markers [7], or by aligning point clouds via registration pipelines [21, 14]. These reduce marker dependence and manual effort but remain sensitive to training data distribution and rely on accurate initial alignment for reliable performance. ModuLoop overcomes these issues by enabling automatic, adaptive calibration without prior data or manual tuning. I Low-Level Control with Workspace Validation via Simulation and Closed-Loop Modular Code Generation I-A Iterative Interaction between LLM and Simulator for Physically Valid Motion Planning To verify the physical feasibility of motion trajectories prior to real-world execution, we employ a simulation-based validation environment in Isaac Sim [17] to check workspace reachability and potential collisions. As shown in Fig. 2, within each iteration of the LLMâsimulator feedback loop, the LLM generates calibration targetsâspecifically, a set of 3D end-effector position candidates: Figure 2: A pipeline where candidate positions are pre-validated in simulation for reachability and collision before being used in the closed-loop modular code generation process. Candidate positions are tested in simulation and classified as reachable, collision, or unreachable. These results are fed back to refine future generations until a sufficient number of reachable positions is obtained. gpt=Gâ(,Lrea,Lcol,Lunr)C_gpt=G(W,L_rea,L_col,L_unr) where Wââ3W ^3 is the robotâs workspace constraint, Lrea,Lcol,LunrL_rea,L_col,L_unr are sets of previously observed reachable, collision, and unreachable positions, respectively, Gâ(â )G(¡) is a function that calls the LLM with prior outcomes as prompts, acting as a coordinate generator that produces candidate end-effector positions. The candidate positions CgptC_gpt are then filtered through inverse kinematics (IK): Cgen=câCgptâŁIâKâ(c)â â C_gen=\câ C_gpt IK(c)â \ where, CgenC_gen represents the subset of position candidates from CgptC_gpt that are kinematically reachable, i.e., those that pass IK filtering. Each câCgencâ C_gen denotes a single 3D end-effector position that satisfies this condition. IK serves both to reject infeasible samples and to provide executable joint trajectories. The simulator evaluates each pose câCgencâ C_gen using a function FsimF_sim, which classifies them into: (Crea,Ccol,Cunr)=Fsimâ(Cgen)(C_rea,C_col,C_unr)=F_sim(C_gen) where: CreaC_rea: positions successfully reached without collisions CcolC_col: positions causing self-collision or collisions with the environment CunrC_unr: positions the robot fails to reach These labeled results are re-fed into the LLM prompt to inform the next generation step, forming a feedback loop. The process repeats until enough reachable positions are collected. Through this process, the framework effectively integrates LLM-guided position generation with physics-based validation, enabling efficient collection of valid and executable robot coordinates. Table I compares three approaches for calculating reachable coordinates in the robot workspace. The first is an analytical method using UR3 DenavitâHartenberg (DH) parameters and joint limits. The second extends this by verifying inverse kinematics (IK) on the physical UR3 robot, yielding results closer to actual execution. The third is our simulation-based calibration method, where candidate coordinates are validated in simulation to assess reachability. Results show that the DH-based method has the lowest accuracy due to its reliance on analytical computation, while adding real-robot verification improves performance. Purely analytical approaches are prone to collision or infeasibility, whereas simulation-based validation provides the most reliable outcomes. In conclusion, the proposed framework integrates LLM-guided candidate generation with simulation-based validation, ensuring that only physically feasible and collision-free coordinates are retained. Simulation is employed exclusively in this validation stage, while the validated coordinates are subsequently used for real robot code generation, thereby enhancing execution fidelity. TABLE I: Comparison of reachable coordinate generation accuracy. Metric Analytical (DH-based) Analytical + Robot Verification Simulation-based Accuracy (%) 56.67 83.33 100.00 I-B ModuLoop for Hand-eye Calibration A closed-loop framework is proposed to automate the calibration process between the robotic arm and the camera using GPT-4o. The methodology consists of two main stages: (1) Modular Code Synthesizer (MCS), where an initial Python script is automatically generated from a natural language prompt, and (2) Closed-Loop Debugger (CLD), in which the script is iteratively refined based on execution results. An overview of the entire workflow is shown in Fig. 3. I-B1 Environment The process begins with the definition of the overarching goal of low-level control. In this case, the objective is to calculate the 3D-to-3D transformation matrix between the robotâs workspace and the cameraâs coordinate system. To achieve this, a structured prompt is provided to GPT, containing essential information such as the robotâs IP address and sensor specifications. In addition, movable coordinates obtained from simulation are also included in the prompt. I-B2 Modular Code Synthesizer The code generation stage systematically decomposes the calibration task and transforms it into automatically executable code. This process consists of three steps. Figure 3: ModuLoop: A LLM-based control framework architecture for autonomously executing low-level robotic tasks from natural language instructions. Given a task, along with essential information such as reachable TCP coordinates (from simulation), sensor specifications, and robot IP, the framework proceeds through six stages. (1) The task is decomposed into sequential code-generation steps. (2) A code block is generated for each step. (3) The individual blocks are integrated into an initial executable script. (4) The code is executed via a subprocess for real-time interaction. (5) Errors are analyzed: syntax errors are handled through direct debugging, while low accuracy triggers debugging-message insertion and further diagnosis. (6) Based on analysis, the code is refined iteratively until both correctness and task accuracy criteria are met. ⢠Task Decomposition & Planning: The LLM decomposes the given calibration task into a series of independent subtasks. For example, to compute the transformation matrix, the procedure may include steps such as moving to predefined coordinates, acquiring sensor data, and performing matrix computation. Each subtask is then passed as a structured prompt to the next LLM instance. ⢠Modular Code Generation: Each subtask is implemented as a Python function or code block. ⢠Code Integration: The generated code blocks are assembled into a coherent executable script while maintaining logical consistency. This script is executed via subprocesses, enabling the system to observe robot motions and process camera data in real time. Input: Initial code C0C_0; Evaluator E; Analyzer A; Refiner R; Max iterations T Output: Refined calibration code CâC^* CâC0Câ C_0; tâ0tâ 0; while t<Tt<T do /* Execute current code */ ((output o, error e)âRun(C)e)â Run(C) if eâ â eâ then /* If error occurs, insert probes and fix */ CdâA.add_probingâ(C,o,e)C_dâ A. add\_probing(C,o,e) (oâ˛,eâ˛)âRunâ(Cd)(o ,e )â Run(C_d) CâR.apply_fixâ(Cd,oâ˛,eâ˛)Câ R. apply\_fix(C_d,o ,e ) else /* If no error, evaluate calibration quality */ result râE.evaluateâ(o)râ E. evaluate(o) if r == âinaccurateâ then /* If inaccurate, diagnose and refine */ dâiâaâgâA.analyze_accuracyâ(C,o)diagâ A. analyze\_accuracy(C,o) CâR.apply_improvementâ(C,o,dâiâaâg)Câ R. apply\_improvement(C,o,diag) else if r == âcalibration successâ then /* If successful, return calibrated code */ return C tât+1tâ t+1; return Failure: calibration did not converge Algorithm 1 Closed-Loop Feedback for Calibration Figure 4: Syntax Error Debugging in Closed-Loop Process. This case illustrates the process of resolving a version mismatch error in OpenCVâs ArUco module. Based on the pre-refined code and its output message, probing code is generated to identify the root cause, and the issue is resolved by refining the code accordingly. Figure 5: Accuracy Improvement Debugging in Closed-Loop Process. This case demonstrates how adding depth alignment and filtering to the camera pipeline improves analysis accuracy. Based on the initial code and its output messages, the system analyzes potential causes of low accuracy and recommends refinement strategies. I-B3 Closed-Loop Debugger The code generation stage provides an initial Python script, but in real robotic environments, unexpected runtime errors or insufficient accuracy may prevent reliable execution. To address these issues, we propose a closed-loop feedback structure that diagnoses execution outcomes and iteratively refines the code (Algorithm 1). In this process, the LLM acts not only as a code generator but also as an active participant in a continuous loop of execution, analysis, and refinement. At each iteration, the current code C is executed in the robotic environment, producing an output o and an error message e. If an error occurs, the Analyzer formulates hypotheses about possible causes and re-executes the code with probing messages inserted to validate these hypotheses. The results are then passed to the Refiner, which applies targeted modifications accordingly. This process is exemplified in Fig. 4, which shows a syntax error case in OpenCVâs ArUco module resolved through probing and refinement. If no runtime error is detected, the Evaluator assesses calibration accuracy based on the output messages. The output includes either âcalibration successâ or âcalibration inaccurate.â If classified as âinaccurate,â the Analyzer diagnoses potential causes, and the Refiner updates the code to improve accuracy. An example of this process is shown in Fig. 5, where calibration accuracy is improved through additional alignment and filtering strategies. This loop continues until one of two termination conditions is met: (1) the calibration reaches the required accuracy, in which case the refined script CâC^* is returned, or (2) the maximum number of iterations T, is reached, in which case the process is considered a failure. Through this mechanism, the LLM functions as a self-correcting agent, capable of iterative adaptation, real-time diagnosis, and execution-driven refinement throughout the calibration process. IV Performance OF Modular Code generation and Closed-Loop Debugger TABLE I: Comparison of calibration performance across different code generation and feedback strategies. The iteration was terminated and deemed unsuccessful after ten attempts. Success Cases: calibration error << 0.02 m. Abbreviations: CaP - Code as Policies, SPG â Single Path Generator, MCS â Modular Code Synthesizer. Model Code Gen Success (%) Accuracy Success (%) Verified Error (Success Cases, m) mean / median Verified Error (All Cases, m) mean / median CaP with Basic Robot Control API - - - - CaP with Full API 90.0 - - 0.038/ 0.039 ProgPrompt 66.67 - - 0.026/ 0.026 SPG - - - - GPT-4o SPG + Debugger 6.67 - - 0.802/ 0.802 SPG + Debugger w/ Probing 10.0 - - 0.485/ 0.062 MCS 23.33 6.67 0.029/ 0.029 0.308/ 0.283 MCS + Debugger 73.33 60.0 0.174/ 0.019 0.074/ 0.014 MCS + Debugger w/ Probing 96.67 86.67 0.048/ 0.011 0.170/ 0.012 GPT-4.1 mini MCS + Debugger w/ Probing 73.33 40.0 0.085/ 0.068 0.337 /0.068 Gemini MCS + Debugger w/ Probing 90.0 73.33 0.064/ 0.064 0.091/ 0.065 IV-A Experimental Setup This experiment was designed to quantitatively evaluate the performance of two key components of the proposed closed-loop calibration framework: MCS and CLD. The objective is to assess their contribution to the robustness of code execution and calibration accuracy. The experiment was conducted using a UR3 robotic manipulator and an Intel RealSense D435i depth camera. An ArUco marker was attached to the robotâs tool center point (TCP), and the camera was positioned 1.3 meters in front of the robot, facing it directly. The robot is controlled through the RTDE (Real-Time Data Exchange) interface, and the control system operates at a frequency of 125 Hz. Communication between the robot and the control system is conducted via TCP/IP over a local network. To assess the effectiveness of the MCS and CLD, we included an additional model to provide a baseline for comparison. This model, referred to as the Single-Path Generator (SPG), generates the entire code in a single pass without any task planning, directly producing code from the given task description in a single prompt. In the feedback stage, two additional comparison models were included. The first is an open-loop model that performs no feedback at all, and the second is a limited closed-loop model without probing. The latter follows the closed-loop structure but does not insert diagnostic code to observe runtime behavior, relying solely on the final execution results as feedback. Thus, it represents a minimally applied feedback loop. All experiments were repeated 30 times to ensure statistical reliability. IV-B Experimental Results Each model was evaluated according to the following three criteria: 1. Successful code generation of without syntax errors 2. Satisfaction of a predefined accuracy threshold (<2<2 cm) 3. Positional error between the robotâs actual TCP position and the transformed position computed from the camera The experimental results are summarized in Table I. The combination of the Modular Code Synthesizer (MCS) and the Closed-Loop Debugger (CLD) achieved the highest success rate and accuracy, whereas the single-prompt approach showed substantially lower performance, indicating that task decomposition is critical for generating structurally and logically complete programs. Figure 6: Comparison of feedback loop iterations between models with and without the proposed Closed-Loop Debugger. Each bar represents the number of refinement iterations required for successful task completion across 16 code generation experiments. As shown in Fig. 6, integrating the CLD with the MCS improved calibration efficiency: Through probing-based debugging, more initial samples were successfully calibrated, and fewer iterations were required for convergence. This demonstrates that error-hypothesis probing accelerates convergence and enhances robustness This study compared ModuLoop with ProgPrompt [20] and Code as Policies (CaP) [16] on the same handâeye calibration task. ProgPrompt is a prompting technique that leverages code structures instead of natural language, where available robot actions and environmental objects are presented as Python functions and lists. By providing example tasks together with these definitions, the LLM is guided to generate new task plans directly in the form of code. Code as Policies is a paradigm where the LLM constructs control code around user-defined APIs, generating programs mainly by invoking provided APIs and, when necessary, defining new ones recursively. In our experiments, we evaluated two conditions. In CaP with Full API, the LLM was given a rich set of APIs covering not only low-level robot and sensor operations but also calibration procedures, along with hints. In contrast, CaP with Basic Robot Control API provided only simple functions for primitive robot control, while essential information as in ModuLoop was supplied for the rest. ProgPrompt and CaP with Full API achieved relatively strong performance but were constrained by their reliance on predefined APIs and the absence of feedback for error correction. ModuLoop, by comparison, does not depend on predefined APIs and directly generates low-level control code, which is iteratively refined through feedback, thereby demonstrating robustness and intuitive applicability even in precise robotic tasks. We further evaluated ModuLoop with different LLM variants. GPT-4.1-mini achieved slightly lower success rates than GPT-4o, but demonstrated advantages in cost and response speed. This indicates that it can serve as a cost- and time-efficient alternative, albeit with reduced reliability. Gemini achieved code-generation and accuracy comparable to GPT-4o, but the derived transformation matrices fell below the required accuracy threshold, indicating limited calibration stability. V Evaluation of the generalizability and practicality of the closed-loop code generation and debugging framework To validate the calibration-derived transformation matrix and assess ModuLoopâs generality, we applied it to a representative pick-and-place task involving object recognition, coordinate alignment, motion planning, and grasp planningâall handled through LLM-based reasoning and code generation. The LLM received natural language instructions, target poses from perception, and motion planning guidelines. Unlike calibration, where closed-loop debugging was feasible, executing pick-and-place trials risked collisions and environmental changes. We therefore adopted static debugging, supplying GPT-4o with a structured evaluation checklist to anticipate and reason about potential failure cases. V-A Object detection for pick and place When a natural language instruction is given, the system executes the task through a multi-stage pipeline. RGB-D images from front and gripper-mounted cameras are processed by the Segment Anything Model [13], which segments object masks and extracts their center coordinates. Each object is assigned a unique label, and the annotated images with the user command are provided to the LLM, which contextually interprets the scene to identify target objects rather than performing simple classification. For example, given âIâm hungry. Can you get me some snacks?â, the LLM selects relevant items such as âsausagesâ or âchocolate,â returning their names and IDs from each view (Fig. 7). The LLM also considers object geometry and context to decide if rotation is needed for precise pick-and-place execution. Figure 7: LLM-based Object Recognition for Pick-and-Place Tasks. RGB-D images from both front and gripper-mounted cameras are processed by the Segment Anything Model (SAM) to segment object masks and extract their center coordinates. The annotated images and the natural-language command are then provided to the LLM, which contextually interprets the scene to identify the target object and determine whether rotation is required for precise manipulation. V-B Quantitative Evaluation of ModuLoop on Pick-and-Place Tasks We evaluated ModuLoop on five pick-and-place tasks of increasing complexity and linguistic difficulty (see Table I), each executed 25 times using three code generation methods: (1) Single-Prompt Code Generation, (2) Modular Code Synthesizer, and (3) ModuLoop. The tasks were specifically structured to assess the systemâs capabilities in perception, reasoning, and control, with increasing complexity. in Fig. 8. shows example scenes of a robotic arm performing each of the five pick-and-place tasks based on natural language instructions. The main characteristics of each task are as follows: ⢠Simple 1: Basic object targeting and orientation control ⢠Moderate 1: Logical reasoning involving multiple objects and inference of color mixing ⢠Moderate 2: Semantic categorization of objects and orientation control ⢠Moderate 3: Distance-based ordering and applying appropriate orientations for identical objects located at different positions ⢠Hard 1: Spatial reasoning, precise relative positioning, and orientation control As shown in Table IV, ModuLoop achieved the highest success rates, particularly in complex tasks (Tasks 3â5) requiring semantic reasoning, spatial understanding, and precise motion planning, outperforming the Single-Prompt baseline. These results underscore the value of structured command decomposition and GPT-based debugging for translating complex manipulation commands into executable code. While this study confirms the feasibility of LLM-based low-level code generation in calibration and pick-and-place tasks, we acknowledge a key limitation: ModuLoop currently relies only on minimal APIs that provide object positions, limiting its applicability to relatively simple tasks. Contact-rich manipulation, such as assembly or opening a drawer, requires richer environmental understandingâincluding perception modules, motion planners, and value map composition provided as APIs as exemplified in works such as VoxPoser [10]. TABLE I: Capabilities required by each task to guide LLM-based code generation. Task ID Task Description Key Capabilities Required per Task (increasing complexity) Class Logical Multi Orientation Sequencing Spatial Reasoning Inference Object Control Reasoning Simple 1 Pick up the sausage on a plate. â Moderate 1 Put paints for coloring the Eggplant into the box. â â Moderate 2 Iâm hungry, can you put some snacks on a plate? â â â Moderate 3 Pick up Mentos in order from farthest to nearest. â â â Hard 1 Move the red block next to the blue block. â â â Figure 8: Pick-and-place tasks performed by a robot based on language instructions. Each task is visualized as a 2Ă2 image sequence, ordered from top-left to bottom-right to show the temporal progression of execution. TABLE IV: Success Rates(%) of Code Generation Methods Across Tasks Task ID Single Modular Code ModuLoop Code Generator Synthesizer Simple 1 60 88 96 Moderate 1 64 80 96 Moderate 2 36 60 92 Moderate 3 12 76 92 Hard 1 24 48 60 VI Conclusion This study verified that pre-trained large language models (LLMs) can autonomously generate and execute low-level robot control code without further training. The ModuLoop pipeline demonstrated high execution accuracy in camera-to-robot calibration and pick-and-place tasks, confirming the feasibility of LLM-based control in physical environments. By positioning the LLM as the central agent in the control process, this work demonstrates the feasibility of using pre-trained LLMs to translate natural language into robot-executable code. Its ability to handle complex control logic without manual programming or task-specific training underscores the approachâs practicality. However, due to the inherent structure of LLMs, generating complete control code introduces noticeable latency, which can limit responsiveness. In future work, we will extend the framework to real-world industrial applications such as process and factory automation, and evaluate its generality and hardware-independence across diverse robotic platforms beyond the UR3. References [1] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich (2023) RT-1: robotics transformer for real-world control at scale. In Robotics: Science and Systems (RSS), Cited by: §I-A. [2] A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al. (2023) Do as i can, not as i say: grounding language in robotic affordances. In Conference on robot learning, p. 287â318. Cited by: §I, §I-A. [3] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. (2020) Language models are few-shot learners. Advances in neural information processing systems 33, p. 1877â1901. Cited by: §I. [4] J. Chen, Y. Mu, Q. Yu, T. Wei, S. Wu, Z. Yuan, Z. Liang, C. Yang, K. Zhang, W. Shao, Y. Qiao, H. Xu, M. Ding, and P. Luo (2024) RoboScript: code generation for free-form manipulation tasks across real and simulation. CoRR abs/2402.14623. External Links: Link, Document, 2402.14623 Cited by: §I-A. [5] M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba (2021) Evaluating large language models trained on code. CoRR abs/2107.03374. External Links: Link, 2107.03374 Cited by: §I-B. [6] D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence (2023) PaLM-e: an embodied multimodal language model. CoRR abs/2303.03378. External Links: Link, Document, 2303.03378 Cited by: §I-A. [7] A. Falisse, S. D. Uhlrich, A. S. Chaudhari, J. L. Hicks, and S. L. Delp (2025) Marker data enhancement for markerless motion capture. IEEE Transactions on Biomedical Engineering. Cited by: §I-C. [8] H. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y. Xie, and C. Lu (2023) AnyGrasp: robust and efficient grasp perception in spatial and temporal domains. IEEE Transactions on Robotics (T-RO). Cited by: §I-A. [9] R. Horaud and F. Dornaika (1995) Hand-eye calibration. The international journal of robotics research 14 (3), p. 195â210. Cited by: §I-C. [10] W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei (2023) VoxPoser: composable 3d value maps for robotic manipulation with language models. In 7th Annual Conference on Robot Learning, External Links: Link Cited by: §I, §V-B. [11] J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim (2025-07) A survey on large language models for code generation. ACM Trans. Softw. Eng. Methodol.. External Links: ISSN 1049-331X, Document Cited by: §I. [12] Y. Jin, D. Li, J. Shi, P. Hao, F. Sun, J. Zhang, B. Fang, et al. (2024) Robotgpt: robot manipulation learning from chatgpt. IEEE Robotics and Automation Letters 9 (3), p. 2543â2550. Cited by: §I-A. [13] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, et al. (2023) Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, p. 4015â4026. Cited by: §V-A. [14] L. Li, X. Yang, R. Wang, and X. Zhang (2024) Automatic robot hand-eye calibration enabled by learning-based 3d vision. Journal of Intelligent & Robotic Systems 110 (3), p. 130. Cited by: §I-C. [15] Y. Li, L. Wang, S. Piao, B. Yang, Z. Li, W. Zeng, and F. Tsung (2024) MCCoder: streamlining motion control with llm-assisted code generation and rigorous verification. CoRR abs/2410.15154. External Links: Link, Document, 2410.15154 Cited by: §I-B. [16] J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng (2023) Code as policies: language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA), p. 9493â9500. Cited by: §I, §I, §I-A, §IV-B. [17] NVIDIA (2021) Isaac Sim. Note: https://developer.nvidia.com/isaac-simAccessed: 2025-04-30 Cited by: §I-A. [18] OpenAI (2024) GPT-4o. Note: https://openai.com/index/gpt-4oAccessed: 2024-04-10 Cited by: §I. [19] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao (2023) Reflexion: language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36, p. 8634â8652. Cited by: §I-B. [20] I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg (2023) Progprompt: generating situated robot task plans using large language models. In 2023 IEEE International Conference on Robotics and Automation (ICRA), p. 11523â11530. Cited by: §I, §I-A, §IV-B. [21] T. Tang, M. Liu, W. Xu, and C. Lu (2024) Kalib: markerless hand-eye calibration with keypoint tracking. arXiv preprint arXiv:2408.10562. Cited by: §I-C. [22] Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang (2024) AutoGen: enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling, Cited by: §I-B. [23] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han (2023-06â09 Nov) RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of The 7th Conference on Robot Learning, J. Tan, M. Toussaint, and K. Darvish (Eds.), Proceedings of Machine Learning Research, Vol. 229, p. 2165â2183. External Links: Link Cited by: §I, §I, §I-A.