Paper deep dive
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
Shudong Liu, Dongyang Chen, Enci Zhang, Jinwei Liang, Zheng Ma, Lewei Lu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/27/2026, 5:15:55 AM
Summary
The paper introduces EASEL, a benchmark designed to evaluate 'dexterous visual tool use' in multimodal agents, defined as fine-grained, closed-loop parameterized visual action. Unlike existing benchmarks that tolerate spatial imprecision, EASEL requires agents to incrementally paint a canvas to match a reference image or follow semantic instructions (annotation, handwriting, path planning) using precise brush parameters. The authors propose EASEL-Data, a 440k-sample curriculum dataset, and train EASEL-9B. Evaluations of 25 models show that current agents struggle with reconstruction similarity (0.40-0.54) and exhibit trajectory instability, though EASEL-9B shows significant improvement over its base model.
Entities (9)
Relation Signals (7)
EASEL â evaluates â dexterous visual tool use
confidence 98% · We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use...
EASEL-9B â trainedon â EASEL-Data
confidence 97% · EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%...
EASEL â includestask â Reference-Guided Visual Reconstruction
confidence 96% · EASEL... adopts reference-guided visual reconstruction as its primary proxy task...
EASEL-9B â finetunedfrom â Qwen3.5-9B
confidence 95% · We fine-tune Qwen3.5-9B (Team 2026) in two stages... EASEL-9B is not intended as an upper bound...
EASEL â includestask â Semantic Tasks
confidence 95% · EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning.
EASEL-Data â derivedfrom â COCO
confidence 85% · Reference pool... single-/dual-/multi-subject crops from COCO...
EASEL-Data â derivedfrom â Tiny-ImageNet
confidence 85% · Reference pool... natural images from Tiny ImageNet...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.
Tags
Links
- Source: https://arxiv.org/abs/2608.25417v1
- Canonical: https://arxiv.org/abs/2608.25417v1
Trouble viewing inline? Open PDF directly â
Full Text
44,579 characters extracted from source content.
Expand or collapse full text
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents Shudong Liu Dongyang Chen Enci Zhang Jinwei Liang Zheng Ma Lewei Lu Abstract Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this spaceâdexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40â0.54), while trajectory diagnostics expose severe closed-loop instabilityâmodels typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models. 1Peking University, 2Tsinghua University, 3SenseTime Research shooo, eczhang@stu.pku.edu.cn, chen-dy25@mails.tsinghua.edu.cn, liangjinwei, mazheng2, luotto@sensetime.com https://github.com/OOOHS/EASEL https://huggingface.co/EASEL-Bench 1 Introduction Figure 1: Approaches to visual task execution. Existing paradigms each avoid precise executionâthrough tolerance margins, understanding-only operation, or delegation to external models. EASEL requires the agent to directly produce parameterized actions through closed-loop refinement. Multimodal models increasingly complete tasks through external tools in complex digital environments. Evaluation has followed suit, expanding beyond answer-only QA toward tool-mediated execution in software engineering (Jimenez et al. 2024), web navigation (Deng et al. 2023; Zhou et al. 2024), desktop GUI use (Xie et al. 2024), and API calling (Patil et al. 2025). Even in the most visual of these settings, however, interaction remains precision-tolerant: actions succeed as long as they fall within a generous bounding region, so precise spatial control is never strictly demanded. Beyond such structured interfaces, recent efforts have extended agent capabilities into open-ended visual domains through visual search (Wu and Xie 2024; Shen et al. 2025; Yu et al. 2025), region annotation (Yang et al. 2023; You et al. 2024; Chen et al. 2023), pixel-level segmentation via specialized decoders (Lai et al. 2024; Zhang et al. 2024; Chen et al. 2024), code-driven visual programs (Gupta and Kembhavi 2023; SurĂs et al. 2023), and generative editing (Zhang et al. 2023; Sheynin et al. 2024; Ye et al. 2026). As Figure 1 illustrates, these approaches either aid understanding without producing persistent visual output, or delegate execution to external segmentation and generative models. Neither paradigm requires the agent to precisely control action parameters that directly form the result. To genuinely manipulate visual content, agents must go beyond high-level delegation to perform fine-grained spatial actions, which demands translating fine-grained visual understanding directly into precise execution parameters. By analogy with dexterous manipulation in robotics, where the manipulatorâs own motor actions directly shape physical outcomes through continuous feedback rather than issuing commands to an external actuator, we term this capability dexterous visual tool use. We operationalize it as fine-grained, closed-loop parameterized visual action: at each step, the model reads pixel-level visual cues to infer explicit tool parameters, those parameters directly govern the visual output, and the updated state drives subsequent refinement. What makes this demanding is not per-step perfection, but the need to continuously read fine-grained visual signals and translate them into corrective parameterized actions, without delegating to an external executor. Parameterized visual action is achievable in restricted settings: neural painting (Huang et al. 2019; Liu et al. 2021) and SVG generation (Li et al. 2020) via task-specific architectures, and SimpleSeg (Song et al. 2026) via direct polygon prediction in general-purpose models. SimpleSeg demonstrates that highly fine-grained, even pixel-level perception is achievable natively within general-purpose MLLMsâbut in a one-shot, static setting. Whether such perception can sustain across multi-step closed-loop execution, where each action changes the visual state and demands fresh spatial reasoning, remains untested; current evaluations center on task completion, where tolerance margins, abstraction, or external models absorb imprecision. Painting offers a natural proxy for this capability: each stroke visibly changes the canvas and shapes the next decision, while parameter errors are not absorbed by a tolerant interface or external generator. Evaluation instances moreover require no manual annotation: reference images can be sourced freely, and stroke trajectories are derived automatically, making the setting easy to scale. Guided by these properties, we introduce EASEL, whose primary task is reference-guided visual reconstruction: the agent observes a reference image and the current canvas, then issues incremental drawing actions, exposing the gap between semantic understanding and geometric precision. EASEL further probes the same parameterized-action interface through annotation, handwriting, and path-planning tasks driven by natural-language instructions; its annotation setting offers, to our knowledge, the first closed-loop evaluation in which general-purpose agents perform segmentation-style action without dedicated segmentation models. To examine whether trajectory-level supervision can improve this capability, we additionally construct EASEL-Data and train EASEL-9B via a two-stage curriculum: first on early-phase strokes, then extending supervision to mid-to-late steps. 1. We propose EASEL, operationalizing dexterous visual tool use as fine-grained, closed-loop parameterized visual action through 110 reference-guided reconstruction samples and 5 instruction-conditioned semantic tasks, covering up to 11,392 interactions per full evaluation. 2. We construct EASEL-Data via a scalable two-stage curriculum pipeline that converts stroke-based rendering trajectories into 440k next-action supervision samples (C1: early reconstruction; C2: light completion), and train EASEL-9B to examine the effect of trajectory supervision on closed-loop visual action. 3. We systematically evaluate 25 representative multimodal agents across both result quality and trajectory quality, revealing two distinct failure modes in closed-loop visual actionâearly saturation and post-peak degradationâand sharp capability boundaries on semantic tasks invisible to reconstruction metrics alone. Figure 2: EASEL overview. (A) At each step, the agent observes reference R and current canvas CtC_t, outputs a parameterized brush_stroke action, and receives the updated canvas as feedback. (B) The full rollout is evaluated along two axes: Result Quality and Trajectory Quality. 2 Related Work 2.1 Agent Evaluation and Visual Tool Use Agent benchmarks evaluate goal-directed execution in interactive environments, covering software engineering (Jimenez et al. 2024), web and GUI navigation (Deng et al. 2023; Zhou et al. 2024; Xie et al. 2024), and long-horizon visual tool use (Su et al. 2026). GUI grounding work such as SeeClick and ScreenSpot-Pro studies screen coordinate prediction (Cheng et al. 2024; Li et al. 2025b), but coordinates typically locate high-tolerance interactive elements and their precision does not directly determine result quality. A separate line of work lets models operate directly on images for understanding and reasoning: through active visual search and region refinement (Wu and Xie 2024; Shen et al. 2025; Zheng et al. 2025; Yu et al. 2025), region annotation and grounding (Yang et al. 2023; You et al. 2024), drawable workspaces (Hu et al. 2024; Li et al. 2025a), and program or tool invocation (Gupta and Kembhavi 2023; SurĂs et al. 2023; Zhang et al. 2025; Hong et al. 2025). Segmentation-capable MLLMs such as LISA, PSALM, and SAM4MLLM further extend to pixel-level annotation through mask decoders (Lai et al. 2024; Zhang et al. 2024; Chen et al. 2024), where segmentation is produced by a specialized interface rather than the agentâs own actions. SimpleSeg instead predicts polygon coordinate sequences directly (Song et al. 2026), instantiating parameterized visual action without decoder delegationâthe closest existing paradigm to dexterous visual tool use, though it operates one-shot without closed-loop refinement. In visual editing and production, instruction-based image editing modifies images from text instructions (Zhang et al. 2023; Sheynin et al. 2024; Liu et al. 2025; Ye et al. 2026), and tool-augmented design benchmarks connect VLMs to real design software (Jeong et al. 2026). These works extend agent capability from âlook and answerâ to âuse visual tools to complete tasks.â EASEL shares this trend but differs in that parameter precision directly determines result quality. 2.2 Parameterized Drawing and Generation Neural painting and stroke-based rendering are closest to EASEL in task form. Learning to Paint and Paint Transformer decompose target images into stroke sequences via RL and feed-forward prediction respectively (Huang et al. 2019; Liu et al. 2021). These works optimize task-specific models for reconstruction quality, whereas EASEL uses reconstruction as a diagnostic lens to evaluate general-purpose agents under a unified tool interface. Draw with Thought reconstructs scientific diagrams as structured code (Cui et al. 2025); SketchAgent drives sketch creation from language descriptions (Vinker et al. 2025). Neither targets closed-loop refinement toward an explicit visual goal, which is EASELâs distinguishing focus. 3 EASEL 3.1 Task Formulation EASEL centers on a reconstruction task; semantic tasks extend the same closed-loop framework to instruction-guided drawing (Figure 2). Reconstruction. Given a reference R and blank canvas C0C_0, at each step t the agent observes (R,Ct)(R,C_t) and outputs a brush_stroke action ata_t in JSON; the renderer updates Ct+1=Renderâ(Ct,at)C_t+1=Render(C_t,a_t). The episode runs for T steps with the goal of making CTC_T as visually close to R as possible. Semantic tasks. The agent observes a task image R and a text instruction l, aiming to produce a drawing that satisfies the instruction l (region outline, handwriting, or maze path). The action space extends to brush_stroke,undo,submit\ brush\_stroke, undo, submit\. The brush appearance is fixed (fully opaque dark-blue, medium width), and the model predicts only the path parameters. In maze tasks, wall-colliding actions are rejected as no-ops but consume one step, with collision feedback passed to subsequent context. Model interface. The default interface is minimal: in reconstruction, the model observes the reference image, the current canvas, and the brush tool schema. Semantic tasks additionally preserve a 3-turn history window. 3.2 Benchmark Construction Category #Samples Resolution Budget Programmatic 15 64Ă64 32 Spatial Geometry 40 128Ă128 96 Abstract/Cartoon 15 128Ă128 96 Natural Images 40 128Ă128 128 Semantic Tasks 5 64â512 64â128 Table 1: EASEL benchmark categories. EASEL includes 110 reconstruction samples and 5 semantic tasks across five categories (Table 1). The four reconstruction categories form a difficulty gradient along geometric structure, color distribution, and semantic prior strength. Programmatic provides the cleanest baseline for tool format and spatial parameterization; Spatial Geometry tests multi-object layout; Abstract/Cartoon weakens semantic priors, forcing reliance on visual evidence; Natural Images is the primary source of difficulty variation. Semantic Tasks are reported separately and include two region-annotation tasks (circle outline and left-apple outline), one handwriting task (write âhelloâ), and two constrained path-planning tasks (simple maze and harder maze path). For annotation and maze tasks, the initial canvas is the task image itself and the model draws over it; for handwriting, the model writes on a blank canvas. Programmatic, maze, and Circle Outline tasks are procedurally synthesized; remaining categories are generated via GPT-image-2. 3.3 Evaluation Metrics Figure 3: EASEL-Data construction pipeline. (1) Reference Pool: 11k reference images spanning four source domains. (2) Vector Trajectory Generation: a stroke-based rendering policy produces a 250-step vector trajectory per reference. (3) Two-stage SFT Construction: C1 densely covers early strokes (0â49, with full retention through stroke 16); C2 spans the full trajectory (0â249), yielding 440k next-action prediction samples. For 10% of references, a reasoning model annotates all corresponding trajectory samples with CoT supervision (44k CoT samples). Overview. EASEL evaluates each rollout along two axes. Result Quality measures whether the final visual output satisfies the target; Trajectory Quality diagnoses how the model uses intermediate feedback over the action sequence. Reconstruction. We compute SSIM, normalized RGB L1L_1 distance, and Edge IoU, and define the similarity StS_t as: St= S_t= 0.5â SSIM \ 0.5·SSIM +0.3â (1âRGBL1) +0.3·(1-RGB_L_1) +0.2â Edge IoU +0.2·Edge IoU (1) Weights encode the intended balance: SSIM primary, RGB fidelity for appearance, Edge IoU for boundary structure. Ranking sensitivity across alternative weightings is reported in the supplementary material. For standard reconstruction, Result Quality is measured by Final Similarity STS_T, the primary ranking metric. Trajectory Quality is measured by four diagnostics: Similarity@50% (SâT/2âS_ T/2 ), Trajectory AUC (1TââtSt 1T _tS_t), Best Similarity (maxtâĄSt _tS_t), and Final-Best Gap (maxtâĄStâST _tS_t-S_T). F-B Gap diagnoses retention of the best achieved state and should be interpreted jointly with Final Similarity: a small gap can indicate either a strong result that was preserved or limited progress with little post-peak loss. Semantic Tasks. For semantic tasks, Result Quality is the task-specific score â[0,1]â[0,1]: (1) outline tasks (Circle Outline, Apple Outline) use the Dice coefficient between the modelâs drawn region and the target mask; (2) handwriting (Hello) uses a multimodal LLM to recognize each letter and compute positional match rate; (3) maze tasks (Simple Maze, Maze) compute the normalized progress of the modelâs path along the reference route. Semantic trajectories also log action traces, invalid action rate, undo/submit behavior, and intermediate canvases for failure analysis. Full semantic task specifications are in the supplementary material. 3.4 Tool Semantics The brush_stroke action comprises 13 continuous normalized parameters: a= a= (xs,ys,xc,yc,xe,yeâpath, ( x_s,y_s,x_c,y_c,x_e,y_e_path, (2) OPENrs,αs,re,αeâradius & opacity,cr,cg,cbâcolor) r_s, _s,r_e, _e_radius \& opacity, c_r,c_g,c_b_color) The path parameters (xs,ys)(x_s,y_s), (xc,yc)(x_c,y_c), (xe,ye)(x_e,y_e) define a quadratic BĂ©zier curve in normalized canvas coordinates. Appearance parameters include per-endpoint radius and opacity (rs,αs,re,αe)(r_s, _s,r_e, _e), linearly interpolated along the stroke, and a shared RGB color (cr,cg,cb)(c_r,c_g,c_b). A quadratic BĂ©zier primitive keeps the schema compact while exposing the continuous path and appearance parameters that constitute the target capability. Semantic tasks provide a reduced-schema setting by fixing appearance and requiring only the six path coordinates. The stroke is composited onto the canvas via alpha blending. The full schema and renderer specification are in the supplementary material. 4 EASEL-Data EASEL-Data construction proceeds in three stages: reference image preparation, vector trajectory generation, and two-stage SFT construction (Figure 3). Reference pool. We prepare 11k reference images from four sources: procedurally generated geometry and text/glyphs, single-/dual-/multi-subject crops from COCO (Lin et al. 2014), and natural images from Tiny ImageNet (Le and Yang 2015), spanning a visual complexity spectrum from simple geometry to complex natural scenes. Detailed category specifications are in the supplementary material. Vector trajectory pool. For each reference image, we generate a 250-step brush_stroke trajectory using a stroke-based painting policy (data policy) (Huang et al. 2019). This 250-step run is the data-native setting; a budget-aligned variant (âB/5â B/5 policy steps, five raster strokes each) serves as a score anchor in Section 5. Two-stage SFT construction. We convert trajectories into next-action prediction: given the reference and current canvas, the model predicts the next stroke, with canvases produced by replaying prior actions. Each supervision target is one observed continuation from the rendering policy. We construct a two-stage curriculum: C1 (Early Reconstruction, 308k samples) retains all steps from strokes 0â15 and samples from strokes 16â49, ensuring every blank-canvas state is covered and training the model to establish coarse structure and color; C2 (Light Completion, 132k samples) spans the full trajectory with decreasing density to cover mid-to-late refinement. Spatial coordinates are quantized to integers in [0,1000][0,1000] to align with MLLM tokenization. A reasoning model additionally generates CoT supervision for 10% of references, yielding 44k CoT samples. Full construction specifications are in the supplementary material. 4.1 EASEL-9B We fine-tune Qwen3.5-9B (Team 2026) in two stages using LoRA (rank 64, alpha 128, all-linear targets). Stage 1 trains on C1 with learning rate 1Ă10â41Ă 10^-4; Stage 2 continues from Stage 1 adapters on C2 with 15% C1 replay at a reduced learning rate 5Ă10â55Ă 10^-5. Both stages use bf16 precision, DeepSpeed ZeRO-2, effective batch size 64 (8 GPUs), max sequence length 8192, AdamW with weight decay 0.1 and cosine schedule. The ViT last four blocks and merger layers receive full gradient updates via modules_to_save. EASEL-9B is not intended as an upper bound on painting quality, but as a diagnostic probe for the effect of the two-stage curriculum. Full training hyperparameters are in the supplementary material. Control / Agent Final Sim.â Edge IoUâ Blank-white canvas 0.446 0.000 Reference mean-color fill 0.534 0.000 Reference dominant-color fill 0.531 0.000 Gemini 3.1 Pro (best MLLM) 0.535 0.053 Data policy, budget-aligned 0.665 0.281 Data policy, native (250-step) 0.715 0.563 Table 2: Reconstruction score anchors. Uniform fills score zero on Edge IoU; the data policy provides a reference point for attainable scores under two step-budget settings. Model Category Breakdown Overall Trajectory Model Prog. Spatial Abs. Nat. Finalâ @50%â AUCâ Bestâ F-B Gapâ Gemini 3.1 Pro 0.598 0.579 0.521 0.472 0.535 0.532 0.534 0.566 0.032 Claude Opus 4.7 0.469 0.531 0.468 0.415 0.472 0.473 0.474 0.506 0.035 Claude Sonnet 4.6 0.439 0.496 0.475 0.390 0.447 0.450 0.451 0.478 0.031 GLM-5V-Turbo 0.367 0.495 0.430 0.382 0.428 0.431 0.433 0.487 0.060 GPT-5.5 0.434 0.467 0.430 0.380 0.426 0.436 0.440 0.499 0.073 qwen3.5-plus 0.412 0.482 0.464 0.367 0.428 0.433 0.434 0.460 0.032 Qwen3.6-27B 0.489 0.476 0.498 0.364 0.440 0.440 0.441 0.454 0.014 Qwen3.5-35B-A3B 0.483 0.490 0.495 0.359 0.442 0.442 0.442 0.451 0.009 Qwen3.5-27B 0.469 0.473 0.486 0.361 0.433 0.433 0.434 0.451 0.018 Qwen3.5-9B 0.473 0.474 0.475 0.359 0.432 0.432 0.433 0.447 0.015 Qwen3.5-4B 0.482 0.484 0.486 0.362 0.440 0.440 0.420 0.448 0.008 Qwen3-VL-32B-Thinking 0.482 0.481 0.490 0.360 0.438 0.438 0.416 0.447 0.008 Qwen3-VL-30B-A3B-Inst. 0.421 0.491 0.469 0.361 0.431 0.433 0.433 0.464 0.032 Qwen3-VL-30B-A3B-Think. 0.416 0.478 0.484 0.355 0.426 0.428 0.410 0.455 0.029 kimi-k2.5 0.459 0.490 0.481 0.374 0.442 0.444 0.445 0.464 0.022 InternVL3-14B 0.472 0.438 0.448 0.344 0.420 0.420 0.420 0.451 0.019 InternVL3-8B 0.483 0.448 0.466 0.346 0.418 0.418 0.420 0.447 0.029 Llama 4 Maverick 0.480 0.475 0.490 0.360 0.437 0.437 0.437 0.445 0.008 Llama 4 Scout 0.473 0.470 0.480 0.355 0.430 0.430 0.430 0.438 0.008 Llama 3.2 Vision 90B 0.460 0.455 0.465 0.348 0.422 0.422 0.422 0.435 0.013 Llama 3.2 Vision 11B 0.455 0.450 0.460 0.345 0.415 0.415 0.415 0.428 0.013 Gemma3-27B 0.470 0.465 0.472 0.352 0.426 0.426 0.426 0.440 0.014 DeepSeek-VL2 0.452 0.445 0.455 0.342 0.414 0.414 0.414 0.428 0.014 MiMo-VL-7B-RL 0.411 0.442 0.444 0.348 0.404 0.406 0.409 0.451 0.047 EASEL-9B (w/o C2) 0.505 0.491 0.486 0.391 0.456 0.455 0.456 0.475 0.020 EASEL-9B 0.495 0.499 0.501 0.390 0.459 0.460 0.459 0.478 0.019 Table 3: EASEL standard reconstruction results and trajectory diagnostics. Final is the primary ranking metric; @50%, AUC, Best, and F-B Gap are trajectory diagnostics defined in Section 3.3. Models are grouped by family; EASEL-9B variants are listed separately as a curriculum ablation. Bold indicates the best result in each column; underline indicates the second best. 5 Experiments We evaluate 25 multimodal agents, including 6 closed-source APIs and 19 open-source models spanning 4Bâ90B parameters selected to cover the current performance frontier. The evaluated model families include proprietary Gemini, Claude, GPT, GLM, and Qwen APIs (Google DeepMind 2026; Anthropic 2026; OpenAI 2026; GLM-V Team 2026; Team 2026), as well as open-source Qwen, InternVL, Llama, Gemma, DeepSeek-VL, MiMo-VL, and Kimi models (Team 2026; Bai et al. 2025; Qwen Team 2026; Zhu et al. 2025; Meta AI 2025; Meta AI 2024; Gemma Team et al. 2025; Wu et al. 2024; LLM-Core-Team Xiaomi 2025; Kimi Team et al. 2026). 5.1 Standard Reconstruction Figure 4: Average similarity over painting progress across reconstruction tasks. Most gains occur within the first 10% of the step budget, after which trajectories plateau or degrade. Model Result Quality Trajectory Model Circleâ Appleâ Helloâ S-Mazeâ H-Mazeâ Maze Inv.â Maze Undo Non-Maze Undo Avg. Submit (%) Gemini 3.1 Pro 0.919 0.783 1.000 0.183 0.007 0.21 87 0 43.5 GPT-5.5 0.785 0.761 1.000 1.000 0.031 0.19 81 0 41.6 Claude Sonnet 4.6 0.521 0.491 0.200 0.157 0.047 0.93 0 1 48.7 Claude Opus 4.7 0.000 0.602 0.200 0.026 0.000 0.87 3 124 87.7 GLM-5V-Turbo 0.484 0.722 0.000 0.000 0.000 1.00 0 0 43.0 Table 4: Semantic task results and trajectory diagnostics. Circle and Apple are outline Dice scores; Hello is letter-level recognition accuracy; S-Maze and H-Maze are normalized path-progress scores. Maze Inv.: fraction of invalid actions on maze tasks; Maze Undo / Non-Maze Undo: undo counts on maze vs. outline+handwriting tasks; Avg. Submit%: mean of per-episode submit percentile (submit step Ă· episode budget Ă100Ă 100), with budget-exhausted episodes counted as 100%. Bold indicates best result per quality column and lowest Maze Inv. Table 3 reports reconstruction results and trajectory diagnostics; Figure 4 plots the average reconstruction curves for the six closed-source models. Score anchors and task solvability. Table 2 anchors the absolute score range. The best agent, Gemini 3.1 Pro (0.535), approximately matches the reference mean-color fill (0.534) in composite similarity, although Gemini captures edge structure (Edge IoU 0.053) while all uniform fills score zero on Edge IoU. The data policy reaches 0.665 under the budget-aligned setting (Edge IoU 0.281) and 0.715 in the data-native setting (Edge IoU 0.563). Both substantially exceed every evaluated agent, confirming that the task is not inherently unsolvable at this score range. The continuous span from uniform fills (Edge IoU 0) to Gemini (Edge IoU 0.053) to the data policy (Edge IoU 0.281â0.563) also confirms that the composite metric retains discriminative range well above current MLLMs, rather than saturating near 0.53. Overall performance. Overall performance is low and clearly stratified. Gemini 3.1 Pro achieves the best Final Similarity (0.535), followed by Claude Opus 4.7 (0.472), while open-source models cluster in the 0.40â0.45 range with small inter-model variance and no clear correlation with parameter count. EASEL-9B (full curriculum) reaches 0.459, ranking third overall and the highest score among non-proprietary models, outperforming its base model Qwen3.5-9B (0.432) by +0.027. Across categories, Natural Images is hardest due to complex textures and fine-grained color distributions, whereas Programmatic is easiest due to strong geometric priors. These results show that EASEL exposes a systematic weakness in closed-loop visual action. Trajectory patterns. Trajectory diagnostics reveal two qualitatively distinct behavioral patterns among closed-source models. Gemini 3.1 Pro, Claude Opus 4.7, and Claude Sonnet 4.6 rise sharply within the first 5% of steps and then plateau: their @50%, AUC, and Final scores are nearly identical, indicating these models effectively stop making meaningful changes after early gains. By contrast, GPT-5.5, GLM-5V-Turbo, and qwen3.5-plus continue acting but degrade: AUC exceeds Final in all three cases, with GPT-5.5 showing the strongest degradation (F-B Gap 0.073; AUC 0.440 vs. Final 0.426). The distinction reflects different local action tendencies: plateau models tend toward conservative outputs once visual gain diminishes, while degrading models continue generating novel actions regardless of canvas state. Across all closed-source models, the effective peak is reached within the first 10% of the step budget, indicating that neither strategy exploits the remaining feedback signal. F-B Gap interpretation. F-B Gap measures retention, not quality, and its interpretation depends on the underlying behavior. The five open-weight models with F-B Gap below 0.01 have exact-repeat rates of 82.3%â92.4%, so their small gaps reflect near-total inaction rather than active quality maintenance. Closed-source plateau models achieve comparable F-B Gaps (†0.035) through selectively conservative actions, whichâunlike verbatim repetitionâimplies genuine sensitivity to canvas state; Final Similarity separately captures the quality of what is being preserved. The gap between closed-source plateau models (0.47â0.54) and open-source near-inaction models (0.40â0.45) suggests stronger models read visual evidence more effectively early on. Yet even top closed-source models plateau well below the data policyâs 0.665â0.715: the bottleneck is not initial perception but sustained, feedback-driven correction across the trajectory. Trajectory supervision. We compare Qwen3.5-9B (base, 0.432) with EASEL-9B trained on C1 only (0.456, +0.024) and the full C1+C2 curriculum (0.459, +0.027 over base). C1 training alone accounts for the majority of the gain, establishing coarse structure and color from early strokes; adding C2 contributes a further +0.003, indicating that mid-to-late completion supervision provides modest but measurable improvement. Final-Best Gap increases slightly from 0.015 (base) to 0.019â0.020 (trained), suggesting that trajectory supervision improves local action quality while long-horizon feedback correction remains limited under the current setup. 5.2 Semantic Task Analysis Figure 5: Best results on each semantic task. Top row: starting canvas. Bottom row: model output (Gemini 3.1 Pro for Circle, Apple, and Hello; GPT-5.5 for S-Maze; Claude Sonnet 4.6 for H-Maze). We analyze five representative models with nontrivial and differentiated semantic behavior (Table 4); models producing near-zero scores on all tasks are omitted. Semantic tasks reveal sharper capability boundaries than reconstruction. Outline and handwriting. Top closed-source models can perform segmentation-style annotation through direct stroke control: Gemini 3.1 Pro and GPT-5.5 lead both outline tasks and achieve perfect handwriting, while weaker models produce partial or zero scores. Claude Opus 4.7 fails on Circle Outline entirely due to persistent draw-undo loopsâaction strategy stability is a prerequisite for tasks requiring cumulative canvas accumulation. Path planning. GPT-5.5 achieves a perfect score on Simple Maze despite weaker reconstruction performance, confirming that reconstruction precision and instruction-following are partially orthogonal dimensions. The harder Maze Path remains unsolved across all models. Action traces. Gemini 3.1 Pro and GPT-5.5 use undo almost exclusively on maze tasks, suggesting targeted collision recovery rather than general uncertainty; GPT-5.5 shows iterative correction on Simple Maze, submitting after 25 undos with a perfect scoreâthe clearest example of closed-loop self-correction in our evaluation. Claude Opus 4.7 inverts this pattern, with 124 non-maze undos and only 3 maze undos, consistent with unstable draw-undo loops that erase useful strokes and prevent accumulation. Claude Sonnet 4.6 almost never uses undo yet produces near-maximal invalid rates on maze tasks, indicating it neither detects nor responds to constraint violationsâa different failure mode in which the model proceeds confidently despite errors. 6 Discussion Scope and proxy validity. EASEL trades breadth for control: the 2D canvas isolates visual goals, parameterized actions, and closed-loop feedback from confounding variables, enabling focused diagnosis of fine-grained visual manipulation and closed-loop correction under a controllable protocol. Applications. In our setting, agents act on visual content through explicit, editable strokes rather than delegating execution to specialized modules or generative systems. The semantic tasks demonstrate breadth within this interface through annotation, handwriting, and path planning. Medical annotation, CAD sketching, diagram authoring, and embodied manipulation motivate the broader capability because they share visual state observation, structured action prediction, and potential closed-loop feedback. Beyond evaluation, the explicit action traces and environment feedback make EASEL compatible with outcome-based optimization, reinforcement learning, and planning methods. Generalization of trajectory supervision. EASEL-9B preserves visual perception after training: on standard perception benchmarks it matches or modestly improves upon the base modelâMMVP +0.7%, POPE +2.0%, HallusionBench +0.6% (Tong et al. 2024; Li et al. 2023; Guan et al. 2024), and on BLINKâs spatial reasoning and counting subtasks +0% and +1.7% respectively (Fu et al. 2024)âsuggesting that closed-loop action training preserves and modestly broadens fine-grained multimodal perception. This raises the possibility that EASEL-style tasks, combined with more targeted supervision design, could serve as a training signal for fine-grained visual perception. 7 Conclusion We present EASEL, a benchmark that operationalizes dexterous visual tool use as precision-critical, closed-loop parameterized visual action. By combining reference-guided reconstruction with instruction-conditioned semantic tasks and trajectory diagnostics, EASEL exposes failures obscured by answer-only QA and high-level tool-use benchmarks: models plateau early, degrade after initial gains, and rarely exploit corrective feedback. A two-stage curriculum (EASEL-9B) improves over the base model by a relative 6.3%, with early-phase training accounting for most of the gain and fine-grained visual perception improved, while long-horizon feedback correction remains limited. References Anthropic (2026) Anthropic Model system cards. Note: https://w.anthropic.com/system-cards Cited by: §5. Bai et al. (2025) S. Bai, Y. Cai, R. Chen, et al. Qwen3-vl technical report. External Links: 2511.21631, Link Cited by: §5. Chen et al. (2023) K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao Shikra: unleashing multimodal llmâs referential dialogue magic. arXiv preprint arXiv:2306.15195. Cited by: §1. Chen et al. (2024) Y. Chen, W. Li, C. Sun, Y. F. Wang, and C. Chen Sam4mllm: enhance multi-modal large language model for referring expression segmentation. In European Conference on Computer Vision, p. 323â340. Cited by: §1, §2.1. Cheng et al. (2024) K. Cheng, Q. Sun, Y. Chu, F. Xu, L. YanTao, J. Zhang, and Z. Wu Seeclick: harnessing gui grounding for advanced visual gui agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 9313â9332. Cited by: §2.1. Cui et al. (2025) Z. Cui, J. Yuan, H. Wang, Y. Li, C. Du, and Z. Ding Draw with thought: unleashing multimodal reasoning for scientific diagram generation. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 5050â5059. Cited by: §2.2. Deng et al. (2023) X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su Mind2web: towards a generalist agent for the web. Advances in Neural Information Processing Systems 36, p. 28091â28114. Cited by: §1, §2.1. Fu et al. (2024) X. Fu, Y. Hu, B. Li, Y. Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W. Ma, and R. Krishna Blink: multimodal large language models can see but not perceive. In European Conference on Computer Vision, p. 148â166. Cited by: §6. Gemma Team et al. (2025) Gemma Team, A. Kamath, J. Ferret, S. Pathak, et al. Gemma 3 technical report. External Links: 2503.19786, Link Cited by: §5. GLM-V Team (2026) GLM-V Team GLM-5v-turbo: toward a native foundation model for multimodal agents. External Links: 2604.26752, Link Cited by: §5. Google DeepMind (2026) Google DeepMind Gemini 3.1 pro model card. Note: https://deepmind.google/models/model-cards/gemini-3-1-pro/ Cited by: §5. Guan et al. (2024) T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, D. Manocha, and T. Zhou HallusionBench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 14375â14385. Cited by: §6. Gupta and Kembhavi (2023) T. Gupta and A. Kembhavi Visual programming: compositional visual reasoning without training. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 14953â14962. Cited by: §1, §2.1. Hong et al. (2025) J. Hong, C. Zhao, C. Zhu, W. Lu, G. Xu, and X. Yu Deepeyesv2: toward agentic multimodal model. arXiv preprint arXiv:2511.05271. Cited by: §2.1. Hu et al. (2024) Y. Hu, W. Shi, X. Fu, D. Roth, M. Ostendorf, L. Zettlemoyer, N. A. Smith, and R. Krishna Visual sketchpad: sketching as a visual chain of thought for multimodal language models. Advances in Neural Information Processing Systems 37, p. 139348â139379. Cited by: §2.1. Huang et al. (2019) Z. Huang, W. Heng, and S. Zhou Learning to paint with model-based deep reinforcement learning. In Proceedings of the IEEE/CVF international conference on computer vision, p. 8709â8718. Cited by: §1, §2.2, §4. Jeong et al. (2026) D. Jeong, S. Byun, K. Son, D. H. Kim, and J. Kim CANVAS: a benchmark for vision-language models on tool-based user interface design. Proceedings of the AAAI Conference on Artificial Intelligence 40 (26), p. 22182â22190. Cited by: §2.1. Jimenez et al. (2024) C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan Swe-bench: can language models resolve real-world github issues?. In International Conference on Learning Representations, Vol. 2024, p. 54107â54157. Cited by: §1, §2.1. Kimi Team et al. (2026) Kimi Team, T. Bai, Y. Bai, Y. Bao, et al. Kimi k2.5: visual agentic intelligence. External Links: 2602.02276, Link Cited by: §5. Lai et al. (2024) X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia Lisa: reasoning segmentation via large language model. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9579â9589. Cited by: §1, §2.1. Le and Yang (2015) Y. Le and X. Yang Tiny imagenet visual recognition challenge. CS 231N 7 (7), p. 3. Cited by: §4. Li et al. (2025a) C. Li, W. Wu, H. Zhang, Y. Xia, S. Mao, L. Dong, I. VuliÄ, and F. Wei Imagine while reasoning in space: multimodal visualization-of-thought. arXiv preprint arXiv:2501.07542. Cited by: §2.1. Li et al. (2025b) K. Li, Z. Meng, H. Lin, Z. Luo, Y. Tian, J. Ma, Z. Huang, and T. Chua Screenspot-pro: gui grounding for professional high-resolution computer use. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 8778â8786. Cited by: §2.1. Li et al. (2020) T. Li, M. LukĂĄÄ, M. Gharbi, and J. Ragan-Kelley Differentiable vector graphics rasterization for editing and learning. ACM Transactions on Graphics (TOG) 39 (6), p. 1â15. Cited by: §1. Li et al. (2023) Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J. Wen Evaluating object hallucination in large vision-language models. In The 2023 Conference on Empirical Methods in Natural Language Processing, External Links: Link Cited by: §6. Lin et al. (2014) T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. DollĂĄr, and C. L. Zitnick Microsoft coco: common objects in context. In European Conference on Computer Vision (ECCV), ZĂŒrich. External Links: Link Cited by: §4. Liu et al. (2025) S. Liu, Y. Han, P. Xing, F. Yin, R. Wang, W. Cheng, J. Liao, Y. Wang, H. Fu, C. Han, et al. Step1x-edit: a practical framework for general image editing. arXiv preprint arXiv:2504.17761. Cited by: §2.1. Liu et al. (2021) S. Liu, T. Lin, D. He, F. Li, R. Deng, X. Li, E. Ding, and H. Wang Paint transformer: feed forward neural painting with stroke prediction. In Proceedings of the IEEE/CVF international conference on computer vision, p. 6598â6607. Cited by: §1, §2.2. LLM-Core-Team Xiaomi (2025) LLM-Core-Team Xiaomi MiMo-vl technical report. External Links: 2506.03569, Link Cited by: §5. Meta AI (2024) Meta AI Llama 3.2 vision model card. Note: Accessed: 2026-05-26 External Links: Link Cited by: §5. Meta AI (2025) Meta AI Llama 4 model card. Note: Accessed: 2026-05-26 External Links: Link Cited by: §5. OpenAI (2026) OpenAI GPT-5.5 system card. Note: https://openai.com/index/gpt-5-5-system-card/ Cited by: §5. Patil et al. (2025) S. G. Patil, H. Mao, F. Yan, C. C. Ji, V. Suresh, I. Stoica, and J. E. Gonzalez The berkeley function calling leaderboard (bfcl): from tool use to agentic evaluation of large language models. In Forty-second International Conference on Machine Learning, Cited by: §1. Qwen Team (2026) Qwen Team Qwen3.6-27b model card. Note: https://huggingface.co/Qwen/Qwen3.6-27BAccessed: 2026-05-26 Cited by: §5. Shen et al. (2025) H. Shen, K. Zhao, T. Zhao, R. Xu, Z. Zhang, M. Zhu, and J. Yin Zoomeye: enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 6613â6629. Cited by: §1, §2.1. Sheynin et al. (2024) S. Sheynin, A. Polyak, U. Singer, Y. Kirstain, A. Zohar, O. Ashual, D. Parikh, and Y. Taigman Emu edit: precise image editing via recognition and generation tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 8871â8879. Cited by: §1, §2.1. Song et al. (2026) T. Song, H. Lu, H. Yang, L. Sui, H. Wu, Z. Zhou, Z. Huang, Y. Bao, Y. Charles, X. Zhou, and L. Wang Towards pixel-level vlm perception via simple points prediction. External Links: 2601.19228, Link Cited by: §1, §2.1. Su et al. (2026) Z. Su, J. Gao, H. Guo, Z. Liu, L. Zhang, X. Geng, S. Huang, P. Xia, G. Jiang, C. Wang, et al. Agentvista: evaluating multimodal agents in ultra-challenging realistic visual scenarios. arXiv preprint arXiv:2602.23166. Cited by: §2.1. SurĂs et al. (2023) D. SurĂs, S. Menon, and C. Vondrick Vipergpt: visual inference via python execution for reasoning. In Proceedings of the IEEE/CVF international conference on computer vision, p. 11888â11898. Cited by: §1, §2.1. Team (2026) Q. Team Qwen3.5: accelerating productivity with native multimodal agents. External Links: Link Cited by: §4.1, §5. Tong et al. (2024) S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie Eyes wide shut? exploring the visual shortcomings of multimodal llms. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9568â9578. Cited by: §6. Vinker et al. (2025) Y. Vinker, T. R. Shaham, K. Zheng, A. Zhao, J. E Fan, and A. Torralba Sketchagent: language-driven sequential sketch generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 23355â23368. Cited by: §2.2. Wu and Xie (2024) P. Wu and S. Xie V?: guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13084â13094. Cited by: §1, §2.1. Wu et al. (2024) Z. Wu, X. Chen, Z. Pan, et al. DeepSeek-vl2: mixture-of-experts vision-language models for advanced multimodal understanding. External Links: 2412.10302, Link Cited by: §5. Xie et al. (2024) T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, et al. Osworld: benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems 37, p. 52040â52094. Cited by: §1, §2.1. Yang et al. (2023) J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441. Cited by: §1, §2.1. Ye et al. (2026) Y. Ye, X. He, Z. Li, S. Yuan, Z. Yan, B. Hou, L. Yuan, et al. Imgedit: a unified image editing dataset and benchmark. Advances in Neural Information Processing Systems 38. Cited by: §1, §2.1. You et al. (2024) H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S. Chang, and Y. Yang Ferret: refer and ground anything anywhere at any granularity. In International Conference on Learning Representations, Vol. 2024, p. 57153â57180. Cited by: §1, §2.1. Yu et al. (2025) X. Yu, D. Guan, and Y. Gu Zoom-refine: boosting high-resolution multimodal understanding via localized zoom and self-refinement. arXiv preprint arXiv:2506.01663. Cited by: §1, §2.1. Zhang et al. (2023) K. Zhang, L. Mo, W. Chen, H. Sun, and Y. Su Magicbrush: a manually annotated dataset for instruction-guided image editing. Advances in Neural Information Processing Systems 36, p. 31428â31449. Cited by: §1, §2.1. Zhang et al. (2025) Y. Zhang, X. Lu, S. Yin, C. Fu, W. Chen, X. Hu, B. Wen, K. Jiang, C. Liu, T. Zhang, et al. Thyme: think beyond images. arXiv preprint arXiv:2508.11630. Cited by: §2.1. Zhang et al. (2024) Z. Zhang, Y. Ma, E. Zhang, and X. Bai Psalm: pixelwise segmentation with large multi-modal model. In European Conference on Computer Vision, p. 74â91. Cited by: §1, §2.1. Zheng et al. (2025) Z. Zheng, M. Yang, J. Hong, C. Zhao, G. Xu, L. Yang, C. Shen, and X. Yu Deepeyes: incentivizing" thinking with images" via reinforcement learning. arXiv preprint arXiv:2505.14362. Cited by: §2.1. Zhou et al. (2024) S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, et al. Webarena: a realistic web environment for building autonomous agents. In International Conference on Learning Representations, Vol. 2024, p. 15585â15606. Cited by: §1, §2.1. Zhu et al. (2025) J. Zhu, W. Wang, Z. Chen, et al. InternVL3: exploring advanced training and test-time recipes for open-source multimodal models. External Links: 2504.10479, Link Cited by: §5.