Paper deep dive
Task-Adaptive Rubrics for GUI Reward Modeling
Tao Xiong, Xavier Hu, Wenkai Wang, Qinzhuo Wu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/29/2026, 4:07:44 AM
Summary
The paper introduces AdaptRubric, a Coarse-to-Fine Rubrics Framework for task-adaptive GUI outcome reward modeling. It addresses limitations in existing reward verifiers by constructing explicit judging criteria through a category-level coarse stage (retrieving reusable task-family criteria) and an instance-level fine stage (generating compact, instruction-specific cues). AdaptRubric outperforms prior reward agents, improving F1 by 3.6 points and yielding a 4.23-point task-success gain in online reinforcement learning.
Entities (8)
Relation Signals (8)
AdaptRubric → improves → F1 Score
confidence 95% · AdaptRubric consistently outperforms prior reward agents, improving F1 by 3.6 points over the baseline average
AdaptRubric → improves → Task Success Rate
confidence 95% · yielding a 4.23-point task-success gain
AdaptRubric → uses → Category-Level Coarse Rubric Retrieval
confidence 92% · AdaptRubric performs category-level coarse rubric retrieval by routing the instruction to a GUI task family
AdaptRubric → uses → Instance-Level Fine Rubric Generation
confidence 92% · then conducts instance-level fine rubric generation to surface compact cues
AdaptRubric → outperforms → ZeroGUI
confidence 90% · AdaptRubric consistently outperforms prior reward agents... improving F1 by 3.6 points over the baseline average
AdaptRubric → outperforms → OS-Themis
confidence 90% · AdaptRubric consistently outperforms prior reward agents... improving F1 by 3.6 points over the baseline average
Qwen3-VL-8B-Instruct → servesas → Reward Verifier
confidence 88% · For AdaptRubric, the reward verifier uses Qwen3-VL-8B-Instruct
GRPO → optimizes → Policy
confidence 85% · the policy is optimized with Group Relative Policy Optimization (GRPO)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success criteria implied by the user instruction. Existing GUI reward verifiers, however, often under-specify how these criteria should be constructed for each task instance. Whether using generic rubric structures or implicit model reasoning, their judging criteria are not sufficiently task-adaptive: they can transfer checks across tasks, overlook concrete constraints in the current instruction, or become overly strict by enforcing unstated requirements. To address this limitation, we propose AdaptRubric, a Coarse-to-Fine Rubrics Framework that constructs task-adaptive judging criteria through a category-level coarse stage and an instance-level fine stage. AdaptRubric performs category-level coarse rubric retrieval by routing the instruction to a GUI task family and retrieving reusable task-family criteria, then conducts instance-level fine rubric generation to surface compact cues for concrete values, scopes, and constraints in the current instruction. Across offline reward evaluation and online reinforcement learning optimization, AdaptRubric consistently outperforms prior reward agents, improving F1 by 3.6 points over the baseline average under a matched image budget and yielding a 4.23-point task-success gain.
Tags
Links
- Source: https://arxiv.org/abs/2608.24174v1
- Canonical: https://arxiv.org/abs/2608.24174v1
Trouble viewing inline? Open PDF directly →
Full Text
100,696 characters extracted from source content.
Expand or collapse full text
Task-Adaptive Rubrics for GUI Reward Modeling Tao Xiong Affiliation: Zhejiang University Xavier Hu Affiliation: Zhejiang University Wenkai Wang Affiliation: Zhejiang University Qinzhuo Wu Affiliation: MiLM Plus, Xiaomi Inc.Correspondence: xiongtao@zju.edu.cn,sy_zhang@zju.edu.cn Changqiao Wu Affiliation: MiLM Plus, Xiaomi Inc.Correspondence: xiongtao@zju.edu.cn,sy_zhang@zju.edu.cn Pengzhi Gao Affiliation: MiLM Plus, Xiaomi Inc.Correspondence: xiongtao@zju.edu.cn,sy_zhang@zju.edu.cn Wei Liu Affiliation: MiLM Plus, Xiaomi Inc.Correspondence: xiongtao@zju.edu.cn,sy_zhang@zju.edu.cn Jian Luan Affiliation: MiLM Plus, Xiaomi Inc.Correspondence: xiongtao@zju.edu.cn,sy_zhang@zju.edu.cn Shengyu Zhang Affiliation: Zhejiang University Abstract Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome rewards by judging whether an executed trajectory satisfies the success criteria implied by the user instruction. Existing GUI reward verifiers, however, often under-specify how these criteria should be constructed for each task instance. Whether using generic rubric structures or implicit model reasoning, their judging criteria are not sufficiently task-adaptive: they can transfer checks across tasks, overlook concrete constraints in the current instruction, or become overly strict by enforcing unstated requirements. To address this limitation, we propose AdaptRubric, a Coarse-to-Fine Rubrics Framework that constructs task-adaptive judging criteria through a category-level coarse stage and an instance-level fine stage. AdaptRubric performs category-level coarse rubric retrieval by routing the instruction to a GUI task family and retrieving reusable task-family criteria, then conducts instance-level fine rubric generation to surface compact cues for concrete values, scopes, and constraints in the current instruction. Across offline reward evaluation and online reinforcement learning optimization, AdaptRubric consistently outperforms prior reward agents, improving F1 by 3.6 points over the baseline average under a matched image budget and yielding a 4.23-point task-success gain. †footnotetext: ‡Corresponding Author 1 Introduction Graphical User Interface (GUI) agents (Hu et al., 2025; Zhang et al., 2025; Liu et al., 2025a; Wang et al., 2024) have emerged as a promising paradigm for automating realistic digital devices. Enhancing GUI agents relies heavily on reliable outcome reward models (ORMs), which can filter high-value trajectories (Xia et al., 2025; Chen et al., 2025b; Xiong et al., 2025) from costly interactions, support benchmark evaluation and inference-time trajectory selection (Rawles et al., 2025; Xie et al., 2024), as well as provide reward signals (Xu et al., 2025; Xu et al., 2026; Li et al., 2026) for reinforcement learning. Figure 1: A motivating example for task-adaptive verification. The trajectory sets a related VS Code option to 50, but misses the requested line-length setting. While baselines verify only the surface value, AdaptRubric constructs coarse-to-fine criteria and correctly identifies the failure. Figure 2: Overview of AdaptRubric. The left panel motivates task-adaptive verification with a concrete trajectory example and the resulting criterion gap. The middle panel illustrates coarse-to-fine rubric construction, where AdaptRubric retrieves a category-level coarse rubric from the rubric bank and augments it with compact instruction-specific cues. The right panel shows the final verification stage, where the fused criterion and selected trajectory context are passed to a VLM verifier for outcome reward prediction. To determine whether a GUI trajectory succeeds, a reward verifier must first identify the success criteria (Gupta et al., 2025; Xie et al., 2026; Gunjal et al., 2025) implied by the instruction. These criteria specify the goal to be achieved, the instance-specific constraints to satisfy, and the visual evidence needed to support the final judgment. Existing GUI reward verifiers construct or obtain these criteria in two common ways. Static-template methods (Yang et al., 2025; Qi et al., 2025; Lai et al., 2025; Wang et al., 2025a; Pan et al., 2024) encode the criteria in a fixed rubric structure, which makes verification stable but keeps the rubric generic. Such a rubric may ask whether the task is completed, but it does not adapt its checks to the current instruction’s target object, required value, operation scope, or output format. Other methods (Li et al., 2026; Dai et al., 2026; Cui et al., 2026) leave the criteria to the model’s implicit reasoning during verification. This gives the verifier more flexibility, but without clear verification boundaries for different task categories, the model may overlook required instruction details or introduce constraints beyond the intended scope of the task. These limitations suggest that GUI reward verification requires task-adaptive criterion construction. The criteria should preserve verification boundaries for different task categories while adapting to the concrete requirements of each instruction instance. This allows the verifier to assess the trajectory against the task’s actual requirements, rather than a generic template or a loosely defined set of implicit checks. We propose AdaptRubric, a framework for task-adaptive criterion construction in GUI outcome reward modeling. AdaptRubric constructs explicit judging criteria through a category-level coarse stage and an instance-level fine stage. The coarse stage first routes each instruction to a GUI task category and retrieves a reusable rubric from a category-level rubric bank, where the rubric specifies verification steps, common pitfalls, and special rules for that category. The fine stage then uses the current instruction and the retrieved rubric to generate compact instance-level rubric cues, turning concrete requirements into task-specific checking items. Finally, AdaptRubric fuses both components into a single criterion supplied to the VLM verifier, so the reward model preserves category-level verification boundaries while adapting to the current instruction. We evaluate AdaptRubric on both offline reward discrimination and online reinforcement learning. On the offline GUI reward benchmark, AdaptRubric achieves 86.7% accuracy and 86.6 F1, improving F1 by 3.6 points over the baseline average under a matched image budget. In online reinforcement learning experiments, using AdaptRubric as the reward verifier yields a 4.23-point absolute gain in task success rate and achieves the best result among compared reward agents. These results show that task-adaptive criterion construction improves trajectory-level reward judgment and transfers to downstream GUI agent optimization. Our contributions are summarized as follows: • We formulate task-adaptive criterion construction for GUI outcome reward modeling, where a verifier must derive explicit judging criteria from the user instruction before assigning reward. • We introduce AdaptRubric, a coarse-to-fine rubric framework that combines category-level rubrics with compact instance-level rubrics into a single task-adaptive criterion for VLM-based reward verification. • We evaluate AdaptRubric on offline reward discrimination and online reinforcement learning, achieving 86.7% accuracy and 86.6 F1 on the offline benchmark and a 4.23-point absolute gain in online task success rate. 2 Related Work GUI Agents for Task Automation. The rapid progress of multimodal large language models (Singh et al., 2026; Anthropic, 2025; Team, 2026) has fundamentally transformed the landscape of autonomous GUI agents. Early GUI agents commonly relied on structured interface representations, such as DOM/HTML trees (Gur et al., 2018; Deng et al., 2023) for web pages and accessibility metadata for web or mobile interfaces (Li et al., 2024). With the emergence of large multimodal language models, later work increasingly shifted toward screenshot-based observations, often augmented with Set-of-Mark-style (Yang et al., 2023) visual annotations to expose clickable regions for visual grounding. To push their capability further, recent work post-trains these agents with supervised fine-tuning (Liu et al., 2025b; Hong et al., 2024) and offline reinforcement learning (Liu et al., 2025b; Liu et al., 2024). However, offline data limits how much an agent can explore beyond the trajectories it has already seen. Online reinforcement learning lets the agent interact with real environments, collect more diverse trajectories, and improve from a reward signal. To provide scalable training data and a reliable reward for both settings, outcome reward modeling is gaining increasing attention. Outcome Reward Modeling for GUI Agents. Reliable outcome reward modeling is critical for GUI agents. A common approach is to use programmatic or rule-based evaluators (Rawles et al., 2025; Xie et al., 2024; Chen et al., 2025a) when the task can be deterministically checked, as in environments with executable assertions or handcrafted reward functions. Such evaluators are precise but costly to design and hard to scale to open-ended GUI tasks. Recent work therefore turns to model-based evaluators that estimate task success from screenshots, action histories, final states, or UI metadata. One line of work directly applies LLM-as-a-judge, feeding the trajectory and the instruction into a single model to obtain a binary or graded reward (Yang et al., 2025; Qi et al., 2025; Lai et al., 2025; Wang et al., 2025a; Pan et al., 2024). Another line improves verification reliability by decomposing the trajectory into milestones, adding verification steps, or introducing process-level rewards (Li et al., 2026; Zheng et al., 2026; Dai et al., 2026; Cui et al., 2026). These methods mainly change how evidence is collected, decomposed, or aggregated during verification. We focus on a complementary issue: before the verifier evaluates any evidence, it needs a task-adaptive judging criterion derived from the instruction. AdaptRubric addresses this by combining category-level rubrics with instance-level rubrics into an explicit criterion for GUI outcome reward verification. 3 Preliminary A GUI agent interacts with an environment through a sequence of visual states and actions. Given a user instruction ℐI and an initial screenshot s1s_1, the agent executes an action ata_t at each step and receives the next screenshot st+1s_t+1 until termination. We denote the completed trajectory as τ=(s1,a1,s2,a2,…,sT,aT,sT+1),τ=(s_1,a_1,s_2,a_2,…,s_T,a_T,s_T+1), (1) where st∈s_t is a GUI screenshot, at∈a_t is the executed action, and sT+1s_T+1 is the terminal screenshot after the final action. An outcome reward model (ORM) maps the instruction and trajectory to a binary success judgment, r^=ℳ(ℐ,τ)∈0,1. r=M(I,τ)∈\0,1\. (2) Task-adaptive outcome reward modeling requires an explicit judging criterion that specifies the success conditions for the current instruction. We denote this criterion by ℛR and write r^=fθ(ℐ,τ′,ℛ)∈0,1, r=f_θ(I,τ ,R)∈\0,1\, (3) where fθf_θ is the verifier, τ′⊆τ τ is the selected trajectory context, and ℛR specifies the success criteria used to judge the trajectory. The goal of task-adaptive criterion construction is to derive ℛR for the current instruction before assigning reward. 4 Method Figure 2 gives a compact overview of AdaptRubric. As shown in the figure, Section 4.1 introduces category-level coarse rubric retrieval, Section 4.2 describes instance-level fine rubric generation, and Section 4.3 presents criterion fusion and reward verification. 4.1 Category-Level Coarse Rubric Retrieval We construct the category-level rubric bank once before evaluation through an LLM-assisted induction procedure. Starting from a development trajectory pool devD_dev collected from existing GUI agent benchmarks, we sample successful and failed trajectories and ask an LLM to summarize the verification dimensions that separate successful completions from failures. We then group these dimensions into broad GUI task categories. Each category rubric is refined on held-out trajectories, and whenever the rubric judgment disagrees with the trajectory label, we ask the LLM to revise it. The resulting taxonomy C contains K=8K=8 categories, namely info_query, create_modify, delete_cleanup, communication, transfer, state_navigation, composite_workflow, and general. The rubric bank is a fixed collection of category-level entries, ℬ=Ec=(mc,Sc)c∈.B=\E_c=(m_c,S_c)\_c . (4) Each entry EcE_c contains metadata mcm_c and a set of natural-language rubric sections ScS_c. The metadata records the category label and prompt-assembly information, while the rubric sections specify verification steps, common pitfalls, special rules, and output format: Sc=Scstep,Scpitfall,Scrule,Scformat,S_c=\S_c^step,S_c^pitfall,S_c^rule,S_c^format\, (5) During verification, AdaptRubric renders ScS_c into the category-level coarse rubric RcR_c. The taxonomy and rubric dimensions are summarized in Appendix A. At inference time, a task router predicts the GUI task category from the instruction, c=gϕ(ℐ),c∈.c=g_φ(I), c . (6) AdaptRubric then retrieves and renders the corresponding bank entry, Ec=Lookup(ℬ,c),Sc=Sections(Ec),Rc=Render(Sc). gatheredE_c=Lookup(B,c), S_c=Sections(E_c),\\ R_c=Render(S_c). gathered (7) We use top-1 lookup and if the router does not identify a specialized category, AdaptRubric falls back to the general entry EgeneralE_ general. 4.2 Instance-Level Fine Rubric Generation The fine stage constructs an instance-level fine rubric for the current instruction. Given the instruction ℐI, the predicted category c, and the retrieved coarse rubric RcR_c, a rubric generator produces Rf=hψ(ℐ,c,Rc),|Rf|≤2,R_f=h_ψ(I,c,R_c), |R_f|≤ 2, (8) where RfR_f is a compact list of fine-grained rubric items. The generator receives the coarse rubric as context, but it is not asked to rewrite RcR_c or fill its sections. Instead, it generates only additional instance-level checks that are grounded in the current instruction. The generator follows three constraints: • Instruction grounding. Each fine rubric item must be supported by an explicit phrase, value, or constraint in the user instruction. • Compactness. RfR_f contains at most two items, so the fine rubric highlights only the most useful instance-level requirements rather than becoming another full checklist. • Abstention. The generator may return Rf=∅R_f= when no reliable fine rubric item is needed. These constraints keep the fine rubric compact and reduce the chance of adding requirements that are not specified by the instruction. Model Ubuntu Mobile Windows macOS Web Overall Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc Prec Rec F1 ZeroGUI Qwen3-VL-4B-Instruct 83.8 84.0 74.5 76.2 80.3 75.6 90.9 78.8 81.1 83.2 82.0 83.2 80.0 81.6 Qwen3-VL-8B-Instruct 84.6 85.0 83.0 84.5 79.3 74.4 94.8 88.2 78.4 80.6 83.3 84.1 81.9 83.0 Qwen3-VL-32B-Instruct 84.6 85.2 83.5 85.0 80.3 75.6 96.1 90.3 82.1 83.5 84.1 84.6 83.1 83.9 Qwen3-VL-235B-A22B-Instruct 86.9 87.3 84.0 85.8 84.5 80.9 96.1 90.3 87.4 88.6 86.7 87.2 85.9 86.5 Qwen3.5-122B-A10B 85.8 86.5 82.4 83.7 83.6 80.2 96.1 90.9 80.0 82.7 84.8 84.1 85.6 84.8 Qwen3.6-27B 86.2 87.5 80.9 83.5 83.6 81.5 94.8 88.9 82.1 84.7 85.0 81.2 90.9 85.8 Gemini 3 Flash 88.5 89.0 80.3 80.6 87.8 85.7 97.4 93.8 87.9 88.3 87.7 89.2 85.7 87.4 Gemini 3.1 Flash-Lite 87.2 87.9 80.3 81.6 85.4 82.3 94.8 87.5 85.3 86.8 86.2 85.9 86.3 86.1 Mean 86.0 86.5 81.1 82.6 83.1 79.5 95.1 88.6 83.0 84.8 85.0 84.9 84.9 84.9 OS-Themis Qwen3-VL-4B-Instruct 72.6 71.4 79.3 78.9 75.1 68.3 84.4 57.1 80.0 81.9 75.5 79.5 68.3 73.5 Qwen3-VL-8B-Instruct 76.2 75.0 84.0 83.7 78.4 72.0 83.1 51.9 81.6 82.4 78.7 84.6 69.9 76.5 Qwen3-VL-32B-Instruct 77.1 74.6 81.9 80.7 76.5 65.3 90.9 77.4 83.7 81.9 79.3 91.6 64.1 75.5 Qwen3-VL-235B-A22B-Instruct 86.4 86.8 93.6 93.7 77.5 69.6 93.5 82.8 91.6 91.9 87.1 90.5 82.7 86.4 Qwen3.5-122B-A10B 85.8 85.8 88.3 88.2 79.8 73.0 88.3 66.7 79.0 75.6 84.5 92.0 75.3 82.8 Qwen3.6-27B 87.9 88.5 93.6 93.9 88.3 86.9 96.1 90.3 92.1 92.1 89.7 90.2 89.0 89.6 Gemini 3 Flash 81.5 81.5 84.0 84.2 86.9 84.4 88.3 71.0 80.5 77.3 82.9 88.1 75.9 81.5 Gemini 3.1 Flash-Lite 82.9 83.3 82.4 82.7 87.3 84.2 90.9 75.9 84.2 83.3 84.1 87.7 79.1 83.2 Mean 81.3 80.9 85.9 85.8 81.2 75.5 89.4 71.6 84.1 83.3 82.7 88.0 75.5 81.1 AdaptRubric Qwen3-VL-4B-Instruct 84.3 85.6 80.9 81.1 83.1 79.5 97.4 94.1 82.1 84.5 84.1 82.8 85.9 84.3 Qwen3-VL-8B-Instruct 83.9 85.1 83.5 84.6 81.7 79.1 97.4 94.1 81.6 83.3 84.0 82.6 85.9 84.2 Qwen3-VL-32B-Instruct 86.5 86.6 85.6 85.9 80.3 75.0 93.5 83.9 86.8 87.6 85.9 89.3 81.3 85.1 Qwen3-VL-235B-A22B-Instruct 86.1 86.7 87.8 88.9 84.5 81.4 94.8 87.5 88.9 90.0 86.9 87.0 86.7 86.8 Qwen3.5-122B-A10B 86.2 86.8 89.9 90.4 81.7 77.2 94.8 88.2 88.9 89.8 86.9 87.8 85.4 86.6 Qwen3.6-27B 87.3 88.5 91.0 91.6 84.5 82.5 93.5 85.7 91.6 92.4 88.3 85.4 92.1 88.7 Gemini 3 Flash 90.4 90.7 91.0 91.4 87.3 84.7 94.8 88.9 94.7 95.0 90.8 92.3 89.0 90.6 Gemini 3.1 Flash-Lite 86.4 87.1 83.0 83.3 85.0 81.6 94.8 87.5 88.9 89.7 86.5 87.2 85.4 86.3 Mean 86.4 87.1 86.6 87.2 83.5 80.1 95.1 88.7 87.9 89.0 86.7 86.8 86.5 86.6 Table 1: Offline reward discrimination results on OGRBench. Each method reports Accuracy (Acc) and F1-score (F1) for each platform; Overall columns include Acc, Precision (Prec), Recall (Rec), and F1. ZeroGUI uses the last-10 image budget. In each column, bold marks the best value across all 3×83×8 method-backbone cells, and underline marks the second-best (computed separately for the Mean rows). 4.3 Criterion Fusion and Reward Verification After obtaining the category-level coarse rubric and the instance-level fine rubric, AdaptRubric assembles the final judging criterion as ℛ=Φ(Rc,Rf)=Rc⊕Rf.R= (R_c,R_f)=R_c R_f. (9) The operator ⊕ keeps RcR_c as the main body of the criterion and appends RfR_f as a separated task-specific block. If the fine stage abstains, the final criterion reduces to the coarse rubric, Φ(Rc,∅)=Rc (R_c, )=R_c. The fused criterion therefore preserves the category-level verification boundary while adding the instruction-specific fine rubric when available. Before verification, AdaptRubric selects a compact trajectory context from the full GUI trajectory, τ′=σ(τ,c),|τ′|≤M.τ =σ(τ,c), |τ |≤ M. (10) The selector keeps screenshot-bearing steps that are useful for outcome judgment, including the initial and final states and frames around high-signal actions such as text entry, submission, or state-changing operations. The selected context, the instruction, and the fused criterion are then passed to the verifier, r^=fθ(ℐ,τ′,ℛ)∈0,1. r=f_θ(I,τ ,R)∈\0,1\. (11) The verifier outputs a binary reward in a parsable format. Thus, AdaptRubric changes the judging criterion and trajectory context supplied to the verifier, while keeping the verifier architecture unchanged. 5 Experiment 5.1 Offline Evaluation We evaluate offline GUI reward discrimination on OmniGUIRewardBench (OGRBench) (Li et al., 2026), a cross-platform benchmark for testing whether a reward verifier can correctly judge trajectory-level task success. (a) Overall Accuracy. (b) Overall F1. Figure 3: Cross-backbone radar comparison on OGRBench under the matched ten-screenshot budget. Each axis corresponds to one Qwen-family judge backbone, and each curve corresponds to one reward verifier. Benchmark. OGRBench contains 1,409 trajectories from five environments: OSWorld (Xie et al., 2024) for Ubuntu, AndroidWorld (Rawles et al., 2025) for mobile Android, (Bonatti et al., 2024) for Windows, macOSArena (Wang et al., 2025b) for macOS, and WebArena-Lite-v2 (Wang et al., 2025b; Zhou et al., 2023) for web tasks. The benchmark is nearly balanced overall, with 700 positive and 709 negative trajectories. Baselines. We compare with six representative GUI reward verification frameworks: DigiRL (Bai et al., 2024), DistRL (Wang et al., 2025a), AndroidGen (Lai et al., 2025), WebRL (Qi et al., 2025), ZeroGUI (Yang et al., 2025), and OS-Themis (Li et al., 2026). To align their trajectory inputs with AdaptRubric, all offline reward verifiers are evaluated under the same ten-screenshot budget. Appendix E details each baseline’s original trajectory-context setting and our unified matched-budget instantiation, and Appendix D analyzes the effect of image budget. For judge backbones, we evaluate open-source Qwen models from the Qwen3-VL (Bai et al., 2025), Qwen3.5 (Team, 2026), and Qwen3.6 (Qwen Team, 2026) series, as well as closed-source Gemini models (Team et al., 2025). Metrics. For offline reward discrimination, we report Accuracy, Precision, Recall, and F1. Main Results. Table 1 reports the main OGRBench comparison with ZeroGUI and OS-Themis across all eight judge backbones. In this setting, AdaptRubric achieves the best average performance, with 86.7% accuracy and 86.6 F1. Compared with the average of the two baselines under the main protocol, AdaptRubric improves the four overall metrics by 3.3 points on average, including 2.9 points in accuracy and 3.6 points in F1. The main gain comes from recall: the baseline average is 80.2%, while AdaptRubric raises recall to 86.5% with a comparable precision of 86.8%. This indicates that task-adaptive criteria help the verifier recover more successful trajectories while still controlling false positives. The improvement is also consistent across platforms, where AdaptRubric achieves the best average F1 on Ubuntu, Android, Windows, macOS, and Web tasks. These results support our core claim that combining category-level rubrics with instance-level fine rubrics yields more reliable trajectory-level reward judgment than generic terminal-state judging or implicit multi-step verification. Figure 3 further compares AdaptRubric with four additional baselines, DigiRL, DistRL, AndroidGen, and WebRL, using the open-source Qwen-family judge backbones. This expanded comparison tests whether the same advantage holds against a broader set of reward-agent prompts under the matched ten-screenshot budget. AdaptRubric achieves the best accuracy on all backbones and the best F1 on five of six backbones, showing that the gain is not tied to a single judge scale. The detailed values in Appendix E show that the expanded-baseline mean F1 of AdaptRubric remains higher than all four additional baselines. 5.2 Online Reinforcement Learning We further evaluate whether the offline reward advantage transfers to downstream GUI agent optimization. Setup. We conduct online RL training using the ClawGUI (Tang et al., 2026) framework on the MobileWorld (Kong et al., 2025) environment. The policy backbone is MAI-UI-8B (Zhou et al., 2025), and the policy is optimized with Group Relative Policy Optimization (GRPO) (Shao et al., 2024). For each task, GRPO samples four rollouts as a group, and the reward agent assigns an outcome reward to each completed trajectory. We keep the training protocol fixed and replace only the reward agent, comparing DigiRL, ZeroGUI, OS-Themis, and AdaptRubric. For AdaptRubric, the reward verifier uses Qwen3-VL-8B-Instruct and observes at most ten trajectory screenshots. All runs use a maximum episode length of 50 steps and train for 2 epochs; other hyperparameters are reported in Appendix C. Main Results. Table 2 reports online reinforcement learning results when different reward agents provide training rewards for MAI-UI-8B. Without an external reward agent, MAI-UI-8B reaches 19.70% task success. Using AdaptRubric as the reward verifier increases success to 23.93%, a 4.23-point absolute gain and the best result among compared reward agents. This suggests that the criterion constructed by AdaptRubric is not only better at offline trajectory discrimination, but also provides a more useful reward signal for policy improvement. Backbone Reward Agent SR (%) Δ MAI-UI-8B – 19.70† – DigiRL 23.08 +3.38+3.38 ZeroGUI 22.22 +2.52+2.52 OS-Themis 22.22 +2.52+2.52 AdaptRubric 23.93±0.8523.93± 0.85 +4.23+4.23 Table 2: Online reinforcement training results on MAI-UI-8B. We report the success rate (SR) after training with different reward agents. † denotes the result from Tang et al. (2026). (a) EarlyStop SR@N (N=1,3,5,7N=1,3,5,7). (b) BestOfN SR@N (N=1,2,4,6,8N=1,2,4,6,8). Figure 4: Test-time scaling results on the heterogeneous AndroidWorld pool of 113 tasks. Method ES@7 BoN@8 Acc. F1 FPR OS-Themis +2.15 +7.97 75.75 72.92 10.81 DigiRL +7.46 +7.97 80.80 81.90 22.71 ZeroGUI +10.11 +12.39 86.55 87.16 15.38 AdaptRubric +11.88 +13.28 88.14 88.41 11.17 Table 3: Heterogeneous-pool offline reward-guided test-time scaling. ES@7 (EarlyStop) and BoN@8 (BestOfN) report success-rate gains over Random in percentage points. Acc., F1, and FPR (false-positive rate) are computed over trajectory-level reward predictions. 5.3 Reward-Guided Test-Time Scaling We test whether reward verification helps select successful trajectories when multiple candidates are sampled for the same task. Setup. We evaluate on 113 AndroidWorld tasks, with ten candidate trajectories sampled for each task. We use this offline candidate pool to simulate online test-time scaling, so that all verifiers select from the same trajectories and the dynamic execution environment remains controlled. To create a discriminative candidate pool, we mix eight rollouts from Qwen3-VL-8B-Instruct and two rollouts from Qwen3-VL-235B-A22B-Instruct. Details are provided in Appendix F. All reward verifiers share Qwen3-VL-32B-Instruct as the judge and consume the same per-task shuffled trajectory stream. We evaluate two selection protocols: EarlyStop@N scans the first N trajectories sequentially and stops at the first one predicted as successful (online, low-latency setting), while BestOfN@N scores all N candidates and returns the highest-scored one (offline, batch setting). All verifiers in this comparison output binary scores, so multiple candidates often share the same score under BestOfN. In this case, we always pick the last tied candidate as the final selection. Both gains are reported relative to Random selection. Overall results. Table 3 summarizes performance under the two selection protocols. AdaptRubric achieves the largest gains over Random, improving EarlyStop@7 by +11.88+11.88 points and BestOfN@8 by +13.28+13.28 points. It also obtains the strongest trajectory-level reward prediction results, with 88.14 accuracy and 88.41 F1. At the same time, its false-positive rate remains low at 11.17, close to the most conservative baseline OS-Themis. EarlyStop: low-latency selection. Figure 4(a) reports EarlyStop SR@N as the per-task budget grows from 1 to 7. AdaptRubric is the strongest method at every budget from N=3N=3 onward. At N=7N=7, it reaches 64.60%64.60\% task success, exceeding ZeroGUI by 1.77 points, DigiRL by 4.42 points, and OS-Themis by 9.73 points. The Oracle upper bound at the same budget is 69.91%69.91\%, so AdaptRubric closes about 69%69\% of the gap between Random selection and the Oracle. Because EarlyStop commits to the first accepted trajectory, it favors verifiers that can avoid false-positive acceptances. BestOfN: batch selection. Figure 4(b) reports BestOfN SR@N for N∈1,2,4,6,8N∈\1,2,4,6,8\. AdaptRubric achieves the best result at N=4N=4 with 64.60%64.60\% task success and at N=8N=8 with 65.49%65.49\% task success, and it ties ZeroGUI at N=6N=6. The only weaker point is N=2N=2, where binary reward outputs often produce ties and the fixed tie-breaking rule has a large effect on the selected trajectory. DigiRL decreases as the candidate set grows from N=4N=4 to N=6N=6, consistent with the difficulty of judging longer visual contexts. Overall, the batch setting amplifies the advantage of AdaptRubric, because reward errors accumulate when the verifier must compare many candidates. Coarse Fine Setting Acc. Δ . F1 Δ 1 ✓ ✓ Full (Coarse + Fine) 86.9 — 86.6 — ✓ × w/o Fine 84.9 −2.0-2.0 84.6 −2.0-2.0 × ✓ w/o Coarse 85.2 −1.7-1.7 84.6 −2.0-2.0 × × w/o Both (image-only) 81.5 −5.4-5.4 80.7 −5.9-5.9 Table 4: Ablation study of rubric construction on Qwen3.5-122B-A10B. Δ . and Δ 1 report the absolute drop relative to the full model. 5.4 Ablation Study Table 4 isolates the contribution of the Coarse and Fine rubrics in AdaptRubric. The full model achieves 86.9 accuracy and 86.6 F1 on Qwen3.5-122B-A10B. Removing either component consistently weakens reward verification. Without the Fine rubric, F1 drops by 2.0 points, and without the Coarse rubric, F1 drops by 2.0 points as well. The largest decline appears when both rubrics are removed, where the verifier relies only on the instruction and trajectory screenshots and F1 falls by 5.9 points. This pattern shows that the two rubrics are complementary, since the Coarse rubric anchors the verifier to task-category boundaries and the Fine rubric adds instruction-specific success requirements. Method Quality LLM Cost Runtime Acc. F1 Calls Calls/traj. Tokens Tok./traj. Min. s/traj. OS-Themis 78.7 76.5 21,490 15.25 243.75M 173.0K 88.8 3.78 ZeroGUI 83.3 83.0 5,608 3.98 100.15M 71.1K 51.7 2.20 DigiRL 78.7 80.8 1,402 1.00 23.46M 16.6K 34.2 1.46 AdaptRubric 84.0 84.2 4,220 3.00 26.98M 19.1K 36.6 1.56 Table 5: Efficiency comparison on OGRBench using Qwen3-VL-8B-Instruct. All methods use the same ten-screenshot trajectory budget. Lowest cost in each column is in bold. 5.5 Efficiency Analysis Table 5 compares OGRBench quality and inference cost using Qwen3-VL-8B-Instruct. AdaptRubric achieves the best accuracy and F1, with 84.0 Acc. and 84.2 F1, while using substantially fewer calls, tokens, and runtime than OS-Themis and ZeroGUI. DigiRL is slightly cheaper, but its F1 remains 3.4 points lower than AdaptRubric. Overall, AdaptRubric provides the strongest quality–cost trade-off for repeated reward verification. 5.6 Case Study Appendix G presents a representative failure case showing why task-adaptive criteria matter. The task asks the agent to add text to the top of a note, but the trajectory inserts the text below existing content. Baselines focus on whether the requested text appears and incorrectly mark the task as successful. AdaptRubric combines a coarse create/modify rubric with fine placement cues, allowing the verifier to check both the edited item and the instruction-specific ordering constraint. 6 Conclusion We presented AdaptRubric, a coarse-to-fine rubric framework for task-adaptive GUI outcome reward modeling. The category-level Coarse rubric anchors verification to task-family success criteria, while the instance-level Fine rubric specifies the task-adaptive criteria needed for the current instruction. Across offline and online experiments, AdaptRubric consistently demonstrates the value of explicitly constructing task-adaptive criteria for success, supporting our core claim for reliable GUI reward modeling across GUI interaction settings. Limitations AdaptRubric improves GUI reward verification by constructing task-adaptive rubrics, but it still has several scope boundaries. First, although our evaluation covers mobile, desktop, and web environments, it does not exhaust all possible applications, interface designs, or user instruction styles. Second, the coarse rubric bank is constructed offline and kept fixed during evaluation to ensure reproducibility and controlled comparison across methods; this makes the protocol stable, but it may not capture every newly emerging application or domain-specific workflow. References Anthropic (2025) Anthropic Introducing claude 4. Note: https://w.anthropic.com/news/claude-4Accessed: 2026-05-25 Cited by: §2. Bai et al. (2024) H. Bai, Y. Zhou, M. Cemri, J. Pan, A. Suhr, S. Levine, and A. Kumar DigiRL: training in-the-wild device-control agents with autonomous reinforcement learning. External Links: 2406.11896, Link Cited by: Table 10, §5.1. Bai et al. (2025) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: §5.1. Bonatti et al. (2024) R. Bonatti, D. Zhao, F. Bonacci, D. Dupont, S. Abdali, Y. Li, Y. Lu, J. Wagle, K. Koishida, A. Bucker, L. Jang, and Z. Hui Windows agent arena: evaluating multi-modal os agents at scale. External Links: 2409.08264, Link Cited by: §5.1. Chen et al. (2025a) Y. Chen, Y. Liu, L. Zhang, P. Gao, J. Luan, and W. Liu STEP: success-rate-aware trajectory-efficient policy optimization. External Links: 2511.13091, Link Cited by: §2. Chen et al. (2025b) Z. Chen, D. Chen, R. Sun, W. Liu, and C. Gan Scaling autonomous agents via automatic reward modeling and planning. External Links: 2502.12130, Link Cited by: §1. Cui et al. (2026) C. Cui, J. Huang, S. Wang, L. Zheng, Q. Kong, and Z. Zeng Agentic reward modeling: verifying gui agent via online proactive interaction. External Links: 2602.00575, Link Cited by: §1, §2. Dai et al. (2026) G. Dai, S. Jiang, T. Cao, Y. Yang, Y. Li, R. Tan, M. Li, and L. Qiu ProRe: a proactive reward system for gui agents via reasoner-actor collaboration. External Links: 2509.21823, Link Cited by: §1, §2. Deng et al. (2023) X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su Mind2Web: towards a generalist agent for the web. External Links: 2306.06070, Link Cited by: §2. Gunjal et al. (2025) A. Gunjal, A. Wang, E. Lau, V. Nath, Y. He, B. Liu, and S. Hendryx Rubrics as rewards: reinforcement learning beyond verifiable domains. External Links: 2507.17746, Link Cited by: §1. Gupta et al. (2025) T. Gupta, S. Shandilya, X. Zhang, R. Madhavan, S. Ghosh, C. Bansal, H. Yao, and S. Rajmohan CARMO: dynamic criteria generation for context-aware reward modelling. External Links: 2410.21545, Link Cited by: §1. Gur et al. (2018) I. Gur, U. Rueckert, A. Faust, and D. Hakkani-Tur Learning to navigate the web. External Links: 1812.09195, Link Cited by: §2. Hong et al. (2024) W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Zhang, J. Li, B. Xu, Y. Dong, M. Ding, and J. Tang CogAgent: a visual language model for gui agents. External Links: 2312.08914, Link Cited by: §2. Hu et al. (2025) X. Hu, T. Xiong, B. Yi, Z. Wei, R. Xiao, Y. Chen, J. Ye, M. Tao, X. Zhou, Z. Zhao, Y. Li, S. Xu, S. Wang, X. Xu, S. Qiao, Z. Wang, K. Kuang, T. Zeng, L. Wang, J. Li, Y. E. Jiang, W. Zhou, G. Wang, K. Yin, Z. Zhao, H. Yang, F. Wu, S. Zhang, and F. Wu OS agents: a survey on mllm-based agents for general computing devices use. External Links: 2508.04482, Link Cited by: §1. Kong et al. (2025) Q. Kong, X. Zhang, Z. Yang, N. Gao, C. Liu, P. Tong, C. Cai, H. Zhou, J. Zhang, L. Chen, Z. Liu, S. Hoi, and Y. Wang MobileWorld: benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments. External Links: 2512.19432, Link Cited by: §5.2. Lai et al. (2025) H. Lai, J. Gao, X. Liu, Y. Xu, S. Zhang, Y. Dong, and J. Tang AndroidGen: building an android language agent under data scarcity. External Links: 2504.19298, Link Cited by: Table 10, §1, §2, §5.1. Li et al. (2024) W. Li, W. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva On the effects of data scale on ui control agents. External Links: 2406.03679, Link Cited by: §2. Li et al. (2026) Z. Li, Z. Wu, Y. Zhao, B. Yang, J. Xie, Z. Liu, Z. Liu, K. Jin, J. Liang, Z. Li, F. Wu, B. Zhou, Z. Wang, and Z. Ding OS-themis: a scalable critic framework for generalist gui rewards. External Links: 2603.19191, Link Cited by: Table 10, §1, §1, §2, §5.1, §5.1. Liu et al. (2024) X. Liu, B. Qin, D. Liang, G. Dong, H. Lai, H. Zhang, H. Zhao, I. L. Iong, J. Sun, J. Wang, J. Gao, J. Shan, K. Liu, S. Zhang, S. Yao, S. Cheng, W. Yao, W. Zhao, X. Liu, X. Liu, X. Chen, X. Yang, Y. Yang, Y. Xu, Y. Yang, Y. Wang, Y. Xu, Z. Qi, Y. Dong, and J. Tang AutoGLM: autonomous foundation agents for guis. External Links: 2411.00820, Link Cited by: §2. Liu et al. (2025a) Y. Liu, P. Li, Z. Wei, C. Xie, X. Hu, X. Xu, S. Zhang, X. Han, H. Yang, and F. Wu InfiGUIAgent: a multimodal generalist gui agent with native reasoning and reflection. External Links: 2501.04575, Link Cited by: §1. Liu et al. (2025b) Y. Liu, P. Li, C. Xie, X. Hu, X. Han, S. Zhang, H. Yang, and F. Wu InfiGUI-r1: advancing multimodal gui agents from reactive actors to deliberative reasoners. External Links: 2504.14239, Link Cited by: §2. Pan et al. (2024) J. Pan, Y. Zhang, N. Tomlin, Y. Zhou, S. Levine, and A. Suhr Autonomous evaluation and refinement of digital agents. External Links: 2404.06474, Link Cited by: §1, §2. Qi et al. (2025) Z. Qi, X. Liu, I. L. Iong, H. Lai, X. Sun, W. Zhao, Y. Yang, X. Yang, J. Sun, S. Yao, T. Zhang, W. Xu, J. Tang, and Y. Dong WebRL: training llm web agents via self-evolving online curriculum reinforcement learning. External Links: 2411.02337, Link Cited by: Table 10, §1, §2, §5.1. Qwen Team (2026) Qwen Team Qwen3.6-27B: flagship-level coding in a 27B dense model. External Links: Link Cited by: §5.1. Rawles et al. (2025) C. Rawles, S. Clinckemaillie, Y. Chang, J. Waltz, G. Lau, M. Fair, A. Li, W. Bishop, W. Li, F. Campbell-Ajala, D. Toyama, R. Berry, D. Tyamagundlu, T. Lillicrap, and O. Riva AndroidWorld: a dynamic benchmarking environment for autonomous agents. External Links: 2405.14573, Link Cited by: §1, §2, §5.1. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo DeepSeekMath: pushing the limits of mathematical reasoning in open language models. External Links: 2402.03300, Link Cited by: §5.2. Singh et al. (2026) A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. El-Kishky, A. McLaughlin, A. Low, A. Ostrow, A. Ananthram, A. Nathan, A. Luo, A. Helyar, A. Madry, A. Efremov, A. Spyra, A. Baker-Whitcomb, A. Beutel, A. Karpenko, A. Makelov, A. Neitz, A. Wei, A. Barr, A. Kirchmeyer, A. Ivanov, A. Christakis, A. Gillespie, A. Tam, A. Bennett, A. Wan, A. Huang, A. M. Sandjideh, A. Yang, A. Kumar, A. Saraiva, A. Vallone, A. Gheorghe, A. G. Garcia, A. Braunstein, A. Liu, A. Schmidt, A. Mereskin, A. Mishchenko, A. Applebaum, A. Rogerson, A. Rajan, A. Wei, A. Kotha, A. Srivastava, A. Agrawal, A. Vijayvergiya, A. Tyra, A. Nair, A. Nayak, B. Eggers, B. Ji, B. Hoover, B. Chen, B. Chen, B. Barak, B. Minaiev, B. Hao, B. Baker, B. Lightcap, B. McKinzie, B. Wang, B. Quinn, B. Fioca, B. Hsu, B. Yang, B. Yu, B. Zhang, B. Brenner, C. R. Zetino, C. Raymond, C. Lugaresi, C. Paz, C. Hudson, C. Whitney, C. Li, C. Chen, C. Cole, C. Voss, C. Ding, C. Shen, C. Huang, C. Colby, C. Hallacy, C. Koch, C. Lu, C. Kaplan, C. Kim, C. Minott-Henriques, C. Frey, C. Yu, C. Czarnecki, C. Reid, C. Wei, C. Decareaux, C. Scheau, C. Zhang, C. Forbes, D. Tang, D. Goldberg, D. Roberts, D. Palmie, D. Kappler, D. Levine, D. Wright, D. Leo, D. Lin, D. Robinson, D. Grabb, D. Chen, D. Lim, D. Salama, D. Bhattacharjee, D. Tsipras, D. Li, D. Yu, D. Strouse, D. Williams, D. Hunn, E. Bayes, E. Arbus, E. Akyurek, E. Y. Le, E. Widmann, E. Yani, E. Proehl, E. Sert, E. Cheung, E. Schwartz, E. Han, E. Jiang, E. Mitchell, E. Sigler, E. Wallace, E. Ritter, E. Kavanaugh, E. Mays, E. Nikishin, F. Li, F. P. Such, F. de Avila Belbute Peres, F. Raso, F. Bekerman, F. Tsimpourlas, F. Chantzis, F. Song, F. Zhang, G. Raila, G. McGrath, G. Briggs, G. Yang, G. Parascandolo, G. Chabot, G. Kim, G. Zhao, G. Valiant, G. Leclerc, H. Salman, H. Wang, H. Sheng, H. Jiang, H. Wang, H. Jin, H. Sikchi, H. Schmidt, H. Aspegren, H. Chen, H. Qiu, H. Lightman, I. Covert, I. Kivlichan, I. Silber, I. Sohl, I. Hammoud, I. Clavera, I. Lan, I. Akkaya, I. Kostrikov, I. Kofman, I. Etinger, I. Singal, J. Hehir, J. Huh, J. Pan, J. Wilczynski, J. Pachocki, J. Lee, J. Quinn, J. Kiros, J. Kalra, J. Samaroo, J. Wang, J. Wolfe, J. Chen, J. Wang, J. Harb, J. Han, J. Wang, J. Zhao, J. Chen, J. Yang, J. Tworek, J. Chand, J. Landon, J. Liang, J. Lin, J. Liu, J. Wang, J. Tang, J. Yin, J. Jang, J. Morris, J. Flynn, J. Ferstad, J. Heidecke, J. Fishbein, J. Hallman, J. Grant, J. Chien, J. Gordon, J. Park, J. Liss, J. Kraaijeveld, J. Guay, J. Mo, J. Lawson, J. McGrath, J. Vendrow, J. Jiao, J. Lee, J. Steele, J. Wang, J. Mao, K. Chen, K. Hayashi, K. Xiao, K. Salahi, K. Wu, K. Sekhri, K. Sharma, K. Singhal, K. Li, K. Nguyen, K. Gu-Lemberg, K. King, K. Liu, K. Stone, K. Yu, K. Ying, K. Georgiev, K. Lim, K. Tirumala, K. Miller, L. Ahmad, L. Lv, L. Clare, L. Fauconnet, L. Itow, L. Yang, L. Romaniuk, L. Anise, L. Byron, L. Pathak, L. Maksin, L. Lo, L. Ho, L. Jing, L. Wu, L. Xiong, L. Mamitsuka, L. Yang, L. McCallum, L. Held, L. Bourgeois, L. Engstrom, L. Kuhn, L. Feuvrier, L. Zhang, L. Switzer, L. Kondraciuk, L. Kaiser, M. Joglekar, M. Singh, M. Shah, M. Stratta, M. Williams, M. Chen, M. Sun, M. Cayton, M. Li, M. Zhang, M. Aljubeh, M. Nichols, M. Haines, M. Schwarzer, M. Gupta, M. Shah, M. Y. Guan, M. Huang, M. Dong, M. Wang, M. Glaese, M. Carroll, M. Lampe, M. Malek, M. Sharman, M. Zhang, M. Wang, M. Pokrass, M. Florian, M. Pavlov, M. Wang, M. Chen, M. Wang, M. Feng, M. Bavarian, M. Lin, M. Abdool, M. Rohaninejad, N. Soto, N. Staudacher, N. LaFontaine, N. Marwell, N. Liu, N. Preston, N. Turley, N. Ansman, N. Blades, N. Pancha, N. Mikhaylin, N. Felix, N. Handa, N. Rai, N. Keskar, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, O. Gleeson, P. Mishkin, P. Lesiewicz, P. Baltescu, P. Belov, P. Zhokhov, P. Pronin, P. Guo, P. Thacker, Q. Liu, Q. Yuan, Q. Liu, R. Dias, R. Puckett, R. Arora, R. T. Mullapudi, R. Gaon, R. Miyara, R. Song, R. Aggarwal, R. Marsan, R. Yemiru, R. Xiong, R. Kshirsagar, R. Nuttall, R. Tsiupa, R. Eldan, R. Wang, R. James, R. Ziv, R. Shu, R. Nigmatullin, S. Jain, S. Talaie, S. Altman, S. Arnesen, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Yoo, S. Heon, S. Ethersmith, S. Grove, S. Taylor, S. Bubeck, S. Banesiu, S. Amdo, S. Zhao, S. Wu, S. Santurkar, S. Zhao, S. R. Chaudhuri, S. Krishnaswamy, Shuaiqi, Xia, S. Cheng, S. Anadkat, S. P. Fishman, S. Tobin, S. Fu, S. Jain, S. Mei, S. Egoian, S. Kim, S. Golden, S. Mah, S. Lin, S. Imm, S. Sharpe, S. Yadlowsky, S. Choudhry, S. Eum, S. Sanjeev, T. Khan, T. Stramer, T. Wang, T. Xin, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Degry, T. Shadwell, T. Fu, T. Gao, T. Garipov, T. Sriskandarajah, T. Sherbakov, T. Korbak, T. Kaftan, T. Hiratsuka, T. Wang, T. Song, T. Zhao, T. Peterson, V. Kharitonov, V. Chernova, V. Kosaraju, V. Kuo, V. Pong, V. Verma, V. Petrov, W. Jiang, W. Zhang, W. Zhou, W. Xie, W. Zhan, W. McCabe, W. DePue, W. Ellsworth, W. Bain, W. Thompson, X. Chen, X. Qi, X. Xiang, X. Shi, Y. Dubois, Y. Yu, Y. Khakbaz, Y. Wu, Y. Qian, Y. T. Lee, Y. Chen, Y. Zhang, Y. Xiong, Y. Tian, Y. Cha, Y. Bai, Y. Yang, Y. Yuan, Y. Li, Y. Zhang, Y. Yang, Y. Jin, Y. Jiang, Y. Wang, Y. Wang, Y. Liu, Z. Stubenvoll, Z. Dou, Z. Wu, and Z. Wang OpenAI gpt-5 system card. External Links: 2601.03267, Link Cited by: §2. Tang et al. (2026) F. Tang, Z. Lu, B. Zhang, W. Lu, J. Xiao, Y. Zhuang, and Y. Shen ClawGUI: a unified framework for training, evaluating, and deploying gui agents. External Links: 2604.11784, Link Cited by: §5.2, Table 2. Team et al. (2025) G. Team, R. Anil, S. Borgeaud, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Silver, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. Barham, T. Hennigan, B. Lee, F. Viola, M. Reynolds, Y. Xu, R. Doherty, E. Collins, C. Meyer, E. Rutherford, E. Moreira, K. Ayoub, M. Goel, J. Krawczyk, C. Du, E. Chi, H. Cheng, E. Ni, P. Shah, P. Kane, B. Chan, M. Faruqui, A. Severyn, H. Lin, Y. Li, Y. Cheng, A. Ittycheriah, M. Mahdieh, M. Chen, P. Sun, D. Tran, S. Bagri, B. Lakshminarayanan, J. Liu, A. Orban, F. Güra, H. Zhou, X. Song, A. Boffy, H. Ganapathy, S. Zheng, H. Choe, Á. Weisz, T. Zhu, Y. Lu, S. Gopal, J. Kahn, M. Kula, J. Pitman, R. Shah, E. Taropa, M. A. Merey, M. Baeuml, Z. Chen, L. E. Shafey, Y. Zhang, O. Sercinoglu, G. Tucker, E. Piqueras, M. Krikun, I. Barr, N. Savinov, I. Danihelka, B. Roelofs, A. White, A. Andreassen, T. von Glehn, L. Yagati, M. Kazemi, L. Gonzalez, M. Khalman, J. Sygnowski, A. Frechette, C. Smith, L. Culp, L. Proleev, Y. Luan, X. Chen, J. Lottes, N. Schucher, F. Lebron, A. Rrustemi, N. Clay, P. Crone, T. Kocisky, J. Zhao, B. Perz, D. Yu, H. Howard, A. Bloniarz, J. W. Rae, H. Lu, L. Sifre, M. Maggioni, F. Alcober, D. Garrette, M. Barnes, S. Thakoor, J. Austin, G. Barth-Maron, W. Wong, R. Joshi, R. Chaabouni, D. Fatiha, A. Ahuja, G. S. Tomar, E. Senter, M. Chadwick, I. Kornakov, N. Attaluri, I. Iturrate, R. Liu, Y. Li, S. Cogan, J. Chen, C. Jia, C. Gu, Q. Zhang, J. Grimstad, A. J. Hartman, X. Garcia, T. S. Pillai, J. Devlin, M. Laskin, D. de Las Casas, D. Valter, C. Tao, L. Blanco, A. P. Badia, D. Reitter, M. Chen, J. Brennan, C. Rivera, S. Brin, S. Iqbal, G. Surita, J. Labanowski, A. Rao, S. Winkler, E. Parisotto, Y. Gu, K. Olszewska, R. Addanki, A. Miech, A. Louis, D. Teplyashin, G. Brown, E. Catt, J. Balaguer, J. Xiang, P. Wang, Z. Ashwood, A. Briukhov, A. Webson, S. Ganapathy, S. Sanghavi, A. Kannan, M. Chang, A. Stjerngren, J. Djolonga, Y. Sun, A. Bapna, M. Aitchison, P. Pejman, H. Michalewski, T. Yu, C. Wang, J. Love, J. Ahn, D. Bloxwich, K. Han, P. Humphreys, T. Sellam, J. Bradbury, V. Godbole, S. Samangooei, B. Damoc, A. Kaskasoli, S. M. R. Arnold, V. Vasudevan, S. Agrawal, J. Riesa, D. Lepikhin, R. Tanburn, S. Srinivasan, H. Lim, S. Hodkinson, P. Shyam, J. Ferret, S. Hand, A. Garg, T. L. Paine, J. Li, Y. Li, M. Giang, A. Neitz, Z. Abbas, S. York, M. Reid, E. Cole, A. Chowdhery, D. Das, D. Rogozińska, V. Nikolaev, P. Sprechmann, Z. Nado, L. Zilka, F. Prost, L. He, M. Monteiro, G. Mishra, C. Welty, J. Newlan, D. Jia, M. Allamanis, C. H. Hu, R. de Liedekerke, J. Gilmer, C. Saroufim, S. Rijhwani, S. Hou, D. Shrivastava, A. Baddepudi, A. Goldin, A. Ozturel, A. Cassirer, Y. Xu, D. Sohn, D. Sachan, R. K. Amplayo, C. Swanson, D. Petrova, S. Narayan, A. Guez, S. Brahma, J. Landon, M. Patel, R. Zhao, K. Villela, L. Wang, W. Jia, M. Rahtz, M. Giménez, L. Yeung, J. Keeling, P. Georgiev, D. Mincu, B. Wu, S. Haykal, R. Saputro, K. Vodrahalli, J. Qin, Z. Cankara, A. Sharma, N. Fernando, W. Hawkins, B. Neyshabur, S. Kim, A. Hutter, P. Agrawal, A. Castro-Ros, G. van den Driessche, T. Wang, F. Yang, S. Chang, P. Komarek, R. McIlroy, M. Lučić, G. Zhang, W. Farhan, M. Sharman, P. Natsev, P. Michel, Y. Bansal, S. Qiao, K. Cao, S. Shakeri, C. Butterfield, J. Chung, P. K. Rubenstein, S. Agrawal, A. Mensch, K. Soparkar, K. Lenc, T. Chung, A. Pope, L. Maggiore, J. Kay, P. Jhakra, S. Wang, J. Maynez, M. Phuong, T. Tobin, A. Tacchetti, M. Trebacz, K. Robinson, Y. Katariya, S. Riedel, P. Bailey, K. Xiao, N. Ghelani, L. Aroyo, A. Slone, N. Houlsby, X. Xiong, Z. Yang, E. Gribovskaya, J. Adler, M. Wirth, L. Lee, M. Li, T. Kagohara, J. Pavagadhi, S. Bridgers, A. Bortsova, S. Ghemawat, Z. Ahmed, T. Liu, R. Powell, V. Bolina, M. Iinuma, P. Zablotskaia, J. Besley, D. Chung, T. Dozat, R. Comanescu, X. Si, J. Greer, G. Su, M. Polacek, R. L. Kaufman, S. Tokumine, H. Hu, E. Buchatskaya, Y. Miao, M. Elhawaty, A. Siddhant, N. Tomasev, J. Xing, C. Greer, H. Miller, S. Ashraf, A. Roy, Z. Zhang, A. Ma, A. Filos, M. Besta, R. Blevins, T. Klimenko, C. Yeh, S. Changpinyo, J. Mu, O. Chang, M. Pajarskas, C. Muir, V. Cohen, C. L. Lan, K. Haridasan, A. Marathe, S. Hansen, S. Douglas, R. Samuel, M. Wang, S. Austin, C. Lan, J. Jiang, J. Chiu, J. A. Lorenzo, L. L. Sjösund, S. Cevey, Z. Gleicher, T. Avrahami, A. Boral, H. Srinivasan, V. Selo, R. May, K. Aisopos, L. Hussenot, L. B. Soares, K. Baumli, M. B. Chang, A. Recasens, B. Caine, A. Pritzel, F. Pavetic, F. Pardo, A. Gergely, J. Frye, V. Ramasesh, D. Horgan, K. Badola, N. Kassner, S. Roy, E. Dyer, V. C. Campos, A. Tomala, Y. Tang, D. E. Badawy, E. White, B. Mustafa, O. Lang, A. Jindal, S. Vikram, Z. Gong, S. Caelles, R. Hemsley, G. Thornton, F. Feng, W. Stokowiec, C. Zheng, P. Thacker, Ç. Ünlü, Z. Zhang, M. Saleh, J. Svensson, M. Bileschi, P. Patil, A. Anand, R. Ring, K. Tsihlas, A. Vezer, M. Selvi, T. Shevlane, M. Rodriguez, T. Kwiatkowski, S. Daruki, K. Rong, A. Dafoe, N. FitzGerald, K. Gu-Lemberg, M. Khan, L. A. Hendricks, M. Pellat, V. Feinberg, J. Cobon-Kerr, T. Sainath, M. Rauh, S. H. Hashemi, R. Ives, Y. Hasson, E. Noland, Y. Cao, N. Byrd, L. Hou, Q. Wang, T. Sottiaux, M. Paganini, J. Lespiau, A. Moufarek, S. Hassan, K. Shivakumar, J. van Amersfoort, A. Mandhane, P. Joshi, A. Goyal, M. Tung, A. Brock, H. Sheahan, V. Misra, C. Li, N. Rakićević, M. Dehghani, F. Liu, S. Mittal, J. Oh, S. Noury, E. Sezener, F. Huot, M. Lamm, N. D. Cao, C. Chen, S. Mudgal, R. Stella, K. Brooks, G. Vasudevan, C. Liu, M. Chain, N. Melinkeri, A. Cohen, V. Wang, K. Seymore, S. Zubkov, R. Goel, S. Yue, S. Krishnakumaran, B. Albert, N. Hurley, M. Sano, A. Mohananey, J. Joughin, E. Filonov, T. Kępa, Y. Eldawy, J. Lim, R. Rishi, S. Badiezadegan, T. Bos, J. Chang, S. Jain, S. G. S. Padmanabhan, S. Puttagunta, K. Krishna, L. Baker, N. Kalb, V. Bedapudi, A. Kurzrok, S. Lei, A. Yu, O. Litvin, X. Zhou, Z. Wu, S. Sobell, A. Siciliano, A. Papir, R. Neale, J. Bragagnolo, T. Toor, T. Chen, V. Anklin, F. Wang, R. Feng, M. Gholami, K. Ling, L. Liu, J. Walter, H. Moghaddam, A. Kishore, J. Adamek, T. Mercado, J. Mallinson, S. Wandekar, S. Cagle, E. Ofek, G. Garrido, C. Lombriser, M. Mukha, B. Sun, H. R. Mohammad, J. Matak, Y. Qian, V. Peswani, P. Janus, Q. Yuan, L. Schelin, O. David, A. Garg, Y. He, O. Duzhyi, A. Älgmyr, T. Lottaz, Q. Li, V. Yadav, L. Xu, A. Chinien, R. Shivanna, A. Chuklin, J. Li, C. Spadine, T. Wolfe, K. Mohamed, S. Das, Z. Dai, K. He, D. von Dincklage, S. Upadhyay, A. Maurya, L. Chi, S. Krause, K. Salama, P. G. Rabinovitch, P. K. R. M, A. Selvan, M. Dektiarev, G. Ghiasi, E. Guven, H. Gupta, B. Liu, D. Sharma, I. H. Shtacher, S. Paul, O. Akerlund, F. Aubet, T. Huang, C. Zhu, E. Zhu, E. Teixeira, M. Fritze, F. Bertolini, L. Marinescu, M. Bölle, D. Paulus, K. Gupta, T. Latkar, M. Chang, J. Sanders, R. Wilson, X. Wu, Y. Tan, L. N. Thiet, T. Doshi, S. Lall, S. Mishra, W. Chen, T. Luong, S. Benjamin, J. Lee, E. Andrejczuk, D. Rabiej, V. Ranjan, K. Styrc, P. Yin, J. Simon, M. R. Harriott, M. Bansal, A. Robsky, G. Bacon, D. Greene, D. Mirylenka, C. Zhou, O. Sarvana, A. Goyal, S. Andermatt, P. Siegler, B. Horn, A. Israel, F. Pongetti, C. ". Chen, M. Selvatici, P. Silva, K. Wang, J. Tolins, K. Guu, R. Yogev, X. Cai, A. Agostini, M. Shah, H. Nguyen, N. Ó. Donnaile, S. Pereira, L. Friso, A. Stambler, A. Kurzrok, C. Kuang, Y. Romanikhin, M. Geller, Z. Yan, K. Jang, C. Lee, W. Fica, E. Malmi, Q. Tan, D. Banica, D. Balle, R. Pham, Y. Huang, D. Avram, H. Shi, J. Singh, C. Hidey, N. Ahuja, P. Saxena, D. Dooley, S. P. Potharaju, E. O’Neill, A. Gokulchandran, R. Foley, K. Zhao, M. Dusenberry, Y. Liu, P. Mehta, R. Kotikalapudi, C. Safranek-Shrader, A. Goodman, J. Kessinger, E. Globen, P. Kolhar, C. Gorgolewski, A. Ibrahim, Y. Song, A. Eichenbaum, T. Brovelli, S. Potluri, P. Lahoti, C. Baetu, A. Ghorbani, C. Chen, A. Crawford, S. Pal, M. Sridhar, P. Gurita, A. Mujika, I. Petrovski, P. Cedoz, C. Li, S. Chen, N. D. Santo, S. Goyal, J. Punjabi, K. Kappaganthu, C. Kwak, P. LV, S. Velury, H. Choudhury, J. Hall, P. Shah, R. Figueira, M. Thomas, M. Lu, T. Zhou, C. Kumar, T. Jurdi, S. Chikkerur, Y. Ma, A. Yu, S. Kwak, V. Ähdel, S. Rajayogam, T. Choma, F. Liu, A. Barua, C. Ji, J. H. Park, V. Hellendoorn, A. Bailey, T. Bilal, H. Zhou, M. Khatir, C. Sutton, W. Rzadkowski, F. Macintosh, R. Vij, K. Shagin, P. Medina, C. Liang, J. Zhou, P. Shah, Y. Bi, A. Dankovics, S. Banga, S. Lehmann, M. Bredesen, Z. Lin, J. E. Hoffmann, J. Lai, R. Chung, K. Yang, N. Balani, A. Bražinskas, A. Sozanschi, M. Hayes, H. F. Alcalde, P. Makarov, W. Chen, A. Stella, L. Snijders, M. Mandl, A. Kärrman, P. Nowak, X. Wu, A. Dyck, K. Vaidyanathan, R. R, J. Mallet, M. Rudominer, E. Johnston, S. Mittal, A. Udathu, J. Christensen, V. Verma, Z. Irving, A. Santucci, G. Elsayed, E. Davoodi, M. Georgiev, I. Tenney, N. Hua, G. Cideron, E. Leurent, M. Alnahlawi, I. Georgescu, N. Wei, I. Zheng, D. Scandinaro, H. Jiang, J. Snoek, M. Sundararajan, X. Wang, Z. Ontiveros, I. Karo, J. Cole, V. Rajashekhar, L. Tumeh, E. Ben-David, R. Jain, J. Uesato, R. Datta, O. Bunyan, S. Wu, J. Zhang, P. Stanczyk, Y. Zhang, D. Steiner, S. Naskar, M. Azzam, M. Johnson, A. Paszke, C. Chiu, J. S. Elias, A. Mohiuddin, F. Muhammad, J. Miao, A. Lee, N. Vieillard, J. Park, J. Zhang, J. Stanway, D. Garmon, A. Karmarkar, Z. Dong, J. Lee, A. Kumar, L. Zhou, J. Evens, W. Isaac, G. Irving, E. Loper, M. Fink, I. Arkatkar, N. Chen, I. Shafran, I. Petrychenko, Z. Chen, J. Jia, A. Levskaya, Z. Zhu, P. Grabowski, Y. Mao, A. Magni, K. Yao, J. Snaider, N. Casagrande, E. Palmer, P. Suganthan, A. Castaño, I. Giannoumis, W. Kim, M. Rybiński, A. Sreevatsa, J. Prendki, D. Soergel, A. Goedeckemeyer, W. Gierke, M. Jafari, M. Gaba, J. Wiesner, D. G. Wright, Y. Wei, H. Vashisht, Y. Kulizhskaya, J. Hoover, M. Le, L. Li, C. Iwuanyanwu, L. Liu, K. Ramirez, A. Khorlin, A. Cui, T. LIN, M. Wu, R. Aguilar, K. Pallo, A. Chakladar, G. Perng, E. A. Abellan, M. Zhang, I. Dasgupta, N. Kushman, I. Penchev, A. Repina, X. Wu, T. van der Weide, P. Ponnapalli, C. Kaplan, J. Simsa, S. Li, O. Dousse, F. Yang, J. Piper, N. Ie, R. Pasumarthi, N. Lintz, A. Vijayakumar, D. Andor, P. Valenzuela, M. Lui, C. Paduraru, D. Peng, K. Lee, S. Zhang, S. Greene, D. D. Nguyen, P. Kurylowicz, C. Hardin, L. Dixon, L. Janzer, K. Choo, Z. Feng, B. Zhang, A. Singhal, D. Du, D. McKinnon, N. Antropova, T. Bolukbasi, O. Keller, D. Reid, D. Finchelstein, M. A. Raad, R. Crocker, P. Hawkins, R. Dadashi, C. Gaffney, K. Franko, A. Bulanova, R. Leblond, S. Chung, H. Askham, L. C. Cobo, K. Xu, F. Fischer, J. Xu, C. Sorokin, C. Alberti, C. Lin, C. Evans, A. Dimitriev, H. Forbes, D. Banarse, Z. Tung, M. Omernick, C. Bishop, R. Sterneck, R. Jain, J. Xia, E. Amid, F. Piccinno, X. Wang, P. Banzal, D. J. Mankowitz, A. Polozov, V. Krakovna, S. Brown, M. Bateni, D. Duan, V. Firoiu, M. Thotakuri, T. Natan, M. Geist, S. tan Girgin, H. Li, J. Ye, O. Roval, R. Tojo, M. Kwong, J. Lee-Thorp, C. Yew, D. Sinopalnikov, S. Ramos, J. Mellor, A. Sharma, K. Wu, D. Miller, N. Sonnerat, D. Vnukov, R. Greig, J. Beattie, E. Caveness, L. Bai, J. Eisenschlos, A. Korchemniy, T. Tsai, M. Jasarevic, W. Kong, P. Dao, Z. Zheng, F. Liu, F. Yang, R. Zhu, T. H. Teh, J. Sanmiya, E. Gladchenko, N. Trdin, D. Toyama, E. Rosen, S. Tavakkol, L. Xue, C. Elkind, O. Woodman, J. Carpenter, G. Papamakarios, R. Kemp, S. Kafle, T. Grunina, R. Sinha, A. Talbert, D. Wu, D. Owusu-Afriyie, C. Du, C. Thornton, J. Pont-Tuset, P. Narayana, J. Li, S. Fatehi, J. Wieting, O. Ajmeri, B. Uria, Y. Ko, L. Knight, A. Héliou, N. Niu, S. Gu, C. Pang, Y. Li, N. Levine, A. Stolovich, R. Santamaria-Fernandez, S. Goenka, W. Yustalim, R. Strudel, A. Elqursh, C. Deck, H. Lee, Z. Li, K. Levin, R. Hoffmann, D. Holtmann-Rice, O. Bachem, S. Arora, C. Koh, S. H. Yeganeh, S. Põder, M. Tariq, Y. Sun, L. Ionita, M. Seyedhosseini, P. Tafti, Z. Liu, A. Gulati, J. Liu, X. Ye, B. Chrzaszcz, L. Wang, N. Sethi, T. Li, B. Brown, S. Singh, W. Fan, A. Parisi, J. Stanton, V. Koverkathu, C. A. Choquette-Choo, Y. Li, T. Lu, A. Ittycheriah, P. Shroff, M. Varadarajan, S. Bahargam, R. Willoughby, D. Gaddy, G. Desjardins, M. Cornero, B. Robenek, B. Mittal, B. Albrecht, A. Shenoy, F. Moiseev, H. Jacobsson, A. Ghaffarkhah, M. Rivière, A. Walton, C. Crepy, A. Parrish, Z. Zhou, C. Farabet, C. Radebaugh, P. Srinivasan, C. van der Salm, A. Fidjeland, S. Scellato, E. Latorre-Chimoto, H. Klimczak-Plucińska, D. Bridson, D. de Cesare, T. Hudson, P. Mendolicchio, L. Walker, A. Morris, M. Mauger, A. Guseynov, A. Reid, S. Odoom, L. Loher, V. Cotruta, M. Yenugula, D. Grewe, A. Petrushkina, T. Duerig, A. Sanchez, S. Yadlowsky, A. Shen, A. Globerson, L. Webb, S. Dua, D. Li, S. Bhupatiraju, D. Hurt, H. Qureshi, A. Agarwal, T. Shani, M. Eyal, A. Khare, S. R. Belle, L. Wang, C. Tekur, M. S. Kale, J. Wei, R. Sang, B. Saeta, T. Liechty, Y. Sun, Y. Zhao, S. Lee, P. Nayak, D. Fritz, M. R. Vuyyuru, J. Aslanides, N. Vyas, M. Wicke, X. Ma, E. Eltyshev, N. Martin, H. Cate, J. Manyika, K. Amiri, Y. Kim, X. Xiong, K. Kang, F. Luisier, N. Tripuraneni, D. Madras, M. Guo, A. Waters, O. Wang, J. Ainslie, J. Baldridge, H. Zhang, G. Pruthi, J. Bauer, F. Yang, R. Mansour, J. Gelman, Y. Xu, G. Polovets, J. Liu, H. Cai, W. Chen, X. Sheng, E. Xue, S. Ozair, C. Angermueller, X. Li, A. Sinha, W. Wang, J. Wiesinger, E. Koukoumidis, Y. Tian, A. Iyer, M. Gurumurthy, M. Goldenson, P. Shah, M. Blake, H. Yu, A. Urbanowicz, J. Palomaki, C. Fernando, K. Durden, H. Mehta, N. Momchev, E. Rahimtoroghi, M. Georgaki, A. Raul, S. Ruder, M. Redshaw, J. Lee, D. Zhou, K. Jalan, D. Li, B. Hechtman, P. Schuh, M. Nasr, K. Milan, V. Mikulik, J. Franco, T. Green, N. Nguyen, J. Kelley, A. Mahendru, A. Hu, J. Howland, B. Vargas, J. Hui, K. Bansal, V. Rao, R. Ghiya, E. Wang, K. Ye, J. M. Sarr, M. M. Preston, M. Elish, S. Li, A. Kaku, J. Gupta, I. Pasupat, D. Juan, M. Someswar, T. M., X. Chen, A. Amini, A. Fabrikant, E. Chu, X. Dong, A. Muthal, S. Buthpitiya, S. Jauhari, N. Hua, U. Khandelwal, A. Hitron, J. Ren, L. Rinaldi, S. Drath, A. Dabush, N. Jiang, H. Godhia, U. Sachs, A. Chen, Y. Fan, H. Taitelbaum, H. Noga, Z. Dai, J. Wang, C. Liang, J. Hamer, C. Ferng, C. Elkind, A. Atias, P. Lee, V. Listík, M. Carlen, J. van de Kerkhof, M. Pikus, K. Zaher, P. Müller, S. Zykova, R. Stefanec, V. Gatsko, C. Hirnschall, A. Sethi, X. F. Xu, C. Ahuja, B. Tsai, A. Stefanoiu, B. Feng, K. Dhandhania, M. Katyal, A. Gupta, A. Parulekar, D. Pitta, J. Zhao, V. Bhatia, Y. Bhavnani, O. Alhadlaq, X. Li, P. Danenberg, D. Tu, A. Pine, V. Filippova, A. Ghosh, B. Limonchik, B. Urala, C. K. Lanka, D. Clive, Y. Sun, E. Li, H. Wu, K. Hongtongsak, I. Li, K. Thakkar, K. Omarov, K. Majmundar, M. Alverson, M. Kucharski, M. Patel, M. Jain, M. Zabelin, P. Pelagatti, R. Kohli, S. Kumar, J. Kim, S. Sankar, V. Shah, L. Ramachandruni, X. Zeng, B. Bariach, L. Weidinger, T. Vu, A. Andreev, A. He, K. Hui, S. Kashem, A. Subramanya, S. Hsiao, D. Hassabis, K. Kavukcuoglu, A. Sadovsky, Q. Le, T. Strohman, Y. Wu, S. Petrov, J. Dean, and O. Vinyals Gemini: a family of highly capable multimodal models. External Links: 2312.11805, Link Cited by: §5.1. Team (2026) Q. Team Qwen3.5: accelerating productivity with native multimodal agents. External Links: Link Cited by: §2, §5.1. Wang et al. (2024) S. Wang, W. Liu, J. Chen, Y. Zhou, W. Gan, X. Zeng, Y. Che, S. Yu, X. Hao, K. Shao, et al. Gui agents with foundation models: a comprehensive survey. arXiv preprint arXiv:2411.04890. Cited by: §1. Wang et al. (2025a) T. Wang, Z. Wu, J. Liu, J. Hao, J. Wang, and K. Shao DistRL: an asynchronous distributed reinforcement learning framework for on-device control agents. External Links: 2410.14803, Link Cited by: Table 10, §1, §2, §5.1. Wang et al. (2025b) X. Wang, Z. Wu, J. Xie, Z. Ding, B. Yang, Z. Li, Z. Liu, Q. Li, X. Dong, Z. Chen, W. Wang, X. Zhao, J. Chen, H. Duan, T. Xie, S. Su, C. Yang, Y. Yu, Y. Huang, Y. Liu, X. Zhang, X. Yue, W. Su, X. Zhu, W. Shen, J. Dai, and W. Wang MMBench-gui: hierarchical multi-platform evaluation framework for gui agents. arXiv preprint arXiv:2507.19478. Cited by: §5.1. Xia et al. (2025) Y. Xia, J. Fan, W. Chen, S. Yan, X. Cong, Z. Zhang, Y. Lu, Y. Lin, Z. Liu, and M. Sun AgentRM: enhancing agent generalization with reward modeling. External Links: 2502.18407, Link Cited by: §1. Xie et al. (2026) L. Xie, S. Huang, Z. Zhang, A. Zou, Y. Zhai, D. Ren, K. Zhang, H. Hu, B. Liu, H. Chen, Z. Liu, and B. Ding Auto-rubric: learning from implicit weights to explicit rubrics for reward modeling. External Links: 2510.17314, Link Cited by: §1. Xie et al. (2024) T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V. Zhong, and T. Yu OSWorld: benchmarking multimodal agents for open-ended tasks in real computer environments. External Links: 2404.07972, Link Cited by: §1, §2, §5.1. Xiong et al. (2025) T. Xiong, X. Hu, Y. Chen, Y. Liu, C. Wu, P. Gao, W. Liu, J. Luan, and S. Zhang GUI-pra: process reward agent for gui tasks. External Links: 2509.23263, Link Cited by: §1. Xu et al. (2026) H. Xu, X. Zhang, H. Liu, J. Wang, Z. Zhu, S. Zhou, X. Hu, F. Gao, J. Cao, Z. Wang, Z. Chen, J. Liao, Q. Zheng, J. Zeng, Z. Xu, S. Bai, J. Lin, J. Zhou, and M. Yan Mobile-agent-v3.5: multi-platform fundamental gui agents. External Links: 2602.16855, Link Cited by: §1. Xu et al. (2025) Y. Xu, X. Liu, X. Liu, J. Fu, H. Zhang, B. Jing, S. Zhang, Y. Wang, W. Zhao, and Y. Dong MobileRL: online agentic reinforcement learning for mobile gui agents. External Links: 2509.18119, Link Cited by: §1. Yang et al. (2025) C. Yang, S. Su, S. Liu, X. Dong, Y. Yu, W. Su, X. Wang, Z. Liu, J. Zhu, H. Li, W. Wang, Y. Qiao, X. Zhu, and J. Dai ZeroGUI: automating online gui learning at zero human cost. External Links: 2505.23762, Link Cited by: Table 10, §1, §2, §5.1. Yang et al. (2023) J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. External Links: 2310.11441, Link Cited by: §2. Zhang et al. (2025) S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu, and G. Wang Instruction tuning for large language models: a survey. External Links: 2308.10792, Link Cited by: §1. Zheng et al. (2026) C. Zheng, X. Mo, X. Ma, Q. Lin, Y. Zhao, J. Zhu, X. Lou, J. Wang, Z. Wang, W. Liu, Z. Zhang, Y. Yu, and W. Zhang Adaptive milestone reward for gui agents. External Links: 2602.11524, Link Cited by: §2. Zhou et al. (2025) H. Zhou, X. Zhang, P. Tong, J. Zhang, L. Chen, Q. Kong, C. Cai, C. Liu, Y. Wang, J. Zhou, and S. Hoi MAI-ui technical report: real-world centric foundation gui agents. External Links: 2512.22047, Link Cited by: §5.2. Zhou et al. (2023) S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, Y. Bisk, D. Fried, U. Alon, et al. WebArena: a realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854. Cited by: §5.1. Appendix A Taxonomy and Rubric Bank Table 6 lists the eight task families in our GUI taxonomy C. For each family, we show its identifier, a brief description, and the key verification dimensions encoded in its static rubric RcR_c. Family Description Key Dimensions info_query Answer from screen Format; content; evidence create_modify Create or edit Explicit properties; completion; consistency delete_cleanup Remove or clear Scope; target; completeness communication Send to recipient Source; recipient; channel; confirmation transfer Move or copy Source; action; destination; fidelity state_navigation Toggle or navigate Final state; toggle; persistence composite_workflow Multiple goals Per-family rubric; all subgoals succeed general Fallback Task logic; action; final state Table 6: GUI task taxonomy used for coarse rubric retrieval. Appendix B Coarse Rubric Bank Construction We construct the coarse rubric bank before evaluation and keep it fixed for all experiments. To build the development pool, we run Qwen3.5-122B-A10B on AndroidWorld and MobileWorld and collect complete trajectories with binary success labels from the environment evaluators. The pool contains 116 trajectories from each benchmark: AndroidWorld has 55 successful and 61 failed trajectories, while MobileWorld has 35 successful and 81 failed trajectories. MobileWorld originally contains 117 tasks in this collection, but one task fails during environment execution and therefore has no usable trajectory. We use Claude Opus 4.6 as the offline rubric-construction model to summarize recurring success criteria and failure patterns, group them into GUI task families, and refine the category rubrics using disagreement cases. The resulting bank contains the eight entries shown in Table 6. Appendix C Online RL Training Details Our online reinforcement learning experiments use ClawGUI-RL with GRPO on MobileWorld. Table 7 summarizes the main training hyperparameters. All compared reward agents use the same policy backbone and training protocol; only the reward verifier is changed. Item Value Training framework ClawGUI-RL Environment MobileWorld Policy backbone MAI-UI-8B Optimization algorithm GRPO Training epochs 2 Number of GPUs 8 Training batch size 8 Rollouts per task 4 History length 3 Maximum episode steps 50 Learning rate 1×10−61×10^-6 KL coefficient 0.01 Rollout temperature 0.7 Validation temperature 0.4 Maximum prompt length 28,000 Maximum response length 512 AdaptRubricverifier backbone Qwen3-VL-8B-Instruct Maximum screenshots for reward verification 10 Table 7: Key hyperparameters for online RL training. Appendix D Image Budget Analysis Trajectory context is an important input factor for GUI reward verification. The OS-Themis evaluation protocol runs ZeroGUI with the final two screenshots, while our main experiments use at most ten trajectory screenshots. To align the image budget, we also evaluate ZeroGUI with its final ten screenshots and use the ten-screenshot setting in subsequent experiments. Table 8 reports both settings. Increasing the image budget improves ZeroGUI mainly by raising recall, which makes our main comparison more conservative. Setting Acc Prec Rec F1 Last 2 screenshots 79.1 84.9 70.7 76.8 Last 10 screenshots 85.0 84.9 84.9 84.9 Table 8: Effect of image budget on ZeroGUI. We report overall performance on OGRBench averaged across eight judge backbones. Table 9 further reports the additional baselines under their original trajectory-context settings. Most of these evaluators were designed around a final screenshot or a small final-state context, while ZeroGUI follows the two-screenshot setting used in prior evaluation. Comparing Table 9 with Table 11 shows that increasing the trajectory context substantially improves several baselines, which motivates the matched-budget setting used in the expanded comparison shown in Figure 3. Model Ubuntu Mobile Windows macOS Web Overall Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc Prec Rec F1 DigiRL Qwen3-VL-4B-Instruct 70.7 71.0 78.7 81.3 76.5 72.5 87.0 70.6 76.3 79.6 74.3 74.1 74.1 74.1 Qwen3-VL-8B-Instruct 73.3 73.8 73.4 76.9 77.9 73.7 88.3 71.0 74.2 77.8 74.9 74.7 75.0 74.8 Qwen3-VL-32B-Instruct 77.9 78.1 80.9 82.5 79.8 74.9 90.9 75.9 81.1 82.9 79.7 81.1 77.1 79.1 Qwen3-VL-235B-A22B-Instruct 77.1 76.5 78.7 80.4 78.4 72.3 90.9 75.9 81.1 82.5 78.8 81.9 73.6 77.5 Qwen3.5-122B-A10B 73.0 69.4 75.5 75.8 78.4 69.3 87.0 61.5 76.8 77.3 75.4 84.4 62.0 71.5 Qwen3.6-27B 73.1 71.0 76.6 78.0 77.0 69.6 92.2 78.6 78.4 79.6 75.9 81.3 67.0 73.5 Mean 74.2 73.3 77.3 79.1 78.0 72.0 89.4 72.2 78.0 80.0 76.5 79.6 71.5 75.1 DistRL Qwen3-VL-4B-Instruct 71.8 73.7 75.0 76.1 72.8 66.7 84.4 68.4 76.3 80.0 73.7 72.6 75.6 74.0 Qwen3-VL-8B-Instruct 79.1 80.1 75.0 77.3 78.9 74.6 89.6 75.0 73.2 77.3 78.3 77.4 79.6 78.5 Qwen3-VL-32B-Instruct 83.9 84.8 78.2 80.4 81.7 77.2 92.2 81.2 83.2 85.6 83.2 82.4 84.1 83.3 Qwen3-VL-235B-A22B-Instruct 80.6 81.0 81.9 83.7 79.8 74.3 90.9 78.8 80.0 82.4 81.1 81.8 79.7 80.8 Qwen3.5-122B-A10B 77.2 75.3 73.4 72.8 77.0 67.5 89.6 71.4 79.5 80.8 77.6 85.1 66.7 74.8 Qwen3.6-27B 80.4 80.2 77.7 77.4 77.9 69.3 88.3 71.0 78.4 79.6 79.8 84.6 72.7 78.2 Mean 78.8 79.2 76.9 78.0 78.0 71.6 89.2 74.3 78.4 81.0 79.0 80.6 76.4 78.3 AndroidGen Qwen3-VL-4B-Instruct 77.3 79.2 68.6 74.9 70.0 69.5 75.3 61.2 77.9 82.1 75.0 70.9 84.4 77.1 Qwen3-VL-8B-Instruct 76.4 75.5 75.5 78.3 70.4 65.2 85.7 70.3 79.5 82.8 76.3 77.2 74.1 75.7 Qwen3-VL-32B-Instruct 76.9 79.0 64.9 74.0 69.5 70.3 74.0 60.0 74.2 79.7 73.7 68.8 86.1 76.5 Qwen3-VL-235B-A22B-Instruct 78.1 79.6 71.8 77.1 73.2 67.8 84.4 71.4 80.5 82.5 77.2 75.1 81.0 77.9 Qwen3.5-122B-A10B 77.1 80.1 70.2 76.5 68.1 66.3 75.3 59.6 74.2 79.7 74.3 69.2 87.1 77.1 Qwen3.6-27B 82.6 83.0 77.1 81.2 74.2 68.2 85.7 66.7 75.8 79.5 79.8 79.1 80.9 79.9 Mean 78.1 79.4 71.4 77.0 70.9 67.9 80.1 64.9 77.0 81.0 76.0 73.4 82.3 77.4 WebRL Qwen3-VL-4B-Instruct 75.6 74.6 75.0 72.8 78.9 72.7 79.2 33.3 78.4 79.0 76.6 82.5 67.1 74.0 Qwen3-VL-8B-Instruct 76.8 75.1 79.8 79.6 70.9 59.2 80.5 34.8 82.1 83.3 77.2 84.1 66.7 74.4 Qwen3-VL-32B-Instruct 81.0 80.7 78.2 77.1 70.0 53.6 87.0 58.3 82.1 83.3 79.4 85.7 70.3 77.2 Qwen3-VL-235B-A22B-Instruct 80.3 82.0 74.5 76.2 77.9 73.1 84.4 62.5 81.1 83.8 79.5 77.7 82.3 79.9 Qwen3.5-122B-A10B 70.2 63.6 70.7 70.9 66.7 43.2 79.2 0.0 79.5 80.8 71.5 83.9 52.7 64.7 Qwen3.6-27B 83.9 85.0 78.7 81.8 81.2 76.2 81.8 53.3 80.0 82.4 82.2 80.9 84.0 82.4 Mean 78.0 76.8 76.1 76.4 74.3 63.0 82.0 40.4 80.5 82.1 77.7 82.5 70.5 75.4 ZeroGUI Qwen3-VL-4B-Instruct 72.3 67.8 75.5 75.8 76.5 67.5 87.0 58.3 80.5 81.2 75.3 85.1 61.0 71.0 Qwen3-VL-8B-Instruct 72.9 68.3 75.0 75.1 77.5 70.4 92.2 76.9 76.3 76.4 75.4 85.1 61.1 71.2 Qwen3-VL-32B-Instruct 74.9 72.1 83.0 84.2 80.8 73.5 88.3 64.0 80.5 81.4 78.4 86.1 67.3 75.5 Qwen3-VL-235B-A22B-Instruct 76.7 74.1 81.4 82.9 78.9 71.7 93.5 81.5 82.1 83.0 79.3 86.6 69.0 76.8 Qwen3.5-122B-A10B 85.8 86.5 82.4 83.7 83.6 80.2 96.1 90.9 80.0 82.7 84.8 84.1 85.6 84.8 Qwen3.6-27B 78.1 78.7 77.7 79.8 77.5 73.3 89.6 73.3 80.5 82.6 78.9 79.2 78.1 78.6 Gemini 3 Flash 80.3 79.6 77.1 76.8 82.2 78.2 90.9 75.9 84.7 85.0 81.3 86.7 73.7 79.7 Gemini 3.1 Flash-Lite 78.0 76.5 77.7 78.4 81.2 75.6 93.5 81.5 79.5 79.6 79.5 86.1 70.0 77.2 Mean 77.4 75.5 78.7 79.6 79.8 73.8 91.4 75.3 80.5 81.5 79.1 84.9 70.7 76.8 Table 9: Additional offline OGRBench results under each baseline’s original trajectory-context setting. Most baselines use only the final screenshot or a small final-state context, while ZeroGUI follows the prior two-screenshot setting. Appendix E Additional Baselines under Matched Image Budget This section provides the detailed numerical results behind the expanded matched-budget comparison summarized in Figure 3. Table 10 summarizes each baseline’s original trajectory-context setting and our matched-budget instantiation. For a controlled comparison, we provide every baseline with the same ten-screenshot trajectory budget used by AdaptRubric and evaluate them with the same Qwen-family judge backbones used in Figure 3. Table 11 reports the full results, including ZeroGUI and OS-Themis for reference. The results show that AdaptRubric maintains the strongest mean overall F1 across the expanded baseline set, indicating that the advantage does not come from comparing only against ZeroGUI and OS-Themis. Baseline Original trajectory context Matched-budget instantiation Verification style ZeroGUI (Yang et al., 2025) Final-state selection; the OS-Themis evaluation protocol uses the last two screenshots. Use the final ten screenshots and the same four-vote majority setting as the main offline comparison. Terminal-state judgment with voting. DigiRL (Bai et al., 2024) Autonomous evaluator prompt that judges task completion from visual evidence, typically centered on the observed screenshot/state. Provide the same ten selected trajectory screenshots as visual evidence. Autonomous trajectory evaluator. DistRL (Wang et al., 2025a) AUTO-EVALUATOR uses the task description with the last screenshot and recent action context. Provide the same ten selected trajectory screenshots and the corresponding trajectory context. Autonomous trajectory evaluator. AndroidGen (Lai et al., 2025) StepCritic-style evaluation uses the task, action history, and final screen to decompose required conditions and check completion. Provide the same ten selected trajectory screenshots together with the action history. Condition-wise task checking. WebRL (Qi et al., 2025) Outcome reward model uses the user intent, action history, and final web state. Provide the same ten selected trajectory screenshots and trajectory history, replacing web-only state evidence with visual GUI evidence for cross-platform OGRBench. Outcome reward modeling. OS-Themis (Li et al., 2026) Milestone-based multi-agent critic that selects and verifies outcome-critical trajectory evidence. Use the same ten-screenshot trajectory budget as AdaptRubric for offline comparison. Structured multi-step evidence verification. Table 10: Baseline context budgets and matched-budget instantiations for the expanded OGRBench comparison. All matched-budget runs use at most ten trajectory screenshots to align with AdaptRubric. Model Ubuntu Mobile Windows macOS Web Overall Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 Acc Prec Rec F1 DigiRL Qwen3-VL-4B-Instruct 77.5 81.2 76.1 80.5 77.0 77.4 88.3 76.9 72.1 77.6 77.1 70.7 92.0 80.0 Qwen3-VL-8B-Instruct 80.0 82.8 76.6 80.4 79.8 78.8 89.6 78.9 70.0 76.2 78.7 73.2 90.1 80.8 Qwen3-VL-32B-Instruct 85.2 86.8 82.4 85.2 82.6 80.6 94.8 88.2 79.5 82.7 84.2 79.6 91.6 85.2 Qwen3-VL-235B-A22B-Instruct 87.9 89.1 80.9 83.9 84.0 82.3 94.8 87.5 82.1 84.5 85.9 81.8 92.3 86.7 Qwen3.5-122B-A10B 84.2 84.7 78.7 81.0 78.4 72.6 90.9 75.9 78.9 81.1 82.3 83.0 80.9 81.9 Qwen3.6-27B 84.2 85.1 79.3 81.7 77.9 71.9 94.8 87.5 80.0 82.2 82.6 82.1 83.1 82.6 Mean 83.2 85.0 79.0 82.1 80.0 77.3 92.2 82.5 77.1 80.7 81.8 78.4 88.3 82.9 DistRL Qwen3-VL-4B-Instruct 77.5 80.6 74.5 77.6 80.3 79.6 87.0 72.2 73.7 78.3 77.5 72.7 87.7 79.5 Qwen3-VL-8B-Instruct 83.8 85.5 79.3 81.7 78.9 78.3 90.9 81.1 75.8 80.0 81.8 77.1 90.0 83.1 Qwen3-VL-32B-Instruct 86.8 88.2 80.9 83.2 83.6 81.9 93.5 84.8 79.5 83.0 84.9 80.5 91.9 85.8 Qwen3-VL-235B-A22B-Instruct 87.3 88.5 81.4 84.0 82.6 80.0 94.8 88.2 84.7 86.9 85.9 82.4 91.0 86.5 Qwen3.5-122B-A10B 82.9 83.1 83.0 84.2 80.3 74.7 96.1 90.3 80.5 81.8 82.9 85.1 79.4 82.2 Qwen3.6-27B 87.0 88.1 79.3 80.6 82.6 79.1 93.5 85.7 81.1 82.7 84.9 83.3 87.0 85.1 Mean 84.2 85.7 79.7 81.9 81.4 78.9 92.6 83.7 79.2 82.1 83.0 80.2 87.8 83.7 AndroidGen Qwen3-VL-4B-Instruct 79.5 82.6 73.4 78.4 71.8 73.7 71.4 54.2 81.1 84.3 77.3 70.9 92.1 80.1 Qwen3-VL-8B-Instruct 84.6 85.7 77.7 81.3 73.2 71.9 87.0 75.0 81.1 84.6 81.6 77.7 88.4 82.7 Qwen3-VL-32B-Instruct 78.4 81.4 65.4 74.3 71.4 73.6 70.1 58.2 71.1 78.3 74.2 67.7 91.9 77.9 Qwen3-VL-235B-A22B-Instruct 82.5 84.3 71.3 76.7 76.5 74.5 89.6 80.0 83.2 85.6 80.6 76.0 88.9 81.9 Qwen3.5-122B-A10B 83.1 85.3 72.3 78.2 78.4 78.3 85.7 74.4 75.8 80.2 80.1 74.0 92.6 82.2 Qwen3.6-27B 86.5 88.1 81.4 84.2 75.6 75.5 94.8 88.9 75.8 80.5 83.2 77.4 93.4 84.7 Mean 82.4 84.6 73.6 78.9 74.5 74.6 83.1 71.8 78.0 82.2 79.5 74.0 91.2 81.6 WebRL Qwen3-VL-4B-Instruct 85.6 86.4 77.1 75.1 80.3 75.9 94.8 86.7 77.4 78.4 83.0 84.8 80.3 82.5 Qwen3-VL-8B-Instruct 82.9 83.2 81.4 81.7 75.1 66.7 87.0 61.5 83.2 84.9 81.8 84.4 77.6 80.9 Qwen3-VL-32B-Instruct 86.6 87.3 80.3 78.9 81.7 76.4 85.7 47.6 86.3 87.4 85.0 87.8 81.0 84.2 Qwen3-VL-235B-A22B-Instruct 86.0 87.7 78.2 80.4 82.2 80.6 90.9 81.1 78.9 82.6 83.7 78.7 92.1 84.9 Qwen3.5-122B-A10B 82.1 81.3 76.1 78.5 68.1 48.5 83.1 31.6 78.9 80.6 78.8 84.6 70.0 76.6 Qwen3.6-27B 86.1 87.8 82.4 85.1 82.6 80.6 92.2 82.4 82.6 85.1 85.0 80.2 92.6 85.9 Mean 84.9 85.6 79.2 80.0 78.3 71.5 89.0 65.1 81.2 83.2 82.9 83.4 82.3 82.5 ZeroGUI Qwen3-VL-4B-Instruct 83.8 84.0 74.5 76.2 80.3 75.6 90.9 78.8 81.1 83.2 82.0 83.2 80.0 81.6 Qwen3-VL-8B-Instruct 84.6 85.0 83.0 84.5 79.3 74.4 94.8 88.2 78.4 80.6 83.3 84.1 81.9 83.0 Qwen3-VL-32B-Instruct 84.6 85.2 83.5 85.0 80.3 75.6 96.1 90.3 82.1 83.5 84.1 84.6 83.1 83.9 Qwen3-VL-235B-A22B-Instruct 86.9 87.3 84.0 85.8 84.5 80.9 96.1 90.3 87.4 88.6 86.7 87.2 85.9 86.5 Qwen3.5-122B-A10B 85.8 86.5 82.4 83.7 83.6 80.2 96.1 90.9 80.0 82.7 84.8 84.1 85.6 84.8 Qwen3.6-27B 86.2 87.5 80.9 83.5 83.6 81.5 94.8 88.9 82.1 84.7 85.0 81.2 90.9 85.8 Mean 85.3 85.9 81.4 83.1 81.9 78.0 94.8 87.9 81.9 83.9 84.3 84.1 84.6 84.3 OS-Themis Qwen3-VL-4B-Instruct 72.6 71.4 79.3 78.9 75.1 68.3 84.4 57.1 80.0 81.9 75.5 79.5 68.3 73.5 Qwen3-VL-8B-Instruct 76.2 75.0 84.0 83.7 78.4 72.0 83.1 51.9 81.6 82.4 78.7 84.6 69.9 76.5 Qwen3-VL-32B-Instruct 77.1 74.6 81.9 80.7 76.5 65.3 90.9 77.4 83.7 81.9 79.3 91.6 64.1 75.5 Qwen3-VL-235B-A22B-Instruct 86.4 86.8 93.6 93.7 77.5 69.6 93.5 82.8 91.6 91.9 87.1 90.5 82.7 86.4 Qwen3.5-122B-A10B 85.8 85.8 88.3 88.2 79.8 73.0 88.3 66.7 79.0 75.6 84.5 92.0 75.3 82.8 Qwen3.6-27B 87.9 88.5 93.6 93.9 88.3 86.9 96.1 90.3 92.1 92.1 89.7 90.2 89.0 89.6 Mean 81.0 80.4 86.8 86.5 79.3 72.5 89.4 71.0 84.7 84.3 82.5 88.1 74.9 80.7 AdaptRubric Qwen3-VL-4B-Instruct 84.3 85.6 80.9 81.1 83.1 79.5 97.4 94.1 82.1 84.5 84.1 82.8 85.9 84.3 Qwen3-VL-8B-Instruct 83.9 85.1 83.5 84.6 81.7 79.1 97.4 94.1 81.6 83.3 84.0 82.6 85.9 84.2 Qwen3-VL-32B-Instruct 86.5 86.6 85.6 85.9 80.3 75.0 93.5 83.9 86.8 87.6 85.9 89.3 81.3 85.1 Qwen3-VL-235B-A22B-Instruct 86.1 86.7 87.8 88.9 84.5 81.4 94.8 87.5 88.9 90.0 86.9 87.0 86.7 86.8 Qwen3.5-122B-A10B 86.2 86.8 89.9 90.4 81.7 77.2 94.8 88.2 88.9 89.8 86.9 87.8 85.4 86.6 Qwen3.6-27B 87.3 88.5 91.0 91.6 84.5 82.5 93.5 85.7 91.6 92.4 88.3 85.4 92.1 88.7 Mean 85.7 86.5 86.5 87.1 82.6 79.1 95.2 88.9 86.6 87.9 86.0 85.8 86.2 86.0 Table 11: Additional offline OGRBench results under the matched ten-screenshot trajectory budget. We include DigiRL, DistRL, AndroidGen, WebRL, ZeroGUI, OS-Themis, and AdaptRubric across multiple judge backbones. Appendix F Reward-Guided Test-Time Scaling Details We evaluate reward-guided trajectory selection on AndroidWorld. The pool contains 113 tasks with complete rollout traces and ground-truth labels from the AndroidWorld evaluator. Instead of executing new rollouts separately for each verifier, we use this pre-collected trajectory pool to simulate online repeated attempts under a controlled environment. This ensures that different verifiers are compared on the same candidate trajectories rather than on different samples produced by a stochastic dynamic environment. For each task, we collect ten independent candidate trajectories: eight generated by Qwen3-VL-8B-Instruct and two generated by Qwen3-VL-235B-A22B-Instruct, all sampled with temperature 0.7. Pool Trials/task All-fail All-pass Mixed Trial SR Oracle@N Headroom 8B-only 8 48 (42.5%) 44 (38.9%) 21 (18.6%) 51.44% 57.52% +6.08 235B-only 2 46 (40.7%) 52 (46.0%) 15 (13.3%) 52.65% 59.29% +6.64 8B+235B 10 32 (28.3%) 29 (25.7%) 52 (46.0%) 51.68% 71.68% +20.00 Table 12: Candidate-pool composition for reward-guided test-time scaling. Mixed tasks contain both successful and failed trajectories, and are the only tasks where reward-guided selection can change the final task success. Headroom is Oracle@N minus Random@1. Why a heterogeneous pool. Table 12 shows why we use a heterogeneous pool rather than a homogeneous single-policy pool. The 8B-only pool is strongly bimodal: 42.5%42.5\% of tasks fail on every trial and 38.9%38.9\% succeed on every trial, leaving only 18.6%18.6\% of tasks where any reward verifier can change the outcome. The 235B-only pool has a similar issue because it contains only two trajectories per task, yielding only 13.3%13.3\% mixed tasks. In these all-success or all-failure cases, the verifier choice is irrelevant and SR@N differences are dominated by pool composition rather than reward quality. The heterogeneous pool raises the mixed-task fraction to 46.0%46.0\% and widens the Oracle-minus-Random headroom from about +6+6 points to +20.00+20.00 points. Relative to the 8B-only pool, it creates 31 additional mixed tasks, all caused by the added 235B trajectories changing an otherwise all-success or all-failure task into a task with both positive and negative candidates. This policy-level diversity makes the reward-guided selection setting informative enough to distinguish verifiers. Per-task shuffling. If the trial order followed the source model ([8B×8, 235B×2]), EarlyStop@N for small N would be dominated by the 8B success rate and a jump to large N would mostly reflect the position of the stronger 235B rollouts rather than the verifier’s judgment. To remove this position prior, we apply a fixed per-task shuffle (seed 4242) so that every position i∈[0,9]i∈[0,9] contains a 235B trajectory with probability ≈20%≈ 20\%. All reward methods score the same shuffled pool and do not receive the source model metadata. We use Qwen3-VL-32B-Instruct as the judge backbone for all compared reward methods. For EarlyStop@N, the verifier scans candidates in order and stops at the first trajectory predicted as successful. If no candidate is predicted as successful within the budget, the protocol falls back to the last candidate in the scanned prefix. For BestOfN@N, the verifier scores all candidates in the prefix and selects a predicted-success trajectory; when binary rewards create ties, we use a deterministic last-candidate tie break. Random@N uniformly selects a candidate from the same prefix and serves as the lower reference, while Oracle@N succeeds whenever the prefix contains at least one truly successful trajectory. The main table reports method gains over Random in percentage points. Appendix G Detailed Case Study Figure 5: Qualitative case study illustrating the importance of task-adaptive criteria. The task requires adding Hello, World! to the top of note_SiFbv.txt in Markor. The trajectory inserts the text, but places it below the original note content, yielding a failed task. Baseline verifiers focus on content presence and incorrectly judge the task as completed. AdaptRubric combines a coarse create/modify rubric with fine placement cues, enabling the verifier to check both the edited item and the instruction-specific requirement that no text appears above the inserted phrase. Figure 5 shows a representative AndroidWorld case where task-adaptive criteria are crucial for avoiding a false positive reward. The user asks the agent to edit note_SiFbv.txt in Markor and add Hello, World! to the top of the note. The ground-truth label is failure. Although the agent opens the correct file and types the target phrase, the final screenshot shows that the original sentence, “Don’t forget to water the plants while I’m away.”, remains above the inserted text. Thus, the trajectory satisfies content presence but violates the positional requirement implied by “to the top of the note”. The baseline verifiers fail because they rely on a surface-level notion of task completion. ZeroGUI states that the final screenshot confirms Hello, world! is now the first line of the note and therefore assigns a successful reward. OS-Themis makes a similar error: its critical evidence claims that Hello, world! appears as the first line, followed by the original content. Both judgments are false positives. They verify that the requested text appears, but they do not correctly verify the spatial or textual ordering of the final note. In other words, they verify the content but miss the placement. AdaptRubric avoids this error by constructing a task-adaptive criterion before making the reward judgment. The coarse stage routes the instruction to the create_modify category, whose rubric bounds verification to the final state of the edited item and the explicitly required properties. This prevents the verifier from judging the task using generic signals such as whether some edit was made. The fine stage then adds instruction-specific cues, including whether Hello, World! appears at the very top of note_SiFbv.txt before any other text. With the fused criterion, the verifier checks both the edited note and the top-placement requirement. It correctly observes that the final state contains the original sentence followed by Hello, world!, so the inserted text is not at the top of the note. The trajectory is therefore assigned a failure reward. The single-component variants further illustrate why both levels are needed. The coarse-only verifier predicts success because it recognizes that the note was modified and that the target text is visible, but the category-level rubric alone is too broad to enforce the exact top-of-note placement. The fine-only verifier receives the top-placement cue, but still misreads the final order and predicts success. Only the fused coarse-to-fine criterion combines a structured final-state check with the instruction-specific positional constraint, leading to the correct failure judgment. Appendix H PROMPTS We provide the prompt-facing components used by AdaptRubric. The rubric bank is loaded once from the cold-start bank, the coarse stage routes each instruction to one of the eight rubric families and retrieves the corresponding bank entry, and the fine stage optionally generates instance-level rubric cues. H.1 Rubric Bank The rubric bank ℬB contains one cold-start RubricSet for each task family. Each entry stores natural-language sections, namely verification_steps, common_pitfalls, special_rules, and output_format. It also stores named rubric_items with criticality labels for structured ablations. The default verifier uses the natural-language sections; the structured rubric_items are kept for the legacy structured ablation. ⬇ RubricSet( id="RSET_INFO_QUERY", categories=["info_query"], description="Read screen content and return a value, count, list, or title.", sections= "verification_steps": [ "Check explicit answer-format constraints, if any.", "Ground content correctness in screenshots from the trajectory.", "Check answer/submission evidence when the task requires an answer." ], "common_pitfalls": [ "Correct content but wrong requested format.", "Answer hallucinated from memory rather than screen evidence.", "Wrong count or list item due to misreading visible content." ], "special_rules": [ "A format constraint is hard only when stated by the instruction.", "The supporting evidence need not appear in the final frame.", "Semantic equivalence is acceptable when no format is specified." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "Answer Format Match", "critical"), ("R2", "Content Correctness", "critical"), ("R3", "Answer Recorded", "default"), ("R4", "Screen-Grounded Source", "auxiliary") ] ) RubricSet( id="RSET_CREATE_MODIFY", categories=["create_modify"], description="Create a new item or modify explicitly requested properties.", sections= "verification_steps": [ "Collect only properties explicitly stated in the instruction.", "Check whether the required properties are reflected in the end state.", "Check that the edit is not left mid-flight." ], "common_pitfalls": [ "Inventing filename, folder, toast, or UI-path requirements.", "Only a subset of explicitly required properties is applied.", "The final frame still shows an uncommitted editor or form." ], "special_rules": [ "Any valid save or commit path counts unless a path is specified.", "Unmentioned properties are ignored." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "Required Properties Set", "critical"), ("R2", "Save / Confirm Action Performed", "critical"), ("R3", "Final State Persisted", "critical"), ("R4", "No Unspecified Constraint Imposed", "default"), ("R5", "Form / Editor Closed After Save", "auxiliary") ] ) RubricSet( id="RSET_DELETE_CLEANUP", categories=["delete_cleanup"], description="Remove items or clear state under specific or global scope.", sections= "verification_steps": [ "Distinguish specific deletion from global cleanup.", "Check final-state absence or zero remaining matches.", "Check that any confirmation dialog has been committed.", "Guard against inferring absence from an unrelated screen." ], "common_pitfalls": [ "Passing after one deletion when the instruction says all.", "Confirmation dialog still visible in the final frame.", "Deleting a similarly named but wrong item." ], "special_rules": [ "Global cleanup requires visibly complete final evidence.", "Post-confirmation return to a list with the item gone is strong evidence." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "Correct Target / Scope Identified", "critical"), ("R2", "Deletion Confirmed", "critical"), ("R3", "Final-State Absence / Count Verified", "critical"), ("R4", "Pre-Deletion Visibility", "auxiliary") ] ) RubricSet( id="RSET_COMMUNICATION", categories=["communication"], description="Send, reply, forward, or share content to a recipient.", sections= "verification_steps": [ "Check source content.", "Check recipient and channel.", "Check post-send evidence instead of treating a draft as sent." ], "common_pitfalls": [ "Compose screen remains visible, so the message may be unsent.", "Wrong recipient or wrong channel.", "Visible sent bubble has different content." ], "special_rules": [ "Thread view with a new bubble is sufficient for messaging apps.", "Literal message text must match up to minor whitespace differences." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "Source Content Correctness", "critical"), ("R2", "Recipient Correctness", "critical"), ("R3", "Send Action Confirmation", "critical"), ("R4", "Channel / App Match", "default"), ("R5", "No Side Effects", "auxiliary") ] ) RubricSet( id="RSET_TRANSFER", categories=["transfer"], description="Move, copy, import, export, clone, unzip, or save content from source to sink.", sections= "verification_steps": [ "Identify source evidence.", "Check destination evidence.", "Check content fidelity at the sink.", "Detect partial or corrupted transfer." ], "common_pitfalls": [ "Source read but destination left empty.", "Only a subset transferred when all was required.", "Manual retype introduces truncation or digit errors." ], "special_rules": [ "Both source-side and sink-side evidence are required.", "Fidelity mismatch is a fail even if the transfer action succeeded." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "Source Read Verified", "critical"), ("R2", "Destination Reached", "critical"), ("R3", "Content Fidelity", "critical"), ("R4", "Correct Transfer Mechanism", "default"), ("R5", "Completeness for all Tasks", "auxiliary") ] ) RubricSet( id="RSET_STATE_NAV", categories=["state_navigation"], description="Navigate to a view, toggle a setting, or configure a preference.", sections= "verification_steps": [ "Check the final-frame state.", "Check persistence of the requested state or view.", "Avoid requiring side effects not requested by the instruction." ], "common_pitfalls": [ "Correct view was only visited mid-trajectory.", "Toggle was tapped but visual state did not change.", "Parent page reached instead of the requested sub-page." ], "special_rules": [ "The final frame is the primary anchor.", "Any UI path is acceptable unless a path is specified." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "Final Screen Match", "critical"), ("R2", "State Persisted", "critical"), ("R3", "Exact Sub-Page Reached", "default"), ("R4", "Toggle Visual State", "auxiliary") ] ) RubricSet( id="RSET_COMPOSITE", categories=["composite_workflow"], description="Two or more independent sub-goals must all be completed.", sections= "verification_steps": [ "Decompose the instruction into independent checkpoints.", "Check evidence for each sub-goal using its underlying family logic.", "Fail when any sub-goal is silently dropped.", "Treat ordering as required only when the instruction makes it required." ], "common_pitfalls": [ "Judging only the first or most visible sub-goal.", "Over-decomposing implementation steps into false sub-goals." ], "special_rules": [ "Composite success is an AND over sub-goals.", "Compositeness does not raise the per-sub-goal pass bar." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "All Sub-goals Identified", "critical"), ("R2", "Each Sub-goal Evidenced", "critical"), ("R3", "No Silent Drop", "critical") ] ) RubricSet( id="RSET_GENERAL", categories=["*"], description="Conservative fallback for tasks outside the specialized families.", sections= "verification_steps": [ "Reconstruct the user’s intended outcome.", "Check end-state evidence.", "Check action-visual consistency.", "Use conservative judgment under genuine ambiguity." ], "common_pitfalls": [ "Intermediate success but final state drifts elsewhere.", "Action trace shows intent but no visual corroboration." ], "special_rules": [ "Prefer EVIDENCE_SUFFICIENT=no when evidence is unclear.", "Do not infer success from a confident-looking answer alone." ], "output_format": OUTPUT_FORMAT_HOLISTIC , rubric_items=[ ("R1", "Task Goal Achieved", "critical"), ("R2", "Final State Consistent", "default"), ("R3", "Action-Visual Consistency", "auxiliary") ] ) Shared meta rule appended to each family: - Evidence anchors guide attention; they do not add requirements beyond the task instruction. - If the final state clearly satisfies the user’s intent but a specific anchor cannot be confirmed, prefer PASS unless the instruction made that anchor mandatory. - Do not fail for properties the instruction did not state, such as filenames, folders, confirmation toasts, specific phrasing, or cosmetic details. H.2 Category-Level Coarse Rubric Retrieval The coarse stage first routes the instruction to a rubric family. In the LLM-routing mode, the implementation uses the following classifier prompt and parses the single-line category output. The retrieved rubric is then selected deterministically: if the category is not general, the retriever returns the unique RubricSet whose categories field contains that category; otherwise it returns the wildcard RSET_GENERAL. ⬇ You are routing a GUI task to one of 8 rubric families. TASK INSTRUCTION: instruction CATEGORIES (pick exactly one): - info_query : the agent must answer a question by reading screen content (a value, a count, a list, a title) - create_modify : the agent creates a NEW item or edits properties of an EXISTING item (note, event, file, drawing, form field, spreadsheet cell) - delete_cleanup : the agent removes items, uninstalls, clears data, gets rid of content - communication : the agent sends / replies / forwards / shares content to a recipient via a messaging channel - transfer : the agent moves content source -> sink (copy, paste, import, export, clone repo, unzip, save-as to a path) - state_navigation : the agent toggles a setting, configures a preference, changes theme/font/default app, navigates to a specific view - composite_workflow : the instruction lists TWO OR MORE INDEPENDENT sub-goals (different apps, different target objects) connected by "and then", "; then", ";". Sequential steps that all accomplish ONE goal are NOT composite. - general : the task genuinely does not fit any specific family above. Use ONLY as a last resort. Decision rules: 1. If two or more categories seem to apply, pick the one matching the PRIMARY outcome the user cares about. 2. Prefer a specific family over general whenever possible. 3. composite_workflow requires independent sub-goals, not just multiple sub-steps of one operation. OUTPUT EXACTLY ONE LINE (no other text, no markdown): CATEGORY: <one of info_query | create_modify | delete_cleanup | communication | transfer | state_navigation | composite_workflow | general> After retrieval, the selected category-level rubric sections are rendered into the verifier prompt. The implementation uses the following prompt skeleton; in the coarse-only setting, additional_attention is empty. ⬇ You are evaluating a GUI task. TASK: instruction COMPILED EVIDENCE: - Agent’s submitted answer: agent_answer - Key text entered by agent: typed_texts - Key actions performed: key_actions ACTION TRACE: action_trace verification_steps common_pitfalls special_rules additional_attention Screenshots (in order) are shown below. End your response with EXACTLY these three lines (do not add any text after): SCORE: [0/1] EVIDENCE_SUFFICIENT: [yes/no] MISSING_EVIDENCE: [source_missing | sink_missing | final_state_ambiguous | answer_format_mismatch | global_count_unverified | subgoal_unverified | none] H.3 Instance-Level Fine Rubric Generation The fine stage generates zero, one, or two instruction-specific rubric cues. It receives the instruction, the routed task family, and the critical items from the retrieved coarse rubric. The generator is explicitly allowed to abstain with NO_HINT; otherwise, the generated cues are injected into additional_attention in the verifier prompt above. ⬇ You are generating optional task-specific rubric cues for a GUI task verifier. The verifier already has a coarse task-type rubric that covers standard verification (action trace, final screenshot, task completion). Your job is to decide whether this specific task needs additional fine-grained attention cues, and if so, generate them. TASK INSTRUCTION: instruction TASK TYPE: category COARSE RUBRIC ITEMS already applied (for reference -- do NOT paraphrase them): rubric_items_block YOUR JOB: Decide whether 0, 1, or 2 task-specific rubric cues would help the verifier. OUTPUT NO_HINT when: - The coarse rubric already covers the task adequately. - Any cue you could write would merely restate the coarse rubric in task-specific words. - The task is straightforward and the main success criterion is obvious from the instruction. - You cannot generate a cue without introducing requirements NOT present in the instruction. OUTPUT 1-2 CUES when the task has: - An explicit constraint that the coarse rubric might miss (a specific value, format, recipient, negative constraint like "do not send", exact date/time, recurrence rule). - A plausible partial-completion trap where the agent might stop one step early. - An answer-format requirement where the verifier might accept wrong structure. RULES for valid cues: 1. Must be grounded in an EXPLICIT phrase/value/constraint from the task instruction. 2. Must distinguish plausible partial completion from actual success. 3. Must NOT require "final screenshot" evidence unless the instruction explicitly asks for a visible final state. The verifier can use action trace, intermediate screenshots, and submitted answers as evidence. 4. Must NOT use over-constraining language: avoid "visibly", "exactly", "every/all" (unless the instruction itself uses these words), "saved/persisted" (unless the task is about saving), "final state must show". 5. Must NOT add requirements absent from the instruction (e.g., "must remain in the app", "must show confirmation dialog", "must prove persistence"). 6. Should specify what evidence source is acceptable: action trace, any screenshot, submitted answer, or final screenshot. 7. Use soft attention phrasing: "Pay attention to whether ...", "Check if ...", "The key distinction is ...", "Evidence of success includes ...". Do NOT start with "Confirm that ..." (too imperative). 8. One sentence each, <= 40 words. OUTPUT FORMAT (exactly one of these two forms, no other text): If no cue needed: NO_HINT: <one-sentence reason> If 1-2 cues: CUE_1: <sentence> CUE_2: <optional second sentence> When cues are available, they are rendered as a soft task-specific attention block rather than as additional hard requirements. ⬇ TASK-SPECIFIC ATTENTION CUES (focus aids -- NOT additional pass/fail items): The bullets below highlight aspects of THIS task that are worth attending to. They guide where to look; they do NOT add pass/fail requirements beyond the task instruction itself, and they do NOT override the META RULE above. - cue_1 - cue_2 CUE USAGE -- REMINDERS THAT OVERRIDE EARLIER WORDING: - A cue you cannot confirm in the screenshots is NOT, by itself, a failure. Base the verdict on whether the OVERALL final state satisfies the user’s intent in the task instruction -- not on per-cue verifiability. - Use EVIDENCE_SUFFICIENT=no only when the OVERALL verdict is uncertain, not when a single cue is unconfirmed but the rest of the evidence agrees. - If the final state clearly satisfies the user’s intent, prefer SCORE: 1 even when one or more cues are visually ambiguous. - A cue CLEARLY CONTRADICTED by visible evidence is still a strong FAIL signal.