Paper deep dive
GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis
Long Zhang, Yuhan Chen, Chaoran Zhang, Wanxia Cao, Kun Huang, Pengzhi Gao, Wei Liu, Jian Luan, Chenliang Li, Lixin Zou
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific environments and struggle to generate diverse data, while existing evaluators either suffer from limited scalability or provide inaccurate and unreliable reward signals. To overcome these challenges, we introduce GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization. Our approach features self-evolving data synthesis, which produces multiple environments through task execution and generates diverse tasks and goal states. Complementing this, a state-anchor mechanism automatically annotates task-relevant UI elements in successful goal states as reference anchors. During RL training, these reference anchors provide accurate, scalable reward signals that substantially enhance efficiency. Extensive evaluations demonstrate that our framework achieves over 90% accuracy on offline trajectory verification and performs closest to rule-based methods. Furthermore, agents trained using our reward framework exhibit strong performance on both AndroidWorld and our constructed benchmark, establishing a scalable approach for GUI agent training.
Tags
Links
- Source: https://arxiv.org/abs/2608.22847v1
- Canonical: https://arxiv.org/abs/2608.22847v1
Trouble viewing inline? Open PDF directly →
Full Text
73,052 characters extracted from source content.
Expand or collapse full text
GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis Long Zhang1 Yuhan Chen2 Chaoran Zhang1 Wanxia Cao2 Kun Huang2 Pengzhi Gao2 Wei Liu2 Jian Luan2 Chenliang Li1 Lixin Zou1 1Wuhan University 2Xiaomi Inc. zlongooo, chaoranzhang, cllee, zoulixin@whu.edu.cn yuhanchen240, caowanxia123, huangkun813@gmail.com gaopengzhi, liuwei40, luanjian@xiaomi.com Thanks: Corresponding author. Abstract Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental issues: current data synthesis methods for GUI Agents rely on specific environments and struggle to generate diverse data, while existing evaluators either suffer from limited scalability or provide inaccurate and unreliable reward signals. To overcome these challenges, we introduce GSAR (Goal-State-Anchor Reward), a RL reward framework that supports scalable task generation and delivers reliable reward signals for stable and efficient policy optimization. Our approach features self-evolving data synthesis, which produces multiple environments through task execution and generates diverse tasks and goal states. Complementing this, a state-anchor mechanism automatically annotates task-relevant UI elements in successful goal states as reference anchors. During RL training, these reference anchors provide accurate, scalable reward signals that substantially enhance efficiency. Extensive evaluations demonstrate that our framework achieves over 90% accuracy on offline trajectory verification and performs closest to rule-based methods. Furthermore, agents trained using our reward framework exhibit strong performance on both AndroidWorld and our constructed benchmark, establishing a scalable approach for GUI agent training. 1 Introduction Graphical User Interface (GUI) agents driven by Vision-Language Models (VLMs) 2; 11 demonstrate human-like competence in understanding and operating mobile device environments 27; 36. As a promising paradigm for automating interactions across mobile and desktop platforms, GUI agents transform raw screen observations into actionable decisions through visual comprehension and functional reasoning. These agents operate without extensive fine-tuning or additional pretraining, highlighting their strong potential for broad real-world applications. Figure 1: Rule-based and Model-based reward judge. Model-based judges may be inaccurate when context is limited or model capacity is insufficient. Reinforcement Learning with Verifiable Rewards (RLVR) 13; 24; 37; 7, building on its success in domains such as mathematics and code reasoning, has recently emerged as an effective paradigm for training GUI agents 1; 32. Scaling RLVR for GUI agents fundamentally relies on a unified training infrastructure: diverse task instantiation and strictly verified reward signals. While Large Language Models can trivially generate a vast array of task descriptions, grounding these open-ended instructions into executable training episodes requires configuring the corresponding initial environments. Currently, this environment setup and data collection process remains highly restricted by manually designed configurations 26; 15. This labor-intensive dependency prevents the automatic generation of truly diverse environment-task pairs, which is the foundational prerequisite for large-scale policy learning. Furthermore, evaluating agent performance across such an open-ended task space introduces an equally complex challenge. As shown in Figure 1, existing reward designs struggle to scale effectively. Rule-based methods require manually written execution functions to verify task completion 22; 31; while precise, they are completely impractical to manually scale alongside automatically generated tasks. Conversely, model-based evaluators attempt to automate outcome assessment 1; 28, but they frequently suffer from instability, hallucination, or insufficient coverage of edge cases. These flaws lead to noisy supervision signals that severely impair policy convergence and degrade agent performance. Together, the inability to automatically instantiate diverse initial environments and the lack of scalable, accurate reward supervision form a systemic bottleneck that constrains the true potential of RL for GUI agents. To address this challenge, we introduce GSAR (Goal-State-Anchor Reward), an RL reward framework that tackles both bottlenecks by enabling scalable task and environment generation and providing reliable reward signals for stable policy optimization. First, we leverage self-evolving data synthesis to generate diverse tasks and trajectories, while progressively evolving the environments through interaction, resulting in task, initial environment, and trajectory triplets in a fully automated manner. Next, the state-anchor mechanism automatically annotates task-relevant UI elements in the final states of successful trajectories, producing task, initial environment, and goal-state-anchor reference triplets. During online training, these references are used to provide accurate and semantically grounded reward signals. Both online and offline experiments demonstrate that our approach, by explicitly anchoring task-relevant UI elements in successful goal states, delivers robust and scalable reward feedback for GUI agent training. In summary, our work makes the following primary contributions: • We propose a self-evolving data synthesis mechanism that generates diverse environments and task trajectories while preserving initial environment snapshots, enabling reproducible and scalable reinforcement learning. • We introduce a scalable Goal-State-Anchor reward framework for GUI agents that overcomes the limitations of prior judges to deliver accurate, generalizable signals for online reinforcement learning. • We analyze the impact of reward accuracy on reinforcement learning and show that reward design critically affects agent performance. 2 Related Works 2.1 Data Synthesis for Mobile GUI Agents Effective training of GUI agents requires diverse tasks that expose models to a wide range of interface elements and interaction patterns necessary for robust and generalizable performance. Following the introduction of sequential GUI data for mobile applications in Rico 6, research has increasingly focused on synthesizing GUI datasets. Several efforts 23; 14; 17 generated data for supervised fine-tuning (SFT) 20 models, but the scalability of manually curated datasets is hindered by costly and inefficient data collection. More recently, automated data synthesis approaches have been proposed to address this limitation. These methods first explore application interfaces via random or heuristic traversal, and then construct tasks based on the explored states, either by reverse-engineering instructions 26 or generating tasks directly from screenshots 30. Large pretrained models (e.g., GPT-4o 11) are then employed to execute these tasks and filter out incomplete or failed trajectories. However, these approaches still depend on manually preconfigured app environments and do not provide reproducible initial states required for stable reinforcement learning of GUI agents. Figure 2: Overview of GSAR. (a) self-evolving data synthesis interacts with mobile environments to modify environment states and automatically generate tasks and trajectories. (b) state-anchor mechanism leverages successful trajectories to identify and anchor task-related UI elements in the goal state, providing more accurate reward signals for training. 2.2 Reward System for GUI Agents Recent research has increasingly integrated reinforcement learning (RL) into the training of GUI agents 39; 3, with the goal of reducing supervision requirements and enhancing generalization and long-term decision-making ability. In RL, reward signals are crucial for guiding policy optimization and shaping learning dynamics 8; 5; 10. Initial reward designs for GUI agents primarily focused on grounding tasks 38; 40, such as click position prediction 18 or intersection over union (IoU) measurement between predicted and target boxes 16. Yet these methods are limited in scope and lack validation on more complex high-level tasks. To support training of high-level GUI agents, subsequent studies have developed rule-based outcome rewards 22; 19 that provide reliable supervision through task-specific verification scripts. In addition, model-based reward functions have been explored for high-level tasks 35; 25, where an external model evaluates successful demonstrations or outcomes, providing rewards with high scalability across tasks. Building upon model-based methods, we propose a reward framework designed to provide more accurate and reliable reward signals for RL training while maintaining scalability. 3 Method In this section, we introduce GSAR (Goal-State-Anchor Reward) framework. An overview of the GSAR pipeline is illustrated in Figure 2. The framework operates through a seamless integration of two core phases. First, through self-evolving data synthesis, the framework progressively generates diverse tasks, execution trajectories, and their corresponding initial environments via interaction-driven evolution. Subsequently, the state-anchor mechanism leverages these trajectories for automatic reward annotation. By anchoring to goal states, this mechanism provides reliable and scalable reward supervision to close the learning loop. 3.1 Self-Evolving Data Synthesis To synthesize data for training and to provide high-quality source data for reward annotation, we propose a self-evolving data synthesis framework that is both automatic and scalable. Unlike static or one-shot data generation pipelines, our approach explicitly models data construction as an evolving process. The key idea is to allow the task distribution, environment states, and trajectories to evolve together over time through continuous interaction with the GUI environment, closely mirroring how real-world applications are explored and used by humans. Task generation, execution, and trajectory collection are all driven by the model itself, enabling continuous expansion of the dataset without manual intervention. The self-evolving data synthesis framework is guided by the following principles: Simulation of Real-World Usage. If an app has not undergone manual configuration or real user interaction, the information contained in a freshly installed app is typically minimal and simplistic. To enable data collection in richer and more diverse environments, we modify the app state by executing multiple tasks. Through repeated rounds of task execution and environment updates, the system progressively evolves the app states. This iterative process produces a wide variety of environments, which subsequently support the collection of more diverse tasks. Automatic Task Generation. We first explore the GUI environment of each app through random interaction, producing an exploration trajectory: T=(s0,a0,s1,a1,…,sn),T=(s_0,a_0,s_1,a_1,…,s_n), where sis_i denotes the environment state represented by the GUI screenshot, and aia_i represents the action taken at step i. Given the explored observations along the trajectory, a Vision-Language Model ℳM generates a set of potential user tasks based on the visual input: i=ℳ(si),i∈0,…,n,Q_i=M(s_i), i∈\0,…,n\, where i=qi1,qi2,…,qikQ_i=\q_i^1,q_i^2,…,q_i^k\ denotes the set of generated tasks conditioned on observation sis_i. For each generated task, we preserve the corresponding environment e as the initial state, forming task–environment pairs: task=(q,e)∣q∈i.D_task=\(q,e) q _i\. The resulting pairs are used to initialize reinforcement learning training episodes. Trajectory Collection and Filtering. After task creation, we employ GPT-4o to execute these tasks within their associated environments in order to obtain trajectory data while simultaneously altering the app states. Due to the inherent limitations of the model, some generated tasks may not be executable under the current GUI context, and certain collected trajectories may contain mistakes or incomplete steps. To mitigate this issue, we adopt the filtering strategy ℱF proposed in OS-Genesis 26 to remove task–trajectory mismatches. As a result, we obtain triplets consisting of task q, initial environment e, and successful trajectory τ: traj=(q,e,τ)∣ℱ(q,τ).D_traj=\(q,e,τ) (q,τ)\. Progression from Simple to Complex. Although numerous tasks can be produced from individual states, we observe that the model struggles to generate tasks that are both sophisticated (e.g., involving multiple sub-goals or strict execution order) and diverse (e.g., similar tasks with varying parameters). To address this limitation, we apply a task complexification process to tasks that were successfully completed in the previous iteration. Given a completed source triplet (qs,es,τs)(q_s,e_s, _s), where qsq_s denotes the source task instruction, ese_s is the initial environment, and τs _s represents the successful execution trajectory, we generate a new task qcomplexifyq_complexify through one of three strategies—inheritance, composition, or rewriting: qcomplexify=ℳinherit(qs,τs),e=eT,ℳmerge(qs,τs),e=es,ℳrewrite(qs,τs),e=es.q_complexify= \ aligned &M_inherit(q_s, _s), e=e_T,\\ &M_merge(q_s, _s), e=e_s,\\ &M_rewrite(q_s, _s), e=e_s. aligned . Where eTe_T denotes the initial environment of the current iteration. ℳinheritM_inherit generates a follow-up task starting from eTe_T, ℳmergeM_merge augments the source task by combining it with additional subtasks starting from the original environment, and ℳrewriteM_rewrite produces a task variant by modifying specific parameters of the source task while preserving its overall structure. Through this process, the task distribution progressively evolves from simple tasks toward more complex and diverse workflows, generating increasingly challenging tasks across iterations. Further details and prompts for self-evolving data synthesis are provided in Appendix B and Appendix G. 3.2 State-Anchor Reward While self-evolving data synthesis enables large-scale task and trajectory generation, effective online RL further requires accurate and scalable reward supervision. To this end, we introduce a state-anchor mechanism that supports automatic annotation with minimal human intervention. Limitations in Action History Based Evaluation. In model-based reward evaluation, action history is often incorporated as contextual information to provide useful signals (e.g., the completion of intermediate steps), which helps the model better assess task progress. However, due to the limitations of current models, evaluators relying on historical context frequently suffer from false positives (FP). This issue is particularly pronounced in GUI scenarios, where visually similar states may correspond to semantically different task outcomes. Compared with rule-based approaches, model-based methods are less reliable because the model lacks explicit references of completed task states, making it difficult to determine whether a task has truly been accomplished. Inspired by this observation, we introduce goal-state references that explicitly represent the expected outcome of a task. Specifically, we annotate key elements within the goal state and use them as reference signals, providing the evaluator with a form of “ground-truth answer” that improves the accuracy of model-based reward estimation. Goal-State Acquisition. We obtain goal states from the self-evolving data synthesis stage by extracting the final screenshot of successful trajectories. Formally, given a successful task–trajectory triplet (q,e,τ)(q,e,τ) where τ=(s0,a0,…,sT)τ=(s_0,a_0,…,s_T), the goal state is defined as the terminal observation: sg=sT.s_g=s_T. For certain difficult tasks, models may fail to complete them during trajectory collection, leading to their removal during filtering. However, such tasks may still be beneficial for reinforcement learning. Therefore, for a subset of these challenging tasks, the final goal states are obtained through manual execution in order to provide reliable completion references. Automatic Reward Annotation via Goal-State Anchoring. After obtaining the goal state, we automatically identify UI elements relevant to the task objective and use them to construct a goal-state-anchor reference. These anchored elements provide a structured representation of the expected task outcome. Specifically, given the goal state sgs_g and its corresponding accessibility tree (a11y tree) AgA_g, we first overlay element indices from AgA_g onto the screenshot and then use ℳM to identify the indices of task-relevant elements: Kg=ℳ(ℐ(sg,Ag),Ag),K_g=M(I(s_g,A_g),A_g), where ℐI denotes the operation of overlaying element indices on the screenshot. The a11y tree provides semantic and spatial information for linking UI elements to their corresponding indices. We then retrieve the bounding boxes of the selected elements from AgA_g and use A to anchor their corresponding regions in the goal-state screenshot: G=(sg,bbox(Agk)∣k∈Kg).G=A (s_g,\bbox(A_g^k) k∈ K_g\ ). The resulting G serves as the goal-state-anchor reference for reliable task completion verification. Importantly, the entire process is human-free, enabling large-scale and continual reward annotation. Goal-State-Anchor Reward for Training. Once the anchored goal state is obtained, it can serve as a reference answer to assist the model in evaluating task completion. In addition, inspired by 12, we incorporate the action history as contextual information during evaluation. By combining the action history with the goal-state-anchor reference, the evaluator is provided with sufficient information to make more accurate judgments about task outcomes. Formally, at time step t, the evaluator receives the current GUI state sts_t, the action history τ0:t _0:t, and the goal-state-anchor label representation G, and outputs a binary reward indicating task completion: yt=ℰ(st,τ0:t,G),rt=1,if yt=1,0,otherwise,y_t=E(s_t, _0:t,G), r_t= cases1,&if y_t=1,\\ 0,&otherwise, cases where ℰE denotes the evaluation model, yt∈0,1y_t∈\0,1\ indicates whether the task goal has been achieved, rtr_t is the reward at step t. During training, this reward mechanism can be applied at any step of task execution to determine whether the desired outcome has been achieved, enabling reliable reward supervision for RL. 4 Experiments Model Data Sources AndroidControl-Low AndroidControl-High GUI-Odyssey TM EM TM EM TM EM OS-Genesis-7B Model 90.7 74.2 65.9 44.4 11.7 3.6 OS-Atlas-7B Human&Model 73.0 67.3 70.4 56.5 91.8* 76.8* Aguvis-7B Human&Model 93.9 89.4 65.6 54.2 26.7 13.5 Qwen2.5-VL-7B – 94.1 85.0 75.1 62.9 59.5 46.3 Ours (aw-app) Model 95.4 90.2 76.1 66.1 65.8 48.4 Ours (extra-app) Model 95.3 89.3 76.3 65.2 65.4 47.6 Table 1: Performance comparison on the AndroidControl and GUI-Odyssey benchmarks. Results are reported without the open action type. *OS-Atlas employs different train/test splits on GUI-Odyssey and is thus not directly comparable. Bold and underline indicate the best and second-best results. 4.1 Experiment Settings Benchmarks. We use two widely used benchmarks, the AndroidControl 14 and GUI-Odyssey 17, to evaluate the quality of our synthesized data. We report Type Match (TM) and Exact Match (EM) as the primary evaluation metrics. To evaluate the effectiveness of our GSAR for trajectory completion verification, we further construct an evaluation set from execution traces in AndroidWorld 22. Specifically, we collect more than 300300 trajectories containing both positive and negative examples, and report Accuracy and F1 score. To evaluate the effectiveness of the overall GSAR framework, we construct a benchmark consisting of 8686 queries with their corresponding initial environment snapshots and annotated reference goal states. Half of the queries are used as the training split. We report the Success Rate (SR) on our self-built benchmark, with results verified through both manual assessment and GSAR. Baselines. To evaluate the quality of the synthesized data, we selected three advanced open-source models, Os-Genesis 26, OS-Atlas 29, and Aguvis 33, for comparison with our fine-tuned model. Moreover, we compare our method with three representative model-based approaches for reward verification from prior work: DigiRL 1, DistRL 28, and StepCritic 12. Additionally, we include two baselines that use only the goal-state screenshot and only the anchored goal-state screenshot as input, denoted as GS-Only and GSA-Only. Training Details. For the supervised fine-tuning (SFT) experiments on synthesized trajectory data, we adopt Qwen2.5-VL-7B 2 as the base model and fine-tune it using LoRA 9. For model-based reward evaluation, we use several vision-language models as judge models, including Qwen3-VL-8B, Qwen3-VL-32B 34, GPT-4o 11, and Gemini-2.5-Pro 4. In the online reinforcement learning (RL) experiments, we adopt UI-TARS-7B-DPO 21 and GUI-Owl-7B 36 as the backbone models, and use GRPO 24 as the optimization algorithm, with Qwen3-VL-32B serving as the judge model. Detailed experimental settings are provided in Appendix C. Method Qwen3-VL-8B Qwen3-VL-32B GPT-4o Gemini-2.5-pro Avg Acc F1 Acc F1 Acc F1 Acc F1 Acc F1 DigiRL 67.0 64.9 71.4 65.9 63.8 56.2 71.1 64.3 68.3 62.8 DistRL 67.9 69.3 80.0 79.7 69.8 66.9 82.5 82.3 75.0 74.5 StepCritic 64.4 74.1 76.8 81.0 81.3 84.0 87.9 88.3 77.6 81.8 GS-Only 84.1 84.2 83.8 83.9 78.1 78.8 86.0 85.9 83.0 83.2 GSA-Only 87.6 88.2 86.4 86.9 81.3 80.9 89.8 90.1 86.3 86.5 GSAR 90.2 91.0 92.1 92.4 91.1 91.3 92.7 92.7 91.5 91.8 Table 2: Offline trajectory evaluation results of different model-based reward methods. 4.2 Main Results We first evaluate the fine-tuned model to validate data quality, then assess the accuracy of GSAR on offline trajectories, and finally verify that the entire pipeline scales effectively to online RL. Evaluation of Synthesized Data Quality. We evaluate the effectiveness and scalability of synthesized data for GUI agent training. The model is fine-tuned on data collected from 20 AndroidWorld apps and further scaled using data from 40 selected open-source apps. As shown in Table 1, fine-tuning on synthesized data leads to consistent improvements across both benchmarks, with notable gains on the EM metric for the low and high splits of AndroidControl (+5.2% and +3.2%, respectively) and on the TM metric for GUI-Odyssey (+6.3%). When scaling to additional open-source apps, the model achieves similarly strong performance, although slightly lower due to distribution differences between datasets. These results indicate that our synthesized data effectively improves model performance and generalizes well across diverse application domains. Additional results on other base models are provided in Appendix D. Offline Evaluation for GSAR. We conduct offline evaluations to assess the ability of GSAR in judging task completion, comparing it with three model-based reward methods as shown in Table 2. As model capabilities continue to improve, the accuracy and F1 scores of the three baseline methods also increase. However, they still lag behind the rule-based approach, whose accuracy and F1 are theoretically 100%. Providing the model with a reference state (GS-Only or GSA-Only) significantly enhances its evaluation capability. GSAR achieves the highest results in evaluating trajectory correctness, being the only method to surpass 90% in both accuracy and F1 for both open-source and closed-source models. This demonstrates that GSAR most closely approximates rule-based evaluation while retaining scalability advantages. RL Training with the Full Framework. (a) UI-TARS-7B-DPO (b) GUI-Owl-7B Figure 3: Success rate on the self-built benchmark. Model indicates that GSAR is used as the evaluation method. To further assess the effectiveness and efficiency of GSAR, we use our self-built benchmark, including triplets (task, initial environment, goal-state-anchor reference), for online RL. As shown in Figure 3, the model trained with GSAR as the reward function achieves a 23.2% improvement on the training split and an 8.1% improvement on the full task set compared to the UI-TARS. A similar trend is observed for GUI-Owl, with gains of 18.6% and 8.2%, respectively. Moreover, incorporating the goal-state-anchor reference substantially enhances task execution over StepCritic, while adding action history also improves training compared to GSA-Only. Verification using GSAR shows similar trends to human evaluation, with some metrics closely matching human judgments. Overall, the proposed pipeline effectively enables scalable RLVR training in GUI agents and improves their performance. (a) Reward on UI-TARS. (b) Reward on GUI-Owl. Figure 4: Online training reward curves of three reward mechanisms on the UI-TARS-7B-DPO and GUI-Owl-7B models. Each curve represents the per-step reward across training steps. 4.3 Ablation Study. We perform ablation experiments on offline verification to assess the contribution of each component. Specifically, we test three variants: (1) w/o history, which removes the action history; (2) w/o anchor, which removes the anchor on the goal-state screenshot; and (3) w/o reference, which removes the goal-state screenshot. Method Qwen3-VL-8B Qwen3-VL-32B Acc F1 Acc F1 w/o reference 64.4 74.1 76.8 81.0 w/o anchor 87.3 88.0 90.5 90.7 w/o history 87.6 88.2 86.4 86.9 GSAR (Ours) 90.2 91.0 92.1 92.4 Table 3: The results of the ablation study. As shown in Table 3, combining the action history with the anchored reference screenshot as input yields the best performance, and removing either component leads to a noticeable degradation. Omitting the reference image leads to substantially lower accuracy compared to the rule-based approach, as the model loses information about the final task state. Excluding the anchor annotations from the completed reference also degrades performance, since VLMs may fail to perceive subtle visual differences between similar incomplete and completed screenshots. Additionally, because some tasks involve critical intermediate steps, discarding the action history removes this contextual information, negatively affecting the final decision. 5 Analysis Experiments 5.1 Annotation Accuracy by Task Category Type Answer Delete Normal Total Count 20 15 47 82 Accuracy 90.0 100.0 89.4 91.5 Table 4: Automatic annotation accuracy of goal-state anchoring across task categories. Table 4 summarizes the accuracy of automatic goal-state annotation via GPT-4o across task categories. The overall annotation accuracy reaches 91.5%. Delete tasks achieve 100% accuracy because the deleted object no longer appears in the goal state, eliminating the need for fine-grained localization; therefore, these cases are treated as fully correct. For Answer and Normal tasks, accuracy is slightly lower, as the goal state may contain multiple small or scattered key elements that the model must individually identify and localize, posing a greater challenge. Overall, the automatic annotation pipeline provides reliable annotations across different task categories and can substantially reduce the need for manual annotation. Figure 5: Comparison of false positives (FP) and false negatives (FN) across different reward designs. StepCritic reduces FN but leaves FP high, GSA-Only reduces FP but some FN remain, and combining both mechanisms balances the reward signal while mitigating both FP and FN errors. 5.2 Understanding the Role of Reward Signals in RL Training We investigate the impact of reward signals on reinforcement learning by comparing the performance of models trained with different reward functions. Although StepCritic exhibits relatively low false negative rates in offline evaluation, it suffers from a notably high false positive rate (Figure 5). As a result, its reward signals during RL training are overly optimistic (Figure 4), leading to policies that fail to consistently translate into strong online performance and producing agents that are less stable and less reliable than ours (Figure 3). This suggests that providing more accurate reward signals can lead to greater gains in RL training. 5.3 Comparison with Rule-based Methods (a) Evaluation results (b) Reward curves Figure 6: Evaluation results on AndroidWorld and online training reward curves of the UI-TARS model. Performance Comparison. We conduct a comparative study between GSAR and rule-based rewards on AndroidWorld 22. We select a subset of 42 tasks as the training set, while evaluation is conducted on the full benchmark. As shown in Figure 6(a), both GSAR and rule-based rewards improve the success rate over the UI-TARS-7B-DPO baseline (26.7%). GSAR increases the performance to 30.2% (+3.5%), while rule-based rewards achieve 32.8% (+6.1%) due to their near-noiseless verification signals. Meanwhile, the reward curves in Fig. 6(b) show that GSAR follows a trend highly consistent with rule-based rewards, demonstrating that it can provide stable optimization signals without relying on manually designed rules, offering better scalability and generalization. Reward Method Average Rollout Time (s) Rule-based 2754.92 GSAR (Model-based) 2342.02 Table 5: Average rollout time per training step for different reward methods. Rollout Efficiency. Both methods perform reward evaluation on a step-wise basis during RL training. GSAR introduces additional latency due to model-based evaluation after each action. However, for most tasks, the latency introduced by model inference is lower than the overhead incurred by repeated ADB-based device access for rule-based verification. As shown in Table 5, GSAR therefore achieves lower overall rollout latency than the rule-based baseline. 6 Conclusion We propose GSAR (Goal-State-Anchor Reward), a framework that integrates self-evolving data synthesis with a state-anchor mechanism. To enhance environmental diversity during data synthesis, we introduce an iterative mechanism that evolves environments through task execution and incorporate task complexification to generate diverse tasks. Furthermore, we leverage final states of successful trajectories to anchor task-relevant UI elements, providing richer context for model-based reward evaluation. Our experiments demonstrate that GSAR achieves over 90% accuracy in trajectory verification and improves performance over the base model in online RL training. Moreover, the reward curve of GSAR closely aligns with the rule-based reward, demonstrating its ability to provide stable optimization signals without relying on handcrafted rules. Overall, GSAR provides a scalable data synthesis paradigm and provides a reliable reward framework for training GUI agents. Limitations Our method addresses the scalability issues of rule-based approaches and the accuracy limitations of model-based approaches. However, it considers only a single final state as the success criterion for each task, whereas a task may be achievable through multiple alternative solutions. Moreover, for question-answering tasks, our method may assign a correct reward before the actual correct answer is produced. For tasks whose completion cannot be fully expressed visually (e.g., deleting certain items), it is difficult for the model to annotate reasonable task-relevant elements, and natural language or other modalities may be necessary to describe the final state. Future work should involve a more fine-grained categorization and handling of task types during data collection and training. Meanwhile, both the automated task execution for trajectory collection and the filtering process are performed by VLMs (e.g., GPT-4o), eliminating the need for manual annotation. However, the quality of the generated trajectories is sometimes constrained by the capabilities of the models. For tasks that cannot be completed automatically, manual execution is still required to obtain the correct goal state if they are to be used for reinforcement learning. We expect that stronger GUI-capable models in the future will further improve the utilization of model-generated data. References Bai et al. (2024) H. Bai, Y. Zhou, M. Cemri, J. Pan, A. Suhr, S. Levine, and A. Kumar DigiRL: training in-the-wild device-control agents with autonomous reinforcement learning. ArXiv abs/2406.11896. External Links: Link Cited by: §1, §1, §4.1. Bai et al. (2025) S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin Qwen2.5-vl technical report. ArXiv abs/2502.13923. External Links: Link Cited by: Appendix D, §1, §4.1. Chen et al. (2025) Y. Chen, Y. Liu, L. Zhang, P. Gao, J. Luan, and W. Liu STEP: success-rate-aware trajectory-efficient policy optimization. ArXiv abs/2511.13091. External Links: Link Cited by: §2.2. Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al. Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Cited by: §E.2, §4.1. DeepSeek-AI et al. (2025) DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang DeepSeek-r1 incentivizes reasoning in llms through reinforcement learning. Nature 645, p. 633 – 638. External Links: Link Cited by: §2.2. Deka et al. (2017) B. Deka, Z. Huang, C. Franzen, J. Hibschman, D. Afergan, Y. Li, J. Nichols, and R. Kumar Rico: a mobile app dataset for building data-driven design applications. Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology. External Links: Link Cited by: §2.1. Feng et al. (2025) L. Feng, Z. Xue, T. Liu, and B. An Group-in-group policy optimization for llm agent training. ArXiv abs/2505.10978. External Links: Link Cited by: §1. Gao et al. (2024) J. Gao, S. Xu, W. Ye, W. Liu, C. He, W. Fu, Z. Mei, G. Wang, and Y. Wu On designing effective rl reward at training time for llm reasoning. ArXiv abs/2410.15115. External Links: Link Cited by: §2.2. Hu et al. (2021) J. E. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, and W. Chen LoRA: low-rank adaptation of large language models. ArXiv abs/2106.09685. External Links: Link Cited by: §4.1. Huang et al. (2025) W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Z. Xu, Y. Hu, and S. Lin Vision-r1: incentivizing reasoning capability in multimodal large language models. ArXiv abs/2503.06749. External Links: Link Cited by: §2.2. Hurst et al. (2024) O. A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, A. Mkadry, A. Baker-Whitcomb, A. Beutel, A. Borzunov, A. Carney, A. Chow, A. Kirillov, A. Nichol, A. Paino, A. Renzin, A. Passos, A. Kirillov, A. Christakis, A. Conneau, A. Kamali, A. Jabri, A. Moyer, A. Tam, A. Crookes, A. Tootoochian, A. Tootoonchian, A. Kumar, A. Vallone, A. Karpathy, A. Braunstein, A. Cann, A. Codispoti, A. Galu, A. Kondrich, A. Tulloch, A. Mishchenko, A. Baek, A. Jiang, A. Pelisse, A. Woodford, A. Gosalia, A. Dhar, A. Pantuliano, A. Nayak, A. Oliver, B. Zoph, B. Ghorbani, B. Leimberger, B. Rossen, B. Sokolowsky, B. Wang, B. Zweig, B. Hoover, B. Samic, B. McGrew, B. Spero, B. Giertler, B. Cheng, B. Lightcap, B. Walkin, B. Quinn, B. Guarraci, B. Hsu, B. Kellogg, B. Eastman, C. Lugaresi, C. L. Wainwright, C. Bassin, C. Hudson, C. Chu, C. Nelson, C. Li, C. J. Shern, C. Conger, C. Barette, C. Voss, C. Ding, C. Lu, C. Zhang, C. Beaumont, C. Hallacy, C. Koch, C. Gibson, C. Kim, C. Choi, C. McLeavey, C. Hesse, C. Fischer, C. Winter, C. Czarnecki, C. Jarvis, C. Wei, C. Koumouzelis, D. Sherburn, D. Kappler, D. Levin, D. Levy, D. Carr, D. Farhi, D. Mély, D. Robinson, D. Sasaki, D. Jin, D. Valladares, D. Tsipras, D. Li, P. D. Nguyen, D. Findlay, E. Oiwoh, E. Wong, E. Asdar, E. Proehl, E. Yang, E. Antonow, E. Kramer, E. Peterson, E. Sigler, E. Wallace, E. Brevdo, E. Mays, F. Khorasani, F. P. Such, F. Raso, F. Zhang, F. von Lohmann, F. Sulit, G. Goh, G. Oden, G. Salmon, G. Starace, G. Brockman, H. Salman, H. Bao, H. Hu, H. Wong, H. Wang, H. Schmidt, H. Whitney, H. Jun, H. Kirchner, H. P. de Oliveira Pinto, H. Ren, H. Chang, H. W. Chung, I. Kivlichan, I. O’Connell, I. Osband, I. Silber, I. Sohl, I. Okuyucu, I. Lan, I. Kostrikov, I. Sutskever, I. Kanitscheider, I. Gulrajani, J. Coxon, J. Menick, J. W. Pachocki, J. Aung, J. Betker, J. Crooks, J. Lennon, J. R. Kiros, J. Leike, J. Park, J. Kwon, J. Phang, J. Teplitz, J. Wei, J. Wolfe, J. Chen, J. Harris, J. Varavva, J. G. Lee, J. Shieh, J. Lin, J. Yu, J. Weng, J. Tang, J. Yu, J. Jang, J. Q. Candela, J. Beutler, J. Landers, J. Parish, J. Heidecke, J. Schulman, J. Lachman, J. Mckay, J. Uesato, J. Ward, J. W. Kim, J. Huizinga, J. Sitkin, J. Kraaijeveld, J. Gross, J. Kaplan, J. Snyder, J. Achiam, J. Jiao, J. Lee, J. Zhuang, J. Harriman, K. Fricke, K. Hayashi, K. Singhal, K. Shi, K. Karthik, K. Wood, K. Rimbach, K. Hsu, K. Nguyen, K. Gu-Lemberg, K. Button, K. Liu, K. Howe, K. Muthukumar, K. Luther, L. Ahmad, L. Kai, L. Itow, L. Workman, L. Pathak, L. Chen, L. Jing, L. Guy, L. Fedus, L. Zhou, L. Mamitsuka, L. Weng, L. McCallum, L. Held, O. Long, L. Feuvrier, L. Zhang, L. Kondraciuk, L. Kaiser, L. Hewitt, L. Metz, L. Doshi, M. Aflak, M. Simens, M. Boyd, M. Thompson, M. Dukhan, M. Chen, M. Gray, M. Hudnall, M. Zhang, M. Aljubeh, M. Litwin, M. Zeng, M. Johnson, M. Shetty, M. M. Gupta, M. Shah, M. A. Yatbaz, M. Yang, M. Zhong, M. Glaese, M. Chen, M. Janner, M. Lampe, M. Petrov, M. Wu, M. Wang, M. Fradin, M. Pokrass, M. Castro, M. Castro, M. Pavlov, M. Brundage, M. Wang, M. Khan, M. Murati, M. Bavarian, M. Lin, M. Yesildal, N. Soto, N. Gimelshein, N. Cone, N. M. Staudacher, N. Summers, N. LaFontaine, N. Chowdhury, N. Ryder, N. Stathas, N. Turley, N. A. Tezak, N. Felix, N. Kudige, N. S. Keskar, N. Deutsch, N. Bundick, N. Puckett, O. Nachum, O. Okelola, O. Boiko, O. Murk, O. Jaffe, O. Watkins, O. Godement, O. Campbell-Moore, P. Chao, P. McMillan, P. Belov, P. Su, P. Bak, P. Bakkum, P. Deng, P. Dolan, P. Hoeschele, P. Welinder, P. Tillet, P. Pronin, P. Tillet, P. Dhariwal, Q. Yuan, R. Dias, R. Lim, R. Arora, R. Troll, R. Lin, R. G. Lopes, R. Puri, R. Miyara, R. H. Leike, R. Gaubert, R. Zamani, R. B. Wang, R. Donnelly, R. Honsby, R. Smith, R. Sahai, R. Ramchandani, R. Huet, R. Carmichael, R. Zellers, R. Chen, R. Chen, R. R. Nigmatullin, R. Cheu, S. Jain, S. Altman, S. Schoenholz, S. Toizer, S. Miserendino, S. Agarwal, S. Culver, S. Ethersmith, S. Gray, S. Grove, S. Metzger, S. Hermani, S. Jain, S. Zhao, S. Wu, S. Jomoto, S. Wu, S. Xia, S. Phene, S. Papay, S. Narayanan, S. Coffey, S. Lee, S. Hall, S. Balaji, T. Broda, T. Stramer, T. Xu, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Cunninghman, T. Degry, T. Dimson, T. Raoux, T. Shadwell, T. Zheng, T. Underwood, T. Markov, T. Sherbakov, T. Rubin, T. Stasi, T. Kaftan, T. Heywood, T. A. Peterson, T. Walters, T. Eloundou, V. Qi, V. Moeller, V. Monaco, V. Kuo, V. Fomenko, W. H. Chang, W. Zheng, W. Zhou, W. Manassra, W. Sheu, W. Zaremba, Y. Patil, Y. Qian, Y. Kim, Y. Cheng, Y. Zhang, Y. He, Y. Zhang, Y. Jin, Y. Dai, and Y. Malkov GPT-4o system card. External Links: Link Cited by: Appendix B, §1, §2.1, §4.1. Lai et al. (2025) H. Lai, J. Gao, X. Liu, Y. Xu, S. Zhang, Y. Dong, and J. Tang AndroidGen: building an android language agent under data scarcity. ArXiv abs/2504.19298. External Links: Link Cited by: §3.2, §4.1. Lambert et al. (2024) N. Lambert, J. D. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, X. Lyu, Y. Gu, S. Malik, V. Graf, J. D. Hwang, J. Yang, R. L. Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y. Wang, P. Dasigi, and H. Hajishirzi TÜlu 3: pushing frontiers in open language model post-training. ArXiv abs/2411.15124. External Links: Link Cited by: §1. Li et al. (2024) W. Li, W. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva On the effects of data scale on ui control agents. Advances in Neural Information Processing Systems 37. External Links: Link Cited by: Appendix A, §2.1, §4.1. Lin et al. (2025) M. Lin, M. Liu, T. Lu, L. Yuan, Y. Liu, H. Xu, Y. Miao, Y. Chao, and Z. Li GUI-rewalk: massive data generation for gui agent via stochastic exploration and intent-aware reasoning. ArXiv abs/2509.15738. External Links: Link Cited by: §1. Liu et al. (2025) Y. Liu, P. Li, C. Xie, X. Hu, X. Han, S. Zhang, H. Yang, and F. Wu InfiGUI-r1: advancing multimodal gui agents from reactive actors to deliberative reasoners. ArXiv abs/2504.14239. External Links: Link Cited by: §2.2. Lu et al. (2024) Q. Lu, W. Shao, Z. Liu, F. Meng, B. Li, B. Chen, S. Huang, K. Zhang, Y. Qiao, and P. Luo GUI odyssey: a comprehensive dataset for cross-app gui navigation on mobile devices. ArXiv abs/2406.08451. External Links: Link Cited by: Appendix A, §2.1, §4.1. Lu et al. (2025) Z. Lu, Y. Chai, Y. Guo, X. Yin, L. Liu, H. Wang, G. Xiong, and H. Li UI-r1: enhancing efficient action prediction of gui agents by reinforcement learning. External Links: Link Cited by: §2.2. Luo et al. (2025) R. Luo, L. Wang, W. He, and X. Xia GUI-r1 : a generalist r1-style vision-language action model for gui agents. ArXiv abs/2504.10458. External Links: Link Cited by: §2.2. Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. E. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. J. Lowe Training language models to follow instructions with human feedback. ArXiv abs/2203.02155. External Links: Link Cited by: §2.1. Qin et al. (2025) Y. Qin, Y. Ye, J. Fang, H. Wang, S. Liang, S. Tian, J. Zhang, J. Li, Y. Li, S. Huang, W. Zhong, K. Li, J. Yang, Y. Miao, W. Lin, L. Liu, X. Jiang, Q. Ma, J. Li, X. Xiao, K. Cai, C. Li, Y. Zheng, C. Jin, C. Li, X. Zhou, M. Wang, H. Chen, Z. Li, H. Yang, H. Liu, F. Lin, T. Peng, X. Liu, and G. Shi UI-tars: pioneering automated gui interaction with native agents. ArXiv abs/2501.12326. External Links: Link Cited by: Appendix D, §4.1. Rawles et al. (2024) C. Rawles, S. Clinckemaillie, Y. Chang, J. Waltz, G. R. Lau, M. Fair, A. Li, W. Bishop, W. Li, F. Campbell-Ajala, D. Toyama, R. Berry, D. Tyamagundlu, T. P. Lillicrap, and O. Riva AndroidWorld: a dynamic benchmarking environment for autonomous agents. ArXiv abs/2405.14573. External Links: Link Cited by: §1, §2.2, §4.1, §5.3. Rawles et al. (2023) C. Rawles, A. Li, D. Rodríguez, O. Riva, and T. P. Lillicrap Android in the wild: a large-scale dataset for android device control. ArXiv abs/2307.10088. External Links: Link Cited by: §2.1. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. K. Li, Y. Wu, and D. Guo DeepSeekMath: pushing the limits of mathematical reasoning in open language models. ArXiv abs/2402.03300. External Links: Link Cited by: §1, §4.1. Shi et al. (2025) Y. Shi, W. Yu, Z. Li, Y. Wang, H. Zhang, N. Liu, H. Mi, and D. Yu MobileGUI-rl: advancing mobile gui agent through reinforcement learning in online environment. ArXiv abs/2507.05720. External Links: Link Cited by: §2.2. Sun et al. (2024) Q. Sun, K. Cheng, Z. Ding, C. Jin, Y. Wang, F. Xu, Z. Wu, C. Jia, L. Chen, Z. Liu, B. Kao, G. Li, J. He, Y. Qiao, and Z. Wu OS-genesis: automating gui agent trajectory construction via reverse task synthesis. ArXiv abs/2412.19723. External Links: Link Cited by: §1, §2.1, §3.1, §4.1. Wang et al. (2025) H. Wang, H. Zou, H. Song, J. Feng, J. Fang, J. Lu, L. Liu, Q. Luo, S. Liang, S. Huang, W. Zhong, Y. Ye, Y. Qin, Y. Xiong, Y. Song, Z. Wu, B. Li, C. Dun, C. Liu, F. Leng, H. Wang, H. Yu, H. Chen, H. Guo, J. Su, J. Huang, K. Shen, K. Shi, L. Yan, P. Zhao, P. Liu, Q. Ye, R. Zheng, W. X. Zhao, W. Heng, W. Huang, W. Wang, X. Qin, Y. Lin, Y. Wu, Z. Chen, Z. Wang, B. Zhong, X. Zhang, X. Li, Y. Li, Z. Zhao, C. Jiang, F. Wu, H. Zhou, J. Pang, L. Han, Q. Ma, S. Liu, S. Cai, W. Fu, X. Liu, Z. Zhang, B. Zhou, G. Li, J. Shi, J. Yang, J. Tang, L. Li, T. Lu, W. Lin, X. Tong, X. Li, Y. Zhang, Y. Miao, Z. Jiang, Z. Li, Z. Zhao, C. Li, D. Ma, F. Lin, G. Zhang, H. Yang, H. Guo, H. Zhu, J. Liu, J. Du, K. Cai, K. Li, L. Yuan, M. Han, M. Wang, S. Guo, T. Cheng, X. Ma, X. Xiao, X. Huang, X. Chen, Y. Du, Y. Chen, Y. Wang, Z. Li, Z. Yang, Z. Zeng, C. Jin, C. Li, H. Chen, H. Chen, J. Chen, Q. Zhao, and G. Shi UI-tars-2 technical report: advancing gui agent with multi-turn reinforcement learning. ArXiv abs/2509.02544. External Links: Link Cited by: §1. Wang et al. (2024) T. Wang, Z. Wu, J. Liu, J. Hao, J. Wang, and K. Shao DistRL: an asynchronous distributed reinforcement learning framework for on-device control agents. ArXiv abs/2410.14803. External Links: Link Cited by: §1, §4.1. Wu et al. (2024) Z. Wu, Z. Wu, F. Xu, Y. Wang, Q. Sun, C. Jia, K. Cheng, Z. Ding, L. Chen, P. P. Liang, and Y. Qiao OS-atlas: a foundation action model for generalist gui agents. ArXiv abs/2410.23218. External Links: Link Cited by: §4.1. Xie et al. (2025) B. Xie, R. Shao, G. Chen, K. Zhou, Y. Li, J. Liu, M. Zhang, and L. Nie GUI-explorer: autonomous exploration and mining of transition-aware knowledge for gui agent. In Annual Meeting of the Association for Computational Linguistics, External Links: Link Cited by: §2.1. Xie et al. (2024) T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, T. J. Hua, Z. Cheng, D. Shin, F. Lei, Y. Liu, Y. Xu, S. Zhou, S. Savarese, C. Xiong, V. Zhong, and T. Yu OSWorld: benchmarking multimodal agents for open-ended tasks in real computer environments. ArXiv abs/2404.07972. External Links: Link Cited by: §1. Xu et al. (2025) Y. Xu, X. Liu, X. Liu, J. Fu, H. Zhang, B. Jing, S. Zhang, Y. Wang, W. Zhao, and Y. Dong MobileRL: online agentic reinforcement learning for mobile gui agents. ArXiv abs/2509.18119. External Links: Link Cited by: §1. Xu et al. (2024) Y. Xu, Z. Wang, J. Wang, D. Lu, T. Xie, A. Saha, D. Sahoo, T. Yu, and C. Xiong Aguvis: unified pure vision agents for autonomous gui interaction. ArXiv abs/2412.04454. External Links: Link Cited by: §4.1. Yang et al. (2025a) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. External Links: Link Cited by: §4.1. Yang et al. (2025b) C. Yang, S. Su, S. Liu, X. Dong, Y. Yu, W. Su, X. Wang, Z. Liu, J. Zhu, H. Li, W. Wang, Y. Qiao, X. Zhu, and J. Dai ZeroGUI: automating online gui learning at zero human cost. ArXiv abs/2505.23762. External Links: Link Cited by: §2.2. Ye et al. (2025) J. Ye, X. Zhang, H. Xu, H. Liu, J. Wang, Z. Zhu, Z. Zheng, F. Gao, J. Cao, Z. Lu, J. Liao, Q. Zheng, F. Huang, J. Zhou, and M. Yan Mobile-agent-v3: fundamental agents for gui automation. ArXiv abs/2508.15144. External Links: Link Cited by: §1, §4.1. Yu et al. (2025) Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y. Tong, C. Zhang, M. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, W. Dai, Y. Song, X. Wei, H. Zhou, J. Liu, W. Ma, Y. Zhang, L. Yan, M. Qiao, Y. Wu, and M. Wang DAPO: an open-source llm reinforcement learning system at scale. ArXiv abs/2503.14476. External Links: Link Cited by: §1. Yuan et al. (2025) X. Yuan, J. Zhang, K. Li, Z. Cai, L. Yao, J. Chen, E. Wang, Q. Hou, J. Chen, P. Jiang, and B. Li Enhancing visual grounding for gui agents via self-evolutionary reinforcement learning. ArXiv abs/2505.12370. External Links: Link Cited by: §2.2. Zhang et al. (2025) Z. Zhang, Y. Lu, Y. Fu, Y. Huo, S. Yang, Y. Wu, H. Si, X. Cong, H. Chen, Y. Lin, J. Xie, W. Zhou, W. Xu, Y. Zhang, Z. Su, Z. Zhai, X. Liu, Y. Mei, J. Xu, H. Tian, C. Wang, C. Chen, Y. Yao, Z. Liu, and M. Sun AgentCPM-gui: building mobile-use agents with reinforcement fine-tuning. ArXiv abs/2506.01391. External Links: Link Cited by: Appendix A, §2.2. Zhou et al. (2025) Y. Zhou, S. Dai, S. Wang, K. Zhou, Q. Jia, and J. Xu GUI-g1: understanding r1-zero-like training for visual grounding in gui agents. ArXiv abs/2505.15810. External Links: Link Cited by: §2.2. Appendix A Benchmarks Here we provide additional details on the benchmarks used to evaluate GSAR. AndroidControl. AndroidControl 14 is a large-scale dataset for human-computer interface control on Android devices, containing 15,283 task demonstrations across 833 apps. Each task includes both high-level and low-level human-generated instructions, covering 14,548 unique tasks. The dataset enables analysis of agent performance on tasks of varying complexity, both within the training domain and out-of-domain. AndroidControl supports studies on how fine-tuning with more data affects performance, particularly highlighting the challenges of generalizing to unseen apps or higher-level tasks. GUI-Odyssey. GUI-Odyssey 17 is a large-scale dataset for cross-app mobile GUI navigation, consisting of 8,334 episodes with an average of 15.3 steps per episode. It spans 6 mobile devices, 212 apps, and 1,357 app combinations. Each step includes detailed semantic reasoning annotations to support models in complex cross-app reasoning and decision-making. The dataset enables training and evaluation of agents capable of long-step, multi-app navigation tasks. Following 39, we removed steps related to the “open_app” action type in the evaluation of these benchmarks. Appendix B Self-Evolving Data Synthesis In this section, we present the algorithmic workflow of the self-evolving data synthesis framework used for generating tasks and collecting trajectories in mobile GUI environments. The framework iteratively collects tasks, evolves the GUI environment, and generates increasingly complex instructions based on previously collected trajectories. The workflow, summarized in Algorithm 1, illustrates how the system produces a diverse and high-quality dataset suitable for reinforcement learning training and goal-state–anchored reward evaluation. Algorithm 1 Self-Evolving Data Synthesis 1: Device serial d, package name p, maximum iterations T, 2: maximum tasks per page MtM_t, maximum breadth MbM_b, maximum depth MdM_d, 3: maximum steps per task SmaxS_ 4: Task set Q and trajectory set D 5: Initialize task set ←∅Q← 6: Initialize trajectory set ←∅D← 7: Initialize previous trajectory set prev←∅D_prev← 8: for iteration t=1…Tt=1… T do 9: Pull environment EtE_t from device 10: Collect candidate tasks t=CollectTask(Et,Mt,Mb,Md)Q_t= CollectTask(E_t,M_t,M_b,M_d) 11: if t>1t>1 and prev≠∅D_prev≠ then 12: for each trajectory τ∈prevτ _prev do 13: if τ is completed and not complexified then 14: Generate complexified tasks ′=ComplexifyTask(τ,Et)Q = ComplexifyTask(τ,E_t) 15: t←t∪′Q_t _t 16: end if 17: end for 18: end if 19: Update task set ←∪tQ _t 20: Push environment EtE_t back to device 21: Execute tasks tQ_t with step limit SmaxS_ 22: Collect resulting trajectories tD_t 23: Update trajectory set ←∪tD _t 24: Set prev←tD_prev _t 25: end for 26: return ,Q,D We use GPT-4o 11 for task generation, task complexification, and task execution. For AndroidWorld apps, we set the maximum iteration to 6. For additional open-source apps, we first allow the model to perform a single random exploration to generate trajectories, and then count the number of unique pages in the trajectory screenshots, mapping this number to a range of 1–6 to determine the maximum number of self-evolving iterations for that app. This approach ensures that apps with more diverse interactions undergo more iterations, while avoiding unnecessary exploration time for simpler apps. Appendix C Details of Experiments Supervised Fine-Tuning. We provide the detailed experimental settings for the supervised fine-tuning (SFT) stage in Table 6, including the training data, model configuration, and key hyperparameters used during training. For all fine-tuned models, the training dataset contains 6,331 steps, covering both low-level and high-level instruction types, where the low-level instructions correspond to the action reasoning output. Hyperparameters All methods Fine-tuning method LoRA LoRA rank (r) 8 LoRA α 16 Train batch size 128 Training epochs 3 Learning rate 1×10−51× 10^-5 LR Scheduler Cosine Warmup Ratio 0.1 GPU numbers 8 Table 6: Training hyperparameters for SFT. Offline Verification for GSAR. The datasets used to evaluate GSAR and other methods are derived from the evaluation trajectories of GUI-Owl-7B and Qwen2.5-VL-7B on the AndroidWorld benchmark. To better assess the differences between methods, we constructed a total of 315 trajectories, including 164 positive samples and 151 negative samples. For each task, the goal state is selected from the last step screenshot and anchored using GPT-4o. To ensure fairness, we strictly follow the experimental settings and prompts described in the original papers when using these baselines. Hyperparameters All methods Train batch size 32 PPO batch size 256 Training steps 80 Max turn 20 Rollout numbers 8 Temperature 0.7 Learning rate 1×10−61× 10^-6 KL coefficient 0.001 Clip ratio 0.2 Table 7: Training hyperparameters for RL. Model Data Sources AndroidControl-Low AndroidControl-High GUI-Odyssey TM EM TM EM TM EM UI-TARS-7B-SFT – 95.05 91.52 79.10 71.21 80.71 66.98 GSAR (aw-app) Model 94.99 91.53 80.33 71.97 81.78 68.01 GSAR (extra-app) Model 95.01 91.48 80.76 72.41 81.90 67.93 UI-TARS-7B-DPO – 92.05 88.74 80.77 73.56 85.29 68.33 GSAR (aw-app) Model 92.82 89.52 81.04 73.85 85.56 71.28 GSAR (extra-app) Model 93.33 90.01 81.28 74.09 85.90 71.15 Table 8: Performance comparison of UI-TARS-7B-SFT and UI-TARS-7B-DPO fine-tuned on our synthesized data, evaluated on the AndroidControl and GUI-Odyssey benchmarks. Results are reported without the open action type. Bold and underline indicate the best and second-best results within each block. Online Reinforcement Learning. The experimental configurations for the reinforcement learning (RL) stage are summarized in Table 7, including the backbone agent, reward evaluation model, optimization algorithm, and other key training parameters used during online RL. The experiments were conducted on two nodes, each equipped with 8×NVIDIA H100 GPUs8×NVIDIA H100 GPUs. To support large-scale environment interactions while decoupling the GUI environment from the training process, we deployed a service of 128 Android emulators for parallel trajectory sampling. Specifically, each of the two machines launched 64 emulators, providing a total of 128 emulators that enable the agent to interact with the environment independently of the training, thereby facilitating efficient training and data collection. Appendix D Supplementary SFT Results To further verify the generalization of our synthesized data across different base models, we fine-tune both UI-TARS-7B-SFT and UI-TARS-7B-DPO 21 on the same training data, as shown in Table 8. Across all three benchmarks, our method performs on par with, and in several cases surpasses, the strong UI-TARS baseline, and scaling to additional open-source apps yields comparable results. We note that the margins of improvement on UI-TARS are narrower than those observed on Qwen2.5-VL-7B 2 (Table 1), likely because UI-TARS is already specialized for GUI tasks and achieves substantially stronger baseline performance, leaving less room for further gains. These results confirm that the quality of our synthesized data is not tied to a specific base model, and the benefits transfer across different model initializations and capabilities. Appendix E Additional Analysis Results E.1 Analyzing the Effectiveness of Self-Evolving Data Synthesis Stage Edge Density Entropy UI Elements Early (Iter 1–2) 0.0240 1.90 22.68 Later (Iter 5–6) 0.0295 2.11 24.56 Table 9: Impact of self-evolving process on GUI page complexity. Task Type Task string length Trajectory length Normal 144.41 7.91 Complexified 338.92 12.08 Table 10: Effect of task complexity on average query string length and trajectory length. Dataset R=1 R=2 R=3 R=4 R=5 R≥4R≥ 4 Ratio aw-app 44 281 214 126 328 45.7% extra-app 110 342 196 83 229 32.5% Table 11: Data utilization statistics of synthesized trajectories. Impact of Self-Evolving Mechanism. To evaluate the effect of the self-evolving mechanism on page complexity, we sample 300 tasks each from early iterations and later iterations, as summarized in Table 9. For each task, one state is randomly sampled from its trajectory to compute page complexity metrics. The results show that pages from later iterations exhibit slightly higher edge density, increased image entropy, and more UI elements on average. These observations suggest that the self-evolving process gradually increases the visual and functional complexity of generated pages, leading to richer and more diverse environments. Effect of Task Complexity. We evaluate the impact of task complexification on a filtered set of 339 samples for each setting, as shown in Table 10. After task complexification, both the average length of natural language task descriptions and the average trajectory length increase, demonstrating the effectiveness of the task complexification process. Data Utilization Analysis We analyze the data utilization rate of trajectories generated by our self-evolving data synthesis pipeline, as shown in Table 11. The aw-app and extra-app datasets contain 993 and 960 tasks, respectively, and trajectories with reward score R≥4R≥ 4 are used for fine-tuning. The results show that a substantial portion of generated trajectories can be effectively utilized for training. The aw-app dataset achieves a higher utilization rate of 45.7%, benefiting from the carefully selected AndroidWorld applications with stable and complete functionalities. In contrast, the extra-app dataset has a lower utilization rate of 32.5%, since it consists of open-source applications, some of which have incomplete functionality or belong to relatively niche usage scenarios. E.2 Offline Evaluation by Task Category Method Delete Answer Normal Overall Acc F1 Acc F1 Acc F1 Acc F1 DigiRL 81.4 80.0 41.4 0.0 79.0 75.5 71.1 64.3 DistRL 86.4 86.2 75.7 80.0 83.9 82.1 82.5 82.3 StepCritic 96.6 96.7 77.1 77.8 89.3 89.7 87.9 88.3 GS-Only 91.5 92.1 77.1 79.0 87.6 86.7 86.0 85.9 GSA-Only 88.1 88.5 85.7 88.1 91.9 91.5 89.8 90.1 GSAR 98.3 98.3 90.0 90.9 91.9 91.5 92.7 92.7 Table 12: Offline trajectory evaluation results of different model-based reward methods across task categories. Table 12 reports classification results by task category using Gemini-2.5-Pro 4. GSAR demonstrates consistently strong performance across all task types, highlighting both the varying difficulty of trajectory completion judgment across categories and the effectiveness of GSAR in handling them. StepCritic outperforms both GSA-Only and GS-Only, indicating that action history plays a critical role in determining completion for Delete type tasks. Notably, GSA-Only performs worse than GS-Only on Delete tasks. This may be attributed to the fact that the added anchor boxes tend to cover large screen regions, causing the model to attend to irrelevant elements and thereby degrading judgment accuracy. Answer tasks require a combination of semantic understanding and goal state reference for accurate judgment. DigiRL yields an F1 of zero on this task type because its frame similarity filter compares the last two screenshots and labels trajectories with minimal visual change as incomplete. Since Answer tasks typically exhibit little visual difference between the final two frames, DigiRL fails to identify any positive samples, resulting in zero true positives for both precision and recall, and consequently an F1 of zero. Comparing GS-Only and GSA-Only, anchoring task-relevant UI elements on the goal state brings substantial gains for Answer tasks. For Normal tasks, successful and incomplete trajectories exhibit distinct differences in key UI elements, allowing strong performance to be achieved even without action history. (a) Broccoli (b) OneTimeAlarm Figure 7: Examples of environment state changes under self-evolving data synthesis. Figure 8: Comparison of UI-TARS-trained agents for task: In Mnemosyne, open the "Write journal notes for 10 minutes" item, set the first entry to "Mood" and the second entry to "Focus", without saving it. Appendix F Case Study F.1 Case Study for Self-Evolving We show examples of the self-evolving mechanism illustrating how it modifies the environment state during the iterative process in Figure 7. It can be clearly observed that the app’s main pages become increasingly diverse during the later stages of evolution, enabling the collection of more complex and varied tasks for subsequent training. F.2 Case Study for Online Performance To further illustrate the impact of different reward methods on agent behavior during online interactions, we present two representative case studies in Figure 8 and Figure 9. One shows a failed trajectory from an agent trained with StepCritic, while the other demonstrates a successful trajectory from an agent trained with GSAR. These examples highlight how using our reward framework enables agents to perform complex mobile GUI tasks more reliably and accurately. Figure 9: Comparison of GUI-Owl-trained agents for task: Add a new player named Emma to the Scoreboard app, and set her scores to 8 points in Round 1, 9 points in Round 2, and 7 points in Round 3. Appendix G Prompts G.1 Prompting Template of Task Generation You are an expert in analyzing mobile app user interfaces and generating clear, purposeful, and executable task instructions. You will receive: - The app name - A current mobile screenshot - UI component information extracted from the screenshot - The action history from app launch to the current screen Your goal is to analyze the current user interface and generate purposeful instructions based on the action history, representing the tasks a human user is likely to perform next. Task types: - Task-Oriented: A series of operations to achieve a specific goal. - Question-Oriented: A series of operations that lead to answering a specific question. Requirements: - Each task must mention the app name and provide a high-level instruction without referencing any specific UI elements. - You can combine multiple simple tasks (from the example tasks) into richer, more complex, and more diverse tasks. - If the task involves entering information, specify the exact content to be input. - DO NOT use words such as ’or’, ’any’, ’e.g.’, or any other expression that makes the task description vague or ambiguous. - DO NOT mention any non-existent files. - DO NOT generate tasks related to user login, account registration, account creation, or authentication flows. - Do not generate any tasks if the screen is a permission dialog, an app initialization popup, or unrelated to the current app. - Generate between 0 and max_tasks tasks depending on the complexity of the screen and action history. Context: - App name: app_name - UI components: accessibility_tree - Action history: action_summary Example tasks: examples Output format (strictly follow this format): your analysis: - A brief explanation of the navigation history, user intent, and key UI elements that affect the tasks. your tasks: 1. Task description 2. Task description 3. Task description … Do not add any extra sections or commentary outside this format. G.2 Prompting Template of Task Inherit You are an expert at generating follow-up mobile app tasks based on an existing completed task. You will receive: - A user task trajectory, including the initial user goal, the last few screenshots, and a sequence of action summaries that indicate the task has already been completed. Your goal is to generate a new task that STARTS FROM the environment state after the previous task is finished. The new task should build upon the existing result and introduce additional steps or subtasks to continue the workflow. Requirements: - The new task must assume the previous task has already been successfully completed. - The task must start from the current environment state shown in the screenshots. - The task should extend the workflow by adding new steps or subtasks, without repeating the original task. - The task must mention the app name and provide a high-level instruction without referencing any specific UI elements. - If the task involves entering information, specify the exact content to be input. - DO NOT use words such as ’or’, ’any’, ’e.g.’, or any other expression that makes the task description vague or ambiguous. - DO NOT mention any non-existent files. Input: - App name: app_name - Previous completed user goal: initial_goal - Action summary: action_summary - Current environment screenshots: <image> Output format (strictly follow this format): Your analysis: - A brief explanation of how the new task continues from the completed state and what new steps are added. Output task: <task> (contain only a single task; if no task needs to be generated, write "No task") </task> Do not add any extra sections or commentary outside this format. G.3 Prompting Template of Task Merge You are an expert at generating more complex mobile app tasks by merging an existing task with additional subtasks. You will receive: - A user task trajectory, including the initial user goal, the last few screenshots, and a sequence of action summaries. Your goal is to generate a NEW combined task that integrates the original task and newly added steps into a single, more complex task. The initial environment of the new task corresponds to the environment state shown in the screenshots. Requirements: - The new task must explicitly include the original task goal and extend it with additional steps or subtasks. - The task should describe a complete, merged workflow as a single high-level user goal. - The task must mention the app name and provide a high-level instruction without referencing any specific UI elements. - If the task involves entering information, specify the exact content to be input. - DO NOT use words such as ’or’, ’any’, ’e.g.’, or any other expression that makes the task description vague or ambiguous. - DO NOT mention any non-existent files. Input: - App name: app_name - Original user goal: initial_goal - Action summary: action_summary - Initial environment screenshots: <image> Output format (strictly follow this format): Your analysis: - A brief explanation of how the original task and new subtasks are merged into a more complex task. Output task: <task> (contain only a single task; if no task needs to be generated, write "No task") </task> Do not add any extra sections or commentary outside this format. G.4 Prompting Template of Task Change You are an expert at generating variant mobile app tasks by modifying parameters in an existing task instruction. You will receive: - A user task trajectory, including the initial user goal, the last few screenshots, and a sequence of action summaries. Your goal is to generate a new task that has the SAME task structure as the original task, but differs by changing specific parameters such as names, values, text content, or identifiers. Requirements: - The new task must keep the same overall workflow and steps as the original task. - Only modify concrete parameters in the task instruction, such as file names, text content, titles, numbers, or labels. - Do NOT add new steps or remove existing steps. - The task must mention the app name and provide a high-level instruction without referencing any specific UI elements. - If the task involves entering information, specify the exact content to be input. - DO NOT use words such as ’or’, ’any’, ’e.g.’, or any other expression that makes the task description vague or ambiguous. - DO NOT mention any non-existent files. Input: - App name: app_name - Original user goal: initial_goal - Action summary: action_summary - Environment screenshots: <image> Output format (strictly follow this format): Your analysis: - A brief explanation of which parameters are modified and how they differ from the original task. Output task: <task> (contain only a single task; if no task needs to be generated, write "No task") </task> Do not add any extra sections or commentary outside this format. G.5 Prompting Template of label anchor You are an expert in GUI understanding. Your task is to analyze the input image and UI Elements, and select the indices of the UI elements that are relevant to the instruction. These selected UI elements should uniquely identify the state after the instruction task has been completed. Input Description: - Instruction: A natural language description of the GUI task to be performed. - UI Elements: The complete list of UI elements extracted from the final-state screen (including all element metadata). - Image: The screenshot of the final state after completing the task, with UI element bounding boxes and indices overlaid to indicate their positions. Requirements: - Compare the image and the UI Elements list, and find UI elements whose content or visual role reflects the completion of the instruction. - Select key UI elements that can indicate the final page state, such as navigation bars, menu titles, tab titles, selected buttons, status indicators, etc. - Avoid selecting UI elements whose state may change dynamically due to time or network content. - If you believe that no UI element in the screen is relevant to the instruction, stop the reasoning and output "No related elements". Now output in the following format: Thinking Process: Your thinking process Output Index: [x, x, …] Input: Instruction: query UI Elements: ui_elements_info