Paper deep dive
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks
Saelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin, Jae Won Cho, Yale Song, Juho Kim
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/31/2026, 1:34:46 AM
Summary
GUIDE is a new benchmark designed to evaluate multimodal LLMs on their ability to understand and assist users in open-ended GUI tasks. Unlike prior benchmarks focusing on task automation, GUIDE emphasizes user intent, behavior state detection, and proactive assistance, utilizing 67.5 hours of screen recordings from 120 novice users across 10 software applications.
Entities (5)
Relation Signals (4)
GUIDE â evaluates â Multimodal Large Language Models
confidence 100% ¡ We evaluate a range of multimodal large language models (MLLMs) on our benchmark
GUIDE â includestask â Behavior State Detection
confidence 100% ¡ GUIDE defines three tasks - (i) Behavior State Detection
GUIDE â includestask â Intent Prediction
confidence 100% ¡ GUIDE defines three tasks - (ii) Intent Prediction
GUIDE â includestask â Help Prediction
confidence 100% ¡ GUIDE defines three tasks - (iii) Help Prediction
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Graphical User Interface (GUI) agents have the potential to assist users in interacting with complex software (e.g., PowerPoint, Photoshop). While prior research has primarily focused on automating user actions through clicks and keystrokes, this paradigm overlooks human intention, where users value the ability to explore, iterate, and refine their ideas while maintaining agency. To move beyond automation and toward collaboration, GUI agents must understand what users are doing and why. We introduce GUIDE (GUI User Intent Detection Evaluation), a benchmark that evaluates AI models on their ability to perceive user behavior, infer intent, and provide assistance in open-ended GUI tasks. GUIDE consists of 67.5 hours of screen recordings from 120 novice user demonstrations with think-aloud narrations, across 10 software. GUIDE defines three tasks - (i) Behavior State Detection, (ii) Intent Prediction, and (iii) Help Prediction that test a model's ability to recognize behavior state, reason about goals, and decide when and how to help. Evaluations across eight state-of-the-art multimodal models reveal that all models struggled, achieving only 44.6% and 55.0% accuracy on behavior state and help prediction. However, providing user context significantly improved the performance, raising help prediction by up to 50.2pp, highlighting the critical role of structured user understanding in effective assistance. Our dataset is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.25864v1
- Canonical: https://arxiv.org/abs/2603.25864v1
Trouble viewing inline? Open PDF directly â
Full Text
92,819 characters extracted from source content.
Expand or collapse full text
GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks Saelyne Yang1,2 Jaesang Yu1 Yi-Hao Peng2 Kevin Qinghong Lin3 Jae Won Cho4 Yale Song5 Juho Kim1,6 1KAIST 2Carnegie Mellon University 3University of Oxford 4Konkuk University 5Google Inc. 6SkillBench Abstract Graphical User Interface (GUI) agents have the potential to assist users in interacting with complex software (e.g., PowerPoint, Photoshop). While prior research has primarily focused on automating user actions through clicks and keystrokes, this paradigm overlooks human intention, where users value the ability to explore, iterate, and refine their ideas while maintaining agency. To move beyond automation and toward collaboration, GUI agents must understand what users are doing and why. We introduce GUIDE (GUI User Intent Detection Evaluation), a benchmark that evaluates AI models on their ability to perceive user behavior, infer intent, and provide assistance in open-ended GUI tasks. GUIDE consists of 67.5 hours of screen recordings from 120 novice user demonstrations with think-aloud narrations, across 10 software. GUIDE defines three tasksâ(i) Behavior State Detection, (i) Intent Prediction, and (i) Help Prediction that test a modelâs ability to recognize behavior state, reason about goals, and decide when and how to help. Evaluations across eight state-of-the-art multimodal models reveal that all models struggled, achieving only 44.6% and 55.0% accuracy on behavior state and help prediction. However, providing user context significantly improved the performance, raising help prediction by up to 50.2p, highlighting the critical role of structured user understanding in effective assistance. Our dataset is available at https://guide-bench.github.io. 1 Introduction Dataset Domain # Video # Video Duration Video Source Primary Goal Evaluation Focus Behavior Intent Help PsTuts [25] 1 - 71.4 h Instructional Videos Action Understanding VideoWebArena [19] 6 74 3.8 h Human-Recorded Tutorials Task Automation VideoGUI [27] 11 178 7.1 h Instructional Videos Task Automation â UI-Vision [33] 83 450 4.8 h Experts Performing Tasks Task Automation AssistGUI [14] 9 100 <8.3 h Instructional Videos Task Automation â WorldGUI [54] 10 611 <30.5 h Instructional Videos Task Automation â GUIDE (Ours) 10 120 67.5 h Novice Usersâ Demonstrations Behavior Understanding â â â Table 1: Comparison of GUIDE with existing GUI video understanding datasets. GUIDE differs from existing benchmarks by (i)(i) collecting screen recordings from novice users, (iâi)(i) capturing how they naturally behave in open-ended tasks with a focus on behavior understanding, and (iâiâi)(i) evaluating systems based on human user needs rather than task automation. Figure 1: An example of the GUIDE benchmark, which jointly models three tasks: Behavior State Detection, Intent Prediction, and Help Prediction, to interpret what the user is doing, aiming to achieve, and whether and what they may need assistance with during open-ended software tasks. Graphical User Interface (GUI) agents hold great promise for supporting users in complex workflows, in mobile [29, 20, 53], web [41, 16, 50, 9, 56], and software application tasks [51, 40, 10]. In creative and analytical tools such as Photoshop or PowerPoint, these agents can automate repetitive subtasks or provide guidance to help users achieve their goals more efficiently. Most existing GUI agents, both in academic research [28, 14, 27, 54] and in commercial services like Microsoft Office Copilot [32] or Figma Make [12], focus on full automation: given a goal, they either execute a sequence of clicks and keystrokes to complete the task or directly generate the desired output. While this approach offers convenience, it overlooks how people actually work. In real-world open-ended workflows, success is not driven solely by efficiencyâuser satisfaction plays an equally critical role. Automated agents assume fixed goals, yet users frequently revise their intentions mid-task. For example, a user may reposition an element multiple times before reverting to the originalâbehavior an automated agent would treat as redundant, but which is essential for forming a preference. Rather than replacing user agency, effective assistance should accelerate exploration while keeping the user in control [21]. Recent work on proactive task assistance takes a more balanced approach [45, 31, 47, 52, 48]. Rather than automate tasks for users, proactive assistants infer a userâs context and intent and deliver timely help. Studies in programming and productivity tools show higher efficiency and satisfaction when a system detects a need and intervenes at the right moment [38, 6, 37, 46]. Yet, the ability to model and track usersâ evolving context remains underexplored in current multimodal systems that power GUI agents. To achieve a truly human-assisting GUI agent, a key ability is to comprehend usersâ cognitive context and intentions to provide appropriate support [17]. In real-world scenarios, users rarely articulate their goals or needs explicitly, making it natural for systems to rely primarily on visual cues from the screen. These user actions often carry semantic structure, such as hovering, undoing, or repeatedly opening menus, that signal intent. However, interpretation remains challenging: similar actions may stem from entirely different intents. For example, repeated undo actions might indicate confusion or deliberate refinement. Without deeper reasoning, assistance based solely on surface-level actions can lead to shallow or misaligned responses. To address this challenge, we present GUIDE (GUI User Intent Detection Evaluation), a benchmark designed to evaluate multimodal LLMs (MLLMs) on their ability to understand and assist users in complex software workflows. GUIDE introduces a three-stage evaluation framework: (1) Understanding the userâs behavioral state to identify their current workflow phase; (2) Reasoning about their underlying intentions and goals; and (3) Assisting by delivering the appropriate form of help at the right moment. We collected 67.5 hours of screen recordings from 120 human demonstrations across 10 widely used applicationsâincluding Photoshop, Figma, PowerPoint, Premiere Pro, and Excelâcovering 40 open-ended tasks designed to elicit natural user behavior. Unlike prior work that primarily targets video understanding from expert-recorded instructional videos on closed-ended tasks [25, 27, 33, 14, 54], our focus is on novice users working on open-ended tasks, with the goal of building collaborative AI systems that assist users during exploration, trial-and-error, and learning. Observing novice workflows allows us to capture authentic moments of confusion, decision-making, and discovery, offering rich opportunities for AI to provide timely, context-aware support. Each session includes both screen recordings and think-aloud narrations that surface the userâs underlying intentions and cognitive states. Building on this dataset, we define three-staged benchmark tasks: First, (i) Behavior State Detection evaluates whether a model can identify the userâs behavioral state, such as exploration or confusion, based solely on visual cues. To support this, we developed a taxonomy of nine user states reflecting diverse cognitive and behavioral phases in open-ended GUI workflows, grouped into four high-level categories: Planning, Execution, Problem-Solving, and Evaluation (Figure 3). This structure aligns with human cognition and interaction theories [4, 34], while introducing finer distinctions tailored to GUI-based task behavior. Next, (i) Intent Prediction targets inference of the userâs immediate goalâwhat they are trying to accomplish in the given moment. The final task, (i) Help Prediction, assesses whether a model can determine 1) whether the user needs assistance or not, and if so, 2) what type of help would be most appropriate, such as explaining a feature, suggesting an alternative, or addressing an error. By leveraging both visual screen recordings and accompanying think-aloud narrations, we automatically generated data for each task, which was subsequently verified through human review for accuracy and consistency. Evaluation across eight state-of-the-art MLLMs reveals that while current models struggle to interpret user behavior and predict underlying intent and help neededâachieving only 44.6% accuracy on behavior state detection and 55.0% on help prediction, performance improves significantly when structured user context is provided. For example, supplying behavioral state and intent information boosted help prediction accuracy by up to 50.2 percent points for the lowest-performing model. Our results suggest a promising path forward: providing different layers of human-grounded context, such as behavioral cues, inferred goals, and temporal history, can lead to more accurate assistance decisions. Our benchmark provides a foundation for training and evaluating the next generation of context-aware, collaborative GUI agents. 2 Related Work 2.1 Video Understanding meets GUI Several benchmarks evaluate video understanding in the context of GUI and software workflows. Early work by Li et al. [25] collected Photoshop tutorial videos to understand screencast videos. More recent datasets span multiple applications and tasks. For example, AssistGUI [14] focuses on automating GUI tasks using an actor-critic agent, serving as a benchmark for task-oriented GUI automation. VideoWebArena [19] evaluates long-horizon multimodal agents on web browsing tasks, emphasizing extended video context and web UI interactions. VideoGUI [27] compiles high-quality instructional screen recordings and introduces a hierarchical model for mapping visual observations to GUI actions. UI-Vision [33] provides a fine-grained desktop UI video benchmark with dense annotations for perception and interaction. Lastly, WorldGUI [54] increases task diversity by allowing arbitrary initial interface states for each task, challenging agents to handle varied starting conditions. These prior benchmarks primarily focus on close-ended tasks with predetermined goals, aiming to replicating expert demonstrations. In contrast, our work targets open-ended GUI workflows with novice users, emphasizing understanding of user intent and context rather than step-by-step replication of actions. This shift toward user-centric evaluation fills a gap not covered by existing GUI video datasets that evaluate task completion or action prediction. 2.2 Collaborative and Proactive Agents While GUI agents that automate interface operations based on a given goal or instruction can be effective, this fully autonomous approach can conflict with the needs of users in creative or analytical environments, where retaining control and exploring alternatives are essential. To address this, recent research has shifted toward assistive GUI agents that collaborate with users by understanding context and offering timely support. Several works have explored inferring user goals and intent in both web [36] and software environments [3, 13, 55] to better align assistance with user needs. For example, Zhao et al. [55] introduce ProactiveVA, a visual analytics agent that monitors user interactions and leverages LLMs to detect when users may be stuck, providing context-sensitive suggestions or guidance. Several recent works in the Human-Computer Interaction (HCI) community explore this shift toward collaboration and contextual support. CowPilot [18] proposes a mixed-initiative framework that enables users to share control with an autonomous web navigation agent, improving efficiency while preserving agency. In programming, proactive assistants like Codellaborator [38] and NeedHelp [6] demonstrate how real-time intervention can aid users when well-timed. Studies on software applications [21] show users prefer AI agents that guide them rather than take over entirely, reinforcing the need for transparency and shared control. ProMemAssist [39] further highlights the benefits of modeling user cognition to deliver timely, non-intrusive support. These findings echo broader discussions on autonomy levels [11] and the importance of aligning agent behavior with human preferences [26, 22]. Our work builds on these insights, evaluating how well current multimodal models can perceive a userâs state and intentions in GUI workflow recordings and decide if and how to assist. By situating the evaluation in real user workflows, we aim to push GUI agents toward true user-aware collaboration. 3 GUIDE Benchmark Figure 2: Overview of the three core tasks in the GUIDE benchmark. (1) User Behavior State Detection identifies the userâs current behavioral mode (e.g., Exploration and Decision-Making). (2) Intent Prediction infers what the user is trying to achieve (e.g., Create a progress bar). (3) Help Prediction determines whether the user needs assistance and, if so, what kind of help is relevant (e.g., Get a guide on how to use text effects). Together, these tasks enable a comprehensive understanding of user behavior and assistance needs in software GUI environments. We evaluate MLLMs on their ability to infer these solely from the visual input, without access to the demonstratorâs narration â a setting that closely reflects real-world use. To develop a benchmark that focuses on understanding and assisting users, we collected demonstrations from novice users. Unlike existing datasets that focus primarily on expert demonstrations or polished instructional videos [25, 27, 33, 14, 54, 30, 44], our dataset captures the authentic challenges and exploratory behaviors that novices exhibit during task completion, serving a crucial role in building collaborative agents. Building on these demonstrations, we propose a suite of tasks designed to evaluate modelsâ capabilities to understand users and provide effective assistance. 3.1 Video Collection We collected 120 demonstrations from novice users across 10 applications spanning five categories: Photo Editing (Photoshop, GIMP), Graphic Design (Figma, Canva), Presentation Design (PowerPoint, Google Slides), Video Editing (Premiere Pro, CapCut), and Data Analysis (Google Sheets, Microsoft Excel). For each application, we designed four open-ended tasks aimed at eliciting natural and diverse user behaviors and approaches (Table B2 in supp.). We chose creative and analytical tools to surface exploratory workflows and variation in problem-solving strategies. Each task was completed by three different users to capture diverse strategies and behaviors. We ensured that each task was flexible enough, while still incorporating elements of challenge. Participants were asked to spend at least 20 minutes per task and meet a few minimal requirements (e.g., inserting a relevant image) to mark it as complete. We recruited 54 novice users of software from Prolific and our institution. Participants were screened based on self-reported expertise to ensure novice-level familiarity with the features in the target application (Mean: 2.8, SD: 1.1, Range: 1â5). During the study, participants worked on the assigned task while recording their screen and keyboard/mouse input events. They were also asked to think aloud and record their voice as they carried out the task, verbalizing what they were doing and their thought process. 3.2 Benchmark Tasks To evaluate a modelâs ability to understand user context and deliver appropriate assistance, we design our benchmark as a unified three-stage framework: Understanding â Reasoning â Assisting. These stages progress from interpreting user behavior to inferring intentions and ultimately providing helpful assistance. Each task corresponds to a distinct level of cognitive inference required for a human-assisting GUI agent to effectively support users in open-ended software workflows. To construct a dataset for task evaluation, we used the Human-AI collaborative method. We first transcribed the think-aloud narration using WhisperX [2], and used the narration as a main source for extracting initial annotations in addition to the video. We employed Gemini-2.5-Pro [15] to first create annotations needed for each task, which were then refined by human annotators. Note that we use narration only as an annotation source to capture usersâ intentions and mental states. The benchmark evaluates vision-only understanding, testing whether models can infer these states solely from visual cues, as in real-world settings with limited access to user speech. Figure 3: Our proposed taxonomy of user behavior states in GUI-based software tasks, organized into four main phases: Planning, Execution, Problem-Solving, and Evaluation. Each phase captures distinct patterns of user cognition and interaction, from initial goal formulation to iterative action, troubleshooting, and reflection. 3.2.1 User Behavior State Detection Description. This task evaluates whether a model can interpret the userâs behavioral context directly from visual cues. Models are asked to classify a video segment into one of nine behavior states in our taxonomy (Figure 3), which spans the full range of cognitive and behavioral processes observed in creative and analytical workflows. We developed the taxonomy through a multi-stage, humanâAI collaborative process [24]. First, three authors iteratively created and consolidated an initial taxonomy over five sessions based on observations of online software task videos. Separately, we prompted Gemini-2.5-Pro to generate a taxonomy from scratch using our collected video dataset, without providing our initial version. We then augmented the human-generated taxonomy by integrating novel categories identified by the LLM. Finally, the combined taxonomy was validated against the entire video dataset to ensure comprehensive coverage and reorganized into the final set of nine distinct states. Our taxonomy aligns with Normanâs Seven Stages of Action [34], mapping Planning, Execution, and Evaluation to goal formation, action, and outcome assessment, and draws on Bloomâs cognitive hierarchy [4] that captures the shifts between operational (Execution) and critical work (Evaluation). Dataset Curation. After constructing the taxonomy, we aligned each video with its corresponding narration segments. For every segment, we annotated the userâs behavior state using Gemini-2.5-Pro according to the taxonomy, prompting the model to produce both a predicted label and its reasoning. Two human annotators recruited from Prolific then verified and refined these annotations, achieving a 96.1% agreement rate. Finally, we uniformly sampled 200 instances from each of the nine classes, resulting in a balanced dataset of 1.8K annotated segments. 3.2.2 Intent Prediction Description. This task evaluates whether a model can reason about the userâs short-term, immediate goal in context. It focuses on identifying what the user aims to achieve within open-ended workflows. Dataset Curation. Using the narration-aligned video segments, we prompted Gemini-2.5-Pro to infer usersâ intention in each segment. The think-aloud narrations often revealed usersâ goals (e.g., âIâm going to align these objectsâ, âIâl try another colorâ). Leveraging this signal, we prompted the model to infer the underlying user intention. After collecting and deduplicating the inferred intents, we further instructed the model to generate three plausible but incorrect alternatives to serve as distractors for the multiple-choice evaluation. The resulting intent annotations and distractors were then validated by the authors, with 88.68% of the data retained, yielding a final set of 1.3K instances. 3.2.3 Help Prediction Description. The final task evaluates whether a model can progress from understanding and reasoning to deciding how to assist. Help Prediction consists of two subtasks: (1) Help Need Detection, a binary classification task that determines whether the user needs help, and (2) Help Content Prediction, which identifies the specific type of help needed, such as explaining a feature or suggesting an alternative. Together, these subtasks assess a modelâs ability to anticipate user needs and recommend appropriate assistance, bridging the gap between perception and actionable support. Dataset Curation. We identified potential help-seeking moments using two complementary signals. First, explicit help-seeking behaviors, such as switching to external resources (e.g., Google, YouTube, ChatGPT), indicated direct attempts to seek guidance. Second, implicit help-seeking cues were extracted from user narration, where they expressed uncertainty or confusion (e.g., âHow do I align this?â, âI canât find Layer Mask.â). Additionally, we included clear no-help-needed moments, where users demonstrated confidence through their narration. Using these signals, Gemini-2.5-Pro was prompted to generate initial annotations for help-need and help-content labels. After deduplication, the model was additionally prompted to generate three plausible but incorrect options for each instance for multiple-choice question evaluation. All annotations and distractors were then reviewed by the authors, resulting in 1K validated instances, with 78.89% of the original data retained. For 12.5% of the retained instances, the segmentâs start or end time was adjusted to exclude explicit visual help signals (e.g., user turning to Google Search) to ensure fair evaluation. Overall, 66% of the instances were labeled as help-needed, while the remaining 34% required no help. 4 Experiments Model (1) Behavior Detection (2) Intent Prediction (3) Help Prediction â + Prev. â + Behavior Help Need Detection Help Content Prediction â + Behv. +Behv.+Intent â + Behv. +Behv.+Intent Gemini-2.5-Flash [15] 36.91 38.19 65.40 66.77 53.64 76.33 78.07 49.53 53.75 78.59 Gemini-2.5-Pro [15] 42.44 43.79 67.80 70.16 69.82 84.73 82.38 52.74 57.03 79.69 GPT-4o-mini [35] 17.65 17.07 60.76 62.19 46.05 78.92 82.26 31.32 42.86 79.84 GPT-4o [35] 36.32 37.24 61.19 62.58 49.69 87.79 87.91 45.95 48.37 79.78 Claude-4.5-Sonnet [1] 44.61 45.63 71.39 72.62 39.49 58.56 59.43 55.00 62.17 82.79 Qwen3-VL-8B [42] 37.97 38.13 62.70 64.03 52.83 70.39 77.36 46.06 50.63 80.11 InternVideo2.5-8B [43] 21.57 27.02 43.79 45.13 34.36 35.35 35.25 23.67 29.15 73.86 InternVL3-8B [57] 22.57 24.90 46.11 46.97 34.94 43.73 46.82 27.03 32.20 72.97 Table 2: Evaluation results on accuracy across (1) Behavior State Detection, (2) Intent Prediction, and (3) Help Prediction. 4.1 Experimental Setup We evaluate a range of multimodal large language models (MLLMs) on our benchmark to assess their ability to understand, reason about, and assist users in open-ended software workflows. Our evaluation includes eight representative MLLMs spanning both proprietary and open-source models: Gemini-2.5-Flash [15], Gemini-2.5-Pro [15], GPT-4o-mini [35], GPT-4o [35], Claude-4.5-Sonnet [1], Qwen3-VL-8B [42], InternVideo2.5-Chat-8B [43], and InternVL3-8B [57]. All models are evaluated in a zero-shot setting using publicly available APIs or checkpoints, without any additional fine-tuning. For each test instance, we uniformly sample 32 frames from the corresponding video segment, providing only visual input (excluding narration audio) to simulate perception based solely on visual cues. To ensure consistency across models, we use standardized prompting templates (Section H). We also prompt models to generate both a predicted label and supporting reasoning, a strategy shown to improve task performance [23]. Our main experiments are conducted in an offline inference setting, where models solve the task given the full video. To approximate real-world proactive assistant scenarios, we additionally evaluate an online setting, where the model receives visual input progressivelyâat 25%, 50%, 75%, and 100% of the segment, we uniformly sample 32 frames from the corresponding prefix for inference. Model Help Need Detection â + Behavior State + Behavior State + Intent Acc Prec Rec F1 Acc Prec Rec F1 Acc Prec Rec F1 Gemini-2.5-Flash [15] 53.64 83.27 36.62 50.87 76.33 97.67 65.47 78.39 78.07 94.56 70.62 80.86 Gemini-2.5-Pro [15] 69.82 76.42 78.09 77.42 84.73 93.61 82.34 87.61 82.38 91.20 80.94 86.76 GPT-4o-mini [35] 46.05 83.03 22.31 35.17 76.73 97.61 66.23 78.92 79.71 97.20 71.29 82.26 GPT-4o [35] 49.69 74.41 35.14 47.73 87.79 95.39 85.53 90.19 87.91 95.12 85.95 90.30 Claude-4.5-Sonnet [1] 39.49 87.69 8.92 16.19 58.56 99.16 37.09 53.99 59.43 99.19 38.44 55.41 Qwen3-VL-8B [42] 52.83 79.86 34.23 47.92 70.39 94.35 58.50 72.22 77.36 95.38 67.56 79.09 InternVideo2.5-8B [43] 34.36 33.33 0.16 0.31 35.35 90.91 1.56 3.07 35.25 83.33 1.56 3.07 InternVL3-8B [57] 34.94 72.73 1.25 2.46 43.73 98.88 15.77 27.20 46.82 98.40 19.22 32.16 Table 3: Results for Help Need Detection on accuracy, precision, recall, and F1-score across three conditions (default, with behavior state, with behavior state and intent). 4.2 Evaluation Tasks (1) Behavior State Detection. This task measures whether a model can identify the userâs behavioral state from a given video segment. We provide each model with clips and ask it to classify them into one of nine taxonomy-defined states. Two configurations are tested: (i) using only the current segment and (i) with prior history, where the model is given the immediately preceding segmentâs behavior state. This is framed as a multi-class classification problem, and performance is evaluated using accuracy. (2) Intent Prediction. This task evaluates a modelâs ability to infer the userâs underlying goal within a given video segment. Models are prompted to predict what the userâs goal in two settings: (i) using only the current segment, and (i) with additional behavior state context, where the model is also given the state label and its definition. We adopt a multiple-choice question (MCQ) format, where the model selects the most likely intent from four candidates. Performance is measured using accuracy. For the default setting (i), to mitigate potential bias, we additionally report multi-binary accuracy (MBAcc) following prior work [5, 8, 7], which evaluates whether the model correctly identifies the ground-truth intent in all three pairwise comparisons against incorrect alternatives. (3) Help Prediction. The final task evaluates whether models can move beyond understanding and reasoning to provide actionable assistance. Given a video segment, models are asked to predict whether the user requires help (Need), and if so, what kind of help would be most appropriate (Content). Help Need Detection is framed as a binary classification task and evaluated using accuracy, precision, recall, and F1-score. Help Content Prediction, similar to Intent Prediction, uses a multiple-choice question (MCQ) format and is evaluated using accuracy and multi-binary accuracy (MBAcc) for the default setting. We test three settings for both tasks: (i) video only, (i) video + behavior state, where the model is given the behavior label and its definition for the current segment, and (i) video + behavior state + intent, where the model additionally receives the identified user intention. These settings progressively assess the modelâs ability to leverage layered user context for meaningful, situation-aware assistance. 4.3 Results Table 2 presents the performance of baseline models on GUIDE across the tasks, with accuracies reported under default and context-augmented settings. Overall, models performed weakest on Behavior State Detection and Help Prediction, with default-setting accuracies peaking at 44.61% and 55.00% for Behavior State Detection and Help Content Prediction, respectively, both from Claude-4.5-Sonnet [1]. While Gemini-2.5-Pro [15] reached nearly 70% accuracy on Help Need Detection, most other models showed substantially lower performance across both Help sub-tasks. Across tasks, we observe that models generally benefit from added behavioral and intent context, with particularly notable improvements in help-related predictions. We report the main findings below. 4.3.1 Behavior State Detection Behavior State Detection remains highly challenging. All models struggled to accurately infer the userâs behavioral state from video segments, underscoring the difficulty of the 9-way classification task. While proprietary models such as Claude-4.5-Sonnet [1] and Gemini-2.5-Pro [15] performed best, no model surpassed 45% accuracy, and most fell below 40%. Models often misinterpret signals of struggle. The most common failure was misclassifying Frustration or Debugging as Performing Actions or Exploration and Decision-Making, as shown in the confusion matrix (Figure C4 in supp.). These errors suggest that models overlook subtle indicators of user difficultyâe.g., repeated clicks, hesitation, or undoing actionsâinstead interpreting them as deliberate progress, revealing a lack of nuanced understanding of user frustration signals. Temporal context shows modest potential. Incorporating the prior behavior state led to small but consistent gains across models. While most improvements were marginal, the largest gain was observed for InternVideo2.5-8B [43] with 5.45 percentage points, suggesting that temporal context holds value and may be more effectively utilized with improved temporal reasoning capabilities. Figure 4: Accuracy trends across the tasks in the online setting, where models are given progressively more of the video segment (25%, 50%, 75%, and 100%). Models show consistent improvement as they see more segments, with Gemini-2.5-Flash [15] and Qwen3-VL-8B [42] showing larger and more consistent gains across all four tasks compared to the smaller open-source models. Model Intent Prediction Help Prediction Acc MBAcc Acc MBAcc Gemini-2.5-Flash [15] 65.40 59.09 49.53 44.69 Gemini-2.5-Pro [15] 67.80 64.34 52.74 45.31 GPT-4o-mini [35] 60.76 50.24 31.32 28.59 GPT-4o [35] 61.19 56.58 45.95 41.25 Claude-4.5-Sonnet [1] 71.39 65.44 55.00 50.78 Qwen3-VL-8B [42] 62.70 58.07 46.06 44.69 InternVideo2.5-8B [43] 43.79 27.98 23.67 18.75 InternVL3-8B [57] 46.11 40.75 27.03 23.75 Table 4: Evaluation of Intent Prediction and Help Content Prediction, with Accuracy (Acc) and Multi-Binary Accuracy (MBAcc). 4.3.2 Intent Prediction Intent Prediction is the most tractable task, but still imperfect. Among the three tasks, models achieved the highest performance on intent prediction, with several surpassing 60% accuracy. However, performance drops under the stricter MBAcc metric, which requires consistent discrimination across all answer pairs. This indicates that while models can often select a plausible intent, they still struggle with reliably identifying the correct one over all distractors (Table 4). Behavior context helps, but only slightly. Incorporating behavior state context (i.e., the userâs behavioral label and definition) consistently improved performance, but the gains were relatively modest across all models. This suggests that while such context may offer useful cues, it does not provide sufficient information on its own or is not yet effectively leveraged by current models for intent inference. 4.3.3 Help Prediction High variance and missed help cases in Need Detection. Table 3 shows the full performance results for Help Need Detection. This subtask exhibited the most variance across models, with F1 scores ranging from 0.31 (InternVideo2.5-8B [43]) to 77.42 (Gemini-2.5-Pro [15]). Notably, recall was particularly low across most modelsâexcept for Gemini-2.5-Pro, all others had recall under 37%. This indicates that many instances where users actually needed help were misclassified as not needing it, echoing similar trends in Behavior State Detection (Section 4.3.1) where models frequently misinterpreted signals of struggle. Behavior context improves Help Need Detection. Providing the userâs behavior state led to consistent and significant improvements in Help Need Detection across all models, with the largest gain observed in GPT-4o [35], with a 42.46-point increase in F1 score. This suggests that context, such as whether a user is exploring or showing signs of frustration, provides strong cues for determining help needs. Help Content Prediction remains challenging, but benefits from intent context. Help Content Prediction proved particularly challenging, with all models struggling and the top accuracy reaching only 55% from Claude-4.5-Sonnet [1], which further declined to around 50% under the stricter MBAcc evaluation. However, incorporating user intent led to substantial improvements across models, with the largest gain in InternVideo2.5-8B [43] at 50.19 percentage points, highlighting the importance of understanding both user state and intent for providing targeted support. 4.3.4 Other Findings Online vs. Offline Setting: models benefit more from temporal context. In our online simulation experiment, where models are given progressively more of the video segment (25%, 50%, 75%, and 100%), we observe consistent performance gains across all four tasks (Figure 4). Gemini-2.5-Flash [15] and Qwen3-VL-8B [42] show larger and more consistent gains across all tasks, compared to the smaller open-source models, indicating a strong ability to integrate growing context into more accurate predictions. These findings suggest that gathering appropriate context over time is crucial for proactive AI assistance, where systems must not only react but also anticipate user needs based on incomplete and evolving information. Outlook for Model Improvements. Together, these results suggest that incorporating structured user context (behavior state and intent) and temporal context consistently improves help prediction. Recent work on agents demonstrates the effectiveness of context engineering via stratified memory, where interaction history is selectively structured rather than treated as a flat sequence [49]. Applying this idea to GUI assistance is a promising direction for better leveraging long-horizon user context. 5 Conclusion We introduced a benchmark for evaluating models in understanding, reasoning about, and assisting users in open-ended GUI-based workflows. Grounded in real-world novice user demonstrations, our tasksâbehavior state detection, intent prediction, and help predictionâcapture core capabilities needed for collaborative GUI agents. Evaluation across state-of-the-art MLLMs revealed that models struggle to interpret nuanced user behavior and accurately infer assistance needed in open-ended scenarios. However, when provided with appropriate user context, models showed consistent improvements, highlighting the value of structured user understanding. Overall, our benchmark provides a foundation for user-aware agents that support human workflows. Acknowledgements This work was supported by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korean government (MSIT) (No. 2021-0-01347, Video Interaction Technologies Using Object-Oriented Video Modeling and No. RS-2024-00443251, Accurate and Safe Multimodal, Multilingual Personalized AI Tutors). References [1] Anthropic (2025) Introducing claude sonnet 4.5. Note: https://w.anthropic.com/news/claude-sonnet-4-5Anthropic News Release, September 29, 2025 Cited by: §4.1, §4.3.1, §4.3.3, §4.3, Table 2, Table 3, Table 4. [2] M. Bain, J. Huh, T. Han, and A. Zisserman (2023) WhisperX: time-accurate speech transcription of long-form audio. INTERSPEECH 2023. Cited by: §3.2. [3] O. Berkovitch, S. Caduri, N. Kahlon, A. Efros, A. Caciularu, and I. Dagan (2025) Identifying user goals from ui trajectories. In Companion Proceedings of the ACM on Web Conference 2025, W â25, New York, NY, USA, p. 2381â2390. External Links: ISBN 9798400713316, Link, Document Cited by: §2.2. [4] B. S. Bloom (1956) Taxonomy of educational objectives: the classification of educational goals. 1st edition, Longman Group. Cited by: §1, §3.2.1. [5] M. Cai, R. Tan, J. Zhang, B. Zou, K. Zhang, F. Yao, F. Zhu, J. Gu, Y. Zhong, Y. Shang, Y. Dou, J. Park, J. Gao, Y. J. Lee, and J. Yang (2024) TemporalBench: towards fine-grained temporal understanding for multimodal video models. arXiv preprint arXiv:2410.10818. Cited by: §A.1.2, §4.2. [6] V. Chen, A. Zhu, S. Zhao, H. Mozannar, D. Sontag, and A. Talwalkar (2025) Need help? designing proactive ai assistants for programming. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI â25, New York, NY, USA. External Links: ISBN 9798400713941, Link, Document Cited by: §1, §2.2. [7] J. Cheng, V. Wang, H. Wang, H. Zhou, Y. Peng, H. Liu, H. Huang, K. Chen, C. Yang, W. Chai, et al. (2025) Tempura: temporal event masked prediction and understanding for reasoning in action. arXiv preprint arXiv:2505.01583. Cited by: §4.2. [8] J. H. Cho, A. Madotto, E. Mavroudi, T. Afouras, T. Nagarajan, M. Maaz, Y. Song, T. Ma, S. Hu, H. Rasheed, P. Sun, P. Huang, D. Bolya, S. Jain, M. Martin, H. Wang, N. Ravi, S. Jain, T. Stark, S. Moon, B. Damavandi, V. Lee, A. Westbury, S. Khan, P. KrähenbĂźhl, P. DollĂĄr, L. Torresani, K. Grauman, and C. Feichtenhofer (2025) PerceptionLM: open-access data and models for detailed visual understanding. arXiv:2504.13180. Cited by: §A.1.2, §4.2. [9] X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su (2023) Mind2Web: towards a generalist agent for the web. External Links: 2306.06070 Cited by: §1. [10] A. Feizi, S. Nayak, X. Jian, K. Q. Lin, K. Li, R. Awal, X. H. LĂš, J. Obando-Ceron, J. A. Rodriguez, N. Chapados, D. Vazquez, A. Romero-Soriano, R. Rabbany, P. Taslakian, C. Pal, S. Gella, and S. Rajeswar (2025) Grounding computer use agents on human demonstrations. External Links: 2511.07332, Link Cited by: §1. [11] K. J. K. Feng, D. W. McDonald, and A. X. Zhang (2025) Levels of autonomy for AI agents. Knight First Amendment Institute â AI and Democratic Freedoms Essay Series. External Links: Link Cited by: §2.2. [12] Figma (2025) Figma make. Note: https://w.figma.com/make/Accessed: 2025-11-14 Cited by: §1. [13] K. Gadhave, J. GĂśrtler, Z. Cutler, C. Nobre, O. Deussen, M. Meyer, J. M. Phillips, and A. Lex (2021) Predicting intent behind selections in scatterplot visualizations. Information Visualization 20 (4), p. 207â228. External Links: Document, Link, https://doi.org/10.1177/14738716211038604 Cited by: §2.2. [14] D. Gao, L. Ji, Z. Bai, M. Ouyang, P. Li, D. Mao, Q. Wu, W. Zhang, P. Wang, X. Guo, H. Wang, L. Zhou, and M. Z. Shou (2024) AssistGUI: task-oriented pc graphical user interface automation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, p. 13289â13298. Note: Benchmarks PC GUI automation with an actor-critic agent framework External Links: Document, Link Cited by: Table 1, §1, §1, §2.1, §3. [15] G. Gemini Team (2025) Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. External Links: Link Cited by: §3.2, Figure 4, Figure 4, §4.1, §4.3.1, §4.3.3, §4.3.4, §4.3, Table 2, Table 2, Table 3, Table 3, Table 4, Table 4. [16] W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding, and J. Tang (2023) CogAgent: a visual language model for gui agents. External Links: 2312.08914 Cited by: §1. [17] E. Horvitz, J. Breese, D. Heckerman, D. Hovel, and K. Rommelse (1998) The lumière project: bayesian user modeling for inferring the goals and needs of software users. In Proceedings of the Fourteenth Conference on Uncertainty in Artificial Intelligence, UAIâ98, San Francisco, CA, USA, p. 256â265. External Links: ISBN 155860555X Cited by: §1. [18] F. Huq, Z. Z. Wang, F. F. Xu, T. Ou, S. Zhou, J. P. Bigham, and G. Neubig (2025-04) CowPilot: a framework for autonomous and human-agent collaborative web navigation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations), N. Dziri, S. (. Ren, and S. Diao (Eds.), Albuquerque, New Mexico, p. 163â172. External Links: Link, Document, ISBN 979-8-89176-191-9 Cited by: §2.2. [19] L. Jang, Y. Li, C. Ding, J. Lin, P. P. Liang, D. Zhao, R. Bonatti, and K. Koishida (2024) VideoWebArena: evaluating long context multimodal agents with video understanding web tasks. External Links: 2410.19100, Link Cited by: Table 1, §2.1. [20] Y. Jang, Y. Song, S. Sohn, L. Logeswaran, T. Luo, D. Kim, K. Bae, and H. Lee (2025) Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §1. [21] A. Khurana, X. Su, A. Y. Wang, and P. K. Chilana (2025) Do it for me vs. do it with me: investigating user perceptions of different paradigms of automation in copilots for feature-rich software. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI â25, New York, NY, USA. External Links: ISBN 9798400713941, Link, Document Cited by: §1, §2.2. [22] A. Khurana, H. Subramonyam, and P. K. Chilana (2024) Why and when llm-based assistants can go wrong: investigating the effectiveness of prompt-based interactions for software help-seeking. In Proceedings of the 29th International Conference on Intelligent User Interfaces, IUI â24, New York, NY, USA, p. 288â303. External Links: ISBN 9798400705083, Link, Document Cited by: §2.2. [23] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa (2022) Large language models are zero-shot reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS â22, Red Hook, NY, USA. External Links: ISBN 9781713871088 Cited by: §4.1. [24] M. Lee, Z. M. Kim, V. Khetan, and D. Kang (2024) Human-ai collaborative taxonomy construction: a case study in profession-specific writing assistants. In Proceedings of the Third Workshop on Intelligent and Interactive Writing Assistants, In2Writing â24, New York, NY, USA, p. 51â57. External Links: ISBN 9798400710315, Link, Document Cited by: §3.2.1. [25] K. Li, C. Fang, Z. Wang, S. Kim, H. Jin, and Y. Fu (2020) Screencast tutorial video understanding. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1, §1, §2.1, §3. [26] T. J. Li, J. Chen, T. M. Mitchell, and B. A. Myers (2020) Towards effective human-ai collaboration in gui-based interactive task learning agents. In CHI 2020 Workshop on Artificial Intelligence for HCI: A Modern Approach (AI4HCI), External Links: Link Cited by: §2.2. [27] K. Q. Lin, L. Li, D. Gao, W. Qinchen, M. Yan, Z. Yang, L. Wang, and M. Z. Shou VideoGUI: a benchmark for gui automation from instructional videos. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, Cited by: Table 1, §1, §1, §2.1, §3. [28] K. Q. Lin, L. Li, D. Gao, Z. Yang, S. Wu, Z. Bai, S. W. Lei, L. Wang, and M. Z. Shou (2025) Showui: one vision-language-action model for gui visual agent. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 19498â19508. Cited by: §1. [29] G. Liu, P. Zhao, L. Liu, Z. Chen, Y. Chai, S. Ren, H. Wang, S. He, and W. Meng (2025) LearnAct: few-shot mobile gui agent with a unified demonstration benchmark. arXiv preprint arXiv:2504.13805. Cited by: §1. [30] D. Lu, Y. Xu, J. Wang, H. Wu, X. Wang, Z. Wang, J. Yang, H. Su, J. Chen, J. Chen, Y. Mao, J. Zhou, J. Lin, B. Hui, and T. Yu (2025) VideoAgentTrek: computer use pretraining from unlabeled videos. External Links: 2510.19488, Link Cited by: §3. [31] Y. Lu, S. Yang, C. Qian, G. Chen, Q. Luo, Y. Wu, H. Wang, X. Cong, Z. Zhang, Y. Lin, W. Liu, Y. Wang, Z. Liu, F. Liu, and M. Sun (2025) Proactive agent: shifting LLM agents from reactive responses to active assistance. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1. [32] Microsoft Corporation (2025) Microsoft copilot. Note: https://copilot.microsoft.com/Accessed: 2025-11-14 Cited by: §1. [33] S. Nayak, X. Jian, K. Q. Lin, J. A. Rodriguez, M. Kalsi, N. Chapados, M. T. Ăzsu, A. Agrawal, D. Vazquez, C. Pal, P. Taslakian, S. Gella, and S. Rajeswar (2025-13â19 Jul) UI-vision: a desktop-centric gui benchmark for visual perception and interaction. In Proceedings of the 42nd International Conference on Machine Learning, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267, p. 45817â45851. Note: Fine-grained desktop GUI benchmark with dense annotations External Links: Link Cited by: Table 1, §1, §2.1, §3. [34] D. A. Norman (1988) The design of everyday things. Basic Books, New York. External Links: ISBN 0-465-06709-3 Cited by: §1, §3.2.1. [35] OpenAI (2025) GPT-4o system card. External Links: Link Cited by: §4.1, §4.3.3, Table 2, Table 2, Table 3, Table 3, Table 4, Table 4. [36] S. V. PAWAR, B. Pedapudi, P. Kaushik, S. Sivaprasad, M. Fritz, and S. Karande (2025) EARL: early intent recognition in GUI tasks using theory of mind. In ICML 2025 Workshop on Computer Use Agents, External Links: Link Cited by: §2.2. [37] Y. Peng, D. Li, J. P. Bigham, and A. Pavel (2025) Morae: proactively pausing ui agents for user choices. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST â25, New York, NY, USA. External Links: ISBN 9798400720376, Link, Document Cited by: §1. [38] K. Pu, D. Lazaro, I. Arawjo, H. Xia, Z. Xiao, T. Grossman, and Y. Chen (2025) Assistance or disruption? exploring and evaluating the design and trade-offs of proactive ai programming support. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI â25, New York, NY, USA. External Links: ISBN 9798400713941, Link, Document Cited by: §1, §2.2. [39] K. Pu, T. Zhang, N. Sendhilnathan, S. Freitag, R. Sodhi, and T. R. Jonker (2025) ProMemAssist: exploring timely proactive assistance through working memory modeling in multi-modal wearable devices. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, UIST â25, New York, NY, USA. External Links: ISBN 9798400720376, Link, Document Cited by: §2.2. [40] C. H. Song, Y. Song, P. Goyal, Y. Su, O. Riva, H. Palangi, and T. Pfister (2026) Watch and Learn: Learning to Use Computers from Online Videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Note: To appear Cited by: §1. [41] Y. Song, K. Thai, C. M. Pham, Y. Chang, M. Nadaf, and M. Iyyer (2025) BEARCUBS: a benchmark for computer-using web agents. External Links: 2503.07919, Link Cited by: §1. [42] Q. Team (2025) Qwen3 technical report. External Links: 2505.09388, Link Cited by: Figure 4, Figure 4, §4.1, §4.3.4, Table 2, Table 3, Table 4. [43] Y. Wang, X. Li, Z. Yan, Y. He, J. Yu, X. Zeng, C. Wang, C. Ma, H. Huang, J. Gao, M. Dou, K. Chen, W. Wang, Y. Qiao, Y. Wang, and L. Wang (2025) InternVideo2.5: empowering video mllms with long and rich context modeling. arXiv preprint arXiv:2501.12386. Cited by: §4.1, §4.3.1, §4.3.3, §4.3.3, Table 2, Table 3, Table 4. [44] Q. Wu, D. Gao, Q. Lin, Z. Wu, and M. Z. Shou (2025) GUI-narrator: detecting and captioning computer gui actions. In Proceedings of the 33rd ACM International Conference on Multimedia, M â25, New York, NY, USA, p. 3683â3692. External Links: ISBN 9798400720352, Link, Document Cited by: §3. [45] S. Wu, M. Galley, B. Peng, H. Cheng, G. Li, Y. Dou, W. Cai, J. Zou, J. Leskovec, and J. Gao (2025) CollabLLM: from passive responders to active collaborators. In International Conference on Machine Learning (ICML), Cited by: §1. [46] S. Wu, M. Galley, B. Peng, H. Cheng, G. Li, Y. Dou, W. Cai, J. Zou, J. Leskovec, and J. Gao (2025) Collabllm: from passive responders to active collaborators. arXiv preprint arXiv:2502.00640. Cited by: §1. [47] B. Yang, L. Xu, L. Zeng, K. Liu, S. Jiang, W. Lu, H. Chen, X. Jiang, G. Xing, and Z. Yan (2025) ContextAgent: context-aware proactive llm agents with open-world sensory perceptions. External Links: 2505.14668, Link Cited by: §1. [48] Q. Yang, H. Li, H. Zhao, X. Yan, J. Ding, F. Xu, and Y. Li (2025) FingerTip 20k: a benchmark for proactive and personalized mobile llm agents. External Links: 2507.21071, Link Cited by: §1. [49] R. Yang, Y. Jiang, Y. Jiang, P. Kargupta, Y. Zhang, and J. Han (2026) Grounding agent memory in contextual intent. External Links: 2601.10702, Link Cited by: §4.3.4. [50] S. Ye, H. Shi, D. Shih, H. Yun, T. Roosta, and T. Shu (2025) RealWebAssist: a benchmark for long-horizon web assistance with real-world users. arXiv preprint arXiv:2504.10445. Cited by: §1. [51] B. Zhang, Z. Shang, Z. Gao, W. Zhang, R. Xie, X. Ma, T. Yuan, X. Wu, S. Zhu, and Q. Li (2025) TongUI: building generalized gui agents by learning from multimodal web tutorials. arXiv preprint arXiv:2504.12679. Note: Introduces the TongUI framework and GUI-Net-1M dataset External Links: Link Cited by: §1. [52] C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y. Sun, C. Zhang, Z. Zhang, A. Liu, S. Zhu, X. Chang, J. Zhang, F. Yin, Y. Liang, and Y. Yang (2024) ProAgent: building proactive cooperative agents with large language models. External Links: 2308.11339, Link Cited by: §1. [53] J. Zhang, J. Wu, T. Yihua, M. Liao, N. Xu, X. Xiao, Z. Wei, and D. Tang (2024-11) Android in the zoo: chain-of-action-thought for GUI agents. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 12016â12031. External Links: Link, Document Cited by: §1. [54] H. H. Zhao, K. Yang, W. Yu, D. Gao, and M. Z. Shou (2025) WorldGUI: an interactive benchmark for desktop gui automation from any starting point. External Links: 2502.08047, Link Cited by: Table 1, §1, §1, §2.1, §3. [55] Y. Zhao, X. Shu, L. Fan, L. Gao, Y. Zhang, and S. Chen (2025) ProactiveVA: proactive visual analytics with llm-based ui agent. External Links: 2507.18165, Link Cited by: §2.2. [56] S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, Y. Bisk, D. Fried, U. Alon, et al. (2023) WebArena: a realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854. External Links: Link Cited by: §1. [57] J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao, et al. (2025) Internvl3: exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479. Cited by: §4.1, Table 2, Table 3, Table 4. Supplementary Material A Detailed Evaluation Metrics In this section, we provide the formal definitions for the evaluation metrics used across our four evaluation tasks: Behavior State Detection, Intent Prediction, Help Need Detection, and Help Content Prediction. Let N denote the total number of test samples in the dataset. For the i-th sample, let yiy_i denote the ground-truth label and y^i y_i denote the modelâs predicted label. â(â )I(¡) denotes the indicator function, which equals 1 if the condition inside is true and 0 otherwise. A.1 Metric Definitions by Task A.1.1 Task 1: Behavior State Detection This task is formulated as a multi-class classification problem where the model must classify a video segment into one of 9 distinct behavioral states. We evaluate performance using standard Accuracy. Accuracy=1Nââi=1Nâ(y^i=yi)Accuracy= 1N _i=1^NI( y_i=y_i) (1) A.1.2 Task 2: Intent Prediction This task is framed as a Multiple-Choice Question (MCQ) task with 4 options (1 correct answer and 3 distractors). We use two metrics: Accuracy. Measures the proportion of instances where the model selects the correct intent option from the four candidates. Accuracy=1Nââi=1Nâ(y^i=yi)Accuracy= 1N _i=1^NI( y_i=y_i) (2) Multi-Binary Accuracy (MBAcc). Following prior work [5, 8], we employ MBAcc to evaluate robustness against distractors. For a given sample i, let yiy_i be the correct option and iâ=ci,1,ci,2,ci,3C_i^-=\c_i,1,c_i,2,c_i,3\ be the set of three incorrect distractor options. The model performs a pairwise comparison function fâ(x,optA,optB)f(x,opt_A,opt_B) which returns the chosen option between A and B. A prediction is considered correct under MBAcc only if the model prefers the ground truth yiy_i over every distractor in iâC_i^-. MBAcc=1Nââi=1N(âcâiââ(fâ(xi,yi,c)=yi))MBAcc= 1N _i=1^N ( _c _i^-I(f(x_i,y_i,c)=y_i) ) (3) A.1.3 Task 3-1: Help Prediction (Need Detection) This sub-task is a binary classification problem (Help Needed vs. Not Needed). We evaluate this using Accuracy, Precision, Recall, and F1-Score. Let TâPTP (True Positive), TâNTN (True Negative), FâPFP (False Positive), and FâNFN (False Negative) denote the classification counts. ⢠Accuracy: The ratio of correctly predicted observations to total observations. Accuracy=TâP+TâNTâP+TâN+FâP+FâNAccuracy= TP+TNTP+TN+FP+FN (4) ⢠Precision: The ratio of correctly predicted positive observations to the total predicted positives. Precision=TâPTâP+FâPPrecision= TPTP+FP (5) ⢠Recall: The ratio of correctly predicted positive observations to the all observations in the actual class. Recall=TâPTâP+FâNRecall= TPTP+FN (6) ⢠F1-Score: The harmonic mean of Precision and Recall. F1=2â Precisionâ RecallPrecision+RecallF1=2¡ Precision¡RecallPrecision+Recall (7) A.1.4 Task 3-2: Help Prediction (Content Prediction) Similar to Intent Prediction, this sub-task is an MCQ task where the model must select the appropriate help content. It is evaluated using Accuracy and Multi-Binary Accuracy (MBAcc). Accuracy. Accuracy=1Nââi=1Nâ(y^i=yi)Accuracy= 1N _i=1^NI( y_i=y_i) (8) Multi-Binary Accuracy (MBAcc). Defined identically to the Intent Prediction task. Let iâC_i^- be the set of incorrect help content options for the i-th sample. MBAcc=1Nââi=1N(âcâiââ(fâ(xi,yi,c)=yi))MBAcc= 1N _i=1^N ( _c _i^-I(f(x_i,y_i,c)=y_i) ) (9) B Dataset Details We provide a comprehensive overview of the GUIDE dataset, detailing its statistical properties, task granularity, and the diverse range of software workflows it encompasses. B.1 Dataset Statistics GUIDE comprises a comprehensive collection of 120 screen recording videos, totaling approximately 67.5 hours of footage. A key characteristic of our dataset is the inclusion of rich verbal narration; as shown in Table B1, think-aloud narration covers 78% of the total video duration, providing high-quality ground truth for annotating user intent and mental states. Variable Value # Videos 120 Total Duration 67.5 hours Avg. Duration 33 min 44 sec Max Duration 1 hour 23 min 50 sec Min Duration 16 min 42 sec Think-Aloud Narration Ratio 78% Task Samples & Granularity (1) Behavior State Detection 1.8K Avg. Segment Length 14.16s (2) Intent Prediction 1.3K Avg. Segment Length 25.40s (3) Help Prediction 1K Avg. Segment Length 25.56s Table B1: Statistics of the GUIDE dataset. The dataset focuses on long-horizon, open-ended workflows. The average video duration is 33 minutes and 44 seconds, with sessions ranging from approximately 16 minutes to over 1 hour and 23 minutes (Figure B1). This extended duration ensures that the dataset captures the full evolution of user tasks, including periods of exploration, struggle, and error recovery. Task Granularity. From these raw videos, we extracted varying numbers of instances for our three evaluation tasks. We collected 1.8K samples for Behavior State Detection, 1.3K samples for Intent Prediction, and 1K samples for Help Prediction. Notably, the average segment length for behavior detection is shorter (14.16s) compared to Intent Prediction (25.4s) and Help Prediction (25.56s). This is because when annotating behavior states from narration-aligned segments, we instructed the model to split the clip if two or more states were identified. Figure B1: Distribution of screen recording video lengths in the dataset. Figure B2: Software categories represented in the dataset. Diversity. To ensure generalizability, the dataset spans a wide variety of software domains. As illustrated in Figure B2, users interacted with diverse applications ranging from creative design tools to analytical software. B.2 Task Composition To ensure our benchmark captures a comprehensive range of user behaviors, we designed a set of 20 open-ended tasks across five distinct software categories: Photo Editing, Graphic Design, Presentation Design, Video Editing, and Data Analysis. Table B2 provides a detailed overview of these categories and their corresponding tasks. Open-Ended Task Design. Unlike rigid, step-by-step tutorials that result in linear behavior, our tasks are designed to be goal-oriented and open-ended. For instance, while we provided users with necessary materials (e.g., raw video clips, images) and suggested specific software features to utilize, we did not prescribe a fixed execution path or a target reference outcome. This semi-structured ambiguity is intentional; it forces users to engage in high-level planning, trial-and-error exploration, and problem-solving. Consequently, this setup naturally elicits the complex behavior statesâsuch as Exploration, Debugging, and Frustrationâthat GUIDE aims to detect. Domain Diversity. The selected software categories cover a broad spectrum of software domains, ensuring comprehensive coverage of diverse GUI workflows. Our dataset spans creative domains (Photo Editing, Graphic Design) that rely on visual manipulation and aesthetic decisions, analytical domains (Data Analysis) focused on data processing and logic, and hybrid tasks like Presentation Design or Video Editing. This variety ensures that our models are evaluated on their ability to generalize across diverse user interfaces, toolsets, and workflow paradigms. Category Software Tasks Photo Editing Photoshop, GIMP 1. Create a composite from two images. 2. Create a bakery logo with a warm, friendly identity. 3. Replace a photoâs background with a custom-designed pattern. 4. Design a movie poster. Graphic Design Figma, Canva 1. Design a mobile sign-up screen for a fictional app. 2. Design a custom 404 error page with a visual and animated element. 3. Design compact profile cards that display personal user details. 4. Design an event poster for a music festival. Presentation Design PowerPoint, Google Slides 1. Create a product pitch deck that highlights the MacBookâs key features. 2. Create an interactive timeline presenting a companyâs history. 3. Create a 5-slide nature-themed shape-masked photo scrapbook. 4. Create a quiz deck with 3 multiple-choice questions. Video Editing Premiere Pro, CapCut 1. Edit a short interview to improve clarity and engagement. 2. Design a creative intro using animated text. 3. Edit a short instructional video to clearly guide a process. 4. Transform a long video into a highly engaging short-form clip. Data Analysis Microsoft Excel, Google Sheets 1. Design a Gantt chart for a mini project. 2. Summarize and visualize responses from a survey. 3. Visualize student performance across subjects. 4. Summarize and visualize product sales by category or region. Table B2: Overview of open-ended tasks across software categories. Each category includes two software applications and four tasks designed to elicit natural and diverse user behaviors. C User Behavior Taxonomy To effectively assist users, an agent must understand not just what the user is doing (e.g., clicking a mouse), but why they are doing it and what their current cognitive and behavior state is. We introduce a hierarchical taxonomy of 9 user behavior states, organized into four high-level phases of the software workflow: Planning, Execution, Problem-Solving, and Evaluation. Table C3 provides detailed definitions and examples for each state. Behavior State Description Examples Planning Task Understanding and Preparation The user is focused on the logistics of the task. This includes interpreting the task, gathering necessary digital assets, and configuring the software environment. Their goal is to set up the conditions needed to begin the work. Reading task instructions, opening required software/files/templates, arranging workspace (resizing windows, organizing directories), downloading images for photo editing. Ideation and Planning The user is engaged in high-level conceptual work. They are brainstorming ideas, outlining the structure of the outcome, or creating a plan for how to approach the task. This often involves creating preliminary, non-final content that serves as a guide. Formulating high-level strategy, creating step lists, sketching rough layouts or wireframes. The output is a plan or outline, not the final polished product. Execution Exploration and Decision-Making The user experiments with different options or features to understand their effects and decide which one to use. This exploratory phase involves deliberate trial and comparison, often pausing forward progress to evaluate alternatives. Applying effects and undoing them, hovering over tools to see what they do, testing multiple font sizes to decide which fits best. Performing Actions The user is confidently using the software to make progress on the task. These actions are purposeful and executed with little hesitation. Typing/deleting text, inserting and resizing images, applying formatting with clear intent, searching for functions to use. Problem-Solving Frustration The user encounters a blocker and shows signs of being stuck, confused, or annoyed. The system may not behave as expected, or the user cannot find a way to perform a desired action, leading to repetitive or unproductive behavior. Sighing, pausing for long periods, undoing repeatedly, clicking unresponsive elements, complaining about slow system behavior. Debugging The user moves beyond frustration and begins to actively investigate the cause of a problem. They form and test hypotheses to diagnose and fix an issue. Testing alternative approaches, undoing recent actions step by step, forming hypotheses about causes, adjusting settings to identify errors. Seeking External Help The user recognizes a gap in their own knowledge and turns to an external resource for assistance or procedural guidance. Switching to a web browser for solutions, opening tutorials/documentation, consulting AI assistants or colleagues, posting questions in forums. Evaluation Waiting and Monitoring The user is in a passive state, waiting for a system-controlled process to complete before continuing their work. They are unable to take meaningful action and typically observe progress indicators. Watching loading bars or spinners, waiting for exports or rendering to complete. Assessment The user intentionally pauses their work to review and evaluate their output. They examine the result for quality, accuracy, or aesthetics. Zooming in/out to inspect fine details, replaying video snippets for review, comparing results to reference images or previous versions. Table C3: Taxonomy of user behavior states in open-ended GUI workflows. C.1 Behavior State Distribution Figure C3 illustrates the overall distribution of user behavior states across the four high-level phases defined in our taxonomy. Table C4 provides a granular breakdown of these states across the specific evaluation tasks. Figure C3: Distribution of user behavior states across Planning, Execution, Problem-Solving, and Evaluation phases across the videos in the dataset. Table C4: Distribution of behavior state labels across the full dataset and specific evaluation tasks. Note that annotated instances used in the evaluation tasks may involve two or more states (e.g., a single segment containing both Debugging and Seeking External Help). Behavior State Detection uniformly sampled 200 instances from each class. Dataset Intent Prediction Help Need Detection Help Content Prediction Behavior State Count (%) Count (%) Count (%) Count (%) Planning Task Understanding and Preparation 1054 8.51% 216 9.85% 103 5.65% 36 2.82% Ideation and Planning 993 8.02% 282 12.86% 84 4.61% 41 3.21% Execution Exploration and Decision-Making 1847 14.92% 289 13.18% 162 8.89% 114 8.92% Performing Actions 3733 30.15% 697 31.80% 474 26.00% 228 17.84% Problem-Solving Frustration 1068 8.63% 131 5.98% 416 22.82% 415 32.47% Debugging 915 7.39% 103 4.70% 259 14.21% 252 19.72% Seeking External Help 649 5.24% 114 5.20% 85 4.66% 83 6.49% Evaluation Waiting and Monitoring 220 1.78% 33 1.51% 17 0.93% 9 0.70% Assessment 1903 15.37% 327 14.92% 223 12.23% 100 7.82% C.2 Error Analysis: Behavior State Detection Figure C4 presents the normalized confusion matrix for Behavior State Detection (Sec. 4.3.1). The results reveal a critical limitation in current MLLMs: a systemic bias toward interpreting interactions as productive execution while failing to recognize signs of struggle or hesitation. While models achieve reasonable accuracy for visually distinct states like Seeking External Help (0.61) and Performing Actions (0.57), they show near-zero capability in detecting Frustration (0.07) and Debugging (0.04). Instead, these negative states are overwhelmingly misclassified as Performing Actions (39% and 43%, respectively) or Exploration and Decision-Making (31% and 29%). This suggests that models perceive the visual activity of a struggling userâsuch as repeated clicking or rapid mouse movementsâas deliberate progress, lacking the temporal understanding to distinguish between trial-and-error and confident execution. Figure C4: Normalized confusion matrix for user behavior state classification. The most common errors occur when Frustration or Debugging is misclassified as Performing Actions or Exploration and Decision-Making. D Screen Recording Video Examples We present qualitative examples to illustrate the richness of the multimodal data in GUIDE. Software: Canva, Task: Design a mobile sign-up screen for a fictional app (4:11) âNope, thatâs not what I wanted to do. Try again. Alright, letâs do a text box.â (4:40) âAnd we need to make this much smaller so it fits there at the top.â (6:50) âProgress bar for⌠what? I donât know. Just do progress bar. But you know what? Letâs try ChatGPT because maybe they can help us.â (10:07) âOh, thatâs cool, okay. So I can create the progress bar using the free elements with the shapes. Alright, so letâs try and do that. Alright. Elements.â (21:01) âOoh, what is this? Ooh, I like that so much better. All right, weâre gonna keep this. Yes, weâre gonna keep this.â (22:05) âNot necessarily sure how I can⌠isnât there a way to like make it a certain size? How do I do that?â Table D5: Example video illustrating the userâs on-screen actions accompanied by think-aloud narration. Software: CapCut, Task: Design a creative intro using animated text. (4:56)âbut I want to make it a bit more dynamic so all right.â (7:55)âoh I think here it should be only one word appearing at a time.â (11:39) âthere are so many effects that I get a little bit too overwhelmed to see so many.â (17:55) âI would like to add additional different animation for this one so I like to keep this.â (18:42) âI am satisfied with the results so I will export this.â (20:14) âMaybe I would like to make it a bit slower.â Table D6: Example video illustrating the userâs on-screen actions accompanied by think-aloud narration. E Benchmark Task Examples E.1 Behavior State Detection Screenshot User Behavior State âOkay, I downloaded it already. Delete my test, so I donât get confused. I have the video.â Software: Premiere Pro Task: Edit a short instructional video to clearly guide a process. Behavior State: Task Understanding and Preparation The user is preparing their digital workspace before starting the editing task. They locate the necessary video file on their desktop and delete a superfluous âtestâ file to prevent confusion. âI would like to just use this design or the white some minimalistic like iOS design. Oh, this one. This one looks good. Okay, letâs justâŚâ Software: Google Slides Task: Create a product pitch deck highlighting a productâs key features. Behavior State: Exploration and Decision-Making The user is actively browsing and comparing different templates, as shown by the scrolling and hovering behavior. The narration (âThis one looks goodâ) confirms they are evaluating options to make a final decision. âOkay, thatâs strange. Thatâs very strange, honestly.â Software: CapCut Task: Design a creative intro using animated text. Behavior State: Frustration The user verbally expresses confusion (âthatâs strangeâ) after the software behaved in an unexpected way. They are momentarily paused, indicating a blocker in their workflow before they decide on a new course of action. (no narration) Software: Google Sheets Task: Summarize and visualize product sales by category or region. Behavior State: Seeking External Help The user is unable to find a feature and turns to ChatGPT for assistance. They type a question clarifying their problem, wait for the response, and then read the provided instructions. Table E7: Example instances for the (1) User Behavior State Detection task, showing screenshots, think-aloud narration, and the corresponding behavior state. E.2 Intent Prediction Screenshot Intent âSo now that I have the frame as a design base, I need to include the input field for name, email.â Software: Canva Task: Design a mobile sign-up screen for a fictional app. Intent: A: Rename the design file to reflect the new project B: Add the required input fields to the design C: Search for a suitable illustration to use as a header D: Resize the canvas to a custom dimension âokay looks perfect, I need to adjust the end date as wellâ Software: Excel Task: Design a Gantt chart for a mini project. Intent: A: Adjust the end date of the chartâs horizontal axis B: Adjust the date interval of the chartâs horizontal axis C: Reverse the order of the chartâs vertical axis D: Adjust the start date of the chartâs horizontal axis âWhen this slot comes, we should put some kind of image here.â Software: Premiere Pro Task: Transform a long video into a short-form clip. Intent: A: Create a new text layer above the existing video track B: Add an image to a specific empty slot in the timeline C: Apply a transition effect to the end of a video clip D: Add a video clip to the end of the current sequence âPaste, paste, paste, paste. Done. Done.â Software: PowerPoint Task: Create a product pitch deck highlighting a productâs key features. Intent: A: Align the logos with the main text boxes. B: Delete the logos from all the slides. C: Duplicate the logos onto the remaining slides. D: Change the color of the logos on all slides. Table E8: Example instances for the (2) Intent Prediction task, showing screenshots, think-aloud narration, and the corresponding intent. E.3 Help Prediction Screenshot Help âWhere could I insert the text? [âŚ] Iâm just going to, because the help function I donât quite understand, but I can see if I can add it. Find it in Google.â Software: GIMP Task: Create a bakery logo with a warm, friendly identity. Help Content: A: how to add another image as a layer B: find the tool to add text C: remove the image background D: add a background color or shape âI think I made a mistake here and I need to rectify this.â Software: Google Slides Task: Create a quiz deck with multiple-choice questions testing sustainability facts Help Content: A: align the answer choice boxes B: how to create a quiz slide template C: how to fix a self-identified audio related error D: add animation to reveal the correct answer âIâl scale it. I just want to scale this up. How do I keep it?â Software: Photoshop Task: Create a composite from two images. Help Content: A: how to use the perspective or warp transform tools B: center the new layer on the canvas C: how to use layer blend modes D: maintain aspect ratio while scaling âSo I believe this is, this is great. I believe itâs just simple.â Software: Canva Task: Design a custom 404 error page with a visual and animated element. Help Need: A: help needed B: no help needed Table E9: Example instances for the (3) Help Prediction task. For the Help Need Detection task, the top three instances illustrate cases labeled as help needed, while the last row shows an instance labeled as no help needed. F Software Task Outcome Examples Figure F5 presents final artifacts produced by participants from the study. These examples highlight the open-ended nature of the assigned tasks. Despite receiving identical high-level instructionsâsuch as âDesign a poster for a music festivalâ or âCreate a friendly bakery logoââusers produced markedly different results in terms of layout, aesthetic style, and complexity. This diversity confirms that the study elicited non-linear, creative workflows rather than fixed execution. (a) Music event poster design in Canva (top) and Figma (bottom). (b) Bakery logo design in GIMP (top) and Photoshop (bottom). Figure F5: Example outcomes of the assigned tasks. The diversity across outputs reflects the open-ended nature of the tasks. G Human Verification Interface Figure G6: Annotation interface for validating and refining LLM-generated behavior-state labels. Annotators reviewed the predicted labels and the associated reasoning, correcting them if inaccurate. Each video segment was independently verified by two external annotators. Figure G7: Before participating in the annotation, annotators completed a quiz phase where they had to correctly classify example video segments. This process ensured that all annotators possessed a solid understanding of the behavior taxonomy and definitions. H Prompts H.1 Taxonomy of User Behavior State Generation Taxonomy of User Behavior State Generation internallinenumbers* # Goal: Create a comprehensive taxonomy of user mental and behavior states by analyzing the video recording and transcript. Integrate visual observations (screen interactions, UI changes, cursor behavior) with audio/verbal cues (tone, hesitations, verbal expressions) and transcript content (exact quotes, semantic meaning). # Analysis Guidelines: internallinenumbers* - VISUAL EVIDENCE: Describe what you see on screen (tool selections, menu interactions, visual feedback, cursor patterns) - AUDIO EVIDENCE: Note tone of voice, hesitations, exclamations, and vocal expressions - TRANSCRIPT EVIDENCE: Extract precise quotes that reveal mental states and intentions - CROSS-REFERENCE: Connect visual actions with verbal expressions to understand user intent and mental state # Output Format (return as JSON): "taxonomy": [ "label": "âŚ", "definition": "âŚ", "evidence": [ "timestamp": "00:01:32", "modality": "visual", "description": "User clicks on the brush tool and immediately switches to eraser", "significance": "Indicates uncertainty or trial-and-error behavior" , "timestamp": "00:01:35", "modality": "audio", "description": "User says âhmm, thatâs not rightâ with a frustrated tone", "significance": "Verbal confirmation of confusion and frustration" , "timestamp": "00:01:35", "modality": "transcript", "quote": "hmm, thatâs not right, let me try something else", "significance": "Shows problem-solving mindset and willingness to iterate" ] , ... ] # Context: Software: SOFTWARE Task: TASK_NAME # Transcript: TRANSCRIPT_JSON # Video Content: SEE THE ATTACHED FILE. Figure H8: Prompt to generate a taxonomy of user behavior states given demonstration videos. H.2 Data Annotation Behavior State Annotation # Goal internallinenumbers* You are given a screen recording video snippet of a user using the software SOFTWARE. Annotate the video with the provided taxonomy of user mental and behavior states. Include the taxonomy label and reasoning that explains why the label is appropriate for the video. # Instructions 1. Your annotation label must be based on the userâs current, on-screen behavior shown in the video. internallinenumbers* 2. Annotate video based on what you see, but if the label is not clear, use the think-aloud narration as auxiliary data to understand the userâs intent or thought process. internallinenumbers* 3. Be aware that the userâs narration may refer to past actions or future plans. Always align your annotation label with the userâs current behavior at that specific moment in the video. # Output Format (JSON) "label": "âŚ" // one of the labels in the taxonomy, "reasoning": "âŚ" // Explanation on why the label is appropriate for the video, # Full Video Context Software: SOFTWARE Task the user is performing in the full screen recording video: TASK_NAME The full transcript of the user: TRANSCRIPT_JSON # Target Video Snippet Context - Video snippet time range: TRANSCRIPT_JSON[narration_index]["start"] - TRANSCRIPT_JSON[narration_index]["end"] Narration sentences of the video snippet: "TRANSCRIPT_JSON[narration_index]["sentence"]" # Taxonomy TAXONOMY # Video Content SEE THE ATTACHED FILE. Figure H9: Prompt to annotate a given video segment based on the taxonomy of user behavior states. Intent Annotation internallinenumbers* You are analyzing a userâs screen recording video and think-aloud narration while they use software (e.g., video editing, design, or spreadsheet tools). # Goal internallinenumbers* For the given video snippet and its corresponding narration segment, infer what the user was aiming to complete or achieve by the end of this segment - their short-term goal or intention. internallinenumbers* - The goal should represent **an outcome or result** that the user was either finishing or actively working on as the segment ends. internallinenumbers* - Focus on **tangible, result-oriented goals** (e.g., "finish trimming the clip," "adjust the image color," "complete text alignment"). internallinenumbers* - Ignore interface-level or purely procedural descriptions (e.g., "click this," "open that," "drag the layer"). - If no clear goal or outcome is expressed or shown, return "no tangible goal". # Output Format (JSON) "original_narration": "<verbatim narration text>", internallinenumbers* "goal": "<concise description of what the user aimed to complete or was completing by the end of this segment, or âno tangible goalâ>", internallinenumbers* "evidence_narration_snippet": "<the exact portion of the narration text that supports this inferred goal>", "reasoning": "<brief explanation of how this goal was inferred based on the narration and video>" # Full Video Context Software: SOFTWARE Task the user is performing in the full screen recording video: TASK_NAME The full transcript of the user: TRANSCRIPT_JSON # Target Video Snippet Context internallinenumbers* - Video snippet time range: TRANSCRIPT_JSON[narration_index]["start"] - TRANSCRIPT_JSON[narration_index]["end"] seconds - Narration during this snippet: "TRANSCRIPT_JSON[narration_index]["sentence"]" # Video Content SEE THE ATTACHED FILE. Figure H10: Prompt used to annotate a given video segment with the userâs intent. Help Annotation internallinenumbers* You are analyzing a userâs screen recording video and think-aloud narration while they use software (e.g., video editing, design, or spreadsheet tools). # Goal internallinenumbers* For the given video snippet and its corresponding narration segment, infer **what kind of help or guidance the user is looking for**. internallinenumbers* - Focus on identifying explicit requests for help or moments where the user seeks information, clarification, or suggestions. - Help-seeking can appear in two main forms: internallinenumbers* 1) **On-Screen Help Behavior:** The user performs on-screen actions to seek help (e.g., opening a web browser, typing a query into a search engine or LLM, viewing online documentation or tutorials). internallinenumbers* 2) **Narration-Based Help Expression:** The user verbally expresses confusion, uncertainty, or asks questions (e.g., "I donât know how to fix this," "Why is this not showing up?", "How do I do this?"). - Ignore casual statements or comments unrelated to problem-solving. internallinenumbers* - If there is no clear indication that the user is seeking help, or if the type of help they need is implicit or ambiguous, return "no help needed". # Output Format (JSON) internallinenumbers* "help_needed": "<concise but meaningful description of the help sought - focus on the underlying need, such as âexplain masking featureâ, âsuggest alternative filterâ, âdebug export errorâ, or âclarify timeline snappingâ", "help_source": "<âscreenâ, ânarrationâ, âbothâ, or ânoneâ>", "evidence_narration_snippet": "<exact portion of narration that indicates help-seeking (if any)>", internallinenumbers* "evidence_screen_behavior": "<description of what was seen on screen that indicates help-seeking (if any)>", internallinenumbers* "reasoning": "<brief explanation of how the need for help was inferred based on the narration or/and on-screen behavior>" # Full Video Context Software: SOFTWARE Task the user is performing in the full screen recording video: TASK_NAME The full transcript of the user: TRANSCRIPT_JSON # Target Video Snippet Context internallinenumbers* - Video snippet time range: TRANSCRIPT_JSON[narration_index]["start"] â TRANSCRIPT_JSON[narration_index]["end"] seconds - Narration during this snippet: "TRANSCRIPT_JSON[narration_index]["sentence"]" # Video Content SEE THE ATTACHED FILE. Figure H11: Prompt used to annotate a given video segment with whether help is needed and, if so, what specific help is required. Filtering On-Screen Help-Seeking Behavior internallinenumbers* You are analyzing a userâs sequence of help-seeking actions while they use SOFTWARE for the task "TASK_NAME". You are given HELP_DATA, a list of chronological records where each record includes: - index: original index of the record - help: userâs inferred help need (e.g., "find a graphic for a password field") - screen_behavior: observed on-screen action - narration: userâs think-aloud narration # Goal internallinenumbers* - Keep only the help entries where the screen_behavior is about using external applications (e.g., Google Search, ChatGPT, Gemini, YouTube,etc.). Note that it doesnât include referring to the task instructions page. - Return the list of segment ids that should be kept as a JSON object. # Output Format (JSON) Return only valid JSON: "kept_segment_ids": [<int>, <int>, ...] # Data HELP_DATA: HELP_JSON Figure H12: Prompt used to filter on-screen help-seeking behavior from segments previously marked as help needed. Filtering Narration-Based Help-Seeking Behavior You are analyzing a userâs sequence of actions while they use SOFTWARE for the task "TASK_NAME". You are given HELP_DATA, a list of chronological records where each record includes: - index: original index of the record - help: userâs inferred help need (e.g., "find a graphic for a password field") - screen_behavior: observed on-screen action - narration: userâs think-aloud narration # Goal Keep only segments where the narration explicitly asks for help or guidance. # Definition of explicit help The narration contains a direct request for help or instruction, for example: - "help me", "i need help", "i want help", "can you help", "please help" - How or where questions about operating the software, such as: "how do i ...", "how can i ...", "how to ...", "where is ...", "which option should i ...", "what should i click", "what does this do". - Requests for instructions or explanation: "show me how to ...", "tell me how to ...", "could someone explain ...", "is there a way to ..." # Exclude the following - Exploration or intent without a help request: "iâm going to try", "let me see", "i will search" internallinenumbers* - Trial and error or self-correction without a request: "no, not that", "okay now i got it", "ah ok", "finally" internallinenumbers* - Uncertainty alone: "maybe", "i think", "not sure" unless followed by a direct question that asks for guidance - Statements addressed to self that do not ask for help: "i need to add text", "iâm looking for an icon" - Generic questions not tied to getting guidance on what to do next # Output Format Return only valid JSON: "kept_segment_ids": [<int>, <int>, ...] # Data HELP_DATA: HELP_JSON Figure H13: Prompt used to filter narration-based help-seeking behavior from segments previously marked as help needed. Filtering No Help Needed You are analyzing a userâs sequence of actions while they use SOFTWARE for the task "TASK_NAME". You are given NO_HELP_DATA, a chronological list of records where each record includes: - index: original segment index - start_time, end_time: time range of the segment - no_help_reasoning: explanation for why the user does not need help - narration: the userâs think-aloud narration # Goal Identify up to 5 segments where it is **explicitly clear** that the user does not need help. These are moments when the user: - Performs actions confidently, smoothly, and intentionally. - Demonstrates clear understanding of what to do next without hesitation or correction. - Speaks in a calm, matter-of-fact tone (e.g., "Now Iâl add text here," "Perfect," "That looks good."). # Exclude the following: internallinenumbers* - **Trial and error:** any sign of experimentation, correction, or rapid alternation (e.g., "No, no, no," "Let me try again," "Okay, that worked.") internallinenumbers* - **Self-resolution after confusion:** phrases like "now I got it," âfinally," "oh, thatâs how," "ah okay," "I see," or any narration showing realization after failure or surprise. internallinenumbers* - **Frustration or emotional reactions:** (e.g., "oh shit," "ugh," "why," "come on") even if followed by success. - **Uncertain or exploratory speech:** "I think," "maybe," "letâs see," "try," "not sure." - **Segments that merely lack confusion** but do not clearly express confidence or mastery. # Rules - Select at most 5 segments that best reflect calm, deliberate, fluent progress. - Preserve the original order of the selected segments. - Return only the segment indices of the kept entries. # Output Format (JSON) Return only valid JSON: "kept_segment_ids": [<int>, <int>, ...] # Data NO_HELP_DATA: NO_HELP_JSON Figure H14: Prompt used to filter clear no-help-needed segments. H.3 Model Evaluation Behavior State Detection # Goal You are given a screen recording video snippet of a user working in SOFTWARE. Classify the video into one of the labels from the provided taxonomy of user mental and behavioral states. Include both the taxonomy label and a reasoning that explains why this label best fits the observed segment. # Instructions 1. Base your classification on the userâs on-screen behavior shown in the video. 2. Provide: - label: one of the taxonomy labels - reasoning: a concise explanation of why this label fits the segment internallinenumbers* 3. Return the output strictly in valid JSON, with keys "label" and "reasoning". Do NOT wrap the JSON in markdown code blocks. Return only the raw JSON object. # Output Format (JSON) "label": "âŚ", "reasoning": "âŚ" # Video Context Software: SOFTWARE Task performed in the full recording: TASK_NAME Start and end times of the snippet (relative to the full recording): start_time - end_time seconds # Taxonomy Descriptions TAXONOMY # Previous Segment Context (*** Optional based on the condition) The user behavior in the immediately preceding segment was previous_label: label_definition # Video Content SEE THE ATTACHED FILE. Figure H15: Prompt used to evaluate the model on the (1) Behavior State Detection task. Intent Prediction # Goal You are given a screen recording snippet of a user working in SOFTWARE. Predict which option best describes the userâs intention during this segment. # Instructions 1. Watch the video and analyze on-screen actions. 2. Select the option (A-D) that best matches the goal of the user trying to achieve. 3. Use the provided behavior context to interpret the goal. (*** Optional based on the condition) 4. Return output in JSON: - label: one of A-D - reasoning: short explanation for your choice 5. Output only a valid JSON object (no Markdown). # Output Format (JSON) "label": "A", "reasoning": "..." # Video Context Software: SOFTWARE Task performed in the full screen recording: TASK_NAME Start and end times of the snippet (relative to the full recording): start_time - end_time seconds # User Behavior Context (*** Optional based on the condition) The following user behavior is identified: label: label_definition. Consider this context when predicting the userâs intention and selecting the most appropriate option. # Options options_text # Video Content SEE THE ATTACHED FILE. Figure H16: Prompt used to evaluate the model on the (2) Intent Prediction task. Help Need Detection # Goal You are given a screen recording video snippet of a user working in SOFTWARE. internallinenumbers* Based on the video content, the userâs intention, and the user behavior context (*** Optional based on the condition), determine if the user needs help or not in this segment. # Instructions 1. Watch the video and observe the userâs behavior and actions. internallinenumbers* 2. Consider the userâs intention and behavior context provided to better understand what they are trying to accomplish and their current state. (*** Optional based on the condition) 3. Determine if the user needs help or not. 4. Provide: - label: "yes" if the user needs help, "no" if they do not need help internallinenumbers* - reasoning: a concise explanation of why the user does or does not need help based on the observed behavior. 5. Output only a valid JSON object (no Markdown). # Output Format (JSON) "label": "yes" | "no", "reasoning": "..." # Video Context Software: SOFTWARE Task performed in the full screen recording: TASK_NAME Start and end times of the snippet (relative to the full recording): start_time - end_time seconds # User Behavior Context (*** Optional based on the condition) The following user behavior is identified in order: label: label_definition. Consider this behavior context when determining if the user needs help. # User Intention (*** Optional based on the condition) The userâs intention or goal during this segment is: intention Consider this intention when determining if the user needs help to achieve this goal. # Video Content SEE THE ATTACHED FILE. Figure H17: Prompt used to evaluate the model on the (3-1) Help Need Detection task. Help Content Detection # Goal You are given a screen recording video snippet of a user working in SOFTWARE. internallinenumbers* Based on the video content, the userâs intention, and the user behavior context (*** Optional based on the condition), predict which option best describes what kind of help or guidance the user is looking for during this segment. # Instructions 1. Watch the video and predict what help the user might need at this segment. internallinenumbers* 2. Consider the userâs intention and behavior context provided to better understand what they are trying to accomplish and their current state. (*** Optional based on the condition) internallinenumbers* 3. Select the option (A, B, C, or D) that best matches what help the user needs to accomplish their intention given their behavior context. 4. Provide: - label: one of the option letters (A, B, C, or D) - reasoning: a concise explanation of why this option best fits the observed segment 5. Output only a valid JSON object (no Markdown). # Output Format (JSON) "label": "A", "reasoning": "..." # Video Context Software: SOFTWARE Task performed in the full screen recording: TASK_NAME Start and end times of the snippet (relative to the full recording): start_time - end_time seconds # User Behavior Context (*** Optional based on the condition) The following user behavior is identified in order: label: label_definition. Consider this behavior context when predicting what help the user might need. # User Intention (*** Optional based on the condition) The userâs intention or goal during this segment is: intention # Options options_text # Video Content SEE THE ATTACHED FILE. Figure H18: Prompt used to evaluate the model on the (3-2) Help Content Prediction task.