Paper deep dive
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/16/2026, 3:08:12 AM
Summary
The paper introduces SPIEval, a human-curated benchmark designed to evaluate Large Language Models (LLMs) as mobile assistants handling scattered personal information across multiple applications. The benchmark comprises 250 tasks grounded in 4,335 records across 10 simulated apps, assessing five cognitive capabilities: reasoning, disambiguation, integration, preference inference, and multi-intent decomposition. Evaluation of nine representative LLMs reveals significant performance gaps, with the best model achieving only 57.3% accuracy. The primary failure mode is inaccurate information localization, where models commit to plausible but incorrect data rather than verifying through retrieval.
Entities (12)
Relation Signals (11)
GPT-5.5 XHigh → achievesaccuracy → 57.3%
confidence 95% · The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy
SPIEval → assessescapability → Reasoning
confidence 95% · SPIEval... grounded in five cognitive capabilities (i.e., reasoning...)
SPIEval → assessescapability → Disambiguation
confidence 95% · SPIEval... grounded in five cognitive capabilities (i.e., ... disambiguation ...)
SPIEval → assessescapability → Integration
confidence 95% · SPIEval... grounded in five cognitive capabilities (i.e., ... integration ...)
SPIEval → assessescapability → Preference Inference
confidence 95% · SPIEval... grounded in five cognitive capabilities (i.e., ... preference inference ...)
SPIEval → assessescapability → Multi-Intent Decomposition
confidence 95% · SPIEval... grounded in five cognitive capabilities (i.e., ... and multi-intent decomposition).
SPIEval → comprises → 4,335 personal records
confidence 95% · spanning 4,335 personal records
SPIEval → comprises → 250 tasks
confidence 95% · SPIEval comprises 250 tasks
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.10692v1
- Canonical: https://arxiv.org/abs/2608.10692v1
Trouble viewing inline? Open PDF directly →
Full Text
78,013 characters extracted from source content.
Expand or collapse full text
Preprint. SPIEVAL: EVALUATING LARGE LANGUAGE MODELS AS MOBILE ASSISTANTS OVER SCATTERED PERSONAL INFORMATION Junjie Ye 1,2 , Zhuohui Sheng 1 , Shaofan Liu 1 , Yulun Zhu 1 , Wenjie Fu 1 , Dingwei Zhu 1 , Ming Zhang 1 , Yujiong Shen 1 , Weichao Wang 2 , Xin Zhao 2 , Shihan Dou 1 , Tao Gui 1 , Qi Zhang 1 , Xuanjing Huang 1 , Pluto Zhou 2 1 Fudan University, 2 Tencent Hunyuan Team jjye23@m.fudan.edu.cn ABSTRACT Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEVAL, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition).SPIEVAL comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi- turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environ- ments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/d atasets/Junjie-Ye/SPIEval. 1INTRODUCTION Due to their excellent instruction-following (Ye et al., 2023; Lou et al., 2024) and tool-use capabilities (Qin et al., 2025; Qu et al., 2025), large language models (LLMs) (Anthropic, 2026; Google, 2026; OpenAI, 2026) are increasingly employed as mobile assistants (Hu et al., 2025; Liu et al., 2025a). In this setting, they must leverage personal information scattered across multiple applications (apps) to complete users’ brief and underspecified instructions. As shown in Figure 1, the instruction “Call my manager” specifies neither the manager’s identity nor the preferred calling method. To fulfill this request, the assistant must proactively identify the manager from meeting records, retrieve the corresponding phone number from the contacts app, and initiate a video call based on the user’s preferences recorded in the notes app. Significant research efforts have been devoted to evaluating LLM-based mobile assistants (Liu et al., 2024; Xie et al., 2024; Xu et al., 2025). Some studies focus on app-operation capabilities. For instance, Trivedi et al. (2024) develop a simulated environment spanning nine apps to evaluate whether models can correctly orchestrate API calls. Other studies investigate the safety risks associated with app operation (Lin et al., 2026), while more recent efforts examine whether models can leverage personal data to generate personalized responses (Mok et al., 2025). 1 arXiv:2608.10692v1 [cs.CL] 11 Aug 2026 Preprint. User Call my manager. Mobile Assistant Who?Whichnumber?How? Meeting open › 1 Contacts open › Notes open › Search apps search_meeting Findmy manager’sname. Daniel Reed FindDaniel’s phonenumber. +1 415-739-2098 search_notes Find thecall preferenceforDaniel. Video Call DR +1 415-739-2098 Video Call... search_contacts 9:419:41 3 2 ✓ Daniel Reed— +1 415-739-2098 · · DaisyMiller— +1 408-552-7810 Dana Tim— +1 650-441-9072 ✓ Onboarding Welcome — Manager Daniel Reed Intro · · Weekly Report — Procurement Department Promotion Meeting — VP → GM ✓ Call Preferences — Prefer Video Calls with Daniel · · Team Directory — Procurement Dept Personal Info — Department · Manager · Hire Date Figure 1: An example of a mobile assistant completing a user instruction by leveraging personal information scattered across multiple apps. To execute the instruction “Call my manager,” the assistant must proactively identify the manager, retrieve the corresponding phone number, and determine the preferred calling method from different apps before placing the call. Despite targeting mobile assistants, these benchmarks primarily evaluate tool use and task execution under settings where the required information is explicitly provided or directly accessible. Specifically, user queries often provide the necessary parameters explicitly (Trivedi et al., 2024; Xie et al., 2024), requiring models to map these parameters to appropriate API calls. Even when certain inputs are not specified in the initial instruction, they are typically returned directly by previous tool calls (Ye et al., 2025; Zhang et al., 2025a). While personalization-oriented benchmarks introduce personal data, access to such information is generally limited to retrieving information from individual documents rather than locating information across multiple apps (Tan et al., 2025). As a result, these benchmarks place limited emphasis on the challenge of scattered personal information, leaving the effectiveness of LLM-based mobile assistants in this setting largely unexplored. To address this gap, we propose SPIEVAL, a human-curated benchmark for evaluating mobile assistants in scenarios with scattered personal information.The benchmark comprises 250 tasks covering five cognitive capabilities essential to this setting (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). These tasks are grounded in 4,335 records spanning 10 apps. To support multi-turn interactions, SPIEVAL provides 21 tools, including 11 retrieval tools and 10 execution tools. Detailed analysis shows that SPIEVAL features diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes, making it a rigorous benchmark for evaluating mobile assistants. We conduct a comprehensive evaluation of nine representative LLMs. The results reveal substantial variation across models and reasoning efforts, indicating considerable room for improvement. Even the strongest model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis shows that 79% of failures arise from inaccurate localization of personal information, as LLMs tend to commit to plausible but incorrect information rather than continue retrieving for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial differences in search efficiency across models. These findings suggest that information localization constitutes the primary bottleneck for current mobile assistants and point to promising directions for future research. 2RELATED WORK LLM-Based Mobile AssistantsMobile assistants represent an important application of LLMs in personalized settings and have attracted widespread attention (Hu et al., 2025; Liu et al., 2025a). Early efforts demonstrate that LLMs can automate multi-step mobile tasks by leveraging commonsense knowledge for action planning (Wen et al., 2024). Subsequent work extends these capabilities to real-world environments, enabling end-to-end task completion from natural language instructions (Guan et al., 2023). More recent systems improve cross-app coordination through multi- agent architectures (Wang et al., 2024; Sun et al., 2025), while self-evolving frameworks enable assistants to accumulate experience and continuously improve through interaction (Liu et al., 2024; Wang et al., 2025). These advances have also driven commercial adoption, with products such as 2 Preprint. §3.2 Five cognitive capabilities User Profile Occupation FamilyMembers Phone Number Payment Account Current Timestamp Retrieve ×11 Per-App ×10Global ×1 Name Apps Tools Create4 Components Linked by Name 1 Design Scattered Info 3 2 QualityControl Apps& ToolsReview Tools: Cross-Validation EnvironmentCheck 3 1 2 Reasoning Find Retrieval Chains Disambiguation Select Valid Records Preference Inference InferPreference Multi-Intent Decomposition Decompose Into Subtasks Integration Aggregate All Sources Instruction Manager = Daniel Reed Name = Daniel Reed Phone = +1 415-739-2098 Execute ×10 Per-App ×10 Expected Outcomes Toolcalls Constructed Environment Apps: Create task Attempt to solve AnswersMust Match Reasoning Path Metting → Contacts → Save InstructionSave my manager’s contact info RecordsTask-Specific App Records Gold AnswerContacts_Create(name, phone) Contacts Voicemail AlarmMeetingAccommodation SMS Notes ScheduleTransportTransactions Fields Mirror Real Apps Complete & Unambiguous Figure 2: Framework of SPIEVAL. Bottom: SPIEVAL comprises 10 commonly used apps, together with 10 execution tools and 11 retrieval tools for accessing records across these apps. All records are associated with a unified user profile. Middle: Each task is designed to evaluate a specific cognitive capability. To ensure quality, the user instruction, personal records, reasoning process, and gold answer are manually constructed and independently verified by at least two annotators. Top: SPIEVAL evaluates five cognitive capabilities essential for handling scattered personal information. Apple Intelligence 1 , Doubao Mobile Assistant 2 , and Honor YOYO Agent 3 bringing intelligent task automation to hundreds of millions of users. As these systems become widely deployed, evaluating their capability to handle real-world personalized tasks becomes increasingly important. In this paper, we focus on the challenge of scattered personal information, where the information required to fulfill a user request is distributed across multiple apps. Evaluation of Mobile Assistants A growing body of work proposes benchmarks for evaluating LLM-based mobile assistants in app-centric settings. Early benchmarks focus on multi-step task execution through tool calls, evaluating whether agents can plan and orchestrate operations within simulated environments (Trivedi et al., 2024; Chen et al., 2025). Subsequent efforts expand the evaluation scope along several dimensions, including dynamic conditions with asynchronous events and temporal constraints (Froger et al., 2026) as well as safety awareness during app operations (Lin et al., 2026). More recently, personalization-oriented benchmarks evaluate models’ capabilities to leverage personal data for generating tailored responses (Mok et al., 2025; Tan et al., 2025). However, these benchmarks either provide all necessary information directly in the user instructions or restrict access to personal information to simple document retrieval. As a result, they do not capture a common real-world setting in which fulfilling a user request requires locating information distributed across multiple apps. To address this gap, we introduce SPIEVAL, a benchmark that evaluates mobile assistants under this challenging setting of scattered personal information. 3SPIEVAL 3.1TASK FORMULATION Given a natural-language user instruction q and a set of apps A = a 1 ,a 2 ,...,a K , where each app a k contains a set of structured personal records R k , the objective of a mobile assistant is to generate a sequence of tool calls that fulfills the instruction. In contrast to fully specified 1 https://w.apple.com/apple-intelligence/ 2 https://o.doubao.com/ 3 https://w.honor.com/cn/magic-os/ 3 Preprint. instructions, q does not explicitly provide all information required to complete the task. The model is equipped with a set of tools T = T retrieve ∪ T exec and must proactively formulate search queries to retrieve relevant records, reason over the retrieved information to infer the required parameters, and invoke execution tools with the inferred arguments. Formally, the model generates a trajectory τ = (c 1 ,r 1 ,c 2 ,r 2 ,...,c n ,r n ), where each c i denotes a tool call and r i the corresponding tool feedback. The trajectory terminates with one or more execution calls. 3.2COGNITIVE CAPABILITIES Figure 2 summarizes five cognitive capabilities required for handling scattered personal information in mobile assistant settings. ReasoningMulti-hop reasoning, a capability extensively studied in other domains (Yang et al., 2018), is also vital for mobile assistants, where fulfilling a request typically requires a series of interdependent retrieval steps.The instruction “Save my manager’s contact information” can illustrate this dependency, since resolving the manager’s identity enables retrieval of the corresponding phone number, which in turn supports the final save action. DisambiguationDisambiguation is a well-known challenge in information retrieval (R ̈ ucker & Akbik, 2025), and becomes even more difficult in mobile assistant settings due to ambiguous and evolving personal data. A possible case involves saving a contact’s phone number when multiple numbers are associated with the same individual across different records or time periods. Completing this task requires identifying the currently valid number by leveraging contextual cues together with record-specific information. IntegrationIn contrast to disambiguation, integration requires aggregating information from all relevant sources (Zhu et al., 2024). Unlike the sequential dependency chains characteristic of multi-hop reasoning, these sources are often independent of one another. The instruction “Save all suppliers from this week’s trip” can exemplify this challenge, as it involves collecting supplier names and phone numbers that may be distributed across SMS messages, notes, and voicemail records. Preference InferencePreference inference is particularly important in personalized scenarios, as users often have habitual preferences that are never explicitly stated (Zhang et al., 2025b). A competent assistant must infer such preferences. This challenge may arise when a user says “Call my wife,” where the model must not only locate the appropriate contact information but also recognize based on prior call history that the user may typically prefer video calls over voice calls. Multi-Intent Decomposition Multi-intent decomposition refers to the capability to decompose a complex instruction into multiple independent subtasks (Liu et al., 2025b). Each subtask involves one of the preceding cognitive capabilities. The instruction “Call Dad and transfer this month’s living expenses” can combine a phone call and a payment transaction, requiring the model to decompose the request and retrieve the corresponding parameters from different sources. 3.3BENCHMARK CONSTRUCTION We construct SPIEVAL through four components, followed by a rigorous quality control process. User Profile Construction To capture realistic mobile assistant usage scenarios, we construct a unified user profile that establishes a consistent identity across all tasks, which is provided to the model as part of the system prompt. 4 The profile specifies a set of core attributes, including the user’s occupation, family members, a personal phone number, a payment account, and the current timestamp. These shared attributes allow user instructions to contain natural references such as “my dad” or “my department manager,” which the assistant must resolve correctly. Furthermore, each task is grounded in a distinct set of app records that introduce task-specific entities and information, such as colleagues, clients, and social contacts, thereby enabling diverse scenarios. Application ConstructionTo reflect the diverse sources of personal information, we simulate 10 mobile apps commonly found on personal devices, including Accommodation, Alarm, Contacts, Meeting, Notes, Schedule, SMS, Transactions, Transport, and Voicemail. Each app is abstracted from its real-world counterpart and defined by a structured schema with domain-specific fields, 4 The complete user profile and system prompt are provided in Appendix A. 4 Preprint. Table 1: Distribution of 357 execution operations across the five cognitive capabilities and ten operation categories in SPIEVAL. Capability Book Accommodation Set Alarm Make Call Save Contact Create Meeting Send Message Create Note Make Payment Create Schedule Book TransportTotal Reasoning555556555551 Disambiguation555555555550 Integration101251757115121397 Preference Inference555555555853 Multi-Intent Decomposition9107891110141216106 Total34372740293436343947357 comprising 8.1 fields on average. 5 In particular, each schema distinguishes required fields (e.g., a contact’s name and phone number) from optional fields (e.g., company or birthday). Consequently, records are often only partially populated, and complete information about an entity may need to be assembled from records distributed across different apps. In addition, related fields are shared across apps, allowing records from different sources to be linked through common attributes such as names, phone numbers, and account numbers. Tool ConstructionTo enable multi-turn interaction and support fine-grained analysis of model behavior, we design 21 tools organized into two complementary categories. One category comprises 11 retrieval tools that mirror the information access mechanisms available on mobile devices. Specifically, each app is equipped with a dedicated retrieval tool, alongside a global one. The per-app tools provide three retrieval modes, which are substring matching, regular expression matching, and fuzzy matching. They also support field-specific targeting and case-sensitivity control for precise information access. In contrast, the global tool performs substring-based search across all apps and returns results annotated with their source apps. All retrieval results are returned in a paginated manner, reflecting real-world search interfaces and requiring models to actively request additional results when necessary. The other category comprises 10 execution tools, one for each app, which serve as the execution endpoints for task completion. Each execution tool is accompanied by detailed parameter specifications, including type annotations and distinctions between required and optional fields, enabling parameter-level evaluation of action correctness. 6 Task ConstructionTo ensure that each task admits a unique and verifiable solution, we adopt a fully manual construction process. For each task, an annotator first selects one of the five cognitive capabilities and designs the underlying information structure, specifying how task-relevant information is distributed across apps and connected through shared attributes. Based on this structure, the annotator constructs four components, including a natural-language user instruction, a set of task-specific app records, a step-by-step reasoning annotation, and a gold answer consisting of the exact execution tool calls with all parameters. The records are populated independently for each task such that all information required for task completion is available, but only through the intended retrieval and reasoning process. Following this procedure, we construct 250 tasks, with 50 tasks for each cognitive capability, grounded in 4,335 records distributed across 10 apps. Quality ControlThe benchmark is constructed over a period of three months by six NLP researchers. To guarantee the reliability and consistency of the benchmark, we implement quality control throughout the entire construction process. During the design of apps and tools, each schema is reviewed by all annotators to ensure that the abstracted fields faithfully reflect real-world app functionality and that tool definitions are complete and unambiguous. For task construction, we employ a cross-validation protocol in which one researcher creates a task and another independently attempts to solve it without access to the gold answer. A task is accepted only if both researchers arrive at the same answer through the intended retrieval process; otherwise, it is revised to eliminate ambiguities or unintended solution paths. In addition, all gold answers are executed against the tool implementation to verify that the corresponding tool calls produce the expected outcomes. The benchmark therefore admits 100% human performance by construction. All personal data used in the benchmark is entirely fictional and does not correspond to any real individual. 5 The detailed schema of each app is provided in Appendix B. 6 The complete schema of all tools is provided in Appendix C. 5 Preprint. 02040 60 80 0 10 20 30 40 mean 34 mean 8.5 Figure 3: Relationship between instruction length and the total number of execution parameters required for each task, with marginal distributions shown alongside. 02040 60 80100 10 13 16 19 22 25 28 31 34 37 40 mean 17.3 records 010203040 50 2 3 4 5 6 7 8 9 10 mean 6.2 apps Figure 4:Distribution of task-relevant records across apps in SPIEVAL.Top: Number of records per instruction. Bottom: Number of apps spanned by records. 3.4DATASET ANALYSIS To demonstrate that SPIEVAL provides a rigorous evaluation of mobile assistants operating over scattered personal information, we analyze the benchmark along five dimensions. 7 Diverse Scenarios Mobile assistants are expected to handle a wide range of requests, requiring them to retrieve records through diverse strategies. A representative benchmark should cover a broad spectrum of tasks to enable systematic evaluation. To this end, each task is designed by jointly considering its required cognitive capability and execution operation. As shown in Table 1, the dataset comprises 357 execution operations spanning 10 operation categories. Each operation category is combined with all five cognitive capabilities, so that all 50 capability-operation pairs are represented, each occurring at least five times. Moreover, even the least frequent operation category appears 27 times, ensuring that no operation is underrepresented. Challenging TasksTasks for mobile assistants require models to interpret ambiguous user instructions, identify appropriate tools, infer the required parameters from the corresponding tool descriptions, and retrieve necessary information from personal data distributed across multiple apps, all of which are captured by SPIEVAL. As shown in Figure 3, instructions in SPIEVAL are highly concise, averaging only 34 characters, while the corresponding execution workflows require an average of 8.47 parameters. Notably, one instruction consists of only 20 characters yet requires 45 execution parameters. Moreover, user instructions explicitly provide none of the required parameters, leaving both tool selection and parameter inference largely implicit. Consequently, solving SPIEVAL requires robust semantic understanding and effective information retrieval, making it representative of the challenging tasks encountered in the real world. Scattered Information Personal information on a mobile phone is stored in structured form across many apps, and task-relevant records are rarely confined to a single app. As a result, fulfilling an instruction requires retrieving and reconciling records from multiple apps. Figure 4 quantifies this distribution. Each instruction in SPIEVAL is associated with an average of 17.3 records spanning 6.2 of the 10 apps, with the most demanding instructions involving up to 39 records. This organization requires models to understand the relationships among apps, identify the relevant records, and integrate evidence from multiple sources to infer the necessary information. Controllable EnvironmentsA fair benchmark should attribute performance differences to the models themselves rather than inconsistencies in the runtime environment. Therefore, SPIEVAL provides a fully controllable environment in which all records are preloaded and all retrieval and execution tools are implemented locally. Whenever a model invokes a tool, the environment validates the tool call against the tool specification and returns an informative error message if the input is invalid. Otherwise, the tool executes deterministically, ensuring that identical inputs always produce identical outputs. This design supports multi-turn interactions while ensuring reproducibility, enabling more reliable comparisons and targeted analyses of model behavior. 7 A systematic comparison between SPIEVAL and existing benchmarks is provided in Appendix D. 6 Preprint. Table 2: Main results on SPIEVAL. Each model is evaluated under its highest and lowest reasoning effort levels, indicated in parentheses. Human accuracy is 100% for all categories by construction. ModelReasoning Disam- Integration PreferenceMulti-Intent Overall biguationInferenceDecomposition GPT-5.5 (xhigh)70.04.372.00.073.30.938.72.532.72.557.31.2 Gemini 3.1 Pro (high)74.04.373.30.974.71.920.04.923.31.953.12.2 Claude Opus 4.8 (max)68.73.458.73.467.31.934.74.132.01.652.30.7 DeepSeek-V4-Pro (max)53.30.962.74.160.04.334.02.822.71.946.50.4 Gemini 3.1 Pro (low)65.31.962.75.060.76.216.74.123.33.845.72.2 Claude Opus 4.8 (none)48.73.440.75.046.03.326.72.519.30.936.31.7 Kimi K2.6 (high)48.71.946.76.643.30.925.32.516.01.636.02.3 GLM-5.2 (max)44.75.044.71.949.30.922.75.214.73.435.21.4 Hy3 (high)44.02.841.31.946.04.326.02.814.01.634.31.9 DeepSeek-V4-Pro (none)36.02.839.31.948.75.028.71.918.70.934.31.4 Qwen3.7-Plus (high)44.02.838.70.946.73.823.33.414.73.833.50.7 Seed-2.1-Pro (high)50.04.944.04.339.32.517.30.916.02.833.31.0 GLM-5.2 (none)37.32.532.04.346.72.519.31.910.72.529.21.1 GPT-5.5 (none)41.35.235.33.430.74.722.02.813.33.828.51.1 Seed-2.1-Pro (minimal)34.72.532.71.927.35.018.74.19.32.524.50.8 Qwen3.7-Plus (none)34.73.427.32.536.01.612.71.98.70.923.90.5 Hy3 (no think)25.32.524.71.923.310.513.35.06.01.618.51.6 Kimi K2.6 (none)17.30.926.02.819.37.511.32.58.00.016.42.0 Verifiable Outcomes A fundamental requirement of any benchmark is that its evaluation reflects model capability. To satisfy this requirement, SPIEVAL avoids the LLM-as-a-judge paradigm, which is susceptible to model bias and output variability, and instead verifies model outputs against human-annotated gold answers. Each gold answer specifies the execution tools required for a task together with their parameter values, enabling automatic evaluation at the parameter level without subjective judgment. Certain parameters may take any valid value, such as the number of alarm repetitions, and are marked as unconstrained, in which case only their types are validated. For the remaining parameters, we annotate every acceptable value, and 7.0% of them admit multiple valid values. For the 76 tasks that require multiple execution tools, the evaluation is invariant to tool invocation order. Together, these design choices provide a consistent, transparent, and reproducible evaluation protocol for comparing different models. 4EXPERIMENTAL SETUP We describe the models, metrics, and implementation details to facilitate reproducibility. Models We evaluate nine representative state-of-the-art LLMs, including Claude Opus 4.8 (An- thropic, 2026), DeepSeek-V4-Pro (DeepSeek-AI et al., 2026), Gemini 3.1 Pro (Google, 2026), GLM-5.2 (Zhipu AI, 2026), GPT-5.5 (OpenAI, 2026), Hy3 (Tencent Hunyuan Team, 2026), Kimi K2.6 (Kimi, 2026), Qwen3.7-Plus (Qwen Team, 2026), and Seed-2.1-Pro (Bytedance Seed, 2026). Metrics To objectively evaluate model performance, we adopt a binary accuracy metric. Because a model may adapt its behavior based on tool feedback, evaluating intermediate steps is neither necessary nor appropriate. Instead, we compare the model’s final execution tool calls against the annotated gold answers. An outcome is considered correct only if the invoked execution tools and all of their parameter values exactly match one of the annotated gold answers. Implementation DetailsFor each model, we evaluate performance under both the highest and lowest available reasoning effort levels. All other hyperparameters are kept at their default values to maximize each model’s performance. Since mobile assistants are subject to latency constraints, we limit each task to a maximum of 50 interaction turns, balancing sufficient interaction with the environment against unnecessary computation. All retrieval results are returned in a paginated manner, with at most five results per page, requiring models to actively request additional pages when needed. To mitigate sampling variability, we conduct three independent runs for each experimental setting and report the mean and standard deviation across runs. 7 Preprint. Reasoning Disambiguation Integration Preference Inference Multi-Intent Decomposition Overall 0 20 40 60 80 100 Paginated Full No-search Figure 5: Accuracy under the standard pag- inated protocol and two controlled settings, averaged over all model configurations. Wrong value 79% Wrong count 13% Missing param 6% Others 2% Figure 6: Distribution of error types from GPT-5.5 (xhigh), Gemini 3.1 Pro (high), and Claude Opus 4.8 (max). 5MAIN RESULTS Table 2 summarizes the performance of LLMs, from which we draw the following observations. Current LLMs struggle with information localization, making SPIEVAL a challenging benchmark. Even GPT-5.5 (xhigh), the strongest evaluated model, achieves an accuracy of only 57.3%. Meanwhile, relatively weaker models, such as Kimi K2.6 (none), achieve only 16.4% accuracy, rendering them largely impractical for this task. To better understand the source of this difficulty, we evaluate models under two additional settings. In the full setting, retrieval tools return all matching records in a single response instead of presenting them page by page. In the no-search setting, all task-relevant records are provided directly in the system prompt, eliminating retrieval altogether. As shown in Figure 5, removing the retrieval process increases the average accuracy from 35.5% to 66.8%, whereas returning all matching records at once improves accuracy only marginally, from 35.5% to 36.0%. These results suggest that the primary bottleneck stems from the capability to formulate effective queries that correctly identify the target record. We further analyze the failure modes of the three strongest models. As shown in Figure 6, 79% of all failures arise from incorrect parameter values, whereas incorrect tool selection is relatively uncommon, and only 6% of failures result from mandatory parameters being left unspecified. These findings suggest that LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. Such overconfident behavior poses a serious reliability risk for mobile assistants, where acting on incorrect personal information may be more harmful than declining to act. LLMs exhibit consistent performance differences across cognitive capabilities. LLMs achieve an average accuracy of around 46% on reasoning, disambiguation, and integration, whereas their average performance on preference inference and multi-intent decomposition is only about half as high. This disparity reflects a fundamental difference between these capabilities. The former primarily requires identifying and combining information explicitly available in personal records, whereas the latter requires inferring information that is not explicitly stated. Figure 5 further distinguishes the sources of difficulty for preference inference and multi-intent decomposition. Under the no-search setting, performance on preference inference reaches 67.0%, representing a 44.1-point improvement over the standard setting. This result suggests that LLMs are capable of inferring user preferences once the relevant evidence is available, but struggle to proactively locate that evidence, a capability that is essential for mobile assistants. By contrast, even with all relevant records directly available, performance on multi-intent decomposition reaches only 45.7%. This finding indicates that its primary bottleneck lies in decomposing complex requests into coordinated subtasks and planning their execution, another fundamental capability for mobile assistants. Increasing reasoning effort improves performance, although the gains vary substantially across models. Across all evaluated models, increasing the reasoning effort consistently improves performance, yielding an average gain of 13.8 points. However, these gains vary considerably across models, ranging from 28.8 points for GPT-5.5 to only 6.0 points for GLM-5.2. This variation suggests that the benefits of additional reasoning depend on a model’s ability to translate extra deliberation into more effective actions. In SPIEVAL, this capability is reflected in how models 8 Preprint. 6 8101214 16 GPT-5.5 Gemini 3.1 Pro Claude Opus 4.8 DeepSeek-V4-Pro Kimi K2.6 GLM-5.2 Hy3 Qwen3.7-Plus Seed-2.1-Pro 9.812.3 6.3 8.7 6.7 8.4 10.413.0 11.4 15.2 12.813.8 9.310.8 9.211.2 8.1 9.5 Failed Solved Figure 7: Average number of retrievals on solved versus failed tasks, reported for each model under its highest reasoning effort. Retrieve modeRetrieve scope Substring 98.5% Regex 0.9% Fuzzy 0.6% Full-text 90.5% Field-specific 9.5% Figure 8: Distributions of retrieval configura- tions across all 126,279 retrieval calls issued by 18 model configurations. use the additional reasoning budget to plan retrieval strategies, verify intermediate results, and reformulate queries when the initial search fails to identify the desired records. Models with smaller improvements either already retrieve information efficiently with minimal reasoning, as exemplified by Gemini 3.1 Pro, or fail to convert additional reasoning into better decisions. These findings suggest that the ability to translate additional reasoning into more effective information localization and retrieval is a key advantage for LLMs intended to function as mobile assistants. 6FURTHER ANALYSIS To better understand the factors underlying model performance, we conduct a more in-depth analysis and draw the following observations. LLMs often fail because they commit to decisions too early. By comparing retrieval behavior on successful and failed tasks, we observe a counterintuitive pattern. As shown in Figure 7, every evaluated model performs fewer retrievals on failed tasks than on successful ones, with the difference ranging from 1.0 retrieval for GLM-5.2 to 3.8 retrievals for Kimi K2.6. Combined with the findings in Figure 6, this pattern suggests that models do not fail because they give up on difficult tasks. Instead, they often stop searching as soon as they encounter a seemingly plausible record and proceed with the requested action without performing additional retrievals to verify the information or distinguish it from competing candidates. These findings suggest that an important direction for future models is to develop their capability to determine whether the available evidence is sufficient and decide when additional retrieval is necessary. LLMs make little use of the advanced retrieval methods provided by the tools. A comprehensive analysis of all 126,279 retrieval calls issued across 18 model configurations, shown in Figure 8, reveals that plain substring queries account for 98.5% of all retrievals. By comparison, regular expressions and fuzzy matching together account for less than 2%, while only 9.5% of retrievals restrict the search to specific fields. One possible explanation is that the models inherit keyword- based retrieval strategies from general-purpose text retrieval rather than learning to exploit the structured retrieval capabilities exposed by the tools. This overwhelming reliance on basic substring matching has practical consequences. Records containing noisy values, lexical variations, or only partial matches often cannot be retrieved through exact substring matching alone. Since advanced methods are already exposed through the retrieval interface, enabling models to use them more effectively represents a promising direction for improving future performance. Different LLMs exhibit distinct strengths in reasoning and retrieval behavior. Gemini 3.1 Pro achieves accuracy above 73% on reasoning, disambiguation, and integration, but its performance drops to 20% on preference inference. GPT-5.5 exhibits the opposite trend, achieving the strongest performance on preference inference and multi-intent decomposition. Meanwhile, Gemini 3.1 Pro and Claude Opus 4.8 perform an average of 7.6 retrievals per task, whereas GPT-5.5 performs 11.2 retrievals per task on average. Despite using only about two-thirds as many retrievals, Gemini 3.1 Pro and Claude Opus 4.8 achieve performance comparable to GPT-5.5, suggesting that they rely on more targeted retrieval strategies. By contrast, GPT-5.5 achieves the highest overall accuracy 9 Preprint. through more comprehensive retrieval. These findings suggest that successful mobile assistants can adopt different strategies, balancing retrieval efficiency against comprehensive information gathering and downstream reasoning. 7CONCLUSION In this paper, we introduce SPIEVAL, a benchmark for evaluating LLMs as mobile assistants that complete tasks by leveraging scattered personal information. The benchmark covers five cognitive capabilities and comprises 4,335 personal records collected from 10 commonly used apps. We evaluate nine representative LLMs and find that current models struggle in this setting, with information localization emerging as the primary bottleneck. Further analyses reveal systematic limitations in retrieval behavior, reasoning strategies, and capability utilization, providing a deeper understanding of current LLM-based mobile assistants. REFERENCES Anthropic. System card: Claude opus 4.8, 2026. URL https://w-cdn.anthropic.com /0b4915911b0d19eca5b5e635c80fef830a37ea.pdf. Bytedance Seed. Seed2.1 model card: Agentic intelligence for productivity, 2026. URL https: //lf3-static.bytednsdoc.com/obj/eden-cn/lapzild-tss/ljhwZthlauk jlkulzlp/seed2.1/Seed2_1_Model_Card.pdf. Jingxuan Chen, Derek Yuen, Bin Xie, Yuhao Yang, Gongwei Chen, Zhihao Wu, Li Yixing, Xurui Zhou, Weiwen Liu, Shuai Wang, Kaiwen Zhou, Rui Shao, Liqiang Nie, Yasheng Wang, Jianye Hao, Jun Wang, and Kun Shao. Spa-bench: a comprehensive benchmark for smartphone agent evaluation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/f orum?id=OZbFRNhpwr. DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyu Hou, Chenhao Xu, Chenze Shao, Chong Ruan, Conner Sun, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Donghao Li, Dongjie Ji, Erhang Li, Fang Wei, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanting Chen, Guoai Cao, Guolai Meng, Guowei Li, Han Yu, Han Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoling Zhang, Haoming Luo, Haoran Wei, Haotian Yuan, Haowei Zhang, Haowen Luo, Haoyu Chen, Haozhe Ji, Hengqing Zhang, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J Yang, JQ Zhu, Jia Luo, Jia Song, Jia Yu, Jialiang Huang, Jialu Cai, Jian Liang, Jiangting Zhou, Jiasheng Ye, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jieyu Yang, Jin Chen, Jin Yan, Jingchang Chen, Jingli Zhou, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jingzi Zhou, Jinhua Zhu, Jiping Yu, Joseph Sun, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junmin Zheng, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Leyi Xia, Li Zhang, Liang Zhao, Lihua Guo, Lingxiao Luo, Linwang Ma, Linyan Zhu, Litong Wang, Liyu Cai, Liyue Zhang, Longhao Chen, MS Di, MY Xu, Max Mei, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Mingxu Zhou, Minmin Han, Ning Wang, Panpan Huang, Panpan Wang, Peixin Cong, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Qiwei Jiang, Rui Tian, Ruifan Xu, Ruijie Lu, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqian Chen, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, Ruyi Chen, SH Liu, Shanghao Lu, Shangmian Sun, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoheng Nie, Shaoqing Wu, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Shuying Yu, Songyang Zhou, Tao Ni, Tao Yun, Tian Jin, Tian Pei, Tian Ye, Tianle Lin, Tianran Ji, Tianyi Cui, Tianyuan Yue, Tingting Yu, Tun Wang, W Zhang, WL Xiao, Wangding Zeng, Wei An, Weilin Zhao, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjing Yao, Wenjun Gao, Wenkai Yang, Wenlve Huang, Wenqing Hou, Wentao Zhang, Wenting Ma, Xi Gao, Xiang He, Xiangwen Wang, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xiaowen Sun, Xiaoxiang Wang, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingchen Liu, Xingkai 10 Preprint. Yu, Xingyou Li, Xinyu Yang, Xinyu Zhang, Xu Chen, Xuanyu Wang, Xuecheng Su, Xueyin Chen, Xuheng Lin, Xuwei Fu, YC Yan, YQ Wang, YW Ma, Yanfeng Luo, Yang Zhang, Yanhong Xu, Yanru Ma, Yanwen Huang, Yao Li, Yao Li, Yao Xu, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Shao, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yijia Wu, Yiliang Xiong, Yiling Ma, Ying He, Ying Tang, Ying Zhou, Yingjia Luo, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiang Zhang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, YuKun Li, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuanhao Li, Yuduan Wang, Yuehan Yang, Yuer Xu, Yuhan Wu, Yuhao Meng, Yuheng Zou, Yukun Zha, Yunfan Xiong, Yupeng Chen, Yuping Lin, Yuqian Cao, Yuqian Wang, Yushun Zhang, Yuting Yan, Yutong Lin, Yuxian Gu, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuxuan Zhou, Yuyang Zhou, Yuzhen Huang, ZF Wu, Zehao Wang, Zehua Zhao, Zehui Ren, Zekai Zhang, Zhangli Sha, Zhe Fu, Zhe Ju, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zheren Gao, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhixuan Chen, Zhiyu Wu, Zhizhou Ren, Zhongyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihua Qu, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Ziyi Wan, Zizheng Pan, and Zongqing Yao. Deepseek-v4: Towards highly efficient million-token context intelligence. CoRR, abs/2606.19348, 2026. doi: 10.48550/ARXIV.2606.19348. URL https://doi.org/10.48550/arXiv.2606.19348. Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Laurenc ̧on, Maxime Lecanu, Kunal Malkan, Dheeraj Mekala, Pierre M ́ enard, Gerard Moreno-Torres Bertran, Ulyana Piterbarg, Mikhail Plekhanov, Mathieu Rita, Andrey Rusakov, Vladislav Vorotilov, Mengjue Wang, Ian Yu, Amine Benhalloum, Gr ́ egoire Mialon, and Thomas Scialom. Gaia2: Benchmarking LLM agents on dynamic and asynchronous environments. CoRR, abs/2602.11964, 2026. doi: 10.48550/ARXIV .2602.11964. URL https://doi.org/10.48550/arXiv.2602.11964. Google. Gemini 3.1 pro model card, 2026. URL https://storage.googleapis.com/d eepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf. Yanchu Guan, Dong Wang, Zhixuan Chu, Shiyu Wang, Feiyue Ni, Ruihua Song, Longfei Li, Jinjie Gu, and Chenyi Zhuang. Intelligent virtual assistants with llm-based process automation. CoRR, abs/2312.06677, 2023. doi: 10.48550/ARXIV.2312.06677. URL https://doi.org/10.4 8550/arXiv.2312.06677. Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, and Fei Wu. OS agents: A survey on mllm-based agents for computer, phone and browser use. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 7436–7465. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.ACL- LONG.369. URL https://doi.org/10.18653/v1/2025.acl-long.369. Kimi. Kimi k2.6: Advancing open-source coding, 2026. URL https://w.kimi.com/blo g/kimi-k2-6. Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao, and Dongliang Xu. Mind the third eye! benchmarking privacy awareness in mllm-powered smartphone agents. In Sven Koenig, Chad Jenkins, and Matthew E. Taylor (eds.), Fortieth AAAI Conference on Artificial Intelligence, Thirty- Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, p. 35626–35634. AAAI Press, 2026. doi: 10.1609/AAAI.V40I42.40874. URL https: //doi.org/10.1609/aaai.v40i42.40874. Guangyi Liu, Pengxiang Zhao, Yaozhen Liang, Liang Liu, Yaxuan Guo, Han Xiao, Weifeng Lin, Yuxiang Chai, Yue Han, Shuai Ren, Hao Wang, Xiaoyu Liang, WenHao Wang, Tianze Wu, Zhengxi Lu, Siheng Chen, LiLinghao, Hao Wang, Guanjing Xiong, Yong Liu, and Hongsheng Li. Llm-powered GUI agents in phone automation: Surveying progress and prospects. Trans. Mach. Learn. Res., 2025, 2025a. URL https://openreview.net/forum?id=yWQqoi1G1K. 11 Preprint. Shuodi Liu, Yingzhuo Liu, Zi Wang, Yusheng Wang, Huijia Wu, Liuyu Xiang, and Zhaofeng He. Select-then-decompose: From empirical analysis to adaptive selection strategy for task decomposition in large language models. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, p. 5454– 5477. Association for Computational Linguistics, 2025b. doi: 10.18653/V1/2025.EMNLP-MAI N.278. URL https://doi.org/10.18653/v1/2025.emnlp-main.278. Xiao Liu, Bo Qin, Dongzhu Liang, Guang Dong, Hanyu Lai, Hanchen Zhang, Hanlin Zhao, Iat Long Iong, Jiadai Sun, Jiaqi Wang, Junjie Gao, Junjun Shan, Kangning Liu, Shudan Zhang, Shuntian Yao, Siyi Cheng, Wentao Yao, Wenyi Zhao, Xinghan Liu, Xinyi Liu, Xinying Chen, Xinyue Yang, Yang Yang, Yifan Xu, Yu Yang, Yujia Wang, Yulin Xu, Zehan Qi, Yuxiao Dong, and Jie Tang. Autoglm: Autonomous foundation agents for guis. CoRR, abs/2411.00820, 2024. doi: 10.48550 /ARXIV.2411.00820. URL https://doi.org/10.48550/arXiv.2411.00820. Renze Lou, Kai Zhang, and Wenpeng Yin. Large language model instruction following: A survey of progresses and challenges. Comput. Linguistics, 50(3):1053–1095, 2024. doi: 10.1162/COLI \ A\00523. URL https://doi.org/10.1162/coli_a_00523. Jisoo Mok, Ik-hwan Kim, Sangkwon Park, and Sungroh Yoon. Exploring the potential of llms as personalized assistants: Dataset, evaluation, and analysis. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 10212–10239. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.ACL- LONG.504. URL https://doi.org/ 10.18653/v1/2025.acl-long.504. OpenAI. Gpt-5.5 system card, 2026. URL https://deploymentsafety.openai.com/ gpt-5-5/gpt-5-5.pdf. Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Xuanhe Zhou, Yufei Huang, Chaojun Xiao, Chi Han, Yi R. Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bowen Li, Ziwei Tang, Jing Yi, Yuzhang Zhu, Zhenning Dai, Lan Yan, Xin Cong, Yaxi Lu, Weilin Zhao, Yuxiang Huang, Junxi Yan, Xu Han, Xian Sun, Dahai Li, Jason Phang, Cheng Yang, Tongshuang Wu, Heng Ji, Guoliang Li, Zhiyuan Liu, and Maosong Sun. Tool learning with foundation models. ACM Comput. Surv., 57(4):101:1–101:40, 2025. doi: 10.1145/3704435. URL https://doi.org/10.1145/3704435. Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong Wen. Tool learning with large language models: a survey. Frontiers Comput. Sci., 19(8): 198343, 2025. doi: 10.1007/S11704-024-40678-2. URL https://doi.org/10.1007/s1 1704-024-40678-2. Qwen Team. Qwen3.7-Plus: Multimodal agent intelligence, May 2026. URL https://qwen.a i/blog?id=qwen3.7-plus. Susanna R ̈ ucker and Alan Akbik.Evaluating design decisions for dual encoder-based entity disambiguation. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 15685–15701. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.ACL-L ONG.764. URL https://doi.org/10.18653/v1/2025.acl-long.764. Jiazheng Sun, Te Yang, Jiayang Niu, Mingxuan Li, Yongyong Lu, Ruimeng Yang, and Xin Peng. Fairy: Interactive mobile assistant to real-world tasks via lmm-based multi-agent. CoRR, abs/2509.20729, 2025. doi: 10.48550/ARXIV.2509.20729. URL https://doi.org/10.4 8550/arXiv.2509.20729. Juntao Tan, Liangwei Yang, Zuxin Liu, Zhiwei Liu, Rithesh R. N., Tulika Manoj Awalgaonkar, Jianguo Zhang, Weiran Yao, Ming Zhu, Shirley Kokane, Silvio Savarese, Huan Wang, Caiming Xiong, and Shelby Heinecke. Personabench: Evaluating AI models on understanding personal 12 Preprint. information through accessing (synthetic) private user data. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, volume ACL 2025 of Findings of ACL, p. 878–893. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.FINDINGS-ACL.49. URL https://doi.org/10.18653/v1/2025 .findings-acl.49. Tencent Hunyuan Team. Introducing hy3, 2026. URL https://hy.tencent.com/researc h/hy3. Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian. Appworld: A controllable world of apps and people for benchmarking interactive coding agents. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11- 16, 2024, p. 16022–16076. Association for Computational Linguistics, 2024. doi: 10.18653/V 1/2024.ACL-LONG.850. URL https://doi.org/10.18653/v1/2024.acl-long. 850. Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (eds.), Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. URL http://pape rs.nips.c/paper_files/paper/2024/hash/0520537ba799d375b8f55232 95c337a-Abstract-Conference.html. Zhenhailong Wang, Haiyang Xu, Junyang Wang, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Heng Ji.Mobile-agent-e: Self-evolving mobile assistant for complex tasks.CoRR, abs/2501.11733, 2025. doi: 10.48550/ARXIV.2501.11733. URL https://doi.org/ 10.48550/arXiv.2501.11733. Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu.Autodroid: Llm-powered task automation in android. In Weisong Shi, Deepak Ganesan, and Nicholas D. Lane (eds.), Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2024, Washington D.C., DC, USA, November 18-22, 2024, p. 543–557. ACM, 2024. doi: 10.1145/3636534.3649379. URL https://doi.org/10.1145/3636534.3649379. Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (eds.), Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. URL http://papers.nips.c/paper_files/paper/2024/hash /5d413e48f84dc61244b6be550f1cd8f5-Abstract-Datasets_and_Benchmar ks_Track.html. Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, and Yuxiao Dong. Androidlab: Training and systematic benchmarking of android autonomous agents. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 2144–2166. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.ACL-L ONG.107. URL https://doi.org/10.18653/v1/2025.acl-long.107. Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning.Hotpotqa:A dataset for diverse, explainable multi-hop 13 Preprint. question answering. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, p. 2369–2380. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1259. URL https://doi.or g/10.18653/v1/d18-1259. Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, Jie Zhou, Siming Chen, Tao Gui, Qi Zhang, and Xuanjing Huang. A comprehensive capability analysis of GPT-3 and GPT-3.5 series models. CoRR, abs/2303.10420, 2023. doi: 10.48550/ARXIV.2303.10420. URL https://doi.org/10.48550/arXiv.2 303.10420. Junjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, and Jiecao Chen. Toolhop: A query-driven benchmark for evaluating large language models in multi-hop tool use. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 2995–3021. Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025.acl-lon g.150/. Chi Zhang, Zhao Yang, Jiaxuan Liu, Yanda Li, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. Appagent: Multimodal agents as smartphone users. In Naomi Yamashita, Vanessa Evers, Koji Yatani, Sharon Xianghua Ding, Bongshin Lee, Marshini Chetty, and Phoebe O. Toups Dugas (eds.), Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI 2025, YokohamaJapan, 26 April 2025- 1 May 2025, p. 70:1–70:20. ACM, 2025a. doi: 10.1145/3706598.3713600. URL https://doi.org/10.1145/3706598.3713600. Zhehao Zhang, Ryan A. Rossi, Branislav Kveton, Yijia Shao, Diyi Yang, Hamed Zamani, Franck Dernoncourt, Joe Barrow, Tong Yu, Sungchul Kim, Ruiyi Zhang, Jiuxiang Gu, Tyler Derr, Hongjie Chen, Junda Wu, Xiang Chen, Zichao Wang, Subrata Mitra, Nedim Lipka, Nesreen K. Ahmed, and Yu Wang. Personalization of large language models: A survey. Trans. Mach. Learn. Res., 2025, 2025b. URL https://openreview.net/forum?id=tf6A9EYMo6. Zhipu AI. Glm-5.2: Built for long-horizon tasks, 2026. URL https://z.ai/blog/glm-5.2. Andrew Zhu, Alyssa Hwang, Liam Dugan, and Chris Callison-Burch. Fanoutqa: A multi-hop, multi-document question answering benchmark for large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 18–37. Association for Computational Linguistics, 2024. doi: 10.18653/V1/20 24.ACL-SHORT.2. URL https://doi.org/10.18653/v1/2024.acl-short.2. 14 Preprint. AUSER PROFILE AND SYSTEM PROMPT As described in Section 3.3, we construct a unified user profile that establishes a consistent identity across all tasks. The profile is provided to the model as part of the system prompt. Table 3 presents the user profile, and Figure 9 shows the complete system prompt template. Table 3: Unified user profile provided as part of the system prompt for all tasks. AttributeChineseEnglish Name 徐艺轩Yixuan Xu Occupation 上海光明坚果公司采购部普通员工 Employee, Procurement Dept., Shanghai Guangming Nut Co. Family Members 妻子谢玲、儿子徐鑫、女儿徐淼 Wife: Ling Xie; Son: Xin Xu; Daughter: Miao Xu Phone Number1571122783715711227837 Payment Account69778899678978900986977889967897890098 Current Timestamp2025-11-01 12:00:00周六2025-11-01 12:00:00 Saturday System Prompt Template Chinese: 你是一个智能手机助手。用户会给你一些指令,你需要通过查询手机中各个应用的记 录来收集必要信息,然后调用相应的工具完成用户的请求。,→ user_profile 你可以使用以下两类工具: 1. ** 搜索工具 ** :用于在手机各应用中检索记录,包括len(app_tools)个应 用内搜索工具(app_list)和1个全局搜索工具(global_tool)。,→ 2. ** 执行工具 ** :用于执行具体操作,包括exec_list。 工作流程: 1. 分析用户意图,判断需要哪些信息 2. 调用搜索工具查找相关记录,可以多次搜索、组合使用不同的搜索工具 3. 根据搜索结果,调用执行工具完成用户请求 4. 如果用户的请求包含多个意图,请并行调用多个执行工具同时完成 注意事项: - 完成用户请求所需的所有信息都可以从手机记录中检索得到,请充分搜索,不要向 用户反问,→ - 搜索结果会按页返回,如果提示还有更多结果,可通过page参数翻页查看 - 如果一次搜索未找到所需信息,尝试换个关键词或换个应用搜索 - 多个搜索之间如果相互独立,可以并行调用 - 最终的执行工具调用应包含尽可能完整的参数信息 - 当所有操作完成后,请回复\[DONE]需求已完成,还有什么可以帮您?" English: You are a mobile assistant. The user will give you instructions. You need to retrieve the necessary information from records across different phone applications before invoking the appropriate tools to fulfill the user's request. ,→ ,→ ,→ ,→ user_profile 15 Preprint. You have access to two categories of tools: 1. ** Retrieval tools ** : used to retrieve records from phone applications, including len(app_tools) app-specific retrieval tools (app_list) and one global retrieval tool (global_tool). ,→ ,→ ,→ 2. ** Execution tools ** : used to perform specific operations, including exec_list.,→ Workflow: 1. Analyze the user's intent and determine what information is needed,→ 2. Call search tools to find relevant records---you may issue multiple retrieval queries and combine different retrieval tools 3. Based on the search results, call execution tools to fulfill the user's request ,→ ,→ ,→ 4. If the user's request contains multiple intents, invoke multiple execution tools in parallel,→ Notes: - All information needed to fulfill the user's request can be retrieved from phone records. Search thoroughly instead of asking the user for clarification. ,→ ,→ - Search results are returned in pages; if more results are available, use the page parameter to view additional pages. ,→ ,→ - If a retrieval does not return the required information, try different keywords or a different application.,→ - Independent retrievals may be performed in parallel. - The final execution tool calls should include all available parameter values whenever possible.,→ - Once all tasks are complete, please reply with "[DONE] Request completed. Is there anything else I can help you with?" ,→ ,→ Figure 9: System prompt template used for experiments. 16 Preprint. BAPPLICATION SCHEMAS As described in Section 3.3, we construct 10 simulated apps. Each app is abstracted from its real- world counterpart and represented by a structured schema containing domain-specific fields, with an average of 8.1 fields per app. Each schema further distinguishes between required and optional fields. Table 4 provides the complete schema of the 10 simulated apps in SPIEVAL. Table 4: Schemas of the 10 simulated apps in SPIEVAL, including required and optional fields. AppRequired FieldsOptional Fields Accommodation Check-in Date, Check-out Date, Contact, Contact Number, Cost, Hotel, Quantity, Room Type, Status – Alarm Alarm Name, Repeat, Repeat Count, Ring Duration, Ring Interval, Status, Time – ContactsName, Phone Number Address, Birthday, Company, Email, ID Number, Note, Relationship Meeting End Date, End Time, Meeting ID, Participants, Reminder, Repeat, Start Date, Start Time, Title Location, Meeting Minutes, Status NotesContentCreation Date, Creation Time, Title Schedule Date, End Time, Reminder, Repeat, Start Time, Title Location SMS Content, Date, Recipient Phone, Status, Time Attachment,RecipientName, Sender Name, Sender Phone Transactions Amount, Date, Item, Payee Account, Payer, Payer Account, Time Note, Payee Transport Contact, Contact Number, Date, De- parture Time, Destination, Estimated Arrival Time, Estimated Cost, Mode, Origin, Status – Voicemail Caller Number, Date, Message Con- tent, Time Caller Name 17 Preprint. CTOOL SCHEMAS As described in Section 3.3, we construct 21 tools, including 11 retrieval tools and 10 execution tools. Each app is associated with one app-specific retrieval tool and one execution tool, while an additional global retrieval tool enables cross-app information access. To provide a detailed description of these tools, we present their specification documents in Figure 10 and Figure 11, respectively. Specification Documents for all Retrieval Tools [ "type": "function", "function": "name": "search_accommodation", "description": "在住宿记录中搜索。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'住宿地点'、'入住日期'、'退宿日 期'、'房型'、'数量'、'联系人'、'联系方 式'、'费用'、'状态'。不指定则搜索所有字 段。" ,→ ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] 18 Preprint. , "type": "function", "function": "name": "search_alarm", "description": "在闹钟中搜索记录。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'时间'、'重复'、'闹钟名'、'响铃时长 (分钟)'、'重复响铃次数'、'响铃间隔时间 (分钟)'、'状态'。不指定则搜索所有字 段。" ,→ ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": "name": "search_contacts", "description": "在通讯录中搜索联系人记录。", "parameters": "type": "object", 19 Preprint. "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'姓名'、'电话号码'、'公司'、'电子邮 件'、'住址'、'生日'、'身份证号'、'与本人 关系'、'备注'。不指定则搜索所有字段。" ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": "name": "search_meeting", "description": "在会议中搜索记录。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", 20 Preprint. "description": "指定搜索的字段名。可选字 段:'标题'、'开始日期'、'结束日期'、'开 始时间'、'结束时间'、'提醒时间'、'重 复'、'参会人'、'会议号'、'地点'、'会议纪 要'、'状态'。不指定则搜索所有字段。" ,→ ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": "name": "search_notes", "description": "在便签中搜索记录。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'标题'、'内容'、'创建日期'、'创建时 间'。不指定则搜索所有字段。" ,→ ,→ , "mode": "type": "string", 21 Preprint. "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": "name": "search_schedule", "description": "在日程中搜索记录。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'标题'、'日期'、'开始时间'、'结束时 间'、'提醒时间'、'重复'、'地点'。不指定 则搜索所有字段。" ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": 22 Preprint. "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": "name": "search_sms", "description": "在短信中搜索记录。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'发件人姓名'、'发件人电话号码'、'收 件人姓名'、'收件人电话号码'、'日期'、'时 间'、'正文'、'附件'、'状态'。不指定则搜 索所有字段。" ,→ ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ 23 Preprint. , "required": [ "pattern" ] , "type": "function", "function": "name": "search_transaction", "description": "在交易记录中搜索。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'项目'、'金额'、'日期'、'时间'、'付 款人'、'付款账号'、'收款人'、'收款账 号'、'备注'。不指定则搜索所有字段。" ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": 24 Preprint. "name": "search_transport", "description": "在交通记录中搜索。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'日期'、'出发时间'、'(预计)到达时 间'、'出发地'、'目的地'、'交通方式'、'联 系人'、'联系方式'、'(预计)费用'、'状 态'。不指定则搜索所有字段。" ,→ ,→ ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": "name": "search_voicemail", "description": "在语音留言中搜索记录。", "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , 25 Preprint. "field": "type": "string", "description": "指定搜索的字段名。可选字 段:'来电姓名'、'来电号码'、'日期'、'时 间'、'留言内容'。不指定则搜索所有字段。" ,→ ,→ , "mode": "type": "string", "description": "匹配模式。fixed:子串包含 匹配(默认);regex:正则表达式匹 配;fuzzy:模糊相似度匹配。", ,→ ,→ "enum": [ "fixed", "regex", "fuzzy" ] , "ignore_case": "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] , "type": "function", "function": "name": "search_phone", "description": "全局搜索手机中所有应用的记录。使用子串包 含匹配,搜索所有字段。可指定field缩小搜索范围。每页返 回5条结果,可通过page参数翻页。", ,→ ,→ "parameters": "type": "object", "properties": "pattern": "type": "string", "description": "搜索内容。" , "field": "type": "string", "description": "指定搜索的字段名(可选)。 不指定则搜索所有字段。不同应用的可用字段 不同,常见字段包括:'姓名'、'电话号 码'、'与本人关系'、'备注'、'正文'、'留言 内容'、'标题'、'内容'等。" ,→ ,→ ,→ ,→ , "ignore_case": 26 Preprint. "type": "boolean", "description": "是否忽略大小写,默认 为true。",→ , "page": "type": "integer", "description": "返回第几页结果,每页最 多5条。默认为1。",→ , "required": [ "pattern" ] ] Figure 10: Specification documents for all retrieval tools. Specification Documents for all Execution Tools [ "type": "function", "function": "name": "Accommodation_Create", "description": "创建一个住宿信息。", "parameters": "type": "object", "properties": "dateCheckIn": "type": "string", "description": "入住日期,日期格式 为Y-M-D(例如:'2025-05-20'表 示2025年5月20日)。" ,→ ,→ , "dateCheckOut": "type": "string", "description": "退宿日期,日期格式 为Y-M-D(例如:'2025-05-20'表 示2025年5月20日)。" ,→ ,→ , "location": "type": "string", "description": "住宿地点。" , "roomType": "type": "string", "description": "房间类型。" , "num": "type": "integer", "description": "房间数量,默认为1。" , "people": 27 Preprint. "type": "array", "description": "入住人。", "items": "type": "string", "description": "入住人姓名。" , "telephone": "type": "string", "description": "联系电话。" , "required": [ "dateCheckIn", "dateCheckOut", "location", "people", "telephone" ] , "type": "function", "function": "name": "Clock_CreateAlarm", "description": "创建闹钟,支持设置闹钟的触发时间、重复规 则等属性",,→ "parameters": "type": "object", "properties": "time": "type": "string", "description": "闹钟触发的具体时间(24小时 内的时间,不包含日期),格式 为H:M:S(例如:'14:30:00'表示下 午2点30分)。" ,→ ,→ ,→ , "repeatDayOfWeek": "type": "array", "description": "重复规则,类型为数组,用于 指定闹钟在哪天重复触发。数字1到7分别表示 每个星期一到星期日重复,数字0表示不重复。 默认不重复。", ,→ ,→ ,→ "items": "type": "integer", "description": "重复规则。数字1到7分别 表示星期一到星期日,数字0表示不重 复。", ,→ ,→ "enum": [ 1, 2, 3, 4, 5, 6, 7, 28 Preprint. 0 ] , "name": "type": "string", "description": "闹钟名称,默认为\闹钟"。" , "alarmDuration": "type": "integer", "description": "响铃时长(分钟),可以是1分 钟,3分钟,5分钟,10分钟。默认为3分 钟。", ,→ ,→ "enum": [ 1, 3, 5, 10 ] , "alarmCount": "type": "integer", "description": "重复响铃次数,可以 是1次,3次,5次,10次。默认为1次。",,→ "enum": [ 1, 3, 5, 10 ] , "alarmInterval": "type": "integer", "description": "响铃间隔(分钟),可以 是5-30分钟。默认为5分钟。",,→ "enum": [ 5, 10, 15, 20, 25, 30 ] , "required": [ "time" ] , "type": "function", "function": "name": "Contacts_Create", "description": "新建一个联系人,并保存到手机通讯录中。", "parameters": 29 Preprint. "type": "object", "properties": "name": "type": "string", "description": "联系人的姓名。" , "telephone": "type": "string", "description": "联系人的电话号码。" , "company": "type": "string", "description": "联系人的公司名称。" , "email": "type": "string", "description": "联系人的电子邮件地址。" , "address": "type": "string", "description": "联系人的地址。" , "birthday": "type": "string", "description": "联系人的生日。" , "identifyCard": "type": "string", "description": "联系人的身份证号码。" , "relationship": "type": "string", "description": "联系人与本人的关系。" , "note": "type": "string", "description": "为联系人添加的备注信息。" , "required": [ "name", "telephone" ] , "type": "function", "function": "name": "Meeting_Create", "description": "预订会议,支持设置时间、地点、参与人、会 议主题等信息",,→ "parameters": "type": "object", "properties": "dateStart": "type": "string", 30 Preprint. "description": "会议的开始日期,日期格式 为Y-M-D(例如:'2025-05-20'表 示2025年5月20日)" ,→ ,→ , "dateEnd": "type": "string", "description": "会议的结束日期,日期格式 为Y-M-D(例如:'2025-05-20'表 示2025年5月20日)" ,→ ,→ , "timeStart": "type": "string", "description": "会议的开始时间,时间格式 为H:M:S(例如:'14:30:00'表示下 午2点30分)" ,→ ,→ , "timeEnd": "type": "string", "description": "会议的结束时间,时间格式 为H:M:S(例如:'14:30:00'表示下 午2点30分)" ,→ ,→ , "timeRemind": "type": "string", "description": "会议提醒的时间。默认为10分 钟前。",,→ "enum": [ "准时", "10分钟前", "30分钟前", "1小时前", "1天前", "1周前" ] , "title": "type": "string", "description": "会议的主题。默认为\无主 题"。",→ , "repeat": "type": "string", "description": "会议重复规则。默认为不重 复。",,→ "enum": [ "不重复", "每天", "每周", "每两周" ] , "location": "type": "array", "description": "会议的地点。", "items": "type": "string", 31 Preprint. "description": "地点名称。默认 为\无"。",→ , "participants": "type": "array", "description": "会议的参与者,默认已包含自 己,只需要填写其他仍需加入的用户即可。",,→ "items": "type": "string", "description": "会议除自己外的参与者姓 名,可以是昵称。",→ , "required": [ "dateStart", "dateEnd", "timeStart", "timeEnd" ] , "type": "function", "function": "name": "TodoList_Create", "description": "新建一个待办事项。", "parameters": "type": "object", "properties": "content": "type": "string", "description": "待办事项的内容。" , "date": "type": "string", "description": "待办事项的具体日期,日期格 式为Y-M-D(例如:'2025-05-20'表 示2025年5月20日)" ,→ ,→ , "time": "type": "string", "description": "待办事项的开始时间,时间格 式为H:M(例如:'14:30'表示下 午2点30分)" ,→ ,→ , "required": [ "content", "date", "time" ] , 32 Preprint. "type": "function", "function": "name": "Schedule_Create", "description": "新建一个日程。", "parameters": "type": "object", "properties": "date": "type": "string", "description": "日程触发的具体日期,日期格 式为Y-M-D(例如:'2025-05-20'表 示2025年5月20日)" ,→ ,→ , "title": "type": "string", "description": "日程的标题,用于标识日程的 主题。默认为\无主题"。",→ , "timeStart": "type": "string", "description": "日程开始的具体时间,时间格 式为H:M:S(例如:'14:30:00'表示下 午2点30分)" ,→ ,→ , "timeEnd": "type": "string", "description": "日程结束的具体时间,时间格 式为H:M:S(例如:'14:30:00'表示下 午2点30分)" ,→ ,→ , "timeRemind": "type": "string", "description": "日程提醒的时间。默认为10分 钟前。",,→ "enum": [ "准时", "10分钟前", "30分钟前", "1小时前", "1天前", "1周前" ] , "repeat": "type": "string", "description": "日程重复规则。默认为不重 复。",,→ "enum": [ "不重复", "每天", "每周", "每月", "每年" ] , 33 Preprint. "location": "type": "string", "description": "日程的地点。默认为\无"。" , "required": [ "date", "timeStart", "timeEnd" ] , "type": "function", "function": "name": "Messages_Send", "description": "向某个电话号码发送短信,包括正文和附 件。",,→ "parameters": "type": "object", "properties": "telephone": "type": "string", "description": "需要接受短信的电话号码。" , "text": "type": "string", "description": "需要发送的短信正文,仅支持 文本信息。",→ , "attach": "type": "array", "description": "需要发送的短信附件,通过指 定文件路径发送,最多支持同时发送10个附 件。", ,→ ,→ "items": "type": "string", "description": "文件路径。" , "required": [ "telephone", "text" ] , "type": "function", "function": "name": "Consumption_Create", "description": "向他人付款。", "parameters": "type": "object", "properties": 34 Preprint. "item": "type": "string", "description": "付款条目,如转账、购物等。" , "amount": "type": "number", "description": "交易金额,精确到小数点后两 位。",→ , "receivingName": "type": "string", "description": "收款人。" , "receivingAccount": "type": "string", "description": "收款账户。" , "note": "type": "string", "description": "备注。" , "required": [ "item", "amount", "receivingName", "receivingAccount" ] , "type": "function", "function": "name": "Transport_Create", "description": "创建一个交通信息。", "parameters": "type": "object", "properties": "date": "type": "string", "description": "出发日期,日期格式 为Y-M-D(例如:'2025-05-20'表 示2025年5月20日)。202" ,→ ,→ , "time": "type": "string", "description": "出发时间,时间格式 为H:M:S(例如:'14:30:00'表示下 午2点30分)。如果是现在出发,可以输 入'now'。" ,→ ,→ ,→ , "start": "type": "string", "description": "出发地。打车时如果是当前位 置,可以输入'current'。",→ , 35 Preprint. "end": "type": "string", "description": "目的地。" , "passenger": "type": "array", "description": "乘车人。", "items": "type": "string", "description": "乘车人姓名。" , "telephone": "type": "string", "description": "乘车人联系电话。" , "transportType": "type": "string", "description": "交通方式。", "enum": [ "汽车", "火车", "飞机" ] , "required": [ "date", "time", "start", "end", "passenger", "telephone", "transportType" ] , "type": "function", "function": "name": "Phone_Call", "description": "给某个电话号码打电话,支持语音电话和视频 电话两种类型。",,→ "parameters": "type": "object", "properties": "telephone": "type": "string", "description": "需要拨打的电话号码。" , "type": "type": "string", "description": "电话类型,枚举值:- VOICE 语音电话 - VIDEO视频电话,默认 为VOICE", ,→ ,→ "enum": [ 36 Preprint. "VOICE", "VIDEO" ] , "required": [ "telephone" ] ] Figure 11: Specification documents for all execution tools. 37 Preprint. DCOMPARISON WITH EXISTING BENCHMARKS As described in Section 3.4, SPIEVAL exhibits five key characteristics, including diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. To further illustrate its strengths, we compare SPIEVAL with existing benchmarks, with the results summarized in Table 5. Table 5: Comparison of SPIEVAL with existing benchmarks. For dialogue- and document-based benchmarks, #Tasks reports the number of question–answer instances and #Apps is not applicable. ✓ and✗ indicate the presence and absence of a systematic benchmark-level property, respectively. Benchmark Basic InformationTask SettingBenchmark Design #Tasks#AppsUnderspecified Scattered Personal Information Proactive Retrieval Action ExecutionVerifiable Human- Curated AppWorld7509✗✓✗ Gaia21,12012✗✓✗ SAPA-Bench7,13850✗ HiCUPID60,000–✗ PersonaBench582–✗✓✗ SPIEVAL25010✓ 38