Paper deep dive
LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks
Hengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen, Qihang Cao, Xianfeng Wang, Yucheng Zhu, Xiongkuo Min, Wei Sun, Dandan Zhu, Guangtao Zhai
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/20/2026, 5:33:11 AM
Summary
The paper introduces LifeEval, a multimodal benchmark designed to evaluate Multimodal Large Language Models (MLLMs) as real-time, task-oriented assistants in egocentric daily life scenarios. It addresses the gap in existing benchmarks by focusing on interactive, adaptive, and human-centered assistance rather than passive video understanding. LifeEval comprises 4,075 high-quality question-answer pairs across 6 core capability dimensions: Static Environment Perception, Dynamic Task Reasoning, Contextual Knowledge Retrieval, Goal-oriented Planning, Safety and Feasibility Assessment, and Multi-turn Interactive Collaboration. The benchmark was constructed using a rigorous pipeline involving video collection from Ego4D, automated generation with Gemini 2.5 Pro, and human verification. Extensive evaluations of 26 state-of-the-art MLLMs reveal significant challenges in real-time, task-oriented assistance.
Entities (25)
Relation Signals (24)
LifeEval → evaluates → MLLMs
confidence 95% · LifeEval, a multimodal benchmark designed to evaluate real-time, task-oriented human–AI collaboration
LifeEval → hascapabilitydimension → Static Environment Perception
confidence 93% · We further decompose task-assistance capability into 6 key dimensions: static environment perception...
LifeEval → hascapabilitydimension → Dynamic Task Reasoning
confidence 93% · We further decompose task-assistance capability into 6 key dimensions: ... dynamic task reasoning...
LifeEval → hascapabilitydimension → Contextual Knowledge Retrieval
confidence 93% · We further decompose task-assistance capability into 6 key dimensions: ... contextual knowledge retrieval...
LifeEval → hascapabilitydimension → Goal-oriented Planning
confidence 93% · We further decompose task-assistance capability into 6 key dimensions: ... goal-oriented planning...
LifeEval → hascapabilitydimension → Safety and Feasibility Assessment
confidence 93% · We further decompose task-assistance capability into 6 key dimensions: ... safety and feasibility assessment...
LifeEval → hascapabilitydimension → Multi-turn Interactive Collaboration
confidence 93% · We further decompose task-assistance capability into 6 key dimensions: ... multi-turn interactive collaboration.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely underexplored. Existing video benchmarks predominantly assess passive understanding through retrospective analysis or isolated perception tasks, failing to capture the interactive and adaptive nature of real-time user assistance. To bridge this gap, we introduce LifeEval, a multimodal benchmark designed to evaluate real-time, task-oriented human-AI collaboration in daily life from an egocentric perspective. LifeEval emphasizes three key aspects: task-oriented holistic evaluation, egocentric real-time perception from continuous first-person streams, and human-assistant collaborative interaction through natural dialogues. Constructed via a rigorous annotation pipeline, the benchmark comprises 4,075 high-quality question-answer pairs across 6 core capability dimensions. Extensive evaluations of 26 state-of-the-art MLLMs on LifeEval reveal substantial challenges in achieving timely, effective and adaptive interaction, highlighting essential directions for advancing human-centered interactive intelligence.
Tags
Links
- Source: https://arxiv.org/abs/2603.00490v1
- Canonical: https://arxiv.org/abs/2603.00490v1
Trouble viewing inline? Open PDF directly →
Full Text
50,222 characters extracted from source content.
Expand or collapse full text
LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks Hengjian Gao 1,2 , Kaiwei Zhang 1 * , Shibo Wang 2 , Mingjie Chen 4 , Qihang Cao 4 , Xianfeng Wang 2 , Yucheng Zhu 2 , Xiongkuo Min 2 , Wei Sun 3 , Dandan Zhu 3 , Guangtao Zhai 1,2 1 Shanghai Artificial Intelligence Laboratory, 2 Shanghai Jiao Tong University 3 East China Normal University, 4 Shanghai University of Electric Power hengjiang, zhangkaiwei@sjtu.edu.cn Abstract The rapid progress of Multimodal Large Language Mod- els (MLLMs) marks a significant step toward artificial gen- eral intelligence, offering great potential for augmenting human capabilities. However, their ability to provide ef- fective assistance in dynamic, real-world environments re- mains largely underexplored. Existing video benchmarks predominantly assess passive understanding through ret- rospective analysis or isolated perception tasks, failing to capture the interactive and adaptive nature of real-time user assistance. To bridge this gap, we introduce LifeE- val, a multimodal benchmark designed to evaluate real- time, task-oriented human–AI collaboration in daily life from an egocentric perspective. LifeEval emphasizes three key aspects: task-oriented holistic evaluation, egocentric real-time perception from continuous first-person streams, and human–assistant collaborative interaction through nat- ural dialogues.Constructed via a rigorous annotation pipeline, the benchmark comprises 4,075 high-quality ques- tion–answer pairs across 6 core capability dimensions. Ex- tensive evaluations of 26 state-of-the-art MLLMs on LifeE- val reveal substantial challenges in achieving timely, ef- fective and adaptive interaction, highlighting essential di- rections for advancing human-centered interactive intelli- gence. 1. Introduction The long-standing ambition of artificial intelligence, achieving Artificial General Intelligence (AGI), has driven a fundamental shift in research priorities [24].Rather than focusing on narrowly defined and isolated tasks, re- cent studies aim to build systems that can understand, learn from, and assist humans in accomplishing complex activi- ties [30, 51]. Recent advances in multimodal large language * Corresponding author. Figure 1. Overview of LifeEval. The benchmark emphasizes real-life interactive scenarios, requiring models to deliver precise, context-aware, and adaptive assistance. models (MLLMs) [52] have opened unprecedented path- ways toward this vision. By integrating visual, linguistic, and other modalities within a unified framework, MLLMs have demonstrated strong cross-modal understanding and reasoning capabilities, encompassing tasks such as image captioning [2] and visual question answering [37]. These developments have laid a solid foundation for more precise, adaptive, and human-centered interaction. Meanwhile, as AI technologies become increasingly em- bedded in daily life through wearable and embodied plat- forms such as AI glasses and first-person cameras, the focus of intelligence is turning toward user-centered and egocen- tric interaction [20]. In such scenarios, an intelligent as- 1 arXiv:2603.00490v1 [cs.AI] 28 Feb 2026 sistant must go beyond perceiving and understanding the surrounding environment. It must stay aligned with the user’s perception, intentions, and ongoing activities in real time. This requires the model to analyze egocentric video streams to infer task goals and progress, recognize poten- tial confusion or errors, and provide context-aware and ac- tionable guidance. Unlike traditional third-person video un- derstanding, where the model functions as a detached ob- server, egocentric interaction requires the model to act as an embedded and proactive participant. This paradigm raises substantially higher requirements for dynamic perception, causal reasoning, and adaptive response, representing a crit- ical step toward the development of truly human-aligned in- telligent systems. To evaluate the video understanding abilities of MLLMs, numerous benchmarks have been developed, covering long- term video memory [23, 36, 54], comprehensive reasoning [19, 50], and egocentric understanding [6, 9]. However, ex- isting benchmarks exhibit two key limitations. First, most benchmarks evaluate models in offline settings using com- plete video clips, emphasizing retrospective analysis rather than real-time perception and response. Second, their evalu- ation largely focus on isolated understanding abilities, with limited connection to concrete user goals and intentions. These gaps make it difficult to assess how well models can operate as embedded assistants that must continually align with a user’s evolving intent and provide actionable guid- ance. In other words, existing research has focused on ”what a model can see and understand”, while paying lit- tle attention to ”what a model can actually do to help”, which is essential for building truly intelligent and useful systems. To bridge this gap, we introduce LifeEval, a novel benchmark designed to systematically evaluate MLLMs as real-time and task-oriented assistants in first-person scenar- ios, as illustrated in Fig. 1. The benchmark features three key characteristics: (1) Task-Oriented Holistic Evalua- tion. Instead of assessing isolated perceptual or reasoning skills, we evaluate the model’s holistic ability to assist users throughout task execution, including state tracking, knowl- edge integration, goal planning, error diagnosis, and ac- tionable guidance. (2) Egocentric Real-Time Perception. Models are exposed to continuous egocentric video streams, which requires robust online inference to capture seman- tically meaningful and task-relevant cues under dynamic real-world conditions. (3) Human–Assistant Collabora- tive Interaction. By simulating realistic human-assistant dialogues, our benchmark incorporates diverse and context- driven user questions, moving beyond predefined templates to reflect open-ended collaboration. Through a carefully designed data collection pipeline, LifeEval includes 4,075 high-quality question–answer pairs, covering both multiple-choice and open-ended for- mats.We further decompose task-assistance capability into 6 key dimensions: static environment perception, dy- namic task reasoning, contextual knowledge retrieval, goal- oriented planning, safety and feasibility assessment and multi-turn interactive collaboration. This structure enables a comprehensive and fine-grained evaluation of MLLMs’ performance as real-world assistants. Based on LifeEval, we conduct extensive benchmark- ing of a variety of state-of-the-art MLLMs, including both open-source and proprietary models. The results reveal significant gaps in current systems’ ability to provide re- liable and efficient assistance in egocentric, real-time set- tings. Through LifeEval, we aim to promote research on human-centered multimodal intelligence, and to advance the development of AI systems that can truly understand human intentions and collaborate effectively with people in solving real-world problems. In summary, our main contributions are: • We propose LifeEval, the first benchmark dedicated to evaluating MLLMs as real-time task assistants from a first-person perspective, marking a shift in focus from passive understanding to active, collaborative interaction. • We design a fine-grained six-dimensional evaluation framework and construct a large-scale, high-quality question-answer bank with detailed rationales via a purpose-built pipeline, ensuring both comprehensiveness and reliability. • We benchmark 26 state-of-the-art MLLMs, revealing their current limitations as intelligent assistants and iden- tifying future directions for advancing human-centered interactive intelligence. 2. Related Works 2.1. General Video Understanding Benchmarks With the rapid advancement of Multimodal Large Language Models (MLLMs) [7, 26, 29, 31, 38], evaluating video un- derstanding capabilities has become an increasingly impor- tant research focus [52]. Early benchmarks [13, 17, 40, 41] primarily concentrated on specific subtasks such as multi- modal retrieval and visual question answering. Subsequent works [35, 39] extended these efforts to temporal reason- ing, and more recent benchmarks [12, 36, 54] have focused on long-form video understanding. Among the more com- prehensive efforts, MMBench-Video [11] introduced a de- tailed taxonomy for fine-grained evaluation across multiple dimensions. Video-MME [12] incorporated diverse tasks like recognition, perception, and reasoning, to form a holis- tic evaluation framework. MVBench [19] innovatively for- mulated spatiotemporal tasks to jointly assess spatial and temporal reasoning within a unified benchmark. Despite these advances, most general video understand- ing benchmarks still lack evaluation of real-time perception 2 Table 1. Comparison of LifeEval with existing video benchmarks. Columns summarize key properties, including dataset size (#QAs, #Clips), egocentric perspective, interactive and real-time settings, task-oriented assistance, multi-turn dialogue, availability of reasoning rationales for each answer, and annotation method. Symbols: ✓ (Yes), ❍ (Partial), ✗ (No). Benchmark#QAs#ClipsEgocentricInteractiveReal-TimeTask-Assist.Multi-TurnRationaleQuestion TypeAnnotation MVBench [19]40003641✗CloseAuto VideoMME [12]2700900✗CloseManual Video-T [50]50001000✗✓Open&CloseAuto&Manual OmniMMI [34]22901121❍✓✗❍✗OpenManual OvO-Bench [25]2814644❍✗❍✗CloseAuto&Manual EgoVQA [10]520520✓✗CloseManual EgoSchema [23]50635063✓✗CloseAuto&Manual EgoPlan-Bench2 [27]13211113✓✗✓✗CloseAuto&Manual EOC-Bench [46]3277656✓✗Open&CloseManual EgoTextVQA [55]70641507✓❍✗OpenAuto&Manual EgoLifeQA [42]60006✓✗✓CloseManual VidEgoThink [6]49933665✓✗❍✗OpenAuto LifeEval (Ours)4075591✓❍✓Open&CloseAuto&Manual and response capabilities, which are essential for interac- tive understanding. Although several recent benchmarks [34, 43] have explored dialogue-based evaluation of stream- ing video understanding, they are not primarily designed for egocentric scenarios. This mismatch with the way humans naturally perceive and interpret the world limits their appli- cability to real-world human–AI collaboration. 2.2. Egocentric Video Understanding Benchmarks Compared with third-person videos, egocentric data more closely reflects human visual experience, capturing how individuals actually see and interact with the world. At the same time, the frequent and dynamic perspective changes introduce unique challenges for visual understand- ing and reasoning [20]. The release of large-scale ego- centric datasets [8, 14, 21, 32] has accelerated progress in this domain, enabling the development of benchmarks that assess diverse aspects of egocentric video understand- ing. EgoVQA [10] and EgoTaskQA [16] focused on video reasoning and comprehension abilities, while EgoSchema [23] emphasized temporal understanding of modern vision and language systems. EgoMemoria [44], HourVideo [3], and EgoMem [56] extended the temporal scope to study long-range memory modeling and cross-temporal semantic consistency. EgoPlan-Bench [5] and EgoPlan-Bench2[27] targeted task planning ability.VidEgoThink[6], EOC- Bench[46] and ECBench[9] introduced systematic multi- dimensional evaluations to quantify embodied cognitive abilities. EgoLifeQA [42] further expands the focus to long- term behavioral patterns and complex social interactions. However, most existing benchmarks primarily empha- size retrospective analysis rather than real-time interac- tive response, and are weakly coupled to explicit user goals or intentions. In contrast, LifeEval shifts the evalu- ation paradigm from passive perception to active collabo- ration, offering a more faithful assessment of dynamic hu- man–assistant interactions, as highlighted in the compari- son in Tab. 1. 3. LifeEval In this section, we first present a hierarchical capability tax- onomy for interactive agents. Then, we introduce the con- struction pipeline and detailed statistics of LifeEval. 3.1. Capability Taxonomy To systematically characterize the competencies required for a multimodal agent to act as an effective collabora- tive assistant, we introduce a hierarchical taxonomy con- sisting of six core capabilities. This taxonomy shifts the perspective from passive video understanding to active, context-aware collaboration, emphasizing the correctness, executability, and adaptability of the model’s responses dur- ing interaction. It progresses from fundamental percep- tual abilities to higher-level interactive reasoning, ensuring a comprehensive evaluation of real-time task assistance, as illustrated in Fig. 2. Static Environment Perception (SEP): The ability to rec- ognize and comprehend scene elements that are spatially or instantaneously observable. It focuses on foundational visual grounding, including object identification, attribute recognition, and scene structure comprehension. This di- mension captures precise and immediate situational aware- ness that serves as the perceptual basis for subsequent rea- soning and collaboration. Dynamic Task Reasoning (DTR): The capability to infer evolving task progress and state transitions through continu- ous video streams. It demands temporal reasoning to mon- itor task evolution, recognize user intentions, infer causal state changes, and predict short-term outcomes. The core of this dimension lies in interpreting the dynamic interplay between the user and the environment over time. Contextual Knowledge Retrieval (CKR): The ability to retrieve and integrate external prior knowledge with the current visual context, including commonsense, domain- 3 Figure 2. Capability dimensions and examples in LifeEval. LifeEval defines six core capability dimensions to comprehensively evaluate a model’s interactive assistance abilities in daily life. The benchmark features both multiple-choice and open-ended questions, each accompanied by detailed reasoning explanations. specific expertise, and procedural manuals. This enables reasoning that goes beyond what is directly visible and al- lows the model to provide users with additional and instruc- tive guidance. Goal-oriented Planning (GP): The ability to provide ef- fective solutions or next-step guidance based on the cur- rent state and the user’s goal. This dimension focuses on the model’s capacity to understand the user’s intention, as- sess the situation, and deliver actionable instructions that directly assist in problem solving or task advancement. Safety and Feasibility Assessment (SFA): The ability to evaluate whether the user’s intended actions or current situ- ation are safe and feasible. This includes identifying poten- tial risks, recognizing contraindications, assessing resource availability, and judging the viability of proposed actions, ensuring that the model can act as a reliable safeguard dur- ing task execution. Multi-turn Interactive Collaboration (MIC): The inte- grative ability to maintain coherent and multi-turn dialogue while supporting task completion. This requires preserv- ing contextual consistency, resolving ambiguities, handling coreferences, and collaboratively guiding the user toward task completion. This dimension integrates and evaluates the effective application of all preceding capabilities within a naturalistic conversational flow. 3.2. Benchmark Construction Video Collection To evaluate task-oriented assistance, we rely on egocentric video, which closely reflects the user’s natural perspective and interactions with the environ- ment. Ego4D [14] is a representative large-scale egocen- tric dataset, containing over 3,600 hours of densely anno- tated videos that spans a wide range of daily activities. Its unprecedented scale and scene diversity make it an ideal source for our benchmark. By leveraging both the native annotations and goal-step annotations [28], we carefully sample a balanced set of videos covering diverse everyday tasks, including but not limited to cooking, cleaning, handicrafts, shopping, laun- dry and device maintenance. This strategy ensures that our benchmark encompasses the broad spectrum of scenarios an intelligent assistant would encounter in real-world settings. In total, we collect 100 varied daily-life scenarios with a total duration of 44.19 hours. Data Curation Pipeline Following the collection of raw video data, we annotate the videos with high-quality ques- tion–answer (QA) pairs in both multiple-choice (MCQ) and open-ended (OEQ) formats, balancing objective evaluation with realistic interactive demands. To achieve cost-effective yet reliable annotation, we adopt a multi-stage pipeline that combines the advanced video reasoning capabilities of Gemini 2.5 Pro [7] with targeted human verification, 4 thereby mitigating potential generator-specific priors. (a) MCQ Generation via Direct Video Prompting. Unlike previous works that rely on dense textual narrations to generate QA pairs using text-only LLMs [5, 6, 19, 23, 27], our approach directly feeds the video content to the video MLLM. This ensures a direct alignment with the raw visual stream, preventing potential information loss or bias introduced by intermediate textual representations. Specif- ically, we first randomly sample 30-second clips from our collected videos to obtain manageable yet informative seg- ments, followed by manual filtering to remove uninforma- tive ones. For each clip, we employ carefully designed prompts tailored to each of our six capability dimensions, guiding Gemini 2.5 Pro to analyze the scene and generate multiple-choice questions. The model is asked to provide the question, correct answer, a compelling reasoning chain, a set of plausible distractors, and the specific timestamp within the clip that grounds the question, ensuring semantic precision and visual faithfulness. (b) Iterative Quality Filtering. To ensure the relia- bility and precision of the generated questions, we apply a rigorous two-stage quality control process. In the first stage, Gemini 2.5 Pro itself serves as an automated re- viewer. Given both the original video clips and the corre- sponding QA pairs, the model is prompted to identify and revise a variety of issues, including misalignment with the defined capability dimension, unclear phrasing, ambiguous or incorrect answers, weak visual grounding, or redundancy across questions. In the second stage, human annotators conduct a thorough manual review to confirm the clarity, accuracy, and contextual relevance of each QA pair, filter- ing out any questions that fail to meet our quality standards. (c) Controlled Difficulty Enhancement. To further strengthen the benchmark’s discriminative capability, we further introduce a difficulty enhancement stage. In this phase, the questions that pass the previous quality filtering stage are fed back into Gemini 2.5 Pro together with the corresponding video. The model is instructed to reformu- late each question to make it more challenging, involving strategies such as requiring multi-step reasoning, incorpo- rating subtle visual cues, or introducing more plausible dis- tractors. The enhanced questions are then reviewed by hu- man annotators to ensure they remain fair, unambiguous, and firmly grounded in the visual content. (d) Open-ended Question Reformulation. The final set of high-quality multiple-choice questions serves as the foundation for generating open-ended ones. We prompt Gemini 2.5 Pro to rephrase each MCQ into a concise short- answer question, ensuring that every reformulated question has a clear and semantically unique answer. To preserve answer consistency and evaluability, we exclude questions that are overly broad or subjective. A final round of manual review is then conducted to verify the quality of all refor- Table 2. Question statistics of LifeEval. StatisticsNumberPercentage Capacities - Static Environment Perception (SEP)66616.34% - Dynamic Task Reasoning (DTR)70817.37% - Contextual Knowledge Retrieval (CKR)77519.02% - Goal-oriented Planning (GP)74418.26% - Safety and Feasibility Assessment (SFA)64615.85% - Multi-turn Interactive Collaboration (MIC)53613.15% Formats - Multiple-Choice Questions206950.77% - Open-Ended Questions200649.23% Total4075100% mulated questions. Through this multi-stage pipeline of generation, filter- ing, enhancement, and reformulation, we finally construct a high-quality benchmark of 4,075 QA pairs, each enriched with reasoning chains and precise temporal grounding. Evaluation Metrics We adopt separate evaluation strate- gies for multiple-choice and open-ended questions to en- sure comprehensive and accurate assessment. For multiple- choice questions, following standard practice [12, 19], we report accuracy by directly matching the model’s response to the ground-truth answer. For open-ended questions, evaluation focuses on the se- mantic correctness of free-form text responses. To achieve this, we use the LLM-as-a-Judge paradigm [4, 53], where GPT-5 [26] is leveraged as our judge model. In our evalu- ation framework, the judge compares the model’s answer with the ground-truth reference and assigns a score on a scale from 0 to 1 in 0.25 increments. This graded scor- ing scheme offers a more nuanced and precise assessment than binary accuracy, effectively capturing partial correct- ness and subtle differences in answer quality. 3.3. Benchmark Statistics LifeEval comprises 4,075 QA pairs across 591 video clips, spanning two question formats and six capability dimen- sions to systematically capture everyday human–AI collab- oration scenarios. As summarized in Tab. 2, the QA pairs are evenly distributed across both question types and ca- pability dimensions, ensuring a balanced and fine-grained assessment of MLLMs’ interactive abilities. To further characterize the benchmark composition, we examine the task goals and linguistic distributions of the QA pairs in Fig. 3. Specifically, we analyze the goal de- scriptions [28] of all video scenes, as visualized in Fig. 3a, which highlights the eight most frequent verbs and their corresponding high-frequency nouns. The results reveal a diverse and balanced coverage of common daily activities, confirming that LifeEval encompasses a broad spectrum of 5 (a) Most frequent verb–noun pairs in task goals (Top 8 verbs).(b) Top 15 most frequent nouns and verbs in questions.(c) Top 15 most frequent nouns and verbs in answers. Figure 3. Task goal distribution and QA vocabulary statistics in LifeEval. realistic and representative real-world scenarios. Moreover, Figs. 3b and 3c presents the overall word frequency distri- bution across all QA pairs. We observe that answers tend to contain more specific and concrete nouns, reflecting their goal of providing precise guidance. In contrast, questions exhibit a broader and more diverse vocabulary, including colloquial high-frequency terms such as “look” and “feel”. This pattern indicates that LifeEval effectively captures the natural conversational style of everyday human–assistant interactions rather than artificially constrained or overly for- mal queries. 4. Experiments 4.1. Experimental Setup To establish comprehensive performance baselines, we evaluate 26 leading MLLMs on LifeEval.Specifically, we evaluate six state-of-the-art proprietary MLLMs: GPT- 5 [26], GPT-4o [15], GPT-5-mini [26], Gemini-2.5-Pro [7], Gemini-2.5-Flash [7], and Grok-4 [38]. For open- source MLLMs, five prominent model families are evalu- ated across multiple parameter scales, including Qwen3-VL [29], InternVL3.5 [31], LLaVA-OneVision [18], LLaVA- NeXT [22], and mPLUG-Owl3 [45], each assessed across multiple parameter scales.Furthermore, we also in- clude four video-specialized models: VideoLLaMA3 [47], LLaVA-Video [49], LongVA [48], and InternVideo2.5 [33]. We adopt a streaming-style evaluation protocol: for each query, the model receives only the visual content within the corresponding temporal window rather than the full video. All models are evaluated in a zero-shot setting with de- fault configurations. For models without native video sup- port, we uniformly sample 8 frames from the same inter- val. Given our emphasis on real-time responsiveness and the brief intervals, this strategy provides representative yet concise visual input while ensuring fair comparison, as val- idated in subsequent analyses. To ensure consistent evalu- ation, we enforce a rule-based output format requiring an- swers within predefined tags. 4.2. Main Results In this section, we provide a detailed comparison and analy- sis. Tab. 3 reports performance across all capability dimen- sions and question types, with all scores linearly normalized to 0–100 for consistent comparison. Proprietary MLLMs. Proprietary MLLMs show consis- tently strong performance, with GPT-5 and Gemini-2.5-Pro achieving the highest overall results. On multiple-choice questions, Gemini-2.5-Pro achieves 85.61% accuracy, out- performing the best open-source MLLM by 13.42% and the best video-specialized LM by 27.19%. For open-ended questions, GPT-5 leads with a score of 72.39%, exceed- ing the best open-source and video-specialized models by 23.15% and 43.48%, respectively. Analysis across capabil- ity dimensions reveals a notable performance gap. These models exhibit high accuracy in Contextual Knowledge Re- trieval (CKR), demonstrating a robust capacity for access- ing and integrating external knowledge. However, their performance declines in Dynamic Task Reasoning (DTR) and Goal-oriented Planning (GP). This contrast underscores a critical weakness in procedural reasoning and adaptive planning, both of which are essential capabilities for intelli- gent task execution. Open-Source MLLMs. Among open-source MLLMs, the Qwen3-VL series achieves the best overall performance. Although LLaVA-OneVision-72B attains the highest accu- racy on multiple-choice questions, its performance on open- ended responses remains less competitive, suggesting lim- itations in generating precise and semantically grounded answers. The LLaVA-NeXT series also benefits from its large parameter scale, delivering stable yet suboptimal re- sults compared with Qwen3-VL. Overall, a substantial per- formance gap persists between open-source and proprietary 6 Table 3. Main evaluation results of representative MLLMs on LifeEval across different capability dimensions and question types. Bold andunderline denote the best and second-best results, respectively. Model Multiple-Choice QuestionsOpen-Ended Questions SEPDTRCKRGPSFAMICOverallSEPDTRCKRGPSFAMICOverall Proprietary MLLMs GPT-5 [26]79.1079.7388.2176.6684.4587.9282.4863.6766.0580.6569.8979.8073.9272.39 GPT-4o [15]68.0670.0081.5467.1177.7482.8774.2346.1548.0861.2352.7264.2359.2955.19 GPT-5-mini [26] 77.3175.6884.3675.0782.0186.2579.8564.9558.2173.1863.6975.7969.1267.44 Gemini-2.5-Flash [7]71.0470.0085.3872.4182.6281.3276.9856.1257.9171.3059.5472.5667.2364.05 Gemini-2.5-Pro [7]83.2879.4691.5481.7089.0289.6885.6165.8660.5874.1665.9477.7572.3169.32 Grok-4 [38] 71.9467.5785.3867.9079.8880.2775.3046.2251.3369.6854.7072.0167.6260.07 Open-Source MLLMs Qwen3-VL-30B-A3B [29]60.3965.1275.7866.0475.8778.0970.0146.9540.5451.6441.7760.6952.3948.61 Qwen3-VL-30B-A3B(thinking) [29] 61.4965.1476.9267.3778.6678.7871.0941.4144.8853.3144.2862.5450.4149.24 Qwen3-VL-8B [29]52.8454.5964.1055.9766.7770.3560.3337.8436.3242.9237.6047.2542.4140.61 Qwen3-VL-8B(thinking) [29]59.1065.9572.0565.2573.4874.4468.1639.5039.2047.3441.6957.0045.7844.96 InternVL3.5-38B [31] 57.0154.8663.3353.3260.9871.7259.6929.3830.1032.0129.0236.4036.4131.99 InternVL3.5-14B [31]55.5257.0361.5450.9353.9667.3857.3826.7425.8928.7722.3436.0134.2428.65 InternVL3.5-8B [31]53.1360.0060.0052.7958.8464.5657.9823.8725.7424.2920.6430.9028.4425.40 InternVL3.5-4B [31]48.0651.8954.8750.4053.0563.3253.2325.4523.2222.0820.9127.7527.7824.27 LLaVA-OneVision-72B [18] 65.3770.0075.9072.4172.5677.5472.1935.0536.2443.9635.7647.2538.4939.48 LLaVA-OneVision-1.5-8B [1] 51.6451.1466.0050.9971.8052.9757.4717.7419.6928.2227.6333.3314.4223.83 LLaVA-OneVision-7B [18] 56.4253.2461.7953.3251.5261.4356.1729.9125.7431.9529.1632.5529.4329.81 mPLUG-Owl3-7B [45]49.5557.0366.4153.8561.2832.9654.5521.3725.6726.7521.5332.1527.3625.66 mPLUG-Owl3-2B [45]34.9341.3549.2341.6447.8728.5641.2214.7311.0911.8813.4211.486.9111.78 LLaVA-NeXT-110B [22] 55.2260.8170.2659.9576.2272.2165.4526.5132.3234.8734.2048.1934.2834.97 LLaVA-NeXT-72B [22]54.9362.1670.7764.1973.7873.2066.2630.8933.5141.4337.1949.0637.2738.24 LLaVA-NeXT-7B [22]40.9044.3242.3140.8552.1353.3145.1611.639.3210.0010.5617.2211.5511.61 Open-Source Video-Specialized LMs VideoLLaMA3-7B [47]49.2550.5462.0550.6658.5453.1354.1323.2621.8225.9123.2331.5314.7123.69 LLaVA-Video-7B [49]54.9357.0363.3356.2356.4063.0758.4226.4426.7828.2527.8633.9628.8228.61 LongVA-7B [48]48.9652.4361.0354.6462.8064.2557.0719.1125.3027.7923.9835.3028.0626.47 InternVideo2.5 Chat8B [33]53.1358.3856.6754.1157.9363.2656.9925.9826.5529.8728.0733.7329.5328.91 models across all evaluated capabilities.This disparity is particularly evident in open-ended question answering, where all open-source models fall short of a 50% accuracy threshold. These results indicate that current open-source MLLMs still lack the robustness required to deliver precise, context-aware, and actionable guidance necessary for effec- tive task-oriented collaboration. Open-SourceVideo-SpecializedLMs. Thevideo- specialized LMs achieve relatively close performance, with LLaVA-Video obtaining the highest score on multiple- choice questions and InternVideo2.5 slightly outperforming others on open-ended questions. When compared to open- source MLLMs of similar parameter size, it becomes evident that the main constraint lies not in architectural specialization, but in the limited model capacity itself. Current parameter scales appear inadequate to support the complex understanding and reasoning required for effective human–AI collaboration in real-world assistive contexts. 4.3. Further Analysis We conduct an in-depth analysis to identify the critical fac- tors influencing model performance as capable task assis- tants, focusing on capability distribution, scaling effects, and input efficiency. Performance Gaps Arise from Reasoning and Interac- tion Challenges Rather Than Scene Variations. Fig. 4 illustrates the performance of representative MLLMs across different capability dimensions and everyday scenarios. Overall, GPT-5 and Gemini-2.5-Pro exhibit consistently strong and well-balanced performance across all evalu- ated aspects, whereas open-source models display more un- even capability distributions. For instance, VideoLLaMA3 shows a notable weakness in Multi-Turn Interactive Collab- oration (MIC), while Qwen3-VL performs relatively poorly in Static Environment Perception (SEP). Besides, all mod- els perform considerably better on multiple-choice ques- tions than on open-ended ones, suggesting that structured options allow models to rely more on recognition, whereas open-ended demand deeper reasoning and generative pre- cision. Across video scenarios, however, the performance gap remains modest, implying that the primary challenge lies not in scene-specific variations but in the higher-level reasoning and interaction demands of collaborative tasks. Scaling Parameters Alone Is Insufficient for Collabo- rative Egocentric Tasks. We visualize the relationship 7 20 40 60 80 20 40 60 80 Figure 4. Performance comparison of representative MLLMs on LifeEval. Left: Evaluation results across six capability dimensions. Right: Evaluation results across five everyday scenarios. Figure 5. Relationship between model parameter scale and performance on LifeEval for open-source MLLM families. Table 4. Performance comparison with varying input frame. Num- bers in the second row of each model denote performance gain rel- ative to 1-frame. # Frames MCQOEQ 1f8f32f1f8f32f GPT-5 80.0382.4882.5167.9472.3970.32 (+2.45)(+2.48)(+4.45)(+2.38) InternVL3.5-8B 56.7257.9857.2624.3825.4026.31 (+1.26)(+0.54)(+1.02)(+1.93) LLaVA-OV-7B 56.6556.1755.3929.1129.8129.22 (−0.48)(−1.26)(+0.70)(+0.11) between model performance and parameter count across various open-source MLLM families in Fig. 5.While model performance generally increases with larger parame- ter counts across MLLM families, notable exceptions reveal the limitations of pure scaling. For example, in multiple- choice tasks, the InternVL3.5 family exhibits a performance drop when scaling from 8B to 148B parameters, with the 38B variant offering no clear improvement.Similarly, LLaVA-NeXT-110B performs worse than its smaller 72B counterpart. These inconsistencies suggest that larger mod- els may overemphasize complex reasoning or hallucinate ir- relevant details, which can be detrimental in egocentric col- laboration scenarios. In contrast, Qwen3-VL demonstrates strong performance even with only 30B parameters, high- lighting that effective alignment with interaction-oriented and context-grounded understanding tasks can outweigh the benefits of mere model scaling. Moderate Frame Sampling Suffices for Real-Time Inter- action. We further investigate the impact of varying input frames numbers on LifeEval, as illustrated in Tab. 4. Per- formance improves when increasing from 1 to 8 frames, but further expanding to 32 frames yields no consistent gains and in some cases even degrades results, suggesting that 8 frames represent a near-saturation point for this bench- mark. The effect is most pronounced for GPT-5, whereas it is relatively less evident for InternVL3.5-8B and LLaVA- OneVision-7B, likely due to their limited ability to process multiple frames. This pattern differs from other benchmarks [11, 12, 46, 48], where denser frame sampling often leads to steady improvements. The difference stems from our fo- cus on real-time interaction and cross-domain knowledge integration beyond temporal perception, thereby reducing the benefit of very dense sampling. Excessively frequent frame inputs offer diminishing informational returns while increasing computational latency. These findings highlight a practical trade-off for efficient multimodal assistants: us- ing a moderate number of frames can reduce inference la- tency while maintaining strong performance, better aligning with real-time interaction demands. 5. Conclusion In this work, we introduce LifeEval, a benchmark for evalu- ating real-time, task-oriented human–AI collaboration from an egocentric perspective. Unlike existing video bench- marks that focus on passive or offline understanding, LifeE- val emphasizes interactive assistance in dynamic, real- world scenarios. With 4,075 QA pairs spanning 6 core capability dimensions, it provides a structured framework for assessing how well MLLMs perceive, reason, and as- sist humans in continuous first-person environments. Ex- tensive evaluations of 26 state-of-the-art MLLMs show that, despite strong visual–language capabilities, current models still struggle with adaptive reasoning, multi-turn interac- tion, and grounded collaboration. The gap between mod- els, along with limited gains from scaling or frame sam- pling, further highlights the challenges of translating static understanding into real-time assistance. LifeEval offers a human-centered evaluation framework for advancing more interactive and context-aware multimodal assistants. 8 6. Acknowledgement This work was supported by the National Key R&D Program of China (2025ZD0124104) in collaboration with Shanghai Artificial Intelligence Laboratory, in part by the China Postdoctoral Science Foundation under Grant 2025M781485, and in part by National Natu- ral Science Foundation of China under Grant 62571324. References [1] Xiang An, Yin Xie, Kaicheng Yang, Wenkang Zhang, Xiuwei Zhao, Zheng Cheng, Yirui Wang, Songcen Xu, Changrui Chen, Chunsheng Wu, Huajie Tan, Chunyuan Li, Jing Yang, Jie Yu, Xiyao Wang, Bin Qin, Yumeng Wang, Zizhen Yan, Ziyong Feng, Ziwei Liu, Bo Li, and Jiankang Deng. Llava-onevision-1.5: Fully open framework for de- mocratized multimodal training. In arxiv, 2025. 7 [2] Shuang Bai and Shan An. A survey on automatic image cap- tion generation. Neurocomputing, 311:291–304, 2018. 1 [3] Keshigeyan Chandrasegaran, Agrim Gupta, Lea M Hadzic, Taran Kota, Jimming He, Crist ́ obal Eyzaguirre, Zane Du- rante, Manling Li, Jiajun Wu, and Li Fei-Fei. Hourvideo: 1-hour video-language understanding. Advances in Neural Information Processing Systems, 37:53168–53197, 2024. 3 [4] Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark. In Forty- first International Conference on Machine Learning, 2024. 5 [5] Yi Chen, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li, Rui Wang, Ruifeng Xu, Ying Shan, and Xihui Liu. Egoplan- bench: Benchmarking egocentric embodied planning with multimodal large language models. CoRR, 2023. 3, 5 [6] Sijie Cheng, Kechen Fang, Yangyang Yu, Sicheng Zhou, Bo- hao Li, Ye Tian, Tingguang Li, Lei Han, and Yang Liu. Vide- gothink: Assessing egocentric video understanding capabili- ties for embodied ai. arXiv preprint arXiv:2410.11623, 2024. 2, 3, 5 [7] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blis- tein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025. 2, 4, 6, 7 [8] Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. The epic-kitchens dataset: Collection, challenges and base- lines. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11):4125–4141, 2020. 3 [9] Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin, Bo- qiang Zhang, Long Li, Liuyi Wang, Qinyang Zeng, Xin Li, and Lidong Bing. Ecbench: Can multi-modal foundation models understand the egocentric world? a holistic embod- ied cognition benchmark. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 24593– 24602, 2025. 2, 3 [10] Chenyou Fan. Egovqa-an egocentric video question answer- ing benchmark dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 3 [11] Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen. Mmbench-video: A long-form multi-shot benchmark for holistic video under- standing. Advances in Neural Information Processing Sys- tems, 37:89098–89124, 2024. 2, 8 [12] Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 24108–24118, 2025. 2, 3, 5, 8 [13] Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michal- ski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The” something something” video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on com- puter vision, pages 5842–5850, 2017. 2 [14] Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 18995–19012, 2022. 3, 4 [15] Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 6, 7 [16] Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang. Egotaskqa: Understanding human tasks in egocentric videos. Advances in Neural Information Processing Systems, 35: 3343–3360, 2022. 3 [17] Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics hu- man action video dataset. arXiv preprint arXiv:1705.06950, 2017. 2 [18] Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 6, 7 [19] Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understand- ing benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195– 22206, 2024. 2, 3, 5 [20] Xiang Li, Heqian Qiu, Lanxiao Wang, Hanwen Zhang, Chenghao Qi, Linfeng Han, Huiyu Xiong, and Hongliang 9 Li. Challenges and trends in egocentric vision: A survey. arXiv preprint arXiv:2503.15275, 2025. 1, 3 [21] Yin Li, Miao Liu, and James M Rehg. In the eye of the be- holder: Gaze and actions in first person video. IEEE trans- actions on pattern analysis and machine intelligence, 45(6): 6731–6747, 2021. 3 [22] Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023. 6, 7 [23] Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long- form video language understanding. Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023. 2, 3, 5 [24] Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg. Levels of agi for op- erationalizing progress on the path to agi. arXiv preprint arXiv:2311.02462, 2023. 1 [25] Junbo Niu, Yifei Li, Ziyang Miao, Chunjiang Ge, Yuanhang Zhou, Qihao He, Xiaoyi Dong, Haodong Duan, Shuangrui Ding, Rui Qian, et al. Ovo-bench: How far is your video- llms from real-world online video understanding? In Pro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025. 3 [26] OpenAI.Gpt-5. https://openai.com/index/ introducing-gpt-5, 2025. 2, 5, 6, 7 [27] Lu Qiu, Yi Chen, Yuying Ge, Yixiao Ge, Ying Shan, and Xihui Liu. Egoplan-bench2: A benchmark for multimodal large language model planning in real-world scenarios. arXiv preprint arXiv:2412.04447, 2024. 3, 5 [28] Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani. Ego4d goal-step: Toward hierarchical understanding of procedural activities. Advances in Neural Information Processing Systems, 36: 38863–38886, 2023. 4, 5 [29] Qwen Team. Qwen3 technical report, 2025. 2, 6, 7 [30] Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, et al. A survey on human-ai teaming with large pre-trained models. arXiv e-prints, pages arXiv–2403, 2024. 1 [31] Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Sheng- long Ye, Jie Shao, et al. Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265, 2025. 2, 6, 7 [32] Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bu- gra Tekin, Felipe Vieira Frujeri, et al. Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 20270–20281, 2023. 3 [33] Yi Wang, Xinhao Li, Ziang Yan, Yinan He, Jiashuo Yu, Xi- angyu Zeng, Chenting Wang, Changlian Ma, Haian Huang, Jianfei Gao, Min Dou, Kai Chen, Wenhai Wang, Yu Qiao, Yali Wang, and Limin Wang. Internvideo2.5: Empower- ing video mllms with long and rich context modeling. arXiv preprint arXiv:2501.12386, 2025. 6, 7 [34] Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu, Dongyan Zhao, and Zilong Zheng.Omnimmi: A comprehensive multi-modal interaction benchmark in streaming video con- texts. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 18925–18935, 2025. 3 [35] Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan. Star: A benchmark for situated reason- ing in real-world videos. arXiv preprint arXiv:2405.09711, 2024. 2 [36] Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long-context interleaved video-language understanding. Advances in Neural Informa- tion Processing Systems, 37:28828–28857, 2024. 2 [37] Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. Visual question answer- ing: A survey of methods and datasets. Computer Vision and Image Understanding, 163:21–40, 2017. 1 [38] xAI. Grok 4. https://x.ai/news/grok-4, 2025. 2, 6, 7 [39] Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 9777–9786, 2021. 2 [40] Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answer- ing via gradually refined attention over appearance and mo- tion. In Proceedings of the 25th ACM international confer- ence on Multimedia, pages 1645–1653, 2017. 2 [41] Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016. 2 [42] Jingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong, Xi- amengwei Zhang, Sicheng Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Ziyue Wang, et al. Egolife: Towards ego- centric life assistant. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 28885–28900, 2025. 3 [43] Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue, Sheng- sheng Qian, Jiahong Wu, Fan Yang, Weiming Dong, and Changsheng Xu.Svbench: A benchmark with tempo- ral multi-turn dialogues for streaming video understanding. arXiv preprint arXiv:2502.10810, 2025. 3 [44] Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen, Zongyu Lin, Yanghao Li, Bowen Zhang, Haoxuan You, Dan Xu, Zhe Gan, et al.Mm-ego: Towards building egocentric multimodal llms for video qa. arXiv preprint arXiv:2410.07177, 2024. 3 [45] Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024. 6, 7 [46] Yuqian Yuan, Ronghao Dang, Long Li, Wentong Li, Dian Jiao, Xin Li, Deli Zhao, Fan Wang, Wenqiao Zhang, Jun 10 Xiao, et al. Eoc-bench: Can mllms identify, recall, and forecast objects in an egocentric world?arXiv preprint arXiv:2506.05287, 2025. 3, 8 [47] Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, et al. Videollama 3: Frontier multi- modal foundation models for image and video understand- ing. arXiv preprint arXiv:2501.13106, 2025. 6, 7 [48] Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024. 6, 7, 8 [49] Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Zi- wei Liu, and Chunyuan Li. Video instruction tuning with synthetic data, 2024. 6, 7 [50] Yuanhan Zhang, Yunice Chew, Yuhao Dong, Aria Leo, Bo Hu, and Ziwei Liu. Towards video thinking test: A holistic benchmark for advanced video reasoning and understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20626–20636, 2025. 2, 3 [51] Zicheng Zhang, Junying Wang, Yijin Guo, Farong Wen, Zijian Chen, Hanqing Wang, Wenzhe Li, Lu Sun, Yingjie Zhou, Jianbo Zhang, Bowen Yan, Ziheng Jia, Jiahao Xiao, Yuan Tian, Xiangyang Zhu, Kaiwei Zhang, Chunyi Li, Xi- aohong Liu, Xiongkuo Min, Qi Jia, and Guangtao Zhai. Aibench: Towards trustworthy evaluation under the 45° law. Displays, page 103255, 2025. 1 [52] Zicheng Zhang, Junying Wang, Farong Wen, Yijin Guo, Xi- angyu Zhao, Xinyu Fang, Shengyuan Ding, Ziheng Jia, Ji- ahao Xiao, Ye Shen, Yushuo Zheng, Xiaorong Zhu, Yalun Wu, Ziheng Jiao, Wei Sun, Zijian Chen, Kaiwei Zhang, Kang Fu, Yuqin Cao, Ming Hu, Yue Zhou, Xuemei Zhou, Jun- tai Cao, Wei Zhou, Jinyu Cao, Ronghui Li, Donghao Zhou, Yuan Tian, Xiangyang Zhu, Chunyi Li, Haoning Wu, Xiao- hong Liu, Junjun He, Yu Zhou, Hui Liu, Lin Zhang, Zesheng Wang, Huiyu Duan, Yingjie Zhou, Xiongkuo Min, Qi Jia, Dongzhan Zhou, Wenlong Zhang, Jiezhang Cao, Xue Yang, Junzhi Yu, Songyang Zhang, Haodong Duan, and Guangtao Zhai. Large multimodal models evaluation: A survey. SCI- ENCE CHINA Information Sciences, 2025. 1, 2 [53] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623, 2023. 5 [54] Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu. Mlvu: A comprehensive benchmark for multi- task long video understanding. arXiv e-prints, pages arXiv– 2406, 2024. 2 [55] Sheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li, Xun Yang, Dan Guo, Meng Wang, Tat-Seng Chua, and Angela Yao. Egotextvqa: Towards egocentric scene-text aware video question answering. In Proceedings of the Computer Vi- sion and Pattern Recognition Conference, pages 3363–3373, 2025. 3 [56] Jialong Zuo, Yongtai Deng, Lingdong Kong, Jingkang Yang, Rui Jin, Yiwei Zhang, Nong Sang, Liang Pan, Ziwei Liu, and Changxin Gao. Videolucy: Deep memory backtracking for long video understanding. arXiv preprint arXiv:2510.12422, 2025. 3 11