Paper deep dive
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela Yao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 4/3/2026, 12:16:06 AM
Summary
The paper introduces MyEgo, a novel diagnostic dataset and benchmark for evaluating Multimodal Large Language Models (MLLMs) on 'ego-grounding'âthe ability to understand, track, and reason about the camera-wearer's identity, actions, and objects in long-form egocentric videos. The study reveals that current state-of-the-art models, including GPT-5 and Qwen3-VL, significantly underperform compared to humans, struggling with long-range temporal memory and distinguishing the camera-wearer from other people or similar objects.
Entities (5)
Relation Signals (3)
MyEgo â evaluates â MLLMs
confidence 95% ¡ MyEgo, the first egocentric VideoQA dataset designed to evaluate MLLMs' ability to understand, remember, and reason about the camera wearer.
GPT-5 â isa â MLLMs
confidence 95% ¡ Top closed- and open-source models (e.g., GPT-5 and Qwen3-VL)
MLLMs â performspoorlyon â MyEgo
confidence 90% ¡ Benchmarking reveals that competitive MLLMs across variants... all struggle on MyEgo.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce MyEgo, the first egocentric VideoQA dataset designed to evaluate MLLMs' ability to understand, remember, and reason about the camera wearer. MyEgo comprises 541 long videos and 5K personalized questions asking about "my things", "my activities", and "my past". Benchmarking reveals that competitive MLLMs across variants, including open-source vs. proprietary, thinking vs. non-thinking, small vs. large scales all struggle on MyEgo. Top closed- and open-source models (e.g., GPT-5 and Qwen3-VL) achieve only~46% and 36% accuracy, trailing human performance by near 40% and 50% respectively. Surprisingly, neither explicit reasoning nor model scaling yield consistent improvements. Models improve when relevant evidence is explicitly provided, but gains drop over time, indicating limitations in tracking and remembering "me" and "my past". These findings collectively highlight the crucial role of ego-grounding and long-range memory in enabling personalized QA in egocentric videos. We hope MyEgo and our analyses catalyze further progress in these areas for egocentric personalized assistance. Data and code are available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2604.01966v1
- Canonical: https://arxiv.org/abs/2604.01966v1
Trouble viewing inline? Open PDF directly â
Full Text
69,406 characters extracted from source content.
Expand or collapse full text
Ego-Grounding for Personalized Question-Answering in Egocentric Videos Junbin Xiao 1,2 * ,Shenglang Zhang 1,2* ,Pengxiang Zhu 2 ,Angela Yao 2 1 University of Science and Technology of China, 2 National University of Singapore junbinxiao@ustc.edu.cn, zsl142857@mail.ustc.edu.cn, angela.yao@nus.edu.sg Abstract We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question- answering requiring ego-grounding - the ability to under- stand the camera-wearer in egocentric videos. To this end, we introduce MyEgo, the first egocentric VideoQA dataset designed to evaluate MLLMsâ ability to understand, remem- ber, and reason about the camera wearer. MyEgo comprises 541 long videos and 5K personalized questions asking about âmy thingsâ, âmy activitiesâ, and âmy pastâ. Bench- marking reveals that competitive MLLMs across variants, including open-source vs. proprietary, thinking vs. non- thinking, small vs. large scales all struggle on MyEgo. Top closed- and open-source models (e.g., GPT-5 and Qwen3- VL) achieve only 46% and 36% accuracy, trailing human performance by near 40% and 50% respectively. Surpris- ingly, neither explicit reasoning nor model scaling yield consistent improvements. Models improve when relevant evidence is explicitly provided, but gains drop over time, indicating limitations in tracking and remembering âmeâ and âmy pastâ. These findings collectively highlight the crucial role of ego-grounding and long-range memory in enabling personalized QA in egocentric videos. We hope MyEgo and our analyses catalyze further progress in these areas for egocentric personalized assistance. 1. Introduction As cameras in smart glasses and other wearable devices become ubiquitous, egocentric videos are emerging as a powerful medium for capturing first-person visual experi- ences [9, 15, 21, 36, 45, 46]. These continuous streams enable personalized assistance and timeline construction, helping users recall what they saw, did or interacted with throughout the day. Achieving this requires ego-grounding: understanding âmeâ, âmy thingsâ, âmy activitiesâ, and âmy pastâ in âmyâ first-person videos. We introduce the notion of ego-grounding as a core pre- * Equal Contribution. Is that my rag? -No, your rag is on the steering wheel. 00:26 01:52 Narratives: In a workshop, you wipe your hands using a red rag after finishing your job. Then, you put it on a steering wheel. You start to clean the working area, and again you want to wipe your hands, but you find there is a red rag on the floor, so you might ask: Is that my rag? ... ... Figure 1. Example of egocentric personalized QA concerning my demands. The narratives are provided for better understanding. requisite for personalized egocentric QA assistants. While related ideas appear in co-reference resolution [18, 39], and personalized VLM [1, 7], none addresses the visual and temporal challenges of grounding first-person references (âIâ, âmyâ) in egocentric videos. Current multimodal large language models (MLLMs) [3, 8, 16, 41, 49, 55] appear promising, due to their strong visual reasoning and long- context capabilities. However, it remains unclear whether current models can actually perform ego-grounding reli- ably. In egocentric video, the camera-wearer is often only partially visible, with hands, arms, ego-motion or brief re- flections, making the task especially challenging. For example, consider Fig. 1; answering the simple ques- tion about âmy ragâ. Humans do such tasks easily, yet current MLLMs often fail in such scenarios, revealing that ego-grounding poses challenges not captured by popular VideoQA and egocentric benchmarks [6, 15, 46]. This ex- ample highlights the broader demands of ego-grounding: models must separate the camera-wearer from nearby peo- ple, track âmy thingsâ through visually ambiguous scenes, and recall interactions that may no longer be visible. Successful ego-grounding therefore requires both spatial discrimination and long-range temporal reasoning. Since these capabilities remain open challenges for state-of-the- art MLLMs [2, 8, 16, 23, 41], we investigate how well cur- rent models can perform ego-grounding for answering per- sonalized questions in practice. To that end, we introduce MyEgo, the first personalized egocentric VideoQA bench- 1 arXiv:2604.01966v1 [cs.CV] 2 Apr 2026 mark that emphasizes ego-grounding for answers in egocen- tric videos. Unlike prior VideoQA datasets, the questions in MyEgo are intentionally diagnostic, probing concepts such as âIâ, âmyâ things, âmyâ activities, and âmyâ past inter- actions. MyEgo comprises 541 long videos (9.2 minutes on average) and 5K manually annotated questions. Each question is tied to a specific timestamp with also the corre- sponding answer moment, enabling controlled evaluation in realistic streaming QA settings. We benchmark a wide range of open-source and propri- etary MLLMs. We find that all models struggle: the state- of-the-art GPT-5 achieves only 46.1% accuracy, which is just over half of human performance at 84.7% accuracy. Beyond low accuracy, we also observe systematic failure modes: models confuse the camera-wearer with nearby people, mix up personal vs. non-personal objects, answer using only the most salient visible evidence, and frequently ignore temporally distant and visually inconspicuous but necessary context. Surprisingly, neither model scaling nor chain-of-thought reasoning yields stable and consistent im- provement. To further investigate these failures, we an- alyze model performances over time and conduct oracle experiments in which models are directly shown key evi- dence frames. Results show that model performances decay rapidly over time, and providing key frames helps. For instance, every MLLM tested correctly identifies âmyâ rag on the steering wheel when the question is posed at 26s in Fig. 1. However, all models answer incorrectly when the same question is asked later, after the rag leaves the scene and another person appears with a similar rag. These failures suggest that current models struggle to main- tain stable identity- and object-level representations over time. Instead of grounding the reference to my rag, they likely fall back on short-term appearance cues, leading to errors once the true referent is no longer visible. These lim- itations are further exacerbated by the fact that most models process only 8â32 frames at a time, restricting long-range temporal integration. Together, our findings show that cur- rent MLLMs lack robust ego-grounding and highlight the need for future improvements in long-term memory, tem- poral tracking, and precise retrieval to support personalized QA in egocentric video. Our contributions are threefold: ⢠The first systematic analysis of personalized question an- swering requiring ego-grounding in MLLMs, evaluating their ability to understand, remember, and reason about the camera-wearer in egocentric videos. ⢠MyEgo, a large-scale, diagnostic dataset and benchmark for egocentric personalized VideoQA. ⢠Detailed analyses and insights that reveal limitations of current MLLMs and guidance towards essential research directions for personalized egocentric AI assistants. 2. Related Works Egocentric VQA. Earlier egocentric VQA focused on ac- tion recognition [12] or high-level task understanding [17], yet with rarely a focus on multiple people interactions or coordination. With the push toward egocentric embodied assistance [14, 30, 31, 42] and the release of large-scale egocentric datasets (e.g., Ego4D [15]), egocentric VQA has attracted growing interest [6, 10, 26, 28, 32, 56]. However, most of this work focuses on general-purpose visual under- standing [6, 28] that mirrors third-person VideoQA settings. Recent efforts [43, 45, 46, 57] study assistive egocen- tric QA aligned with user demands. Yet none specifically target the core challenge of personalized reasoning across extended spatial and temporal contexts. Our work fills this gap by providing the first diagnostic study of personalized QA assistance requiring ego-grounding in MLLMs, by fo- cusing on âmeâ (the camera wearer) and âmyâ interacted objects among multiple people and similar distractors. Multimodal LLMs. Existing MLLMs primarily address short or third-person video understanding [20, 25, 27, 50]. More recent efforts [3, 10, 11, 25, 34, 37, 47, 53] incorpo- rate long-form and egocentric videos for first-person QA, e.g., via extended temporal context windows, hierarchical modeling, memory and retrieval-based modeling. Despite these advances, prior work emphasizes general egocentric comprehension [6, 28] rather than personalized referencing. Ego-grounding requires resolving first-person pronouns and user-specific object references (âIâ, âmy ragâ, âmy col- leagueâ) over time. Such capabilities are largely unexam- ined in current MLLMs and our work is the first systematic evaluation of these personalized reasoning abilities. Ego-cues and Personalized Understanding. Egocen- tric research has extensively explored first-person (ego) cues such as gaze prediction [19, 22, 24], hand-object in- teractions [13, 44], and action intent [39]. These cues pro- vide rich information, but prior works treat them as sig- nals for task or activity understanding [17, 33], not as per- sistent identity anchors. In addition, existing personalized vision-language systems [1, 7, 38, 48] adapt to specific users through additional visual exemplars or textual pro- files. In contrast, we aim to reinforcing MLLMs to study personalization directly from past egocentric video itself. This requires a form of visual memory and identity tracking under-explored in prior personalization research. 3. MyEgo Dataset 3.1. Task Definition We formally define egocentric personalized VideoQA as the task of answering questions that require grounding first- person references in streaming egocentric video.Such questions fall into two categories. The first involves disam- biguating âmyâ actions, attributes, or belongings from those 2 Q: What am I holding? GT: A fork. Multiple-choices options: A. A wooden spoon. B. A bottle of oil. C. A sausage. D. A pan handle. E. A fork. question moment: 05:04 answer moment: 05:04 Q: What color are my game pieces? GT: Green. Multiple-choices options: A. Green. B. Black. C. Yellow. D. Red. E. Blue. question moment: 06:59 answer moment: 06:31 05:04 06:2106:3106:59 (a) (b) Figure 2. MyEgo examples. To answer my questions, the models must understand (a) which hand is mine, (b) memorize and track the green chess pieces I used (to be distinguished with the red ones of another person, and the blue ones which are not used). of other people in the scene (see Fig. 1, and Fig. 2(a)). The second involves identifying the specific object the camera wearer interacted with among multiple visually similar in- stances of the same category (challenging distractors), such as the chess example in Fig. 2(b). 3.2. Dataset Construction Video Collection: Our videos are sourced from three pub- lic egocentric video datasets: Ego4D [15], EgoLife [46], and CASTEL2024 [36]. The footage of EgoLife and CAS- TEL2024 already involve multiple people so we addition- ally remove single-person videos from Ego4D. EgoLife dis- tributes their videos in 30-second short clips for easy down- loading. Also, each frame contains a dynamic timestamp watermark at the upper-left corner. Thus, we construct long videos by concatenating sequential clips from the same recording into approximately 10-minute segments. For wa- termark, we apply a black mask (blended into the back- ground) to remove it in each frame. The raw CASTEL2024 videos include long stretches of non-informative content, so we trim them into continuous, activity-focused clips rang- ing from 6 to 20 minutes. After processing, we obtain 541 valid videos of averaging 9.2 minutes, with 182 videos from Ego4D, 257 from EgoLife, and 102 from CASTEL2024. QA Annotation: Our preliminary trial shows that gen- erating high-quality, personalized QA pairs for our bench- mark was non-trivial. Automatic curation using MLLMs such as GPT-5/Gemini-2.5 Pro were insufficient, as the models often failed to capture the context-specific and per- sonalized nature of our queries. We therefore manually an- notated all QA pairs to ensure they accurately reflect the distinctions between actions and objects associated with the camera wearer versus other people in the videos. We recruited and trained 10 university students to manu- ally annotate the QA pairs based on the video content. An- notators were requested to adhere to the following key prin- ciples to ensure meaningful and challenging questions: ⢠Egocentric: Question must be framed from the camera wearerâs first-person perspective to simulate a direct, per- sonal inquiry. ⢠Personalized: The content must be personalized to high- light distinctions between the camera wearerâs actions or objects and those of others, compelling the model to en- gage in personalized reasoning to first determine, which objects or actions are associated with the camera-wearer, before arriving at a correct answer. ⢠Visual Answer: The answers to the questions should be concise and be visible in the videos. The corresponding ground-truth answers are also annotated by the same annotator to ensure their correctness. For each QA pair, we additionally mark the question moment when a question is posed and the answer moment when the vi- sual evidence for the answer occurs. The answer moment is always no later than the corresponding question moment. After a manual annotation and check, we obtain 5,012 first- person view, personalized QA pairs for open-ended task. To facilitate more standardized evaluation, we further augment our open-ended QA pairs for a multiple-choice (MC) setting. We feed Gemini-2.5 Pro with the video, ques- tion and ground-truth answer, and design a prompt to cre- ate 4 plausible distractors by requiring that each option de- scribes something verifiably present in the video. Specifi- cally, the prompt prioritized generating distractors that were temporally relevant (appearing at question moment or an- swer moment) or contextually confusing (e.g., an action per- formed by some others versus âmeâ). As a fallback, any incorrect event or object appearing in the video is deemed acceptable. For the âyes/noâ questions, we regard them as 2-option multiple-choice QA task. After an initial manual screening for duplicate or invalid options, we add a filtering step to limit data biases. Fol- lowing [51], we provide only the video frames and answer 3 action 30.3% object 59.4% others 10.3% (a) Question type. [0,1)[1,2)[2,4)[4,6)[6,8) [8,inf) Time (min) 0 200 400 600 800 1000 1200 Frequency (b) Question moment and answer moment. Question Moment Answer Moment [0,1)[1,2)[2,3)[3,5) [5,10) [10,20)[20,40) [40,inf) Time Delta (s) 0 200 400 600 800 1000 1200 Frequency 18.0% 11.5% 14.3% 14.5% 15.3% 10.1% 6.1% 10.2% Current 29.4% Previous 70.6% (c) Time difference. 02004006008001000 Count no yes E D C B A 504 452 802 801 812 820 821 Binary 5 options (d) Multiple choices. (e) Word cloud of questions.(f) Word cloud of answers. Figure 3. Statistic analysis of MyEgo choices as input, explicitly excluding the question, to both Gemini-2.5 Pro and GPT-5 and identify QA pairs which are solvable by both models. This step filters out simple instances where the correct answer is guessable from the video and options alone. For these rejected QA pairs, we manually refined the distractors based on the videos to en- sure a sufficient level of difficulty and specificity on ego- grounding for identifying the correct answers. After generation and pre-automatic filtering, we invite 2 students (together with authors) to carefully inspect and refine through the whole annotations. Our final dataset con- tains 5,012 questions, each available in both open-ended and multiple-choice tasks, with the latter comprising 953 and 4,059 for 2- and 5-option QA pairs respectively. 3.3. Dataset Analysis 3.3.1. Dataset Statistics Fig. 3 presents key statistics of MyEgo. It features 5,012 QA pairs distributed across 541 videos. The videos are on average 9.2 minutes long with an average of 9.3 questions on each. The queries in the questions can be categorized se- mantically as Object, Action and Others (Fig. 3(a)). Tem- porally, the question and answer moments are distributed throughout the video (Fig. 3(b)), with varying temporal sep- arations (Fig. 3(c)). Questions are considered âCurrentâ if the ground truth answer moment is either concurrent or within 2 seconds of the question. Questions with answers further back in the history is considered âPreviousâ. The average time difference between question and answer mo- ments is about 20 seconds. This average difference already presents a significant challenge, demanding models to re- call specific details and maintain a consistent understand- ing of the âmeâ concept over time. For the multiple-choice (MC) setting, the order of correct choices is randomized to ensure fairness and mitigate positional bias (see Fig. 3(d)). Fig. 3(e) and Fig. 3(f) visualize frequent words in the ques- tions and answers respectively. It shows that the ques- tions strongly feature first-person pronouns (âmyâ,âmeâ) and action keywords (âtakeâ, âputâ). Answer keywords are dominated by adjectives describing objective attributes, like colours and locations. 3.3.2. Dataset Comparison We compare MyEgo with several egocentric video QA benchmarks in Tab. 1. As highlighted, MyEgo features personalized questions that emphasizes ego-grounding and tracking the camera wearers and their interactions in com- plicated environments where multiple people and similar- looking objects present. Prior datasets, such as EgoSchema [28], EgoMemoria [47], and EgoThink [6], primarily fo- cus on general-purpose ego-vision understanding. While QAEgo4D [4] and EgoLifeQA [46] emphasize episodic memory, they do not require linking the recalled moment to help ego-ground the camera wearer in the current scene, which is underscored in MyEgo. Specialized datasets like EgoTextVQA [57] and EgoBlind [43] have concentrated on specific functionalities like understanding scene text or assisting the visually impaired in egocentric manner. MyEgo shares similar goal in egocentric assistance, but dif- fers in the emphasis on ego-grounding for personalized QA in long video stream. It specifically tests the model capacity to analyze, understand, and remember the camera wearerâs 4 Table 1. Benchmark comparison. MyEgo highlights ego-grounding and memorizing the camera wearers from ego video stream where multiple people and multiple similar objects present. MP: questions require distinguishing ego-person from other persons. MO: questions require distinguishing objected associated with ego-person from other similar-looking objects. OE/MC: Open-Ended/Multi-Choice. Benchmark#Q#VAve. LenChallengesMPMOTask QAEgo4D [4]1,8541668.2 minEpisodic MemoryâOE EgoSchema[28]5,0635,0633 minLong Ego Video UnderstandingâMC EgoThink[6]700595-General Ego Vision UnderstandingâOE EgoMemoria[47]7,0266290.5 to 60 minGeneral Ego Video UnderstandingâMC EgolifeQA [46]6,000644.3 hMultimodal Episodic MemoryâMC EgoTextVQA [57]7,0641,5071.7 minEgo Scene-Text UnderstandingâMC EgoBlind [43]5,3111,39240sBlind AssistanceâOE MyEgo (Ours)5,0125419.2 minPersonalized UnderstandingâOE/MC trajectory and intent over time, making it better reflect the demands of personalized embodied QA assistants. 4. Experiments 4.1. Experimental Setups Evaluation. For open-ended (OE) QA, we prompt GPT-5 mini [29] as an evaluator to give a binary judgment (âyesâ or ânoâ) on whether a modelâs response matches the ground truth (GT) answer, to obtain the Accuracy (0-100%, the per- centage of âyesâ answers in evaluation). Meanwhile, we obtain the match Score (0-5, 5 indicates an exact match.) that signals the match extent of two answers. After refin- ing the evaluation prompt for GPT-5 mini, we achieve an agreement rate of 94% with human judgment (see Supple- mentary), demonstrating a reasonable automatic evaluation mechanism. For more stable evaluation, we also report QA accuracy on multiple-choice (MC) setting. Specifically, we provide candidate answers for each question, and prompt the models to output the selected answer. Detailed prompts for QA and evaluation are given in the Supplementary. Models. We analyze both popular closed-source and open-source MLLMs. For closed-source ones, we evalu- ate Gemini-2.5 Pro [8] and GPT-5 [29]. For open-source models, we collect them from three groups: 1) General- purpose understanding: Qwen2.5-VL series [3], Qwen3- VL series [2], InternVL2.5-8B [5], Intern3-VL series [58], Intern3.5-VL series [41], LLaVA-OneVision [20], LLaVA- Video [55], MiniCPM-V 4.5 [49], 2) Long video under- standing: LongVA [54], LongVU [37], and 3) Streaming video understanding: Flash-VStream [52] and Dispider [35]. Additionally, to benchmark human performance, we evaluated two university students who were not involved in annotation on a random subset of 300 samples (6% of the full dataset). Related results are presented in Tab. 2 and Supplementary Tab. 6. Implementation. For each question, we uniformly sam- ple a fixed number of video frames up to the question times- tamp, with a maximum rate of 1 fps. While other models de- fault to 32 frames, LongVA and LongVU are supplied with 128 frames. Furthermore, we explicitly prompt the model to adopt a first-person perspective, focusing on the camera wearerâs (my) objects and actions. 4.2. Main Results Tab. 2 and Fig. 4 present the performances of diverse mod- els on MyEgo. We summarize the following points. General Observations. Both open- and closed-source models lag behind human performance by 33%âź55%, un- derscoring the challenging nature of MyEgo. GPT-5 gener- ally outperforms others in terms of overall accuracy, though no single model consistently excels across all categories. Among the open-source models, InternVL3 wins in MCQA while Qwen3-VL champions in OEQA. Noteworthy, all models, especially for the InternVL series of models, ex- hibit substantial performance drops when transferred from the multiple-choice setting to open-ended QA, suggesting that much of their multiple-choice success may rely on answer-choice shortcuts rather than engaging in faithful multimodal reasoning. However, MC questions, although simpler than OE ones, are still quite challenging and most models only achieve chance-level (50%) performances on binary (MC-2) questions. This is because our distractor an- swers are specially curated to be misleading if without truly grounding the questioner (camera wearer) in the video. Additionally, In the Supplementary, we analyze MLLMs of different parameter sizes and find that larger models do not consistently outperform the smaller ones. Even within the same model family, smaller variants (e.g. 4B) can achieve competitive or even superior performance. This ob- servation indicates that general-purpose model scaling can- not solve the challenge in MyEgo. Failure Instances. Unsurprisingly, questions whose an- swers are not found in the question moments are more chal- lenging (âPreviousâ vs. âCurrentâ category), suggesting that the models struggle with grounding and tracking the camera wearers and their associated things. For instance, Gemini-2.5 Pro [8] and Qwen3-VL-8B-Instruct [2] merely 5 Table 2. Evaluation results on MyEgo dataset. We present the performance on both multiple-choice (MC) and open-ended (OE) tasks. MCQA includes binary (MC-2) and five-option (MC-5) QA. For OEQA, we analyze the results across different categories and temporal locality. Cur.: question whose answer can be found in the current moment (within 2s). Pre.: question whose answer is located in the past video content. 480P: resize the height to be 480. MethodsRes. Multiple-ChoiceOpen-Ended MC-2 MC-5Cur.Pre.Avg.Act.Obj. OthersCur.Pre.Avg. (Acc/Score) Human-95.192.193.492.392.782.484.891.284.085.084.7 Closed-source Models GPT-5 [29]480P66.453.756.9 55.856.150.0 44.643.151.1 44.046.1 / 2.5 Gemini-2.5 Pro [8]720P61.845.549.048.448.640.240.247.742.440.340.9/ 2.2 Open-source Models InternVL2.5-8B [5]448 2 53.136.641.3 39.139.827.8 23.124.127.2 23.524.5 / 1.4 InternVL3.5-8B-Instruct [41]448 2 54.636.640.5 39.840.031.9 28.631.930.3 29.729.9 / 1.6 LongVA [54]336 2 56.930.940.2 34.035.833.4 31.332.434.6 31.032.1 / 1.7 InternVL3.5-8B-Thinking [41]448 2 54.537.741.8 40.540.932.5 32.233.432.5 32.432.4 / 1.8 LongVU [37]ori.53.332.838.0 36.236.733.3 31.734.633.2 32.232.5 / 1.8 LLaVA-OneVision [20]384 2 53.534.540.9 36.938.134.5 31.840.535.7 32.633.5 / 1.9 MiniCPM-V 4.5 [49]ori.53.836.141.4 38.739.536.2 32.139.035.3 33.534.0 / 1.9 InternVL3-8B [58]448 2 54.538.442.4 41.041.434.7 33.338.634.7 34.134.3 / 1.8 Qwen2.5-VL-7B [3]ori.55.132.340.8 34.936.635.4 33.734.837.4 33.034.3 / 1.8 Qwen3-VL-8B-Thinking [2]ori.56.231.538.9 35.136.236.4 33.135.136.6 33.434.3 / 1.8 LLaVA-Video [55]ori.54.836.043.0 38.239.635.2 34.040.037.4 33.935.0 / 1.9 Qwen3-VL-8B-Instruct [2]ori.55.036.641.4 39.540.138.5 35.535.737.4 36.036.4 / 2.0 describe question-irrelevant objects presented in current or past moment (see Fig. 4, left), while LLaVA-Video [55] and InternVL3.5 [41] series tend to summarize the global events in specific frames (see Fig. 4, right) without linking them to the camera wearerâs trajectory, leading to incorrect answers. Counter-intuitive Findings.Unexpectedly, âthink- ingâ models such as Qwen3-VL-8B-Thinking [2] and InternVL3.5-8B-Thinking offer no significant accuracy gains. This result conflicts with the observations in exist- ing research [23], which shows that MLLMs benefit from long-CoT reasoning on general video understanding task, highlighting the unique challenge of ego-grounding for per- sonalized understanding in MyEgo. Models using more in- put frames (e.g., LongVA [54] and LongVU [37] with 128 vs. the default 32) also show no improvement on either Current or Previous questions. We speculate that model- ing more frames will introduce increased noises that in turn will contaminate the often only partially visible ego cues that are key to identify the camera wearers. Finally, scaling up model parameters only provides a marginal yet unstable accuracy increase (see Supplementary). 4.3. Controlled Analyses Memorizing key information is essential for our task. We investigate a question moment and answer moment aware sampling strategy (Q&A strategy), where we uni- formly sample 8 frames from the Âą1.5s intervals surround- ing both the answer moment and question moment, and con- catenate them to form the final 16-frame input. If the two Table 3. Effects of answer moment as input. Uni.: uniform sam- pling. Q&A: key frames at QA moments as input. ModelsInputAcc@Cur.Acc@Pre. Qwen3-VL-8B-Think [2] Uni. Q&A 38.4 41.3â 2.9 32.0 41.1â 9.1 Qwen3-VL-8B-Instruct [2] Uni. Q&A 36.9 46.3â 9.4 35.5 41.4â 5.9 InternVL3-8B [58] Uni. Q&A 35.6 41.3â 5.7 34.8 40.4â 5.6 LLaVA-Video [55] Uni. Q&A 36.3 37.7â 1.4 33.7 41.6â 7.9 Qwen2.5-VL-7B [3] Uni. Q&A 37.7 43.4â 5.7 33.2 42.4â 9.2 Gemini-2.5 Pro [8] Uni. Q&A 42.4 49.3â 6.9 40.3 51.5â 11.2 intervals overlap, we combine them into a single interval from which 16 frames are uniformly sampled. Tab. 3 show that taking such a strategy significantly improves the mod- elsâ performance, though with fewer frames input. The improvement is more dramatic for âPreviousâ questions, where Gemini-2.5 pro [8] and Qwen2.5-VL-7B [3] gain 11.2% and 9.2% respectively. It is likely because correctly answering âPreviousâ questions requires accurately locat- ing and distinguishing key information from tens of sec- onds or even minutes ago, while uniform sampling may miss this information or introducing too much irrelevant 6 [Open-Ended] Which one is my bottle? GT: The one in your hand. Gemini-2.5 Pro: You are holding a white bottle with a white cap. [Acc: No] LLaVA-Video: The bottle is on the left side of the frame. [Acc: No] Qwen3-VL-8B-Thinking: Got it... the user is holding a bottle... But need tocheckthelast frame... </think>The bottle in your hand. [Acc: Yes] InternVL3.5-8B-Instruct: I don't have visibility of your bottle. [Acc: No] LongVU: Your bottle is located on the table. [Acc: No] Qwen3-VL-8B-Instruct: On the floor near the claw machine. [Acc: No] MiniCPM-V 4.5: The bottle in your left hand is yours. [Acc: Yes] 11.0s 33.9s 36.9s 66.6s68.8s 225.7s 406.2s [Multiple-Choice] Is that my stuff toy? Options: A. Yes. It's a prize you won from the coin pusher game. B. No. It's your friend's. (GT) C. Yes. It's a prize you won from the claw machine. D. No. It's the man in green jacket 's. E. Yes. It's with you to the arcade. LongVA, LongVU: A GPT-5: B LLaVA-Video, InternVL3.5, Qwen3-VL: C Figure 4. Result visualization. [0,1)[1,2)[2,4)[4,6)[6,8)[8,inf) Question Moment (min) 25 30 35 40 45 50 55 60 Accuracy 58.6 54.5 48.2 41.5 48.3 38.0 50.7 36.2 49.1 37.0 48.5 35.8 51.4 47.1 42.2 39.8 41.5 31.0 41.9 30.6 43.0 33.9 32.7 30.7 Models Gemini-2.5 Qwen3-VL Input Strategy Uniform Q&A moment Figure 5. Accuracy distribution along the video stream. Accuracy drops along time with uniform sampling but remains relatively sta- ble with GT moment input. background information. Interestingly, accuracy on âCur- rentâ questions also improves by 1.4% to 9.4%, which uni- form sampling should already provide sufficient informa- tion to answer. This may be because the Q&A input pro- vides a larger number of key frames compared to uniform sampling (16 versus ~1). Furthermore, information from irrelevant and redundant frames in uniform sampling can interfere with the modelâs judgment. We also analyze the distribution of accuracy along the video stream under both uniform sampling and GT-moment sampling strategies in Fig. 5.Intuitively, with a fixed number of uniformly sampled frames, models often trade off key details for long-ranged modeling, thus leading to lower accuracy as ego cues are often partially visible. Our experimental results are consistent with this intuition: model accuracy is the highest for questions asked within the first minute and declines as the video stream progressed, with the lowest performance for questions posed after 8163264 Number of frames 28 30 32 34 36 Accuracy 28.7 33.5 33.133.1 34.7 34.3 34.5 34.9 (a) Uniform sampling (OE). 18163248 Number of frames 38 40 42 44 46 48 Accuracy (41.8) (37.5) 38.5 45.3 46.1 46.2 44.6 38.1 45.2 44.9 43.7 42.5 (b) Sampling at 1fps backward from the question moment (MC). InternVL3LLaVA-Video Figure 6. Effects with different number of frame inputs. The dashed lines in (b) are the results of uniformly sampling 32 frames. 8 minutes, which highlights the inherent challenge of long-ranged video QA in MyEgo. Crucially, the Q&A strategy consistently outperforms the uniform sampling method across every time bin for both models, once again demonstrating the importance of key frames grounding. Overall, the improvements suggest that enabling the model to distinguish key, self-related information, filter out irrel- evant visual clutter, and focus on memorizing egocentric details, especially in long-form video QA scenarios, is promising for future exploration. More frames do not always bring better perfor- mance. We feed InternVL3-8B [58] and LLaVA-Video [55] with a varying number of frames, ranging from 8 to 64, as shown in Fig. 6(a). For open-ended QA, the per- formance of InternVL3-8B peaks with 16 frames and then slightly declines, while only reaches 28.7% with 8 frames. LLaVA-Videoâs accuracy remains relatively stable, staying within the range of 34-35%. These suggest that adding 7 Table 4. Effects of different prompt modification. We test sev- eral modelsâ performance with a prompt where cues of "personal- ization" is removed ("Remove") and questions where first-person pronouns are replaced with "the camera wearer" ("Enhanced"), re- spectively. We select the thinking version of InternVL3.5-8B here. ModelsInputOEMC InternVL3.5-8B [41] Original Enhanced Remove 33.1 33.2â 0.1 31.7â 1.4 39.5 41.1â 1.6 38.5â 1.0 InternVL3-8B [58] Original Enhanced Remove 35.0 34.7â 0.3 33.7â 1.3 41.8 40.4â 1.4 41.4â 0.4 LLaVA-Video [55] Original Enhanced Remove 34.5 35.9â 1.4 35.3â 0.8 37.5 38.1â 0.6 37.5â 0.0 more frames alone under uniform sampling strategy does not boost performance. Furthermore, inspired by the nature of human asking questions shortly after an event, and by the observation that in 90% of our QA pairs the time difference between question moment and answer moment is less than 40 seconds (see Fig. 3), we explored a backward sampling strategy. Specifically, for multiple-choice setting, we sam- ple 1 to 48 frames backward from the question moment at 1 fps. Fig. 6(b) shows that such a strategy surpasses the base- line accuracy achieved by uniformly sampling (indicated by the dashed lines), though with much fewer frames. How- ever, the benefits are not linear. The most significant im- provement occurs when increasing the input frames from 1 to 8, yielding accuracy improvement of 6.8% and 7.1% on InternVL3 and LLaVA-Video, respectively. Then the accu- racy enters a plateau, with only minor fluctuations or even a slight decrease, possibly due to redundant frames that do not necessarily for better understanding. These findings un- derscore that the relevance of visual information is more critical than quantity, inspiring future study into more intel- ligent moment detection. Personalization-aware prompting has a small but measurable impact. To evaluate the impact of different prompts on model performance, we introduce two system- atic modifications to the personalized prompt. Specifically, we test an âEnhancedâ prompt, where first-person pronouns (e.g., âIâ, âmyâ) in the questions are replaced with âthe camera wearer (âs)â, and a âRemoveâ prompt, which re- moves personalization cues entirely (see Fig. 7). Results are concluded in Tab. 4. The âEnhancedâ strategy gener- ally improves performance, but not always. For instance, while LLaVA-Video [55] sees a consistent improvement of over 1% in both OE and MC tasks, the performance of InternVL3.5-8B [41] remains largely unchanged. Con- versely, removing personalization predominantly degrades [Remove Personalized Prompt]: You are a first-person AI assistant ... Key Instructions: 1. First-Person Context: All references to 'I', 'me', or 'my' are about the user. You must distinguish between the user's actions/objects and those of other people visible in the video. 2. Grounding: Base your answers primarily on the visual evidence in the video... [Enhanced Question]: Q: Is the woman holding my the camera wearer's phone? A: No. 09:0609:25 09:33 Figure 7. Example prompts of (top) removing the personalized cues and (bottom) enhancing the clarification by reminding that âIâ refers to the camera wearer. performance, as seen in the notable 1.4-1.6% accuracy drop of the InternVL3.5-8B model. However, its impact can be nuanced, occasionally resulting in slight performance gains for some models. These findings suggest that the models are not drastically sensitive to these specific prompt alterations, but a clear pattern indicates that enhancing the prompt (i.e., reminding that âIâ refers to the camera wearer in the ques- tion) is generally helpful. 5. Conclusion We studied whether MLLMs can faithfully understand âmeâ, remember âmyâ past interactions, and track âmyâ context over extended spatial and temporal scales.We introduced MyEgo, the first egocentric VideoQA dataset designed for personalized user questions requiring ego- groundingâdistinguishing me and my objects from others. Using MyEgo, we conducted a comprehensive evaluation of modern MLLMs and found that all models struggle, and even scaling and âthinkingâ offer only limited gains. Mod- els perform better when key moments are provided or lie close to the question time, yet large gaps remain compared with human performance, and their accuracy drops sharply once the answer leaves the current view. This indicates that while models can partially understand me, they fail to truly remember me. These findings highlight the need for stronger long-range memory in the short term and genuinely personalized reasoning in the long term. We hope MyEgo and related analyses help drive progress toward these goals. 8 Acknowledgements This research is supported by the Ministry of Education, Singapore, under its MOE Academic Research Fund Tier 2 (MOE-T2EP20125-0037). References [1] Yuval Alaluf, Elad Richardson, Sergey Tulyakov, Kfir Aber- man, and Daniel Cohen-Or. Myvlm: Personalizing vlms for user-specific queries. In ECCV, pages 73â91. Springer, 2024. 1, 2 [2] Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhao- hai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Jun- yang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan Liu, Dunjie Lu, Ruilin Luo, Chenxu Lv, Rui Men, Lingchen Meng, Xuancheng Ren, Xingzhang Ren, Sibo Song, Yuchong Sun, Jun Tang, Jianhong Tu, Jian- qiang Wan, Peng Wang, Pengfei Wang, Qiuyue Wang, Yux- uan Wang, Tianbao Xie, Yiheng Xu, Haiyang Xu, Jin Xu, Zhibo Yang, Mingkun Yang, Jianxin Yang, An Yang, Bowen Yu, Fei Zhang, Hang Zhang, Xi Zhang, Bo Zheng, Humen Zhong, Jingren Zhou, Fan Zhou, Jing Zhou, Yuanzhi Zhu, and Ke Zhu. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. 1, 5, 6, 13 [3] Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025. 1, 2, 5, 6, 13 [4] Leonard Bärmann and Alex Waibel. Where did i leave my keys? â episodic-memory-based question answering on egocentric videos. In 2022 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition Workshops (CVPRW), pages 1559â1567, 2022. 4, 5, 12 [5] Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Science China Information Sciences, 67(12):220101, 2024. 5, 6, 13 [6] Sijie Cheng, Zhicheng Guo, Jingwen Wu, Kechen Fang, Peng Li, Huaping Liu, and Yang Liu. Egothink: Evalu- ating first-person perspective thinking capability of vision- language models. In CVPR, pages 14291â14302, 2024. 1, 2, 4, 5 [7] Niv Cohen, Rinon Gal, Eli A Meirom, Gal Chechik, and Yuval Atzmon. âthis is my unicorn, fluffyâ: Personaliz- ing frozen vision-language representations. In ECCV, pages 558â577. Springer, 2022. 1, 2 [8] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blis- tein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025. 1, 5, 6, 12, 13 [9] Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Pro- ceedings of the ECCV (ECCV), pages 720â736, 2018. 1 [10] Shangzhe Di and Weidi Xie. Grounded question-answering in long egocentric videos. In CVPR, pages 12934â12943, 2024. 2 [11] Shangzhe Di, Zhelun Yu, Guanghao Zhang, Haoyuan Li, Tao Zhong, Hao Cheng, Bolin Li, Wanggui He, Fangxun Shu, and Hao Jiang. Streaming video question-answering with in-context video kv-cache retrieval.arXiv preprint arXiv:2503.00540, 2025. 2 [12] Chenyou Fan. Egovqa-an egocentric video question answer- ing benchmark dataset. In CVPR Workshops, pages 0â0, 2019. 2 [13] Zicong Fan, Takehiko Ohkawa, Linlin Yang, Nie Lin, Zhis- han Zhou, Shihao Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Xue Zhang, et al. Benchmarks and challenges in pose estimation for egocentric hand interactions with objects. In ECCV, pages 428â448. Springer, 2024. 2 [14] Gabriele Goletto, Tushar Nagarajan, Giuseppe Averta, and Dima Damen. Amego: Active memory from long egocentric videos. In ECCV, pages 92â110. Springer, 2024. 2 [15] Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In CVPR, pages 18995â19012, 2022. 1, 2, 3, 12 [16] Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 1 [17] Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang. Egotaskqa: Understanding human tasks in egocentric videos. NeurIPS, 35:3343â3360, 2022. 2 [18] Shuhei Kurita, Naoki Katsura, and Eri Onami. Refego: Re- ferring expression comprehension dataset from first-person perception of ego4d. In CVPR, pages 15214â15224, 2023. 1 [19] Bolin Lai, Miao Liu, Fiona Ryan, and James M Rehg. In the eye of transformer: Globalâlocal correlation for egocentric gaze estimation and beyond. International Journal of Com- puter Vision, 132(3):854â871, 2024. 2 [20] Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Zi- wei Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 2, 5, 6, 13 [21] Junlong Li, Huaiyuan Xu, Sijie Cheng, Kejun Wu, Kim-Hui Yap, Lap-Pui Chau, and Yi Wang. Building egocentric pro- cedural ai assistant: Methods, benchmarks, and challenges. arXiv preprint arXiv:2511.13261, 2025. 1 [22] Lei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang, Hongkai Yu, Chen Lv, Jianru Xue, and Tat-Seng Chua. Causal-entity reflected egocentric traffic accident video syn- thesis. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 11208â11218, 2025. 2 [23] Xinhao Li, Ziang Yan, Desen Meng, Lu Dong, Xiangyu Zeng, Yinan He, Yali Wang, Yu Qiao, Yi Wang, and 9 Limin Wang.Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning.arXiv preprint arXiv:2504.06958, 2025. 1, 6 [24] Yin Li, Alireza Fathi, and James M Rehg. Learning to predict gaze in egocentric video. In Proceedings of the IEEE inter- national conference on computer vision, pages 3216â3223, 2013. 2 [25] Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual repre- sentation by alignment before projection. In EMNLP, pages 5971â5984, 2024. 2 [26] Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z Xu, Difei Gao, Rong-Cheng Tu, Wen- zhe Zhao, Weijie Kong, et al. Egocentric video-language pretraining. NeurIPS, 35:7575â7586, 2022. 2 [27] Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fa- had Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023. 2 [28] Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long- form video language understanding.In NeurIPS, pages 46212â46244, 2023. 2, 4, 5 [29] OpenAI. Gpt-5 system card. https://cdn.openai.com/gpt-5- system-card.pdf, 2025. 5, 6, 12, 13 [30] Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Sid- dhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi. An outlook into the fu- ture of egocentric vision. International Journal of Computer Vision, 132(11):4880â4936, 2024. 2 [31] Chiara Plizzari, Shubham Goel, Toby Perrett, Jacob Chalk, Angjoo Kanazawa, and Dima Damen. Spatial cognition from egocentric video: Out of sight, not out of mind. In 3DV, pages 1211â1221. IEEE, 2025. 2 [32] ShramanPramanick,YaleSong,SayanNag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang.Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In CVPR, pages 5285â5297, 2023. 2 [33] Zhaobo Qi, Shuhui Wang, Chi Su, Li Su, Qingming Huang, and Qi Tian. Self-regulated learning for egocentric video activity anticipation. IEEE transactions on pattern analysis and machine intelligence, 45(6):6715â6730, 2021. 2 [34] Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuan- grui Ding, Dahua Lin, and Jiaqi Wang. Streaming long video understanding with large language models. Advances in Neu- ral Information Processing Systems, 37:119336â119360, 2024. 2 [35] Rui Qian, Shuangrui Ding, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, and Jiaqi Wang. Dispider: Enabling video llms with active real-time interaction via dis- entangled perception, decision, and reaction. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 24045â24055, 2025. 5, 13 [36] Luca Rossetto, Werner Bailer, Duc-Tien Dang-Nguyen, Gra- ham Healy, BjĂśrn ĂĂłr JĂłnsson, Onanong Kongmeesub, Hoang-Bao Le, Stevan Rudinac, Klaus SchĂśffmann, Flo- rian Spiess, et al.The castle 2024 dataset: Advanc- ing the art of multimodal understanding.arXiv preprint arXiv:2503.17116, 2025. 1, 3, 12 [37] Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Bal- akrishnan Varadarajan, Florian Bordes, et al. Longvu: Spa- tiotemporal adaptive compression for long video-language understanding. In Forty-second International Conference on Machine Learning, 2025. 2, 5, 6, 13 [38] Yufei Shi, Weilong Yan, Gang Xu, Yumeng Li, Yucheng Chen, Zhenxi Li, Fei Yu, Ming Li, and Si Yong Yeo. Pvchat: Personalized video chat with one-shot learning. In CVPR, pages 23321â23331, 2025. 2 [39] Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li, Arjun Akula, and Angela Yao. Visual intention grounding for egocentric assistants. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2512â 2522, 2025. 1, 2 [40] Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Jun- yang Lin. Qwen2-vl: Enhancing vision-language modelâs perception of the world at any resolution, 2024. 13 [41] Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. Internvl3. 5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265, 2025. 1, 5, 6, 8, 12, 13 [42] Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bu- gra Tekin, Felipe Vieira Frujeri, et al. Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world. In CVPR, pages 20270â20281, 2023. 2 [43] Junbin Xiao, Nanxin Huang, Hao Qiu, Zhulin Tao, Xun Yang, Richang Hong, Meng Wang, and Angela Yao. Egob- lind: Towards egocentric visual assistance for the blind peo- ple. NeurIPS, 2025. 2, 4, 5 [44] Yue Xu, Yong-Lu Li, Zhemin Huang, Michael Xu Liu, Cewu Lu, Yu-Wing Tai, and Chi-Keung Tang. Egopca: A new framework for egocentric hand-object interaction under- standing. In CVPR, pages 5273â5284, 2023. 2 [45] Jiaqi Yan, Ruilong Ren, Jingren Liu, Shuning Xu, Ling Wang, Yiheng Wang, Yun Wang, Long Zhang, Xiangyu Chen, Changzhi Sun, et al.Teleego:Benchmark- ing egocentric ai assistants in the wild.arXiv preprint arXiv:2510.23981, 2025. 1, 2 [46] Jingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong, Xi- amengwei Zhang, Sicheng Zhang, Pengyun Wang, Zitang Zhou, Binzhu Xie, Ziyue Wang, et al. Egolife: Towards ego- centric life assistant. In CVPR, pages 28885â28900, 2025. 1, 2, 3, 4, 5, 12 [47] Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen, Zongyu Lin, Yanghao Li, Bowen Zhang, Haoxuan You, Dan Xu, Zhe Gan, et al.Mm-ego: Towards building egocentric multimodal llms for video qa. arXiv preprint arXiv:2410.07177, 2024. 2, 4, 5 10 [48] Chun-Hsiao Yeh, Bryan Russell, Josef Sivic, Fabian Caba Heilbron, and Simon Jenni.Meta-personalizing vision- language models to find named instances in video. In CVPR, pages 19123â19132, 2023. 2 [49] Tianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang, Wenshuo Ma, Zhihui He, Tianchi Cai, Weize Chen, Yuxiang Huang, Yuanqian Zhao, et al. Minicpm-v 4.5: Cooking effi- cient mllms via architecture, data, and training recipe. arXiv preprint arXiv:2509.18154, 2025. 1, 5, 6, 13 [50] Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video un- derstanding. arXiv preprint arXiv:2306.02858, 2023. 2 [51] Hao Zhang, Chen Li, and Basura Fernando. Mitigating easy option bias in multiple-choice question answering, 2025. 3 [52] Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, and Xiaojie Jin. Flash-vstream: Efficient real-time un- derstanding for long video streams. ICCV, 2025. 5, 13 [53] Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024. 2, 13 [54] Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024. 5, 6, 13 [55] Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Zi- wei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024. 1, 5, 6, 7, 8, 12, 13 [56] Yue Zhao, Ishan Misra, Philipp KrähenbĂźhl, and Rohit Gird- har. Learning video representations from large language models. In CVPR, pages 6586â6597, 2023. 2 [57] Sheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li, Xun Yang, Dan Guo, Meng Wang, Tat-Seng Chua, and Angela Yao. Egotextvqa: Towards egocentric scene-text aware video question answering. In CVPR, pages 3363â3373, 2025. 2, 4, 5 [58] Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shen- glong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, et al. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479, 2025. 5, 6, 7, 8, 13 11 Ego-Grounding for Personalized Question-Answering in Egocentric Videos Supplementary Material A. MyEgo Dataset In this part, we will introduce additional details about the construction of MyEgo dataset, including the video source datasets and the multiple-choice generation pipeline. Our dataset and related resources are available at https://github.com/Ryougetsu3606/MyEgo. A.1. Video Source Introduction Ego4D [15] offers 3,670 hours of daily activity video span- ning diverse scenarios: household, outdoor, workplace, and etc, which is captured by 931 unique camera wearers from 74 worldwide locations. We select 182 videos from QAEgo4D [4] subset, with each video involving two or more people engaging in the same scene or event. EgoLife [46] consists of 44-hour egocentric videos record- ing a week of shared living experience of six young vol- unteers. To better raise questions related to multi-person scenarios, we further select clips containing rich group ac- tivities, including videos from all six volunteers on the same day (Day 3), as well as footage from a single volunteer (A4) spanning four days (Day 3 to Day 6). We then concatenate the sequential short clips into 257 longer ones, with each about 10 minutes. Castle2024 [36] is a large-scale, multimodal dataset de- signed for advancing study in lifelogging, human activity analysis and multimodal retrieval. The entire set contains over 600 hours of videos from 15 time-aligned recorders, including 10 participants and 5 fixed cameras. We select a 30-hour subset of videos and segment them into clips of 6â20 minutes. We additionally filter out clips that lack multiple people or contain incomplete recording data, and eventually obtain 102 valid videos. A.2. Multiple Choice Generation Tab. 7 presents the prompt used for distractor generation. After obtaining the generated distractors, we first refine the options to produce concise expressions consistent with the style of the correct answers, preventing answer leakage due to formatting discrepancies. To further limit option bias, we manually revise QAs that both Gemini-2.5 Pro [8] and GPT-5 [29] can correctly answer even without access to the video or question. For human participation, we particularly ensure the negative options be misleading unless the models truly understand the video contents available at the question moments and correctly ground âmeâ, âmy thingsâ, and âmy pastâ in the egocentric scene. Table 5. Question categories and examples of MyEgo. CategoryNumber/Ratio Question Examples Action1,519/30.3% What am I holding? Did I give her my power bank? Where did I put my pan? Object2,975/59.4% Is this my rag? Where is my screwdriver? What color is my helmet? Others518/10.3% Is this bowl the same as before? Did I get more than 200 scores? Human Yes No GPT NoYes 35 1 5 59 Figure 8. Agreement between human and GPT evaluation on 100 instances. A.3. Question Categories For better evaluation and analysis, we categorize the ques- tions into 3 groups: 1) Action: questions focus on distin- guishing my actions from those of other people nearby, 2) Object: questions focus on distinguishing my objects from other similar objects in the scene, and 3) Others: other ques- tions that are not covered by the previous two categories. Examples of questions for each category are presented in Tab. 5. Additionally, we split the questions according to whether their answers are visible at the question moments: 1) Current: answer is visible at the question moment, and 2) Previous: answer is not visible at the question moment and demands retrieving past video contents. B. Evaluation Details Prompt Details. Tab. 8 shows the prompt we used to test the model performances. Following the official inference code, we add a thinking system prompt for InternVL3.5- Thinking series [41], and insert a time instruction right after the video frames for LLaVA-Video [55]. 12 Table 6. Performances across all selected models (results based on 30% data). Dispider performs only MC-QA. Methods Multiple-ChoiceOpen-Ended MC-2MC-5Cur.Pre.Avg.ActionObjectOthersCur.Pre.Avg. Closed-source Models GPT-5 [29]64.554.257.555.956.350.044.643.151.144.046.1 Gemini-2.5 Pro [8]60.345.149.947.548.240.240.247.742.440.340.9 General Open-source Models Parameters > 10B Qwen3-VL-32B-Instruct [2]58.335.444.138.440.041.436.437.840.037.338.1 Qwen3-VL-32B-Thinking [2]55.237.944.340.041.340.135.535.840.635.436.9 InternVL3.5-38B-Thinking [41]54.641.843.844.644.435.534.837.835.235.335.3 Qwen2.5-VL-32B-Instruct [3]49.331.442.931.935.039.733.630.536.834.535.1 InternVL3.5-38B-Instruct [41]56.339.041.942.742.535.034.137.834.334.934.7 InternVL3-38B [58] 51.341.944.143.743.835.732.539.133.334.534.1 4B <Parameters⤠10B Qwen2-VL-7B-Instruct [40]54.329.636.533.934.634.636.335.135.835.635.7 Qwen3-VL-8B-Instruct [58]55.735.544.437.739.637.035.628.536.334.935.3 InternVL3-8B [58]53.638.841.042.141.831.936.336.435.634.835.0 Qwen2.5-VL-7B-Instruct [3]53.332.543.634.036.734.635.031.137.733.234.5 LLaVA-Video [55]52.333.842.735.537.533.933.939.736.333.734.5 Qwen3-VL-8B-Thinking [2]55.331.240.834.236.034.134.131.838.432.033.9 MiniCPM-V 4.5 [49]50.335.841.237.838.735.232.237.134.333.333.6 LongVU [37]56.031.940.335.336.731.933.136.430.834.033.1 InternVL3.5-8B-Thinking [41] 52.036.342.738.239.531.133.934.433.832.833.1 LLaVA-OV-7B [20]49.732.740.834.336.131.132.438.433.832.132.6 LongVA [54]56.328.639.632.134.233.931.029.834.730.531.7 InternVL3.5-8B-Instruct [41]55.635.541.238.939.632.426.827.231.527.328.5 InternVL2.5-8B [5]50.335.341.237.238.327.521.919.227.221.823.3 Parameters⤠4B Qwen3-VL-4B-Instruct [2]55.735.043.637.439.236.135.635.837.435.135.8 Qwen3-VL-4B-Thinking [2]55.329.639.133.134.737.033.529.837.233.034.2 InternVL3.5-4B-Thinking [58]48.336.841.738.139.132.831.035.832.431.832.0 InternVL3.5-4B-Instruct [58]49.031.136.734.034.728.630.730.530.629.930.1 Qwen2.5-VL-3B-Instruct [3] 54.028.738.831.833.829.329.830.528.130.429.7 Memory-Enhanced Streaming QA Methods Flash-VStream (Qwen2-VL-7B) [52]51.331.136.334.835.231.929.935.131.530.831.0 Dispider (Qwen2-VL-7B) [35]45.531.231.933.132.7------ Automatic Evaluation. We feed the prompt in Tab. 9 to GPT-5 mini [29] to evaluate whether the responses from a tested model (e.g., Gemini 2.5 Pro) match the ground-truth answers. After two rounds of manual review and refinement of the evaluation prompt, we achieve an agreement rate of 94% with human judgment on 100 instances (Fig. 8). C. More Experiment Analyses Tab. 6 reports the performance of all selected models on MyEgo. For efficiency, we evaluate on a randomly sampled 30% subset (1,500 instances).We further in- clude two recent memory models Dispider [35] and Flash- VStream [52] that are specialized for long streaming VideoQA for comparison. By comparing among models of different parameter scales (e.g., 4Bâ 8Bâ 32Bâ 38B), we find that larger models do not consistently outperform the smaller ones. Even within the same model family, smaller variants (e.g., 4B) can achieve competitive or even superior performance, as seen in the InternVL3 series. Additionally, the two long Video-LLMs, LongVA [53] and LongVU [37], do not exhibit outstanding results despite processing substantially more frames (128 vs. 32 of other general models). Similarly, the two long streaming QA methods also fail to achieve their desired level of performance. Surprisingly, the rela- 13 [Open-Ended] What color is my bike? GT: Blue. Qwen2.5-VL-7B: Your bike is black. [Acc: No] LLaVA-Video: The bike is blue. [Acc: Yes] Qwen3-VL-8B-Thinking: Got it ... In the frames, the bike has a blue color. So the answer is blue.</think>Blue. [Acc: Yes] InternVL3.5-8B-Instruct: blue [Acc: Yes] LongVU: Your bike is black. [Acc: No] Qwen3-VL-8B-Instruct: You bike is blue. [Acc: Yes] MiniCPM-V 4.5: Blue. [Acc: Yes] 00:09 01:50 02:41 06:31 06:35 [Open-Ended] What color is my helmet? GT: Black. Qwen2.5-VL-7B: The helmet is blue. [Acc: No] LLaVA-Video: Black. [Acc: Yes] Qwen3-VL-8B-Thinking: Got it...a blue helmet attached to the bike frame. ... So the helmet is blue...</think>Blue. [Acc: No] LLaVA-OV: Red [Acc: No] LongVU: Your helmet is black. [Acc: Yes] Qwen3-VL-8B-Instruct: You helmet is black. [Acc: Yes] MiniCPM-V 4.5: Black. [Acc: Yes] [Open-Ended] Did I get more than 200 scores? GT: Yes. You have one 100, one 20 and two 60. GPT-5: No. Only three balls totaling about 110, not 200+. [Acc: No] LLaVA-Video: Yes, you got more than 200 scores. [Acc: Yes] Qwen3-VL-8B-Thinking: ... The wheel has ...the maximum is 100...</think>No, ... max 100 score, not 200+. [Acc: No] LLaVA-OV: Yes. [Acc: Yes] LongVU: You did achieve more than 200 scores. [Acc: Yes] Qwen3-VL-8B-Instruct: No, the highest score is 100. [Acc: No] MiniCPM-V 4.5: No, you did not get more than 200. [Acc: No] [Open-Ended] Who won the Dart Ball game? GT: You. Because the total score of your balls is the highest Qwen2.5-VL-7B: The person in the black shirt. [Acc: No] LLaVA-Video: The person in the white shirt. [Acc: No] Qwen3-VL-8B-Thinking: ... there's a scene with a wheel ... no specific info on Dart Ball... </think> not shown. [Acc: No] LLaVA-OV: The person in the blue hoodie [Acc: No]] LongVU: The person wearing a black shirt. [Acc: No] Qwen3-VL-8B-Instruct: Not shown. [Acc: No] MiniCPM-V 4.5: Not shown. [Acc: No] 06:02 06:47 07:02 07:09 07:34 07:44 07:22 [Open-Ended] Where is my screwdriver? GT: On the table. Qwen2.5-VL-7B : On the workbench behind you. [Acc: No] LLaVA-Video: The screwdriver is on the workbench. [Acc: No] Qwen3-VL-8B-Thinking: ... using tools....</think>Yes, you hold it in your right hand near the handlebars. [Acc: No] LLaVA-OV: The screwdriver is on the workbench. [Acc: No] LongVU: Your screwdriver can be found on the table. [Acc: Yes] Qwen3-VL-8B-Instruct: On the workbench, near the tools. [Acc: No] MiniCPM-V 4.5: On the workbench. [Acc: No] [Multiple-Choice] What am I holding? Options: A. Nothing (GT) B. Not shown C. A bicycle tire. D. A cassette. E. A wrench. Qwen3-VL-8B-Thinking, Gemini-2.5 Pro: E Qwen3-VL-8B-Instruct, LLaVA-OV, InternVL3, GPT-5: C 00:44 00:54 00:56 06:17 06:20 01:04 Figure 9. Result visualization on MyEgo. Models struggle to correctly answer even simple questions, indicating their severe deficiency in performing ego-grounding in videos to answer personalized questions from the camera wearers. Top: The camera wearer is cycling alongside another man whose bike and helmet differ in color from his. Middle: The camera wearer is playing a dart-ball game with others and achieves the highest score. Bottom: The camera wearer is repairing a bicycle with another man, repeatedly putting down and picking up his screwdriver. 14 Table 7. Prompts for Gemini-2.5 Pro to generate distracting answers in multi-choice QA. You are an expert in egocentric video analysis and a creative multiple-choice question designer. Your task is to generate a complete set of multiple-choice options for each question, which includes refining the correct answer and creating four plausible distractors. **IMPORTANT INSTRUCTIONS:** 1. Analyze the Video: Carefully analyze the provided video. The timestamps âquestion_momentâ and âan- swer_momentâ are critical. 2. Create Multiple-Choice Options: For each question, you will generate a single list of options. The **very first option** in your list MUST be the refined correct answer, followed by exactly four distractors. 3. Generate Distractors: Following the refined answer, create four distractors. Each should be formatted as â"distrac- tor (Criterion)"â and ideally meet one of the following criteria: **Criterion A (Contextual Confusion)**: The item is present and visible during the âanswer_momentâ but is not the correct answer. **Criterion B (Environmental Objects)**: The item is present in the video but is not being used or worn by the person recording. **Criterion C (Temporal Misdirection)**: The item appears *after* the âquestion_momentâ but is not the correct answer for that specific question. **Criterion D (Attentional Decoy)**: The item is present and visible during the âquestion_momentâ to distract the viewerâs attention. Additionally, the phrasing of the distractors should be as similar as possible to the correct answer. 4. Fallback Rule: If a distractor cannot meet the above criteria, it must at least be an object or action present in the video and be a plausible, yet incorrect, answer. 5. Special Case for "Yes/No": If the answer is purely "yes" or "no" (i.e., without any additional explanation), you MUST return a list with **only two options**: the correct answer formatted as "Yes (Refined Answer)" or "No (Refined Answer)", and the opposite option formatted as "No (Distractor)" or "Yes (Distractor)". Otherwise, you should return distractor answers with explanation. 6. Input Format: You will receive the list of question-answer pairs in the following JSON format: question: QUESTION answer: ANSWER question_moment: QUESTION_MOMENT answer_moment: ANSWER_MOMENT 7. REQUIRED Output Format: You MUST return your response as a single, valid JSON array (a list of lists). The first item in each inner list must be the refined answer. **Example Output:** [ "ground truth (Refined Answer)", "distractor A (Contextual Confusion)", "distractor B (Environmental Objects)", "distractor C (Temporal Misdirection)", "distractor D (Attentional Decoy)" ] Please now generate the refined answers and distractors for the provided video and QA pairs. tively old model Qwen2-VL leads the models of 7B sizes in open-ended QA, surpassing many recent systems. The above analyses collectively suggest that the explored prob- lem of ego-grounding and personalized understanding, de- spite with significant practical value for personal assistance, is largely overlooked in existing technique evolution, and highlight the importance of our benchmark and analyses to- wards advancements in these fields. By further analyzing performance across different ques- tion categories. We find that most models perform better on âActionâ (vs. âObjectâ) questions. This indicates that identifying âwhat I am doingâ is easier than disambiguat- 15 Table 8. Prompts for models to answer questions. TaskGeneral Prompt Open-endedYou are a first-person AI assistant integrated into a head-mounted camera. Your primary mission is to answer questions from the user (the camera wearer) about their own actions, objects, and environment as seen through your lens. You should answer directly to the user. The question is asked at the moment of the last frame in the video. Key Instructions: 1. First-Person Context: All references to âIâ, âmeâ, or âmyâ are about the user. You must distinguish between the userâs actions/objects and those of other people visible in the video. 2. Grounding: Base your answers primarily on the visual evidence in the video. If the information is not present or cannot be reasonably inferred from the video, state that it is not shown. 3. For any question that requires a âYesâ or âNoâ response, you MUST follow it with a brief explanation for your reasoning. 4. All responses must be direct and clear. VIDEO_CONTENT. Question: QUESTION. Multiple-choiceYou are a first-person AI assistant integrated into a head-mounted camera. Your primary mission is to answer questions from the user (the camera wearer) about their own actions, objects, and environment as seen through your lens. You should answer directly to the user. The question is asked at the moment of the last frame in the video. Key Instructions: 1. All references to âIâ, âmeâ, or âmyâ are about the user. You must distinguish between the userâs actions/objects and those of other people visible in the video. 2. Base your answers primarily on the visual evidence in the video. VIDEO_CONTENT. Question: QUESTION. Options: 2/5 OPTIONS. There is only one correct option. Please only response with the letter of the correct option. Removed personalized cuesYou are a first-person AI assistant integrated into a head-mounted camera. Your primary mission is to answer questions from the user (the camera wearer) about their own actions, objects, and environment. You should answer directly to the user. The question is asked at the moment of the last frame in the video. Key Instructions: 1. Grounding: Base your answers primarily on the visual evidence in the video. If the information is not present or cannot be reasonably inferred from the video, state that it is not shown. 2. For any question that requires a âYesâ or âNoâ response, you MUST follow it with a brief explanation for your reasoning. 3. All responses must be direct and clear. VIDEO_CONTENT. Question: QUESTION. Table 9. Prompts for GPT-5 mini to evaluate open-ended answers. You are an intelligent chatbot designed for evaluating the correctness of generative outputs for question-answer pairs. Your task is to compare the predicted answer with the correct answer and determine if they match meaningfully. Hereâs how you can accomplish the task: â â â ##INSTRUCTIONS: - Focus on meaningful matches: Assess whether the predicted answer and the correct answer have a meaningful match, not just literal word-for-word matches. - Criteria for Correctness: The predicted answer is considered correct if it reasonably matches the standard answer, recognizing that synonyms or varied expressions that convey the same or similar meaning are acceptable. - If the predicted answerâs yes/no conclusion conflicts with the correct answer, it is incorrect. - The predicted answer is considered correct if it contains the core descriptive information and does not contradict the correct answer, even if some non-critical details are missing. - Flexibility in Evaluation: Use judgment to decide if variations in the predicted answer still correctly address the question, even if they do not directly replicate the correct answerâs phrasing. Please evaluate the following question-answer pair: Question: QUESTION Correct Answer: GT-ANSWER Predicted Answer: MODEL-RESPONSE Provide your evaluation result only as a yes/no and score where the score is an integer value between 0 and 5, with 5 indicating the highest meaningful match. Please generate the response in the form of a valid JSON string with keys âpredâ and âscoreâ. For example: "pred": "yes", "score": 5, "pred": "no", "score": 1. 16 ing âmy thingâ from those of others in egocentric videos. The findings may depart from our common understanding about video object and action recognition. We visualize some examples in Fig. 9 for better understanding of results. Moreover, by comparing between âCurrentâ and âPreviousâ questions, we observe a clear performance gap of almost all models: They perform better in answering questions whose answers are visible at the question moments. This suggests that the models are capable of better ground the camera wears and their belongings in the question moments, but such strength drops sharply when the answers are out of scene. 17