Paper deep dive
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
Yunzhe Wang, Runhui Xu, Kexin Zheng, Tianyi Zhang, Jayavibhav Niranjan Kogundi, Soham Hans, Volkan Ustun
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/26/2026, 2:17:12 AM
Summary
GameplayQA is a benchmarking framework designed to evaluate the perception and reasoning capabilities of Multimodal Large Language Models (MLLMs) in 3D virtual environments. It utilizes a triadic entity decomposition (Self, Other, World) and a structured distractor taxonomy to diagnose model hallucinations in temporal, cross-video, and agent-role attribution tasks, featuring 2.4K QA pairs derived from densely annotated multiplayer gameplay videos.
Entities (4)
Relation Signals (3)
GameplayQA â usestaxonomy â Self-Other-World
confidence 98% ¡ structured around a triadic system of Self, Other Agents, and the World
GameplayQA â evaluates â Multimodal Large Language Models
confidence 95% ¡ Evaluation of frontier MLLMs reveals a substantial gap from human performance
GameplayQA â diagnoses â Hallucination
confidence 92% ¡ accompanied by a structured distractor taxonomy that enables fine-grained analysis of where models hallucinate.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct entities, and reason about concurrent multi-agent behaviors from a first-person perspective, capabilities that existing benchmarks do not adequately evaluate. We introduce GameplayQA, a framework for evaluating agentic-centric perception and reasoning through video understanding. Specifically, we densely annotate multiplayer 3D gameplay videos at 1.22 labels/second, with time-synced, concurrent captions of states, actions, and events structured around a triadic system of Self, Other Agents, and the World, a natural decomposition for multi-agent environments. From these annotations, we refined 2.4K diagnostic QA pairs organized into three levels of cognitive complexity, accompanied by a structured distractor taxonomy that enables fine-grained analysis of where models hallucinate. Evaluation of frontier MLLMs reveals a substantial gap from human performance, with common failures in temporal and cross-video grounding, agent-role attribution, and handling the decision density of the game. We hope GameplayQA stimulates future research at the intersection of embodied AI, agentic perception, and world modeling.
Tags
Links
- Source: https://arxiv.org/abs/2603.24329v1
- Canonical: https://arxiv.org/abs/2603.24329v1
Trouble viewing inline? Open PDF directly â
Full Text
87,153 characters extracted from source content.
Expand or collapse full text
GAMEPLAYQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents Yunzhe Wang, Runhui Xu, Kexin Zheng, Tianyi Zhang, Jayavibhav Niranjan Kogundi, Soham Hans, Volkan Ustun University of Southern California yunzhewa, runhuixu, kexinzhe, tzhang62, jniranja, sohamhan, ustun@usc.edu https://hats-ict.github.io/gameplayqa/ Abstract Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute actions to the correct entities, and reason about concur- rent multi-agent behaviors from a first-person perspective, capabilities that existing bench- marks do not adequately evaluate. We intro- duce GAMEPLAYQA, a framework for eval- uating agentic-centric perception and reason- ing through video understanding. Specifically, we densely annotate multiplayer 3D game- play videos at 1.22 labels/second, with time- synced, concurrent captions of states, actions, and events structured around a triadic system of Self, Other Agents, and the World, a natural decomposition for multi-agent environments. From these annotations, we refined 2.4K di- agnostic QA pairs organized into three lev- els of cognitive complexity, accompanied by a structured distractor taxonomy that enables fine-grained analysis of where models halluci- nate. Evaluation of frontier MLLMs reveals a substantial gap from human performance, with common failures in temporal and cross-video grounding, agent-role attribution, and handling the decision density of the game. We hope GAMEPLAYQA stimulates future research at the intersection of embodied AI, agentic per- ception, and world modeling. 1 Introduction Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in advanced reasoning, multimodality, and agency (Comanici et al., 2025; Achiam et al., 2023; Anthropic, 2025; Bai et al., 2025), position- ing them as promising decision-making backbones for autonomous agents in Robotics (Zitkovich et al., 2023; Gemini-Robotics-Team et al., 2025), Com- puter Use (He et al., 2024; Zhang et al., 2025) and 3D virtual agents (Raad et al., 2024; Bolton et al., GameplayQA Question Taxonomy Question FormContext TargetEntity TypesNum VideoDistractors Identification Absent Intent Count Ordering Time Localization POV Identification Single Entity Timestamp Refer Cross-Entity Refer Cross-POV Refer Self State Self Action Other State Other Action World Object World Event Lexical Scene Temporal Role Cross-Video Single Video Multi-Video Figure 1: Question taxonomy of GAMEPLAYQA. Ques- tions are organized along two axes: Entity (Self, Other, World) and Temporal Nature (Action/State for agents, Object/Event for world), yielding six primitive label types. These primitives compose into 15 task categories across three cognitive levels: single-reference percep- tion (L1), temporal reasoning (L2), and cross-video understanding (L3). See Sec. 3.1 and Table 2 for details. 2025; Yue et al., 2026). These applications require perception capabilities beyond passive scene de- scription. Drawing on perspectives from embodied cognition and multi-agent reasoning (Hernandez- Leal et al., 2019), we identify three core require- ments for agentic perception in goal-directed envi- ronments: (1) dense state-action tracking: cap- turing rapid transitions in the agentâs own states and actions; (2) other-agent modeling: reason- ing about the behaviors and intentions of other au- tonomous entities; and (3) environment ground- ing: tracking persistent and transient elements of the shared world. However, current video understanding bench- marks are ill-equipped to diagnose these agentic requirements for three primary reasons. First, the majority of existing evaluation sets suffer from a lack of embodiment and agency grounding (Ma- jumdar et al., 2024; Yang et al., 2025; Dang et al., 2025); they are often composed of slow-paced, pas- sive observations that lack the high-frequency state transitions and dense decision-making loops re- 1 arXiv:2603.24329v1 [cs.CL] 25 Mar 2026 1. Dense Timeline Captioning 3. Combinatorial QA Generation 4. Quality Assurance SA S WO 00:0000:1000:2000:4000:30 Interacting with a villagerChopping woodFollowing a villager outsideLooting a chest Location is inside a stone towerlocation is exteriorLocation is a wooden house A ladder A torch Some crops An oak tree Torch x 2A chickenA wooden chest WE "New Recipes Unlocked!" notificationVillager becomes "Apprentice" ... 2. Hallucination-Inducing Distractors Interacting with a villagerAttacking a villager Lexical Looking for loots in a stone tower Scene [00:00-00:07] Chopping Wood Temporal ... 5. Benchmark Models 6. Hallucination Analysis Human Evaluation Q: When ... Language Priors Filter Answer If always guess right, remove Q Single Correct Answer Correct Timeline Correct Label Type Lexical Scene Temporal Role Cross-POV Q: When a ladder appears, which action is the POV player performing? A. Interacting with a villager B. Attacking a villager C. Looking for loots in a stone tower D. Chopping wood Template Code: WO2SA-IDENT // Answer // Lexical // Scene // Temporal Figure 2: Overview of GAMEPLAYQA. Gameplay videos undergo (1) dense multi-track temporal captioning on 6 types of target entities (Sec. 3.2), (2) captioning includes negative labels for hallucination-inducing distractors, and (3) QA pairs are generated through a combinatorial template-based algorithm (Sec. 3.3). After (4) quality assurance (Sec. 3.4), the benchmark enables (5) model evaluation with (6) fine-grained hallucination analysis (Sec. 4.2). quired to stress-test a modelâs understanding of intentional action. Second, these benchmarks are largely not hallucination-diagnosable, providing global performance metrics while lacking the gran- ular, multi-faceted annotation needed to identify whether a failure stems from temporal misinterpre- tation, object fabrication, or role confusion (Bai et al., 2024; Seth et al., 2025; Tu et al., 2025). Fi- nally, current protocols exhibit a significant lack of multi-video understanding (Peng et al., 2025), focusing almost exclusively on single-viewpoint perception. Multi-video understanding is impor- tant in domains such as sports analytics leveraging various camera angles and autonomous driving re- quiring information fusion from multiple surround cameras. In esports and gaming, cross-POV syn- chronization and collective reasoning, skills that are fundamental to interpreting multi-agent collab- oration in interactive 3D spaces (Long et al., 2024; Savva et al., 2026), are also crucial. To bridge this gap, we introduce GAME- PLAYQA, a comprehensive benchmarking frame- work, not merely a static evaluation set, but an end- to-end pipeline encompassing structured annota- tion protocols, automated question generation, and diagnosable error analysis, designed to evaluate the cognitive foundations of agency in 3D virtual environments. We utilize 3D gameplay as a high- density âcognitive sandboxâ where states and con- sequences are deterministic and decision-making is fast-paced. We meticulously annotate synchro- nized gameplay videos from 9 multiplayer com- mercial games at a decision density ofĎâ 1.22 labels/second (Eq. 1), using a timeline-based dense captioning mechanism structured around a Selfâ OtherâWorld entity decomposition. This tripar- tite schema combined with the properties of 3D gameplay directly addresses the three core agen- tic requirements identified above: Self captures the POV agentâs own states and actions for dense state-action tracking; Other models external agentsâ behaviors and intentions; and World grounds per- ception in persistent environmental elements and transient events (Fig. 5). Leveraging these annotations, we propose a com- binatorial template-based algorithm that generates 2.4K QA pairs organized into a multi-faceted tax- onomy spanning three cognitive levels: (1) basic perception, (2) temporal reasoning, and (3) cross- video understanding. The algorithm initially pro- duces 400K candidate pairs and we downsample to 4K to enforce balanced category coverage before quality assurance yields the final set. A key innova- tion is our structured distractor taxonomy: by cate- gorizing incorrect options as lexical, temporal, or role-based confusions, we can systematically diag- nose model hallucination through multiple-choice questions. Evaluation of state-of-the-art MLLMs 2 OA-IDENT (L1: Action Recognition) Which of the following actions did the NPC horse perform in the video? A. running across the grassB. grazing in the grass C. putting down a crafting tableD. standing still in the grass SA-INTENT (L2: Intent Recognition) What is the primary reason the POV player places down the torch? A. to illuminate the surroundingsB. to mark the way back C. to clear up inventory spaceD. to prevent mobs from spawning WO-COUNT (L1: Static Object Count) How many Rowa Fruit bushes are visible in this scene? A. 2 bushesB. 3 bushesC. 4 bushes D. 5 bushes SA-ABSENT (L2: Absent Recognition) Which action did the POV player not perform during the video? A. getting into the parked white carB. running down the stairs C. deploying the parachute after a failed jump D. inspecting the container's ceiling OA-COUNT (L2: Occurrence Count) How many times did the teammate throws flashbang in total? A. 1B. 2C. 3 D. 4 OS-TIME (L2: Time Localization) When was the enemy ship shown as destroyed/exploding? A. Between 00:10 and 00:12B. Between 00:15 and 00:17 C. Between 00:04 and 00:06 TR2OA-EXIST-Temporal (L2: Timestamp Referring) Between 00:00 and 00:04, was the other player looting the downed guard's body? A. TrueB. False MIX-ORDER (L2: Single-Video Ordering) Which of the following sequences of events is correct? A. 1. The teammate runs toward the burning wreckage, 2. The POV player slides on the ground, 3. The enemy (Harvester/Machine) attacks the players, 4. The POV player shoots into the smoke B. 1. The enemy (Harvester/Machine) attacks the players, 2. The POV player slides on the ground, 3. The POV player shoots into the smoke, 4. The teammate runs toward the burning wreckage C. 1. The teammate runs toward the burning wreckage, 2. The enemy (Harvester/Machine) attacks the players, 3. The POV player slides on the ground, 4. The POV player shoots into the smoke D. 1. The enemy (Harvester/Machine) attacks the players, 2. The teammate runs toward the burning wreckage, 3. The POV player shoots into the smoke, 4. The POV player slides on the ground V4-SA2V1SA-IDENT (L3: Synchronization-Referring) While the player in Video 4 was throwing a grenade, what was the player in Video 1 doing at the same time? A. throwing a smoke grenade B. inspecting the combat knife C. running while holding a knife D. hiding inside smokes <Video 1> <Video 2> <Video 3> <Video 4> SA-POV-ID (L3: POV Identification) The POV Player in which video enters the ambulance? A. Video 1B. Video 2C. Video 3 <Video 1> <Video 2> <Video 3> SA-ORDER-MV (L3: Cross-Video Ordering) Which of the following sequences is the correct chronological order of events across Video 1 and Video 2? A. 1. The POV player in Video 1 ascends the elevator shaft using a zipline, 2. The POV player in Video 2 runs through the dark tunnel, 3. The POV player in Video 2 ascends the elevator shaft via zipline, 4. The POV player in Video 1 runs toward the outside area B. 1. The POV player in Video 1 ascends the elevator shaft using a zipline, 2. The POV player in Video 2 ascends the elevator shaft via zipline, 3. The POV player in Video 2 runs through the dark tunnel, 4. The POV player in Video 1 runs toward the outside area C. 1. The POV player in Video 2 ascends the elevator shaft via zipline, 2. The POV player in Video 2 runs through the dark tunnel, 3. The POV player in Video 1 ascends the elevator shaft using a zipline, 4. The POV player in Video 1 runs toward the outside area D. 1. The POV player in Video 2 runs through the dark tunnel, 2. The POV player in Video 1 runs toward the outside area, 3. The POV player in Video 2 ascends the elevator shaft via zipline, 4. The POV player in Video 1 ascends the elevator shaft using a zipline <Video 1> <Video 2> SA-IDENT (L1: Action Recognition) Which action did the POV player perform in the video? A. swinging a club at an enemyB. cutting trees C. running away from enemy D. building a wooden workbench WO-IDENT (L1: Object Recognition) Which of these objects appeared in the video? A. a spherical storage tankB. a flat landing pad C. a resupply stationD. an industrial complex Cross-Domain (Car Collision) While the driver ahead was turning on their right turn signal, which description best fits the POV driver's state? A. motion is acceleratingB. windshield wipers are on C. motion is slowing downD. direction is moving backward Cross-Domain (Ego Human) While the POV person in Video 2 was placing a green block onto the wall structure, what did the man in the white t-shirt do? A. read a paper manualB. placed two green blocks C. picked up a blue blockD. placed a red block <Video 1> <Video 2> <Video 3> Figure 3: Example questions from GAMEPLAYQA across different question codes and cognitive levels. Each example shows video frames paired with the corresponding QA pair, illustrating the progression from basic perception (L1) to temporal reasoning (L2) to cross-video understanding (L3). Additional cross-domain examples from car collision and egocentric human activity videos demonstrate the generalizability of the framework. reveals a performance gap against human, with models struggling when: (1) the game is fast-paced and decision-dense, (2) questions concern other agents or entities rather than the egocentric player, and (3) cross-video understanding and temporal grounding over long horizons are required. In summary, our contributions are threefold: â˘We introduce an end-to-end benchmarking framework with structured taxonomy, anno- tation schema, combinatorial QA generation, and diagnosable error analysis, enabling re- producible evaluation pipelines that can scale to new games and domains. ⢠We release a benchmark of 2.4K QA pairs from 9 multiplayer games with synchronized multi-POV videos, filling a critical gap in eval- uating the dense, multi-agent perception re- quired for embodied AI. â˘Benchmarking frontier MLLMs against hu- man evaluation reveals a performance gap, with fine-grained diagnostic analysis through structured distractors revealing that models struggle with fast-paced decision-dense sce- narios, other-agent modeling, cross-video syn- chronization grounding, and temporal reason- ing over long horizons. 2 Related Work Multimodal Large Language Models Recent progress in MLLMs has significantly expanded the ability of AI systems to perceive and reason over visual inputs (Comanici et al., 2025; Achiam et al., 2023; Anthropic, 2025; Bai et al., 2023). Many recent MLLMs have been proposed to be video- native for video understanding (Cheng et al., 2024; Comanici et al., 2025; Li et al., 2024b). These sys- tems can process extended visual streams; however, 3 BenchmarkDomainAgent-CentricMulti-POVDiagnosticAnnotation#Q#Vid MVBench (Li et al., 2024a)GeneralââAuto4,0004,000 LongVideoBench (Wu et al., 2024)GeneralââHuman6,6783,763 MVU-Eval (Peng et al., 2025)GeneralâââHuman, Auto1,8244,959 MovieQA (Tapaswi et al., 2016)MovieâHuman15k408 TVQA (Lei et al., 2018)TV ShowsâHuman152k21.8k MarioQA (Mun et al., 2017)GameplayââAuto188k13 hrs PhysGame (Cao et al., 2024)Game GlitchesââHuman880880 VideoGameQA-Bench (Taesiri et al., 2025)Game GlitchesââHuman, Auto4,7861.2k Ego4D (Grauman et al., 2022)Daily/EgoâââHuman74k3.67k hrs EgoSchema (Mangalam et al., 2023)Daily/EgoâââHuman5.1k5.1k EgoIllusion (Seth et al., 2025)Daily/EgoâââHuman, Auto8k1.4k GameplayQA (Ours)GameplayâHuman, Auto2,365100 Table 1: Comparison of relevant video understanding benchmarks. Context ScopeTaskDescriptionExample Question Codes#QDur. Single Reference (L1, 469) Action RecognitionIdentify or verify existence of self & othersâ actions SA-IDENT, OA-EXIST, ...16210.0s State RecognitionIdentify or verify existence of self & othersâ states S-IDENT, OS-EXIST, ...14710.1s Object RecognitionIdentify or verify existence of world objects in scene WO-IDENT, WO-EXIST709.3s Event RecognitionIdentify world events occurring in the environment WE-IDENT618.4s Static Object CountCount static objects present in the scene WO-COUNT2921.3s Temporal (L2, 1383) Cross-Entity ReferringLink one entity to another (X2Y reasoning) SA2S-IDENT, WO2S-EXIST, ...42323.0s Timestamp ReferringGiven time range [t1-t2], identify what entity exists TR2S-IDENT, TR2SA-IDENT, ...8124.3s Time LocalizationLocate exact timestamp when an event occurred SA-TIME, WE-TIME, ...28128.4s Absence RecognitionIdentify actions/states that did not occur over a timespan SA-ABSENT, S-ABSENT, ...19538.9s Occurrence CountCount how many times an action/event happened SA-COUNT, OA-COUNT, WE-COUNT7526.4s OrderingDetermine temporal order sequence of actions/events SA-ORDER, OA-ORDER, MIX-ORDER18032.6s Intent IdentificationIdentify underlying intent or goal behind actions SA-INTENT, OA-INTENT14823.0s Cross-Video (L3, 513) Sync-ReferringLink corresponding entities across synchronized videos V1-SA2-V2OA, V1-WO2-V2S, ...20794.7s Cross-Video OrderingDetermine event order sequence across multiple videos SA-ORDER-MV, MIX-ORDER-MV, ...117110.0s POV IdentificationIdentify who performed what action in which video SA-POV-ID, OA-POV-ID, ...18991.6s Table 2: Definition of 15 categories of cross-entity referring reasoning questions in GAMEPLAYQA, including the number of questions and average video duration. they remain prone to hallucination, including fabri- cating objects, misinterpreting temporal dynamics, and confusing causal relationships (Bai et al., 2024; He et al., 2025; Tu et al., 2025; Seth et al., 2025). Video Understanding Benchmarks Video un- derstanding benchmarks have evolved from early action recognition datasets toward evaluations emphasizing temporal reasoning, spatial ground- ing, and long-context comprehension. General video QA benchmarks such as MVBench (Li et al., 2024a), LongVideoBench (Wu et al., 2024), Video-MME (Fu et al., 2025), and MVU- Eval (Peng et al., 2025) assess multimodal mod- els on fine-grained temporal perception and multi- step inference. Domain-specific benchmarks tar- get narrative understanding in movies and TV shows (Tapaswi et al., 2016; Lei et al., 2018). Egocentric benchmarks including Ego4D (Grau- man et al., 2022), EgoSchema (Mangalam et al., 2023), ECBench (Dang et al., 2025), and EgoIllu- sion (Seth et al., 2025) evaluate first-person video understanding and hallucination detection. Em- bodied QA benchmarks such as OpenEQA (Ma- jumdar et al., 2024) and EmbodiedBench (Yang et al., 2025) ground reasoning in physical environ- ments. In the video game domain, MarioQA (Mun et al., 2017) pioneered event-centric QA on 2D platformer videos, while recent works explored the feasibility of using MLLMs to detect video game graphics glitches, including GlitchBench (Taesiri et al., 2024), VideoGameQA-Bench (Taesiri et al., 2025), and PhysGame (Cao et al., 2024). 3 The GameplayQA Framework We collected 3D multiplayer gameplay footage from 9 commercial games spanning diverse genres (see Appendix C for the full game list). Videos were sourced from YouTube, Twitch streams, and existing datasets (Wang et al., 2025). For games re- quiring synchronized multi-POV footage, we iden- tified groups of streamers who played together in the same match and downloaded their individual recordings, then manually aligned them to con- struct temporally synchronized multi-video sets. This section details how we obtain the bench- mark from these raw videos: defining a ques- tion taxonomy (Section 3.1), annotating via time- line captioning on synchronized multi-POV videos (Section 3.2), generating QA pairs through a com- 4 binatorial template-based algorithm (Section 3.3), and applying quality assurance procedures (Sec- tion 3.4). The final benchmark contains 2.4K QA pairs, generated from 2709 caption true labels and 1586 distractor labels. 3.1 Question Taxonomy Our question taxonomy (Figure 1) is built upon a six-primitive label system that categorizes observ- able events along two axes: Agent (Self, Other, World) and Temporal Nature (Action/State for agents, Object/Event for world). Entity Types. We organize perception in inter- active 3D environments around three entity cate- gories (Figure 5): Self (the POV agent), Other (external entities such as teammates, enemies, or NPCs), and World (the shared environment). This SelfâOtherâWorld decomposition naturally aligns with multi-agent reinforcement learning frame- works and agent-based modeling paradigms (Sut- ton et al., 1998; Busoniu et al., 2008), where agents must simultaneously track their own state, model other agentsâ behaviors, and respond to environ- mental dynamics (Illustration in Fig. 5). For each entity category, we distinguish between dynamic and static properties: Self-Action (SA) captures what the player does (shooting, jumping, reload- ing), while Self-State (S) captures the playerâs condition (health, ammo, equipped weapon). Sim- ilarly, Other-Action (OA) and Other-State (OS) track other agents. The World category is divided into World-Object (WO), referring to static or in- teractive items such as supply crates and vehicles, and World-Event (WE), which includes dynamic events like explosions or game notifications. This labeling system enables hallucination analysis of model error rates by entity type (see Sec. 4.2). Task Categories.We organize questions into 15 task categories across three cognitive levels; ques- tion examples, category sizes, and average video durations are summarized in Table 2. Level 1 (Sin- gle Reference) tests basic perception: recognizing actions, states, objects, and events within a single video segment. These tasks include action recogni- tion (e.g., âWhat did the player do?â), state recog- nition (e.g., âWhat was the playerâs health?â), ob- ject recognition, event recognition, and static object counting. Level 2 (Temporal) introduces tempo- ral reasoning that requires grounding answers to specific time windows. Tasks include cross-entity referring (e.g., âWhen the player jumped, what was their health?â), timestamp referring, time local- ization, absence recognition (identifying what did not occur), occurrence counting, temporal ordering, and intent identification. Level 3 (Cross-Video) extends reasoning across synchronized multi-POV footage, testing sync-referring (e.g., âWhen POV1 was reloading, what did POV2 do?â), cross-video ordering, and POV identification. This hierarchy progressively tests from basic perception to com- plex multi-perspective temporal reasoning. Figure 3 provides typical example questions covering the task categories. Distractor Taxonomy. A key contribution of GAMEPLAYQA is its structured distractor taxon- omy, which enables fine-grained diagnosis of why models hallucinate. We categorize incorrect op- tions by their relationship to the ground truth. Lexi- cal distractors are text-based variants of the correct option, generated by changing the subject, using antonyms, or altering object attributes. Scene dis- tractors are vision-based options listing plausible events that did not actually occur in the video. Tem- poral distractors refer to events that did happen, but outside the queried time window. Role distractors swap the agent attribution (e.g., attributing other agentsâ actions to the POV player). Cross-Video distractors refer to events from other synchronized videos, applicable only to multi-video questions. By analyzing the error rates for each distractor type, we can pinpoint failure modes in temporal ground- ing, agent attribution, or semantic understanding. 3.2 Multi-Video Timeline Captioning We employ dense multi-track timeline caption- ing where each of the six entity types (SA, S, OA, OS, WO, WE) is treated as an independent annotation track (See Figure 7 and Figure 8 for screenshots of labeling interface). Labels within and across tracks can overlap temporally, enabling concurrent event capture (e.g., a player action (SA) occurring while their health state (S) changes dur- ing a world event (WE)). Figure 2 visualizes this process, where the object label âa ladderâ is tem- porally referred to ask a question regarding the playerâs action at the same time. For multi-POV videos, we synchronize timelines across perspec- tives, enabling cross-video temporal alignment. Decision Density. We operationalize decision density as the temporal frequency of semantic la- bels such as actions, states, and events that consti- tute the necessary information stream for an agentâs 5 planning and reaction loop. Formally, we define the density metric Ď as: Ď = N labels T seconds (1) Across our benchmark, 2,709 true labels span a total of 2,219.41 seconds of annotated footage, yieldingĎâ 1.22labels/second. Table 10 (Ap- pendix C) shows the per-type breakdown, reflecting the predominance of self-centric observations in first-person gameplay. This high-frequency label- ing regime sets GAMEPLAYQA apart from passive video benchmarks and underscores the inherent difficulty of temporal grounding tasks in our exper- iments. The annotation process follows a two-stage human-in-the-loop workflow. In the first stage, Gemini-3-Pro generates candidate labels (3,632 pre- dictions) and distractors (1,678 predictions). Four graduate student annotators then verify and refine these candidates: 31.1% of predicted labels were deleted, 42.7% were edited (with 61.9% requir- ing caption changes and 42.2% requiring temporal boundary adjustments), and 26.2% were accepted without modification. Additionally, 7.6% of the final label set were added entirely by annotators to capture events missed by the model. In the second stage, a separate annotator reviews all labels, mak- ing further adjustments to approximately 12% of labels. Detailed annotation protocol and annotator statistics are provided in Appendix E. 3.3 Combinatorial QA Generation We generate questions through a combinatorial template-based algorithm that instantiates ques- tion templates by systematically combining veri- fied labels across five orthogonal dimensions: num- ber of videos, context target, entity type, distrac- tor type, and question form, as summarized in Ta- ble 2 and Table 7. For each combination, the al- gorithm selects a ground-truth label as the correct answer and populates the remaining options with distractors drawn from the corresponding distrac- tor pool, enabling fine-grained diagnosis of model failure modes. Complete templates are listed in Appendix F. Optionally, an LLM paraphrasing step is applied to reword the templated questions into more natural phrasing without altering their mean- ing or answer. The algorithm initially produces 399,214 candi- date QA pairs. Sync-Referring, Cross-Entity Re- ferring, Timestamp Referring, and Ordering types dominate due to their combinatorial nature, so we strategically downsample to 4K questions to en- force balanced category coverage and avoid long- tail bias. After quality assurance described in Sec- tion 3.4, this yields the final 2,365 gold-standard pairs. 3.4 Quality Assurance Language Prior Filtering.Template-based gen- eration can introduce language priors that allow models to guess answers without visual grounding. To mitigate this, we apply a blind filtering proce- dure: for each generated question, we query Gemini- 3-Flash withk = 3trials using only the question text (no video). Questions where the model con- sistently achieves high accuracy are flagged as po- tentially biased and removed from the benchmark. This ensures that remaining questions require gen- uine video understanding rather than exploiting statistical regularities in question phrasing. Human Evaluation.To validate generation qual- ity, we evenly sampled 120 questions covering all question types for human evaluation. Annotators assessed two criteria: (1) the video contains ex- actly one correct answer among the options, and (2) the question adheres to the semantics defined by its question code (e.g., an IDENT question truly requires identification). For questions where an- notators disagreed, we held discussion meetings to reach consensus; when no agreement could be reached, we resolved through majority voting. Dur- ing this process, 8% of questions were flagged as faulty due to issues such as excessive similarity between multiple options or misaligned temporal boundaries, which is consistent with the annota- tion error propagation discussed in our limitations (Section 5). 4 Experiments We evaluate both open-source and proprietary MLLMs. Open-source: Qwen3-VL Series(Bai et al., 2025), Gemma 3 Series (Team et al., 2025).Proprietary: GPT-5 Series (OpenAI, 2025), Claude 4.5 (Anthropic, 2025) (Sonnet, Haiku), Gemini Series (Comanici et al., 2025), and Seed 1.6 (Guo et al., 2025). Evaluation Setup. We evaluate all models in a zero-shot setting using accuracy as the metric. For video-native models (Gemini, Seed), we input the entire video directly. For frame-based models, we 6 ModelAll L1 (Single Ref.)L2 (Temporal)L3 (Cross-Video) ActRec StaRec ObjRec EvtRecSOCX-Ent TsRef TimLoc AbsRec OccCnt Order Intent SyncRef X-VOrd POV-ID Human80.580.080.0100.075.0100.084.2100.076.983.362.577.857.188.977.8100.0 Proprietary Models GPT-567.079.070.770.068.948.371.670.445.986.262.778.354.172.060.754.0 GPT-5 Mini62.770.467.368.672.134.567.666.747.079.033.372.850.072.043.658.7 GPT-5 Nano49.361.760.570.072.137.957.765.433.564.617.335.041.949.829.142.9 Gemini 2.5 Pro 71.369.168.070.080.334.574.577.865.182.138.782.859.581.065.085.7 Gemini 3 Flash 68.271.665.375.768.924.170.780.264.483.632.078.962.876.352.160.3 Gemini 2.5 Flash63.769.859.271.472.131.065.069.160.576.934.772.260.872.950.451.3 Claude 4.5 Sonnet 51.362.349.770.065.648.357.950.634.968.241.342.261.547.830.846.0 Claude 4.5 Haiku41.846.952.460.060.751.741.853.126.053.324.036.146.641.129.938.6 Seed 1.661.875.963.3 77.173.851.770.465.444.178.542.769.460.157.041.947.6 Seed 1.6 Flash56.566.956.172.174.650.065.567.930.968.838.441.563.161.748.244.2 Open-Source Models Qwen3 VL 235B63.871.059.970.0 80.355.268.676.554.480.050.772.863.566.731.649.2 Qwen3 VL 30B60.868.560.574.382.058.665.277.847.779.565.366.756.855.130.847.1 Qwen3 VL 8B57.868.556.574.372.1 62.163.675.346.373.352.057.264.948.327.445.5 Gemma 3 27B48.055.654.458.660.744.857.464.229.266.232.028.350.746.429.946.0 Gemma 3 12B43.753.148.365.759.031.052.554.326.754.99.327.250.050.224.839.7 Gemma 3 4B42.946.942.964.363.924.149.654.326.057.49.327.246.658.523.937.6 All Models56.964.858.469.570.443.062.566.842.772.036.555.555.859.838.849.6 Table 3: Model performance across task categories.Gold= best,Silver= second best. L1: ActRec (Action Recognition), StaRec (State Recognition), ObjRec (Object Recognition), EvtRec (Event Recognition), SOC (Static Object Count). L2: X-Ent (Cross-Entity Referring), TsRef (Timestamp Referring), TimLoc (Time Localization), AbsRec (Absence Recognition), OccCnt (Occurrence Count), Order (Ordering), Intent (Intent Identification). L3: SyncRef (Sync-Referring), X-VOrd (Cross-Video Ordering), POV-ID (POV Identification). sample frames at 1 FPS up to 32 frames; for videos longer than 32 seconds, we uniformly sample 32 frames across the duration. Videos are resized such that the longer side is 720p while preserving aspect ratio. Although models are instructed to output a single letter, they sometimes produce full sen- tences or explanations; we use GPT-5-mini as an LLM judge to extract the selected option. Detailed inference settings are in Appendix B; evaluation prompt templates in Appendix D. 4.1 Main Results Table 3 summarizes model performance across all task categories. Among all models evaluated, Gem- ini 2.5 Pro attains the highest overall accuracy (71.3%), followed by Gemini 3 Flash (68.2%) and GPT-5 (67.0%), yet a substantial gap to human per- formance (80.5%) persists. We highlight two key findings below. Consistent degradation across cognitive levels. Averaged across all models, accuracy drops steadily from L1 Single-Reference (61.2%) to L2 Temporal (56.0%) to L3 Cross-Video (49.4%). This trend validates that the three-level hierarchy of GAME- PLAYQA successfully stratifies task difficulty, with temporal grounding and multi-POV reasoning re- maining substantially more challenging than basic visual perception. Counting and Cross-Video Ordering are the hardest tasks.Two tasks emerge as clear bottle- necks. Occurrence Count (OccCnt) averages only 36.5% across models, making it the hardest L2 task. This suggests that tracking event recurrences over time, which demands sustained temporal at- tention across frames, remains beyond the reach of current models. Cross-Video Ordering (X-VOrd) averages 38.8%, the lowest among L3 tasks, with several models dropping to around 30%, indicating severe difficulty in aligning temporal events across perspectives. Together, these results suggest that precise temporal tracking, whether within a single video or across multiple perspectives, remains a fundamental weakness of current video-language architectures. 4.2 Error Source Analysis We conduct a fine-grained error analysis to iden- tify systematic failure modes by entity category. Table 4 reveals that World-Object (WO) recog- nition is the easiest category (62.0% aggregate), while recognizing Other Agents proves substan- tially harder, for example Other-Action (OA) at 54.0% and Other-State (OS) at 55.4% represent an 8-point gap compared to world objects. This sug- gests MLLMs struggle with other agent attribution in multi-agent scenes. We further plot error rates along four criteria: dis- 7 ModelAllSASSOAOSWOWE Proprietary Models Gemini 2.5 Pro 71.375.273.365.672.071.774.0 Gemini 3 Flash68.270.573.065.664.870.568.4 GPT-567.066.073.560.867.071.669.1 Gemini 2.5 Flash63.765.166.660.860.167.664.9 GPT-5 Mini62.764.468.657.660.466.863.8 Seed 1.661.860.365.857.662.367.065.6 Seed 1.6 Flash56.556.959.356.749.663.063.5 Claude 4.5 Sonnet51.351.653.050.946.957.853.6 GPT-5 Nano49.347.159.947.750.355.951.4 Claude 4.5 Haiku41.841.039.338.144.044.145.7 Open-Source Models Qwen3 VL 235B63.863.667.459.258.870.870.2 Qwen3 VL 30B60.857.662.757.160.166.267.8 Qwen3 VL 8B57.853.759.956.856.365.963.4 Gemma 3 27B48.048.453.548.347.854.649.7 Gemma 3 12B43.745.251.442.144.747.346.4 Gemma 3 4B42.943.449.643.240.352.247.7 All Models56.856.561.054.055.462.060.2 Table 4: Model performance by entity category.Gold = best,Silver= second best. SA: Self-Action, S: Self-State, OA: Other-Action, OS: Other-State, WO: World-Object, WE: World-Event. tractor type in EXIST questions (True/False), game name, video length, and number of synchronized videos. Figure 4 reveals three key insights. First, models are primarily confused by cross-video and temporal distractors, while scene distractors are the easiest, indicating that models handle static visual input better than temporal and cross-video reason- ing. Second, game pace strongly predicts difficulty: competitive shooters with rapid state transitions (Counter-Strike, Battlefield, Apex Legends) rank as the top three hardest games compared to slower exploration titles, validating that decision-dense en- vironments pose fundamentally harder challenges. Third, both temporal extent and multi-POV com- plexity compound errors, as longer clips and ad- ditional synchronized perspectives each generally degrade performance monotonically. 4.3 Language Prior and Temporal Ablation To disentangle the contributions of visual ground- ing and temporal reasoning, we conduct ablation studies on GPT-5-mini under three degraded input conditions: (1) No Video, where only the ques- tion text is provided; (2) Random Frame, where a single randomly selected frame replaces the full video; and (3) Shuffled Frames, where the original frames are presented in random order. Results are shown in Table 5. The No Video condition drops accuracy the largest amount (33.3%), confirming that GAME- PLAYQA requires genuine visual grounding and 010203040 Error Rate (%) scene role lexical temporal cross-video 6.5% 12.2% 14.0% 35.0% 39.7% Error Rate by Distractor Type 01020304050 Error Rate (%) Cyberpunk 2077 Minecraft No Man's Sky Valheim Elden Ring ARC Raiders Apex Legends Battlefield 6 Counter-Strike 2 30.5% 35.6% 38.4% 39.3% 41.7% 44.1% 44.7% 47.1% 49.7% Error Rate by Game 01020304050 Error Rate (%) 30-60s 15-30s 5-15s 0-5s 44.6% 46.4% 38.3% 35.8% Error Rate by Video Length 0204060 Error Rate (%) 5 videos 4 videos 3 videos 2 videos 62.3% 56.9% 51.2% 40.3% Error Rate by Number of Videos Figure 4: Error rate analysis across four dimensions. Top-left: Cross-video and temporal distractors cause the most errors. Top-right: Fast-paced shooters (CS2, Battlefield) are hardest. Bottom-left: Error increases with video length. Bottom-right: Error scales with number of synchronized videos. ConditionAllL1L2L3 Baseline (Full)62.767.261.960.6 No Video29.436.029.124.2 Random Frame41.752.940.933.7 Shuffled Frames54.863.152.653.4 Table 5: Ablation study on GPT-5-mini with degraded visual inputs. Performance drops indicate the contribu- tion of video content and temporal ordering. cannot be solved by language priors alone. The Random Frame condition recovers only partial per- formance (+12.3% over No Video), indicating that static visual content provides useful context but cannot substitute for temporal dynamics. Shuf- fled Frames achieves near-baseline L1 performance (63.1% vs. 67.2%) but degrades substantially on L2 (52.6% vs. 61.9%) and L3 (53.4% vs. 60.6%), showing that temporal ordering is critical for rea- soning tasks but less so for basic perception. 4.4 Cross-Domain Generalization To validate that our framework generalizes be- yond gameplay to broader single-agent and multi- agent ego-centric settings, we conducted a small- scale transfer experiment by applying the identi- cal benchmarking pipeline to two real-world do- mains: (1) dashcam collision videos from the Nexar dataset (Moura et al., 2025), and (2) syn- chronized ego-centric videos of humans collabo- ratively assembling Lego from the Ego-Humans benchmark (Khirodkar et al., 2023). The only pipeline adjustment required was renaming the de- 8 ModelAll L1 (Single Ref.)L2 (Temporal)L3 (Cross-Video) ActRec StaRec ObjRec EvtRec X-Ent TsRef TimLoc AbsRec OccCnt Order Intent SyncRef X-VOrd POV-ID Gemini 2.5 Pro66.261.186.783.3100.062.080.052.672.240.050.066.776.725.055.6 Gemini 2.5 Flash 62.072.273.366.780.064.080.036.861.140.083.355.670.050.038.9 GPT-5 Mini61.055.666.7100.080.066.070.063.277.820.066.766.753.350.027.8 Qwen3 VL 235B 59.250.0 73.3100.0100.066.080.036.872.220.050.066.750.025.044.4 Table 6: Cross-domain generalization results on real-world ego-centric videos (autonomous driving + multi-human collaboration, 213 questions).Gold= best,Silver= second best. Column headers follow the same conventions as Table 3. fault actor from âplayerâ to domain-appropriate labels such as âpersonâ or âdriverâ. Across 4 videos (âź113 seconds total), our auto- mated pipeline generated 5,463 initial questions; following the same downsampling and quality- assurance protocol as the main benchmark, we pro- duced a test set of 213 questions at a label density ofĎ = 0.50labels/second, lower than gameplay (Ď = 1.22), reflecting the slower decision pace of real-world activities. Table 6 shows that performance trends mirror those of the main benchmark: Gemini 2.5 Pro leads overall (66.2%), models degrade progres- sively from L1 to L3, and Occurrence Count and Cross-Video Ordering remain the hardest tasks. The lower label density confirms that real-world videos progress at a slower decision pace than gameplay, yet the relative difficulty ordering across models and task categories is preserved. These results demonstrate that our benchmarking frame- work generalizes to real-world spatiotemporal tasks with only minimal domain-specific adjustments. 5 Conclusion We presented GAMEPLAYQA, an end-to-end benchmarking framework that uses densely anno- tated, synchronized multi-POV gameplay videos to evaluate agentic perception in decision-dense 3D environments. Built on a SelfâOtherâWorld entity decomposition and a three-level cognitive hierarchy, the framework refines 2.4K diagnos- tic QA pairs whose structured distractors pinpoint where models hallucinate. Evaluation of 16 fron- tier MLLMs reveals steady performance degrada- tion from basic perception to temporal reasoning to cross-video understanding, with models partic- ularly failing on other-agent attribution, temporal grounding, and fast-paced decision-dense scenar- ios. Cross-domain experiments on autonomous driving and ego-centric human collaboration con- firm that the pipeline generalizes with minimal adaptation, preserving difficulty and model rank- ings across domains. We hope GAMEPLAYQA drives progress toward MLLMs capable of reliable perception and reasoning in dynamic, multi-agent worlds. Limitations While GAMEPLAYQA includes intent identifica- tion as a proxy for understanding goal-directed be- havior, our framework does not cover decision rea- soning questions such as âWhat is the best action to take at this moment?â. Answering such ques- tions would require estimating expected rewards or action values from raw video observations, which is a capability that remains an open research chal- lenge, as it demands not only perception but also learning implicit reward structures and world dy- namics from uncurated, in-the-wild footage. Addi- tionally, intent identification is inherently more sub- jective than recognizing physical actions or states, which occasionally results in multiple defensible answers; during human evaluation, approximately 8% of questions were flagged as having ambiguous ground-truth labels (Section 3.4). Nevertheless, be- cause anticipating intent is a critical capability for planning agents, we believe this category provides vital diagnostic signal and should be preserved. Annotation Cost and Error Propagation. The dense labeling process underlying GAMEPLAYQA is extremely labor-intensive and susceptible to hu- man error. Annotators must track over 100 la- bels and distractors per video, repeatedly watching the same footage while switching cognitive focus across different entity types and temporal windows. On average, labeling a 30-second video clip re- quires 25â35 minutes. More critically, the combina- torial QA generation algorithm reuses labels across multiple questions, meaning that a single labeling error, whether in timestamp boundaries, entity type, or description content, can propagate to multiple erroneous questions. However, because the QA 9 generation algorithm is deterministic and template- based, perfectly accurate annotations would yield perfectly correct questions; i.e. the source of bench- mark noise is annotation error, not algorithmic er- ror. While our quality assurance procedures mit- igate this risk, some annotation noise inevitably remains in the benchmark. Acknowledgments The authors acknowledge the use of Large Lan- guage Models for assistance with proofreading and grammar checking. All content was reviewed, edited, and approved by the human authors, who take full responsibility for the final manuscript. The project or effort depicted was or is sponsored by the U.S. Army Combat Capabilities Development Command â Soldier Centers under contract number W912CG-24-D-0001. The content of the informa- tion does not necessarily reflect the position or the policy of the Government, and no official endorse- ment should be inferred. References Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 techni- cal report. arXiv preprint arXiv:2303.08774. Anthropic. 2025. System card:claude sonnet 4.5. Tech- nical report, Anthropic. Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, and 1 others. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhi- fang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025. Qwen3-vl technical report. Preprint, arXiv:2511.21631. Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou. 2024. Hallucination of multimodal large language models: A survey. arXiv preprint arXiv:2404.18930. Adrian Bolton, Alexander Lerchner, Alexandra Cordell, Alexandre Moufarek, Andrew Bolt, Andrew Lampinen, Anna Mitenkova, Arne Olav Hallingstad, Bojan Vujatovic, Bonnie Li, and 1 others. 2025. Sima 2: A generalist embodied agent for virtual worlds. arXiv preprint arXiv:2512.04797. Lucian Busoniu, Robert Babuska, and Bart De Schut- ter. 2008. A comprehensive survey of multiagent reinforcement learning. IEEE Transactions on Sys- tems, Man, and Cybernetics, Part C (Applications and Reviews), 38(2):156â172. Meng Cao, Haoran Tang, Haoze Zhao, Hangyu Guo, Jiaheng Liu, Ge Zhang, Ruyang Liu, Qiang Sun, Ian Reid, and Xiaodan Liang. 2024. Physgame: Uncov- ering physical commonsense violations in gameplay videos. arXiv preprint arXiv:2412.01800. Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and 1 others. 2024. Videol- lama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476. Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Mar- cel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin, Boqiang Zhang, Long Li, Liuyi Wang, Qinyang Zeng, Xin Li, and Lidong Bing. 2025. Ecbench: Can multi- modal foundation models understand the egocentric world? a holistic embodied cognition benchmark. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 24593â24602. Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, and 1 oth- ers. 2025. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 24108â 24118. Gemini-Robotics-Team, Abbas Abdolmaleki, Saminda Abeyruwan, Joshua Ainslie, Jean-Baptiste Alayrac, Montserrat Gonzalez Arenas, Ashwin Balakrishna, Nathan Batchelor, Alex Bewley, Jeff Bingham, and 1 others. 2025. Gemini robotics 1.5: Pushing the frontier of generalist robots with advanced embod- ied reasoning, thinking, and motion transfer. arXiv preprint arXiv:2510.03342. Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, and 1 others. 2022. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 18995â19012. Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng, Guang Shi, Haobin Chen, Haoqi Fan, Jian Wang, Jianyu Jiang, Jiawei Wang, and 1 others. 2025. Seed1. 5-vl technical report.arXiv preprint arXiv:2505.07062. 10 Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. Webvoyager: Building an end-to- end web agent with large multimodal models. arXiv preprint arXiv:2401.13919. Yixiao He, Haifeng Sun, Pengfei Ren, Jingyu Wang, Huazheng Wang, Qi Qi, Zirui Zhuang, and Jing Wang. 2025. Evaluating and mitigating object hallu- cination in large vision-language models: Can they still see removed objects? In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Hu- man Language Technologies (Volume 1: Long Pa- pers), pages 6841â6858. Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor. 2019. A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems, 33(6):750â797. Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard Newcombe, Minh Vo, and Kris Kitani. 2023. Ego- humans: An ego-centric 3d multi-human benchmark. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 19807â19819. Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. 2018. Tvqa: Localized, compositional video ques- tion answering. In Proceedings of the 2018 con- ference on empirical methods in natural language processing, pages 1369â1379. Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao. 2024a.Mvbench:A comprehensive multi- modal video understanding benchmark. Preprint, arXiv:2311.17005. Yanwei Li, Chengyao Wang, and Jiaya Jia. 2024b. Llama-vid: An image is worth 2 tokens in large lan- guage models. In European Conference on Computer Vision, pages 323â340. Springer. Qian Long, Zhi Li, Ran Gong, Ying Nian Wu, Demetri Terzopoulos, and Xiaofeng Gao. 2024. Teamcraft: A benchmark for multi-modal multi-agent systems in minecraft. arXiv preprint arXiv:2412.05255. Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, and 1 others. 2024. Openeqa: Embodied question answering in the era of foundation mod- els. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16488â16498. Karttikeya Mangalam, Raiymbek Akshulakov, and Ji- tendra Malik. 2023. Egoschema: A diagnostic bench- mark for very long-form video language understand- ing. Advances in Neural Information Processing Systems, 36:46212â46244. Daniel Moura, Shizhan Zhu, and Orly Zvitia. 2025. Nexar dashcam collision prediction dataset and chal- lenge. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2583â2591. Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han. 2017. Marioqa: Answering questions by watching gameplay videos. In Proceedings of the IEEE International Conference on Computer Vision, pages 2867â2875. OpenAI. 2025. Gpt-5 system card. Technical report, OpenAI. Tianhao Peng, Haochen Wang, Yuanxing Zhang, Zekun Wang, Zili Wang, Ge Zhang, Jian Yang, Shihao Li, Yanghai Wang, Xintao Wang, and 1 others. 2025. Mvu-eval: Towards multi-video understand- ing evaluation for multimodal llms. arXiv preprint arXiv:2511.07250. Maria Abi Raad, Arun Ahuja, Catarina Barros, Fred- eric Besse, Andrew Bolt, Adrian Bolton, Bethanie Brownfield, Gavin Buttimore, Max Cant, Sarah Chakera, and 1 others. 2024. Scaling instructable agents across many simulated worlds. arXiv preprint arXiv:2404.10179. Georgy Savva, Oscar Michel, Daohan Lu, Suppakit Waiwitlikhit, Timothy Meehan, Dhairya Mishra, Sri- vats Poddar, Jack Lu, and Saining Xie. 2026. So- laris: Building a multiplayer video world model in minecraft. arXiv preprint arXiv:2602.22208. Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvaku- mar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, and Dinesh Manocha. 2025. Egoillusion: Benchmarking hallu- cinations in egocentric video understanding. In Pro- ceedings of the 2025 Conference on Empirical Meth- ods in Natural Language Processing, pages 28449â 28468. Richard S Sutton, Andrew G Barto, and 1 others. 1998. Reinforcement learning: An introduction, volume 1. MIT press Cambridge. Mohammad Reza Taesiri, Tianjun Feng, Cor-Paul Beze- mer, and Anh Nguyen. 2024. Glitchbench: Can large multimodal models detect video game glitches? In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 22444â 22455. Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zadtootaghaj, Nabajeet Barman, and Cor-Paul Beze- mer. 2025. Videogameqa-bench: Evaluating vision- language models for video game quality assurance. In Advances in Neural Information Processing Sys- tems (NeurIPS). Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. Movieqa: Understanding stories in movies through question-answering. In Proceedings of the 11 IEEE conference on computer vision and pattern recognition, pages 4631â4640. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre RamĂŠ, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, and 197 others. 2025. Gemma 3 technical report. Preprint, arXiv:2503.19786. Yahan Tu, Rui Hu, and Jitao Sang. 2025. Ode: Open- set evaluation of hallucinations in multimodal large language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 19836â19845. Yunzhe Wang, Soham Hans, and Volkan Ustun. 2025. X-ego: Acquiring team-level tactical situ- ational awareness via cross-egocentric contrastive video representation learning.arXiv preprint arXiv:2510.19150. Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. 2024. Longvideobench: A benchmark for long- context interleaved video-language understanding. Advances in Neural Information Processing Systems, 37:28828â28857. Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, and 1 others. 2025. Embodiedbench: Compre- hensive benchmarking multi-modal large language models for vision-driven embodied agents. arXiv preprint arXiv:2502.09560. Yuguang Yue, Irakli Salia, Samuel Hunt, Chris Green, Wenzhe Shi, and Jonathan J Hunt. 2026. Scaling behavior cloning improves causal reasoning: An open model for real-time video game playing. arXiv preprint arXiv:2601.04575. Chaoyun Zhang, Liqun Li, He Huang, Chiming Ni, Bo Qiao, Si Qin, Yu Kang, Minghua Ma, Qingwei Lin, Saravan Rajmohan, and 1 others. 2025. UFO 3 : Weaving the digital agent galaxy. arXiv preprint arXiv:2511.11332. Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, and 1 others. 2023. Rt-2: Vision-language-action models transfer web knowl- edge to robotic control. In Conference on Robot Learning, pages 2165â2183. PMLR. 12 A GameplayQA Question Taxonomy Figure 5: Illustration of the SelfâOtherâWorld frame- work. A first-person player (Self) perceives a teammate (Other) issuing a warning, set against the surrounding game environment (World). These three perspectives define the entity types used in our question taxonomy. Table 7 summarizes the GameplayQA question taxonomy. Each question is defined by five orthog- onal dimensions. For each category, we provide a short description and representative example ques- tions. Figure 5 illustrates the SelfâOtherâWorld concept. B Model Details and Inference Settings This section provides details about the inference settings and providers used to run the benchmark. Table 8 summarizes the inference configurations for all tested models. Video Input Strategy.For video-native models (Gemini-2.5-Pro, Gemini-3-Flash, Gemini-2.5-Flash), we input the entire video directly without frame sampling. For non-video-native models (GPT, Claude, Qwen, Gemma), we sample frames at 1 FPS up to a default of 32 frames; for videos longer than 32 seconds, we uniformly sample 32 frames across the duration. All videos are resized such that the longer side is 720p while preserving aspect ratio. Special Cases. Qwen 30B and 235B: Due to provider (Fireworks) limitations, we cap the to- tal number of frames at 30 for these models. For multi-video questions, the 30 frames are evenly split across all input videos (e.g., 10 frames per video for 3 synchronized videos). Seed 1.6 and Seed 1.6 Flash: Although these models support video-native input, we experienced high instability when calling the API with raw video. We therefore resort to 32-frame sampling for these models to ensure consistent evaluation. Gemini 3-Pro: We did not benchmark Gemini-3- Pro due to strict rate limits (250 API calls per day) at the time this paper was written, which made full benchmark evaluation infeasible. Reasoning Effort Settings.For models that sup- port configurable reasoning modes, we use the de- fault reasoning effort settings provided by each API. Table 9 summarizes the default reasoning modes for each model family. C Dataset Statistics C.1 Games Selection and Data Source GameplayQA includes footage from 9 commer- cially released games spanning diverse genres: â˘Single-POV Games: Minecraft, Apex Leg- ends, No Manâs Sky, Elden Ring, Cyberpunk 2077, Valheim â˘Multi-POV Synchronized Games: Counter- Strike 2 (Wang et al., 2025), Battlefield 6, Arc Raiders For multi-POV games, synchronized footage was obtained either from existing datasets or by identifying groups of Twitch streamers who played together in the same match and manually aligning their individual recordings. C.2 Label Distribution Table 10 reports the per-type breakdown of the 2,709 true labels annotated across 2,219.41 seconds of footage, corresponding to a decision density of Ďâ 1.22 labels/second. Label TypeCountShare Self-Action (SA)65824.3% Self-State (S)72926.9% Other-Action (OA)1605.9% Other-State (OS)1907.0% World-Event (WE)41715.4% World-Object (WO)55520.5% Total2,709100% Table 10: Per-type label distribution (Ďâ 1.22labels/s over 2,219.41 s of annotated footage). Figure 6 provides a broader view of the dataset regarding the distribution of question opening phrases, the question codes distribution by count and task type, and word distribution by entity type. 13 runs shoots running knife grenade times rifle shooting weapon ship smoke distant scope throws inventory cover throwing switches onto behind sprinting sprints building across bombsite ground area attacks enemies climbs aims using angle sights market attack equips along open stone menu enemy stairs target firing uses boss build path high site fence blaze cave mines side loot water check corridor pickaxe away torch cannon drone via coal fire create items fires loots moves near two chests log flask find map around ore kill search large logs axe surface ran shot line oak item clear gain house leaves dark zip small use avoid looks digs eats flies scans flank first player active weapon health ammo low count critically decreasing empty location item inventory menu interface knife located shield open full stone tool inside depleted equipped hunger rifle taking damage hammer shows magazine container oxygen completely pistol armor increasing backpack bar status gas broken shipping stamina axe load user house crafting vision bullets ability recovering peacekeeper level hearts slots thermal torch ultimate fire downed mask looting assault outside pickaxe shovel village firing obscured renegade integrity notification breath ground 20/20 bubbles venator rounds character encumbered holding diamond underground tunnel top rooftop ammunition sniper team players alive high drumsticks interior decreases shanks kv9 100 money $2300 standing balcony posture prone uncharged body sword fists/unarmed hullcracker unarmed enhanced knockdown battle walkway hemlok burst rune remaining apartments connector ready table destroyed workbench station reduced indicates ak-47 desert eagle exterior ash flashbang showing screen explosive grenade campfire kill positioned critical tutorial tactical storm catcher flesh recovers walking purple new hangar combat blue lava talon fade cold ship wet eco smg 6/54 fists bow 4/54 half 1/10 none cloaked wreckage room raw plain mp9 side alley svk kalĂŠ mid teammate runs boss shoots running shooting around stairs attacks smoke times weapon reviving corner jumps ability bombsite attack firing ultimate throwing ground building jumping heavy hill distance middle player squad slams past throws peeks slam peeking fallen wide driving back air corridor teammates site repeatedly swing bomb deploying platform standing vehicle cover knock across wreckage guard damage launches turns leaps deploys starts field container right slamming lasers seer dome attacking players horse bodies body launching area deal objective near fence still threw grass grenade melee fire ahead path squatting onto sprinting ran activating defensive projectile behind defend attempts vault swings rears kills warehouse falls force break military fired void gravity lift turret scout arrow team control exit create shield away pot spikes high ally loot serve menu wild deer fight cubes test perimeter axe kill point wall flail torch raven strike rock arc star volley arrows bed teammate alive health status shield player damage downed taking active broken vehicle enemy team players ally location knocked full shroud boss dead inside view parachuting ground thermal visibility count magma cube eliminated critically combat armor weapon disconnected harvester scope critical decreasing ammo position low enemies located stunned level visible connector tsoonami high running ahead corridor nearby injured piloting squadmate open small torch guard active/alive large firing outside greyling ash revived gold blinded idle staggered hit stairs energy alerted engaging reviving hugin tapy) doorway flash molotov self-reviving far away showed mode attack red taken request night room site truck soldier good around 80% assist reloading engine bow eyes near stealth equip lock type tank rounds grenade 233 baked item heavy alert shows 00:05 via close rune riot foot dome roof field load fire sword wooden large stone metal red chest crates building blue green tank ground yellow beacon crate box burning campfire small white ore supply smoke industrial set truck wall door stack vehicle turret drop lava table death armored sign respawn cloud pool two potatoes bushes palm station stairs military plants labeled villager fire site grace deer cows coal pieces hay bales trees barrels ladders benches container mechanical items dropped tower staircase three tree four six water orange shield painted block bed trident creeper body kit plant path oak item planet leaves growing arrow turtle mobile golden cargo bot pad rusty loot med walls sugar cane gun ring mill tall cave fences log altar lectern directional arrows iron cobblestone vase straw lever lying fuel care moon signs cart message notification times smoke objective red marker sky menu damage ground explosion wall fire blood new warning interface numbers weapon opens occurred intel flashbang molotov case burns occurs text location squads remain splatter effect feed bullet white explodes kill sparks impacts explosions sun grenade exchange timer rise distance blue screen incoming plumes died golden map added black guard tower glare drop loot stone round flies skill spirit score level occur trail bar death box shield item dust arriving rises starts press status title ring dies well met beam alert max air flare pulse yellow wipe clouds puffs resin sand base archives fall falls hay bales arc rain heavy tree lost hits cane reviving Self ActionSelf StateOther Action Other StateWorld ObjectWorld Event which while when why at how did between what can of video action object option the the did was did what the many the 00:00 action the the the the these shows did is best player pov enemy teammate pov player enemy the the the time moment times pov and did pov pov to following objects events actions the the not describes in player player's was was was player's player in pov enemy teammate pov enemy pov teammate did was the did player player 00:02 00:04 00:03 the the player's player interval SA-POV ID 189 SA-INTENT 123 WE-ORDER 63 MIX-ORDER 56 SA-ORDER 53 WO-TIME 48 WE-TIME 47 OS-TIME 47 OA-TIME 47 SA-TIME 47 S-TIME 45 WE-COUNT 44 S-IDENT 42 S-ABSENT 42 MIX-ORDER MV 41 WO-IDENT 37 WE-ABSENT 35 SA-IDENT 35 OS-IDENT 34 WO-ABSENT 33 SA-ORDER MV 32 SA-ABSENT 31 OA-IDENT 30 WO-COUNT 29 WE-ORDER MV 28 OS-ABSENT 28 WE-IDENT 27 OA-ABSENT 26 OA-INTENT 25 SA-COUNT 22 WO-EXIST True 16 OA-ORDER MV 16 WE-EXIST True 15 OS-EXIST True 15 OA-EXIST True 14 SA-EXIST True 14 OA-EXIST Temporal 13 S-EXIST True 12 SA-EXIST Role 11 OS-EXIST Temporal 10 OA-EXIST Scene 10 R2S-ABSEN 9 WE-EXIST Temporal 9 OA-COUNT 9 OA2S-IDENT 9 SA-EXIST Temporal 8 SA-EXIST Scene 8 OA-ORDER 8 WE2WO-IDEN 8 E2WO-ABSEN 8 WE-EXIST Lexical 8 TR2S-IDENT 8 E2S-ABSEN 8 A2S-ABSEN 8 OA2SA-IDENT 7 OS-EXIST Role 7 WO-EXIST Temporal 7 O2OS-ABSEN 7 TR2OS-IDENT 7 R2OS-ABSEN 7 TR2WO-IDENT 7 OS-EXIST Lexical 7 WO-EXIST Lexical 7 V2-SA2V3 SA-IDENT 7 O2S-ABSEN 7 OA2WE-IDENT 7 WE2S-IDENT 7 WE2SA-IDENT 7 WO2OA-IDEN 7 SA2WO-IDENT 7 OA-EXIST Role 7 R2WE-IDENT S2OA-IDENT WO2WE-IDEN S2SA-ABSEN OA-EXIST Lexical S2S-ABSEN SA-EXIST Lexical A2WE-ABSEN A2S-ABSEN S-EXIST Role O2WE-ABSEN S2WO-IDENT A2WO-IDEN S2OS-IDENT R2WE-ABSEN A2WO-ABSEN O2SA-ABSEN WO2S-IDENT WO2SA-IDENT WO2OS-IDENT V2-SA2V1 SA-IDENT V2-S2V3-SA EXIST-True V1-SA2V4 SA-IDENT A2SA-ABSEN S2WE-IDENT S2SA-IDENT S2WO-ABSEN R2WO-ABSEN V1-SA2V2 SA-IDENT V1-SA2V3 SA-IDENT S2OS-ABSEN R2SA-ABSEN A2WE-ABSEN S-EXIST Temporal OS2SA-IDENT S2WO-ABSEN S2WO-IDENT OS2S-IDENT A2WE-IDENT E2OS-ABSEN E2SA-ABSEN WE2OA-IDENT OS2WE-IDENT A2OS-ABSEN O2OA-ABSEN A2OS-IDENT OS-EXIST Scene V2-SA2V1 S-IDENT V1-S2V3-SA EXIST-True V1-S2V2-S EXIST-True V2-SA2V3 OA-IDENT V4-SA2V3 SA-IDENT R2OA-IDENT S2SA-EXIST True A2OS-ABSEN OA2S-EXIST True A2OA-IDENT S2SA-ABSEN R2SA-IDENT S-EXIST Lexical OS2SA-EXIST Lexical WE2OS-IDENT S2WE-ABSEN L1 â Single Reference Action Recognition Event Recognition Object Recognition State Recognition Static Object Count L2 â Temporal / Relational Absence Recognition Cross-Entity Referring Intent Identification Occurrence Count Ordering Time Localization Timestamp Referring L3 â Cross-Video Cross-Video Ordering POV Identification Sync-Referring Figure 6: Top left: Distribution of the first four words of the questions. The arc length indicates the frequency. Top right: Distribution of question codes, categorized by count and task type (using color for cognitive level). Bottom: Word cloud visualizing question terms by entity type. 14 D Prompt Templates Single-Video Question Answering Watch the video carefully and answer the following multiple choice question: <frame_1> <frame_2> ... <frame_32> Q: <question> <options> Please select the correct answer from the options. Answer with the letter directly. Your answer: Multi-Video Question Answering Watch the video carefully and answer the following multiple choice question: The following are 32 frames of the Video 1: <frame_1> <frame_2> ... <frame_32> The following are 32 frames of the Video 2: <frame_1> <frame_2> ... <frame_32> ... Q: <question> <options> Please select the correct answer from the options. Answer with the letter directly. Your answer: LLM as a Judge (Extract Selected Option) You judge which option a model selected for a multiple choice question. The question was: <question> Available options are: <options> The modelâs response was: <model_output> Your task is to determine which option the model selected. Look for: â˘Explicit mention of a letter (e.g., "A", "B", "C", "D") â˘The model stating or implying a spe- cific choice â˘The response content matching one of the available options If the model clearly selected one of the op- tions, return the corresponding letter. If the modelâs response is empty, an er- ror, unclear, or does not make a definitive choice, return "X". Language Prior Filtering You are answering multiple choice ques- tions about video game footage. You have NOT seen the video. Based only on the question and options provided, select the most likely answer. You must respond with ONLY a single letter (A, B, C, or D). E Annotation Protocol E.1 Annotator Demographics and Expertise Our annotation team consisted of 5 graduate stu- dents (ages 21â31; 3 male, 2 female), assigned to 4 labeler and 2 evaluator roles, with one participant serving in both capacities to ensure cross-stage consistency. Annotators were graduate student co- authors recruited internally for this study and were not financially compensated. All annotators were informed of the purpose of the data collection and consented to their annotations being used for re- search and potential public release. General Gaming Experience. To confirm that annotators were qualified to interpret complex game states, we surveyed their overall gaming habits and game-specific familiarity before the an- notation process. ⢠Video game play frequency in the past 100 days: 60% (3 participants) play regularly (3â5 times/week); 40% (2 participants) play occa- sionally (1â2 times/week). â˘Years of video game experience: 60% (3 par- ticipants) have 8+ years; 20% (1 participant) 3â8 years; 20% (1 participant) 1â3 years. Game-Specific Familiarity. Table 11 reports each annotatorâs self-reported familiarity with the benchmark titles. GameExpertRegular CasualLowNone (>300h)(>30h)(>5h)(>0h) Counter-Strike11210 Minecraft11300 Apex Legends02030 Battlefield 602030 Arc Raiders01040 Cyberpunk 207701130 Elden Ring01121 Valheim00131 Table 11: Annotator familiarity with each game. Values indicate the number of annotators (out of 5) at each familiarity level. Hours denote minimum play time. 15 E.2 Annotation Interface We developed a custom web-based annotation tool supporting both single-video and multi-video la- beling. The annotation starts with Gemini-3-Pro- generated timeline captions on the video about the six target entity types (SA, S, OA, OS, WO, WE) and distractor candidates for lexical semantic dis- tractors and scene distractors. Human annotators then verify each individual caption and distractor regarding type, content, and timeline. The annota- tion tool is publicly available: â˘SourceCode:https://github.com/ wangyz1999/sync-video-label â˘Demo Video:https://w.youtube.com/ watch?v=PKedELJ4XT0 â˘Live Demo:https://sync-video-label. vercel.app/ Figure 7 shows an example of the interface in the single-video setting in the game Valheim, where the player is combating a Boar. Figure 8 shows an example of the interface in the multi-video setting in the game Arc Raiders, where an explosion in the sky is synchronously captured by all three videos. E.3 Annotation Instructions The annotation process follows a structured work- flow designed to ensure temporal precision and semantic accuracy. Annotators begin by importing synchronized video instances (single or multi-POV) into the annotation interface. The process proceeds through three sequential phases: label generation, verification, and question preview. Label Generation and Verification. Annota- tors initiate automated caption generation using Gemini-3-Pro, which produces candidate true labels, lexical distractors, and scene distractors for each video segment. For each generated true label, an- notators verify three criteria: (1) the described event actually occurred, (2) the temporal bound- aries[t start ,t end ]accurately delimit the event, and (3) the assigned entity type (SA, S, OA, OS, WO, WE) correctly categorizes the label. Tem- poral boundaries are adjusted when necessary to ensure precise alignment with observable events. For lexical distractors, annotators confirm that the described events do not occur during the specified video segment. For scene distractors, verification ensures non-occurrence across the entire video du- ration. Incorrectly generated labels are not immedi- ately discarded; instead, we first consider whether they can be repurposed as distractors. Overall, 31.1% of Gemini-3-Pro-predicted labels were deleted as incorrect or irreparable, while 42.7% were edited to meet quality standards. Of these edited labels, 61.9% required caption text corrections (e.g., fixing entity names, action de- scriptions, or semantic precision) and 42.2% re- quired temporal boundary adjustments to better align[t start ,t end ]with the observable event. The remaining 26.2% of predicted labels were accepted without modification. Additionally, 7.6% of the final label set were added entirely by annotators to capture events missed by the model. In the second stage, a separate evaluator reviewed all labels and made further adjustments to approximately 12% of labels, primarily to enforce cross-video consistency and resolve edge cases. Count Labeling. Count questions require spe- cial handling: for action and event types (SA, OA, WE), annotators mark multiple temporally distinct segments sharing the same caption name. For ob- ject types (WO), annotators specify the quantity directly in the label metadata. Quality Control. Annotators systematically re- view labels by filtering by entity type, verifying each category independently. Ambiguous labels, those where truth value cannot be reliably deter- mined, are removed to maintain benchmark in- tegrity. The interface supports efficient navigation through keyboard shortcuts and contextual menus, enabling rapid verification of temporal alignment and semantic correctness. Once all labels and dis- tractors are verified, annotators proceed to the ques- tion preview phase, where generated questions are validated against the verified label set. F Question Templates This section provides the complete list of ques- tion templates for all three levels used for ques- tion generation. Each template is identified by a code combining the entity type and question form. Placeholders:otherrefers to other play- ers (teammate, enemy, NPC);captionrefers to specific action/state/object/event descriptions; refCaptionrefers to the anchor entity descrip- tion;timestamprefers to a formatted time range (e.g., [00:01 to 00:12]). 16 Figure 7: Screenshot of the annotation interface in the single-video setting, shown for the game Valheim. Annotators label actions, states, events, and entities on a synchronized timeline. Figure 8: Screenshot of the annotation interface in the multi-video setting, shown for the game Arc Raiders. Three synchronized video perspectives are displayed in parallel, with aligned timelines enabling annotators to capture cross-video events and temporal relations. The same explosion in the sky is captured by all three videos. 17 DimensionCategoryDescriptionExample Question Number of Videos Single VideoInputs a single videoWhat action did the player perform? Multi-VideoInputs multiple videos When POV1 player was reloading, which ac- tion did POV2 player perform? Context Target Summative(Single-Video Only) Requires aggrega- tion over a temporal segment. Which of the following best summarizes the playerâs actions during the video? Timestamp Referring Refers to a specific moment or interval At [02:45 - 02:52], what action did the player perform? Target Entity Referring Refers to the temporal moment of the 6 entity types. When the player was reloading, what action did the teammate perform? Cross-Video Referring (Multi-Video Only) Explicitly references a video index. Which POV player reload their weapon? Entity Type Self-ActionAction performed by the POV player e.g., shooting, reloading, using item, etc Which action did the POV player performed? Self-StateState or status of the POV player e.g., health, inventory, equipped weapon, etc What was the playerâs health status at that time? Other-ActionAction performed by another player e.g., teammate, enemy, NPC, etc What action did the enemy perform during the fight? Other-StateState of another playerWas the teammate downed during the en- counter? World-ObjectEnvironment objects or landmarks e.g., tree, supply crate, building, cars, etc How many supply crates are visible in the scene? World-EventEnvironmental or system-level events or notifications e.g., explosion, enemy downed, achievement notification, etc Did an explosion occur during this segment? Distractor Type LexicalTextually similar but incorrect descrip- tions Did the player reload instead of switching weapons? ScenePlausible but nonexistent eventsDid a vehicle explode in this area? TemporalReal events outside the question contextDid the explosion occur before the firefight? RoleCorrect event but wrong agentDid the enemy reload their weapon? IntentAlternative plausible motivations Did the player reload to prepare for a long fight? Question Form IdentificationSelect the correct answer from optionsWhich of the following actions occurred? ExistenceBinary true or false questionDid an explosion occur in this clip? AbsentIdentify what did not occurWhich event did not happen during healing? IntentAsk why an action was performedWhy did the player reload their weapon? CountAsk for quantities or frequenciesHow many enemies appeared in the scene? Table 7: GameplayQA question taxonomy. Each question is defined by a combination of dimensions, enabling systematic coverage of perception, temporal reasoning, and cross-video reasoning. 18 Model NameVersionInference ProviderPlatform GPT-5gpt-5OpenAIOpenAI GPT-5-minigpt-5-miniOpenAIOpenAI GPT-5-nanogpt-5-nanoOpenAIOpenAI Gemini-2.5-Progemini-2.5-proGoogleGoogle AI Studio Gemini-3-Flashgemini-3-flashGoogleGoogle AI Studio Gemini-2.5-Flashgemini-2.5-flashGoogleGoogle AI Studio Claude-4.5-Sonnetclaude-4.5-sonnetAmazon BedrockOpenRouter Claude-4.5-Haikuclaude-4.5-haikuAmazon BedrockOpenRouter Seed-1.6bytedance-seed/seed-1.6SeedOpenRouter Seed-1.6-Flashbytedance-seed/seed-1.6-flashSeedOpenRouter Qwen-3-VL-235Bqwen/qwen3-vl-235b-a22b-instructFireworksOpenRouter Qwen-3-VL-30Bqwen/qwen3-vl-30b-a3b-instructFireworksOpenRouter Qwen-3-VL-8Bqwen/qwen3-vl-8b-instructAlibabaOpenRouter Gemma-3-27Bgoogle/gemma-3-27b-itChutesOpenRouter Gemma-3-12Bgoogle/gemma-3-12b-itChutesOpenRouter Gemma-3-4Bgoogle/gemma-3-4b-itChutesOpenRouter Table 8: Inference configurations for tested models. All models are evaluated using the listed inference providers. Model NameDefault Reasoning Mode GPT-5Balanced (medium) GPT-5-miniBalanced (medium) GPT-5-nanoBalanced (medium) Gemini-2.5-ProHigh (dynamic) Gemini-3-FlashHigh (dynamic) Gemini-2.5-FlashHigh (dynamic) Claude-4.5-SonnetStandard (extended thinking off) Claude-4.5-HaikuStandard (extended thinking off) Table 9: Default reasoning effort for tested models. All models are evaluated using their default reasoning settings. 19 FormEntityCodeTemplate IDENT SA SA-IDENTWhich of the following actions did the POV player perform during the video? S S-IDENTWhich of the following best describes the POV playerâs state in the video? OA OA-IDENTWhich of the following actions did other perform during the video? OS OS-IDENTWhich of the following best describes otherâs state in the video? WO WO-IDENTWhich of the following objects appeared in the video? WE WE-IDENTWhich of the following event occurred in the video? EXIST SA SA-EXISTDid the POV player perform the action: "caption"? S S-EXISTCan you describe the POV playerâs state as: "caption"? OA OA-EXISTDid the other perform the action: "caption"? OS OS-EXISTCan you describe the otherâs state as: "caption"? WO WO-EXISTDid the object "caption" appear in the video? WE WE-EXISTDid the event "caption" occur in the video? ABSENT SA SA-ABSENTWhich action did the POV player NOT perform? S S-ABSENTWhich of the following states does not describe the POV playerâs state? OA OA-ABSENTWhich action did the other NOT perform? OS OS-ABSENTWhich of the following does not describe the otherâs state? WO WO-ABSENTWhich objects is NOT present in the scene? WE WE-ABSENTWhich of the following events did NOT occur in the video? COUNT SA SA-COUNTHow many times did the POV player perform the action: "caption"? OA OA-COUNTHow many times did the other perform the action: "caption"? WO WO-COUNTHow many caption are there in the scene? WE WE-COUNTHow many times did the event "caption" occur in the video? INTENT SA SA-INTENTWhy did the POV player perform the action: "caption"? OA OA-INTENTWhy did the other perform the action: "caption"? Table 12: Level 1 (perception) question templates. Entity types: SA (Self-Action), S (Self-State), OA (Other- Action), OS (Other-State), WO (World-Object), WE (World-Event). 20 FormRefâ AnsCodeTemplate IDENT SAâ S SA2S-IDENTWhen the POV player was performing the action: ârefCaptionâ, which of the following best describes their state? SAâ OA SA2OA-IDENTWhen the POV player was performing the action: ârefCaptionâ, which of the following actions did other perform? Sâ SA S2SA-IDENTWhen the POV playerâs ârefCaptionâ, which of the following actions did they perform? OAâ S OA2S-IDENTWhenotherwas performing the action: ârefCaptionâ, which of the following best describes the POV playerâs state? WOâ SA WO2SA-IDENT At the moment when the object ârefCaptionâ appeared, which of the following actions did the POV player perform? WEâ OA WE2OA-IDENT At the moment when the event ârefCaptionâ occurred, which of the following actions did other perform? EXIST SAâ S SA2S-EXISTWhen the POV player was performing the action: ârefCaptionâ, can you describe their state as: âcaptionâ? SAâ OA SA2OA-EXISTWhen the POV player was performing the action: ârefCaptionâ, did other perform the action: âcaptionâ? Sâ SA S2SA-EXISTWhen the POV playerâs ârefCaptionâ, did they perform the action: âcaptionâ? OAâ S OA2S-EXISTWhenotherwas performing the action: ârefCaptionâ, can you describe the POV playerâs state as: âcaptionâ? WOâ SA WO2SA-EXISTAt the moment when the object ârefCaptionâ appeared, did the POV player perform the action: âcaptionâ? WEâ OA WE2OA-EXISTAt the moment when the event ârefCaptionâ occurred, did other perform the action: âcaptionâ? ABSENT SAâ S SA2S-ABSENTWhen the POV player was performing the action: ârefCaptionâ, which of the following does NOT describe their state? SAâ OA SA2OA-ABSENT When the POV player was performing the action: ârefCaptionâ, which action did other NOT perform? Sâ SA S2SA-ABSENTWhen the POV playerâs ârefCaptionâ, which action did they NOT perform? OAâ S OA2S-ABSENTWhenotherwas performing the action: ârefCaptionâ, which of the following does NOT describe the POV playerâs state? WOâ SA WO2SA-ABSENTAt the moment when the object ârefCaptionâ appeared, which action did the POV player NOT perform? WEâ OA WE2OA-ABSENT At the moment when the event ârefCaptionâ occurred, which action did other NOT perform? Table 13: Level 2 entity-reference question templates (representative examples; full set covers all 30 RefâAns pairs across SA, S, OA, OS, WO, WE). Code format:Ref2Ans-Form. Each EXIST question is further instantiated with distractor subtypes: True, Lexical, Scene, Temporal, and Role. 21 FormAnsCodeTemplate IDENT SA TR2SA-IDENTDuringtimestamp, which of the following actions did the POV player perform? S TR2S-IDENT Duringtimestamp, which of the following best describes the POV playerâs state? OA TR2OA-IDENTDuringtimestamp, which of the following actions didotherper- form? OS TR2OS-IDENT Duringtimestamp, which of the following best describesotherâs state? WO TR2WO-IDENTDuring timestamp, which of the following objects appeared? WE TR2WE-IDENTDuring timestamp, which of the following events occurred? EXIST SA TR2SA-EXISTDuringtimestamp, did the POV player perform the action: âcaptionâ? S TR2S-EXIST Duringtimestamp, can you describe the POV playerâs state as: âcaptionâ? OA TR2OA-EXISTDuring timestamp, did other perform the action: âcaptionâ? OS TR2OS-EXIST Duringtimestamp, can you describeotherâs state as: âcaptionâ? WO TR2WO-EXISTDuring timestamp, did the object âcaptionâ appear? WE TR2WE-EXISTDuring timestamp, did the event âcaptionâ occur? ABSENT SA TR2SA-ABSENTDuring timestamp, which action did the POV player NOT perform? S TR2S-ABSENT Duringtimestamp, which of the following does NOT describe the POV playerâs state? OA TR2OA-ABSENTDuring timestamp, which action did other NOT perform? OS TR2OS-ABSENT Duringtimestamp, which of the following does NOT describe otherâs state? WO TR2WO-ABSENTDuring timestamp, which object did NOT appear? WE TR2WE-ABSENTDuring timestamp, which of the following events did NOT occur? Table 14: Level 2 timestamp-reference (TR) question templates. The placeholdertimestampis replaced by a formatted time range such as [00:01 to 00:12]. Code format: TR2Ans-Form. 22 TypeRefâ AnsCodeTemplate Cross-Video Reference (V1-Ref2V2-Ans-Form) IDENT SAâ SA V1-SA2V2-SA-IDENTWhenPOV1playerwasperformingtheaction: ârefCaptionâ, which of the following actions did POV2 player perform at the same time? SAâ S V1-SA2V2-S-IDENTWhenPOV1playerwasperformingtheaction: ârefCaptionâ, which of the following best describes POV2 playerâs state at the same time? OAâ SA V1-OA2V2-SA-IDENTWhenrefOtherwasperformingtheaction: ârefCaptionâ in POV1, which of the following ac- tions did POV2 player perform at the same time? WEâ WO V1-WE2V2-WO-IDENTAt the moment when the event ârefCaptionâ occurred in POV1, which of the following objects appeared in POV2 at the same time? EXIST SAâ SA V1-SA2V2-SA-EXISTWhenPOV1playerwasperformingtheaction: ârefCaptionâ, did POV2 player perform the action: âcaptionâ at the same time? SAâ OA V1-SA2V2-OA-EXISTWhenPOV1playerwasperformingtheaction: ârefCaptionâ,didotherperformtheaction: âcaptionâ in POV2 at the same time? OSâ WE V1-OS2V2-WE-EXISTWhenrefOtherâs ârefCaptionâ in POV1, did the event âcaptionâ occur in POV2 at the same time? WOâ S V1-WO2V2-S-EXISTWhen the object ârefCaptionâ appeared in POV1, was POV2 playerâs âcaptionâ at the same time? POV Identity (POV-ID): Which video corresponds to the player who did X? POV-ID SA SA-POV-ID Which video corresponds to the player who performed the action: âcaptionâ? S S-POV-IDWhich video corresponds to the player whose âcaptionâ? OA OA-POV-IDWhich video showsotherperforming the action: âcaptionâ? OS OS-POV-IDWhich video shows other whose âcaptionâ? WO WO-POV-IDWhich video shows the object âcaptionâ? WE WE-POV-IDWhich video shows the event âcaptionâ? Temporal Ordering (ORDER): Which happened first? ORDERSA SA-ORDERWhich of the following actions happened first? Table 15: Level 3 (cross-video) question templates. Cross-video reference templates cover all 36 RefâAns pairs (6 reference typesĂ6 answer types) in both IDENT and EXIST forms; representative examples are shown.V1/V2 are replaced by actual video indices at generation time.refOtherrefers to the other player in the reference video. POV-ID answer options are video numbers; ORDER options are formatted as âThe POV player in VideoXis [action]â. 23