Paper deep dive
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/14/2026, 6:23:02 AM
Summary
The paper introduces NARU, a benchmark for evaluating Multimodal Large Language Models (MLLMs) on narrative evolution and cultural nuance understanding in Japanese extreme long-form video. NARU comprises 1,481 questions derived from 155 videos totaling 146.8 hours, covering four narrative dimensions (character evolution, sequential flow, plot progression, thematic development) and five cultural dimensions (aizuchi, reading the air, subtext, cultural context, sentiment analysis). The authors propose a hierarchical memory-based annotation pipeline to construct the dataset and evaluate eight model configurations, revealing significant limitations in long-range narrative integration and culturally grounded reasoning.
Entities (8)
Relation Signals (7)
NARU â contains â 1,481 questions
confidence 98% ¡ NARU consists of 1,481 questions grounded in 155 videos
NARU â evaluates â MLLMs
confidence 95% ¡ NARU, a multimodal benchmark designed to evaluate MLLMs on extreme long-form video understanding
NARU â covers â Cultural Understanding
confidence 92% ¡ (2) Native Cultural Understanding, which assesses comprehension of Japanese conversational nuances
NARU â covers â Narrative Intelligence
confidence 92% ¡ NARU focuses on two core competencies... (1) Narrative Intelligence
NARU â uses â Hierarchical Memory-Based Annotation Pipeline
confidence 90% ¡ we propose a hierarchical memory-based annotation pipeline... To construct the benchmark at this scale
Kuuki wo Yomu â ispartof â Cultural Understanding
confidence 85% ¡ C.2: Kuuki wo Yomu(Shared Situational Understanding)... This category evaluates whether a model can construct a situation-level account
LongVideoBench â comparedto â NARU
confidence 80% ¡ Existing benchmarks have advanced the evaluation... LongVideoBench... evaluate long-range retrieval... NARU connects these research directions
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.
Tags
Links
- Source: https://arxiv.org/abs/2608.13210v1
- Canonical: https://arxiv.org/abs/2608.13210v1
Trouble viewing inline? Open PDF directly â
Full Text
59,855 characters extracted from source content.
Expand or collapse full text
JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20261 NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video Yuheng Huang*, Jianlang Chen*, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao and Lei Ma,Member, IEEE AbstractâLong-form video understanding encom- passes tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpret- ing social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabili- ties jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totalling 146.8 hours, span- ning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these per- sistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video. Index Termsâbenchmark, cultural understanding, long-form video understanding, multimodal large lan- guage models, video question answering I. Introduction Recent advances in Multimodal Large Language Models (MLLMs) have substantially improved their abilities in video understanding, supporting accurate captioning and visual question answering for short video clips and tem- poral events [1], [2], [3]. These capabilities have further enabled a growing range of new downstream applications, Yuheng Huang and Jianlang Chen contributed equally to this work. Yuheng Huang, Hua Qi and Lei Ma are with The University of Tokyo, Tokyo, Japan (e-mail: yuhenghuang42@g.ecc.u-tokyo.ac.jp; qi-hua@g.ecc.u-tokyo.ac.jp; ma.lei@acm.org). Jianlang Chen is with Kyushu University, Fukuoka, Japan (e-mail: chen.jianlang.396@s.kyushu-u.ac.jp). Jiayang Song is with Macau University of Science and Technology, Macau, China (e-mail: jiayang.song@ieee.org). Aza Kai, Vincent Markert, and Edison Marrese-Taylor are with Infinimind Japan Inc., Japan (e-mail: kai@infinimind.io; vin- cent@infinimind.io; edison@infinimind.io). Edison Marrese-Taylor is also with The University of Tokyo, Japan. Jianjun Zhao is with Kyushu University, Fukuoka, Japan (e-mail: zhao@ait.kyushu-u.ac.jp) Lei Ma is also with University of Alberta, Edmonton, Canada. such as short-form content retrieval, clip-level summariza- tion, and open-domain video commentary generation [4], [5]. However, in addition to short video clips, many real- world scenarios involve long-form video content [6], such as full-length films and television episodes, livestream archives, and documentaries, where users may seek a holistic understanding across the entire video rather than isolated facts from individual fragments. Tasks such as long-form video summarization [7], content analysis [8], [9], culturally aware content moderation [10], [11], and me- dia archiving [12] require models to integrate information across extended time spans and reason over evolving social and narrative contexts. Crucially, a videoâs length alone does not dictate the need for global context. For example, hours-long contin- uous surveillance footage may support tasks that can be solved through localized detection or retrieval. In contrast, films, television programs, documentaries, and other so- cially situated media often contain events whose signifi- cance emerges only through their relationships to earlier events, evolving participants, and accumulated social con- text. We refer to such content ascontext-rich long-form video. Understanding such content imposes two coupled demands. First, models need to preserve narrative co- herence by tracking entities, events, causal dependencies, and character development across the video. Second, they are required to interpret brief but consequential moments whose significance depends on accumulated social context rather than literal audiovisual content alone. Existing benchmarks have advanced the evaluation of these capabilities from various aspects. LongVideoBench [13] and LVBench [14] evaluate long-range retrieval, temporal grounding, and general reasoning over extended videos, while recent narrative- oriented benchmarks such as VRBench [15] and StoryVideoQA [16] target multi-step reasoning and deep storyline comprehension. Multilingual benchmarks such as ViMUL-Bench [17] further broaden video evaluation across languages and cultural categories. Nevertheless, these directions remain largely separate: existing benchmarks do not jointly examine sustained narrative comprehension and culturally situated implicit reasoning in context-rich long-form video. Japanese long-form media is both a practically conse- quential domain and a scientifically informative setting arXiv:2608.13210v1 [cs.CV] 13 Aug 2026 JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20262 A new dream Challenges and Growth Success and New Start 0:12:46 ... ... 1:04:321:56:07 1:10:34 ... ... 1:43:261:45:16 The Guest enjoys and would like to stay longer The host offers âanother cup of coffeeâ In Kyoto, this actually means âyou need to leaveâ (a) Narrative Intelligence: The model is asked to summarize a characterâs storyline over a long video. A new dream Challenges and Growth Success and New Start 0:12:46 ... ... 1:04:321:56:07 1:10:34 ... ... 1:43:261:45:16 The Guest enjoys and would like to stay longer The host offers âanother cup of coffeeâ In Kyoto, this actually means âyou need to leaveâ (b) Cultural Understanding: The model is asked to infer the hostâs implicit intent when offering another cup of coffee to the guest. 1 Fig. 1: Example scenarios fromNARU. for context-rich video understanding. Japan supports a substantial video-content ecosystem, with its domestic market projected to reach approximately 630 billion yen in fiscal year 2025 [18]. This scale motivates evaluating whether MLLMs can understand Japanese media. More importantly, Japanese communication is characterized by a high-context cultural style, in which meaning is often conveyed implicitly through ambient atmosphere (kuuki wo yomu,, or âreading the airâ), conversa- tional backchannels (aizuchi,), and culturally shared expectations. Understanding these interactions requires models to integrate diffuse cues across extended temporal spans and to maintain a coherent interpretation of evolv- ing interpersonal and narrative dynamics.NARUthere- fore uses Japanese long-form media as a focused setting for evaluating context-rich video understanding, rather than treating temporal duration, narrative reasoning, and cultural knowledge as independent capabilities. To address these gaps, we introduceNARU(NARrative and Cultural Understanding), a multimodal benchmark designed to evaluate MLLMs on extreme long-form video understanding within a high-context cultural framework. NARUfocuses on two core competencies, as demonstrated in Fig. 1: (1) Narrative Intelligence, which measures a modelâs ability to track story evolution across long videos with a duration of 30 to 240 minutes, and (2) Native Cultural Understanding, which assesses comprehension of Japanese conversational nuances, social dynamics, and 1 In certain situations in Japan, such offers are conventionally used as an indirect and polite way to signal the end of a visit. Understand- ing the hostâs true intent requires awareness of the context and the surrounding social atmosphere, beyond the literal semantic content. implicit communicative signals. Constructing a benchmark of this scale exposes fun- damental limitations in traditional dataset creation paradigms. Existing high-quality benchmarks rely heavily on manual annotation pipelines, which are effective for short clips but become unscalable and cognitively unreli- able when applied to extremely long videos in culturally specific domains. To address these challenges, we introduce a hierarchical annotation-to-QA synthesis pipeline for long-form video. The pipeline decomposes videos into temporal segments, maintains cross-segment narrative continuity, and sup- ports multi-level annotations ranging from fine-grained events to high-level narrative and social analysis. This approach enables scalable, context-aware ground truth construction that is infeasible with conventional manual annotation alone. Using this pipeline, we build a high- quality long-form video understanding benchmark, fol- lowed by expert verification to ensure annotation relia- bility and consistency. In summary, this paper makes the following contributions: â˘Problem Formulation: We distinguish context-rich long-form video from video that is merely long in du- ration and formulate its understanding as the joint problem of maintaining narrative state and interpreting culturally situated implicit meaning. â˘Benchmark Development: We introduceNARU, a large-scale benchmark comprising 155 Japanese long- form videos totaling 146.8 hours and 1,481 multiple- choice questions, designed to systematically evaluate Narrative Intelligence and Native Cultural Understand- ing in extreme long-form settings. The final items are verified by a total of 68 native Japanese annotators to ensure video grounding and cultural fidelity. â˘Technical Contribution: We introduce a hierarchi- cal annotation-to-QA synthesis pipeline that constructs temporally coherent annotations for hours-long videos and uses them to generate context-dependent question answer pairs. â˘Systematic Evaluation: We benchmark SOTA MLLMs on NARU , establishing baselines for context- rich Japanese long-form video understanding across narrative and culturally implicit reasoning dimensions. The benchmark data, evaluation resources, and addi- tional details are available on our project website [19]. I. Related Work Video question-answering benchmarks initially focused on short clips, evaluating local event recognition, action understanding, and temporal reasoning [20], [21], [22], [23]. Recent benchmarks extend evaluation to movies, television programs, documentaries, and other long-form videos [8], [24], [25], [26], [13], [14], [27]. These efforts substantially broaden temporal coverage, evaluating capabilities such as evidence retrieval, temporal grounding, event understand- ing, summarization, and multi-timescale integration. How- ever, extended temporal coverage alone does not capture JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20263 whether a model can maintain an evolving interpretation of narrative and social context. To bridge this gap, story-focused benchmarks examine complementary aspects of narrative comprehension. Dra- maQA introduced hierarchical, character-centered ques- tion answering, while SCVBench evaluates story-centric temporal reasoning [28], [29]. More recent benchmarks further address multi-step reasoning, entity persistence, distributed evidence, and long-range storyline comprehen- sion [15], [30], [31], [16]. While these efforts successfully move beyond localized event recognition, they primarily formulate narrative understanding through structural plot elements, entity continuity, temporal relations, and ex- plicit causal dependencies. In parallel, social and cultural reasoning present an- other crucial layer of video understanding. Prior work has explored this through multimodal social reasoning, geographically diverse visual knowledge, and multilingual video evaluation [32], [33], [34], [17]. Although these bench- marks provide vital coverage of social interaction and cultural knowledge, they rarely examine how culturally implicit meanings emerge and evolve across an extended narrative arc.NARUconnects these research directions throughcontext-rich long-form video understanding. Its narrative dimension requires models to track characters, events, causal relations, and plot development across ex- tended temporal horizons, while its cultural dimension requires them to infer socially implicit meanings from evolving interpersonal context. I. NARU Benchmark Construction While Japanese long videos can serve as an ideal testbed for evaluating MLLMsâ abilities, we face a dual dilemma in constructing a high-quality benchmark from them. On one hand, relying solely on manual annotations is impractically costly and unscalable. Namely, annotating extremely long videos requires native Japanese experts to watch hours of content, track complex narrative arcs, and identify subtle cultural cues. This process can incur high cognitive load and time costs. Additionally, while auto- mated MLLM pipelines offer scalability, they are currently hindered by the context-window bottleneck. Even SOTA models struggle to process hours of continuous video without losing fidelity, often hallucinating details or failing to maintain narrative consistency over long durations [14]. To bridge the gap, we propose a hierarchical memory- based annotation pipeline. The system performs short- term chunk processing with long-term narrative history to analyze and annotate multimodal information across extended timelines. This pipeline provides experts with abundant contextual information, substantially reducing routine annotation effort and allowing them to focus on identifying and resolving more complex, high-impact cases. We describe the construction process ofNARU through Sec. I-A to Sec. I-D. A. Capability Taxonomy To evaluate long-form video understanding beyond mere information retrieval, we organizeNARUaround two complementary capabilities. First,narrative intelli- gencemeasures a modelâs ability to track and synthesize evolving events, entities, and concepts across the extended runtime. Second,cultural understandingevaluates how well a model interprets implicit, socially situated meanings that transcend literal audiovisual signals. Rather than relying on a single existing framework, we bridge theories of event cognition and high-context communication to construct subcategories with distinct inference targets. Narrative Intelligence (N) To ground our evaluation of narrative structure, we draw on Event Segmentation Theory (EST) [35]. EST posits that humans perceive continuous experience as a structured sequence of discrete events, maintaining inter- nal event models that are updated at situational bound- aries. Therefore, we distinguish betweenlocal coherence, which captures consistency across adjacent or tightly linked events, andglobal coherence, which requires integrating evidence distributed throughout the video into a unified interpretation. We operationalize these two levels of narrative structure through four subcategories: â˘N.1: Character/Entity Evolution(Local Coherence): The ability to track a recurring character or entity across temporally separated segments and infer how its state evolves over time. Relevant states include goals, beliefs, emotions, roles, relationships, and circumstances. This category evaluates whether a model can maintain a consistent representation of an entity while continuously updating it in response to new events, rather than treating each appearance independently. â˘N.2: Sequential/Topical Flow(Local Coherence): The ability to reconstruct how a video progresses from one event, scene, or topic to the next. This includes recognizing temporal order, identifying transitions and topic shifts, and relating each segment to its immediate context. Different from N.1, which centers on entity evolution, N.2 evaluates the continuity and organization of the narrative itself. â˘N.3: Plot/Conflict Progression(Global Coherence): The ability to trace the evolution of goals, obstacles, con- flicts, and consequential decisions across multiple events. This includes identifying initiating conditions, major turning points, causal consequences, and eventual reso- lutions or shifts in stakes. Rather than simply recovering temporal order, this category evaluates whether a model can integrate dispersed events into a coherent account of how the central narrative unfolds. â˘N.4: Idea/Thematic Development(Global Coherence): The ability to synthesize the higher-level ideas, argu- ments, values, or themes developed throughout a video and explain how they are introduced, supported, refined, or reinterpreted over time. Unlike N.3, which focuses on narrative progression, this category captures forms of global coherence organized around conceptual or JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20264 Video CollectionMLLM-centric AnnotationMQA Generation API Program Video Screening (10,000+ videos) Japanese Themed Long Duration Video Quality Filtering (8,018 videos) â˘Integrity â˘Uniqueness â˘Continuity â˘Interactivity 155 videos (~147 hours) Narrative Intelligence (N.1â N.4) Culture Understanding (C.1â C.5) Step 1-1: Chunking Step 1-2: Semantic Segmentation Step 1-3: Task-Oriented Annotation â˘VisualSampling â˘Chunk Merge Chapter-level Segments â˘Narrative Module: Entities/Events/... â˘Culture Module: Cues/ Subtext/... Taxonomy-aligned Evidence Step 2-1: Capability-Conditioned Formulation â˘Video â˘Segments â˘Annotation MLLM Question Stem MCQ (1 Key + 3 Distractors) A B C D Step 2-2: Iterative Debiasing Refinement â˘Blind Solver Agent â˘Diagnostic Agent â˘Revision Agent Textual Shortcuts Elimination Quality Control: Verification by 68 Native Experts Final NARU Benchmark (1,481 QA Items) Fig. 2: Workflow overview ofNARU. Video collection and filtering (Sec. I-B) select 155 long videos from the candidate set. A MLLM-centric pipeline (Sec. I-C1) is used to produce taxonomy-aligned evidence for narrative intelligence and cultural understanding annotation. Multiple-choice questions (Sec. I-C2) are subsequently generated and refined through a multi-agent pipeline. Verification (Sec. I-D) by 68 native Japanese experts yields the final 1,481 QA items. thematic development, particularly in documentaries, interviews, and discussion-oriented content. Together, N.1 and N.2 assess whether a model can maintain a stable account ofwho or what is involvedand how the content proceeds. N.3 and N.4 test whether it can inferwhy the events matterwithin the narrative andwhat broader meaning they jointly express. Cultural Understanding (C) For cultural understanding, we focus on interactions in which meaning depends on shared social assumptions and contextual cues beyond explicit verbal content [36]. We organize the taxonomy according to an evidence-to- interpretation process. C.1 evaluates the interpretation of an observable conversational signal. C.2, C.3, and C.5 evaluate latent social meaning at three distinct levels: the shared situation, an individual speakerâs intent, and the participantsâ affective states, respectively. C.4 focuses on the culturally specific background knowledge that can support interpretation at any of these levels. â˘C.1: Aizuchi(Interactional Signalling):Aizuchiare short listener utterances in Japanese conversation that sig- nal attention, understanding, or engagement without necessarily indicating agreement (roughly like the En- glish expressionsuh-huhandyep) [37]. This cate- gory assesses a modelâs ability to interpret Japanese backchannels using their lexical, prosodic, temporal, and interactional context. The same backchannel can indicate agreement, surprise, emotional alignment, or an attempt to manage turn-taking [38]. We aim to under- stand whether a model can recover the conversational function of such phatic expressions, rather than equating its surface form with a fixed meaning. â˘C.2: Kuuki wo Yomu(Shared Situational Understand- ing):Kuuki wo Yomu, literallyread the air,refers to inferring the unspoken social atmosphere and interac- tional norms from contextual and interpersonal cues, including group consensus, interpersonal tension, and expectations about appropriate behavior, from partici- pantsâ actions. This category evaluates whether a model can construct a situation-level account of what the group collectively recognizes and how that understanding af- fects participantsâ behavior accordingly. â˘C.3: Subtext Interpretation(Speaker Intent): The ability to infer what a particular speaker intends to communi- cate beyond, or in contrast to, the literal meaning of an utterance. This includes indirect requests, euphemism, irony, socially restrained expression, and distinctions betweentatemae(public-facing expression) andhonne (private intent). Different from C.2, which concerns the shared atmosphere or norms of the overall situation, C.3 targets the latent communicative intention behind a specific speakerâs expression. â˘C.4: Cultural Context Recognition(Cultural Ground- ing): The ability to recognize culturally specific objects, practices, social conventions, and historical references and to explain their significance within the current scene. This category evaluates the background knowl- edge needed to understand the specific meaning behind an observed action or reference, rather than merely checking if the model can visually identify it. â˘C.5: Sentiment Analysis(Affective and Interpersonal Dynamics): The ability of leveraging evidence to infer participantsâ affective states and interpersonal attitudes, and to track how they evolve throughout an interaction. Evidence may include speech content, facial expressions, and body language. This category encompasses social norms, such as concealed discomfort, restrained frustra- tion, or emerging aďŹinity, rather than only explicit emo- tional expressions or coarse positivenegative sentiment. Different from C.2 and C.3, this category evaluates how participants feel and relate to one another. B. Dataset Construction Acquisition and Pre-filtering.To build a broad candidate pool, we searched across 16 YouTube-defined upload categories, continuing retrieval until reaching over 100,000 unique videos after deduplication. We then applied an ASR-based spoken-language identification pipeline, filtering for Japanese-language content, which yielded 51,643 videos. Finally, following the long-video JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20265 criterion established by LVBench [14], we retained videos with durations of at least 30 minutes. This resulted in a final set of 8,018 candidate videos for subsequent content screening and diversity-aware selection. Video Quality Filtering.A key challenge in eval- uating long-context multimodal models is ensuring that long duration reflects meaningful progression rather than redundancy (e.g., videos with looping contents or disor- dered clip compilations). Specifically, we required each video to exhibit temporally ordered changes in at least one benchmark-relevant dimension, such as character or entity states, events or actions, scenes or activities, topics or arguments, or interpersonal dynamics. We excluded repet- itive or looping content, such as Rainstorm Sounds for Relaxing, in which different temporal portions are largely interchangeable with no evolving semantic information. To translate this requirement into a concrete video selection process and ensure alignment with the capabilities defined in our taxonomy, two authors proficient in Japanese man- ually screened the candidate videos against four criteria: 1)Visual Integrity:We excluded videos dominated by static imagery or obstructive watermarks because they provide insuďŹicient evolving visual evidence for long- video understanding. 2)Temporal Semantic Progression:The reviewers in- spected the content at the 25%, 50%, 75%, and 100% timestamps and scrubbed the timeline around each an- chor point. At each location, they identified the active characters or entities, ongoing events or activities, and current topic or argument. A video was retained only when these observations exhibited temporally ordered changes in at least one dimension and the changes formed a connected sequence of events, topics, or ar- guments. We excluded videos whose sampled portions were repetitive or interchangeable, as well as compila- tions of unrelated clips. Incremental Semantic Diversity Sampling.To pre- vent over-representation of certain domains and main- tain broad coverage of cultural contexts, we applied an automated sampling strategy guided by semantic diver- sity. Starting from 30 selected seed videos, we used a greedy, embedding-based selection process. A new video was added only if the cosine similarity between its ti- tle and description and those of all previously selected samples was below a fixed thresholdĎ. This procedure ensures that each video contributes new semantic con- tent, expanding the range of scenarios and perspectives represented inNARU. The full algorithm is described on our website [19]. Finally, we selected 155 videos from the candidate set, corresponding to approximately 146.8 hours of content. C. MLLM-Driven QA Generation Given the selected videos and capability taxonomy, we construct candidate questionâanswer pairs through a hierarchical annotation-to-QA synthesispipeline. Direct generation from an hours-long video would require a single model invocation to preserve fine-grained evidence, inte- grate narrative relations, and interpret culturally implicit cues. We therefore separate benchmark construction into two stages. First, theannotation stagetransforms each video into a temporally connected hierarchy of chunk- and segment-level evidence and enriches this representation with taxonomy-aligned narrative and cultural annotations. Second, thequestionâanswer pair generation stage uses the resulting annotations to synthesize multiple- choice questions and reduce text-only shortcuts through video-blind diagnosis and targeted revision. 1)Annotation:This process converts a long-form video into a structured representation using an MLLM- centric, multi-layer memory pipeline based on chunking, segmentation, and task-oriented annotation. Step 1-1: Chunking.Following prior long-video an- notation protocols that use five-minute clips as units for dense narration [39], [40], we partition each video into approximately five-minute chunks. We treat this duration as a practical local processing unit rather than a semantic boundary. When transcripts are available, boundaries are aligned to the nearest transcript endpoint after five min- utes to avoid splitting dialogue; otherwise, fixed-duration cuts are used. An MLLM processes these chunks sequen- tially to produce schema-constrained JSON records con- taining chunk summaries, entity lists, timestamped events, and closed captions (including dialogue, on-screen text, non-speech audio, and vocal tone). To maintain entity consistency and track cross-boundary events, the model is given a textual recap of preceding chunks with relative timestamps whenever it generates a new record. Finally, chunk records are unified into a global timeline by shifting timestamps to video-relative coordinates, deduplicating entities by identifier, and sorting events chronologically. Step 1-2: Semantic Segmentation.Because chunk boundaries reflect processing constraints rather than nar- rative structure, we perform a second pass over the full video using lower-rate visual sampling and the merged Step 1-1 annotation as reference. Motivated by Segmented Discourse Representation Theory [41], this pass identifies contiguous, chapter-level segments with a coherent topic or narrative function. Each segment records its temporal span, summary, detailed description, and content type. A boundary is introduced at a substantive thematic shift or a change in content function, such as a transition between the main program, ignoring routine camera cuts and speaker turns. The resulting segments are combined with the video-level entities, events, and closed captions to provide the structured context for the taxonomy-aligned annotation in Step 1-3. Step 1-3: Task-Oriented Annotation.Steps 1â2 record what occurs in the video and organize this evi- dence over time, but they do not explicitly capture the higher-level narrative and cultural relations defined in our benchmark taxonomy. We therefore introduce two complementary MLLM-based annotation modules w.r.t. to the narrative intelligence and cultural understanding capabilities mentioned in Sec. I-A. Thenarrative anno- JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20266 tation modulereasons over the video-level representation to assign functional roles to characters and organize events into coherent narrative threads. Each thread identifies its central entities, highlights consequential events, and summarizes their causal progression, providing structured evidence for questions on four subjects in narrative intelli- gence (N.1-N.4). Thecultural annotation moduleanalyzes each semantic segment and extracts timestamped evidence for the five cultural dimensions (C.1-C.5). These segment- level annotations are subsequently integrated along the video timeline. Together, the two modules transform a gen- eral video representation into taxonomy-aligned evidence for subsequent question generation. 2)Question-Answer Pair Generation:Using the taxonomy-aligned annotations, we employ an MLLM to generate candidate four-option multiple-choice questions (MCQs). We organize this process into two steps: (1) capability-conditioned formulation, which associates each question with a taxonomy category and supporting ev- idence; (2) video-blind diagnosis and targeted revision, which identifies and repairs text-only shortcuts. All gener- ated items remain candidates until the human verification described in Sec. I-D. Step 2-1: Capability-Conditioned Formulation. The entry point for generation is the selection of a cognitive task grounded in the annotations provided in Sec. I-C1. For each evaluation category, the MLLM re- ceives the original video, the segment-level representation, the relevant narrative or cultural annotation, and task- specific instructions and examples. Then, it is required to provide a question stem, a proposed correct answer option and a supporting record identifying the question rationale and its relevant entities, events, temporal segments, and audio cues. Simultaneously, the model generates three can- didate distractor options. Distractor design is critical since implausible options or differences in length, specificity, emotional tone, and abstraction can reveal the correct answer without requiring the video (text-only shortcut). We therefore require all options in each MCQ to have comparable granularity, polarity, length, and syntactic complexity. Meanwhile, each distractor must also be plau- sible within the videoâs setting while remaining incorrect w.r.t. its evidence. These constraints provide an initial safeguard against option-level shortcuts before the video- blind screening in Step 2-2. Step 2-2: Iterative Debiasing Refinement.Recent work shows that MLLMs can answer visual questions confidently without visual input, creating an illusion of grounded understanding [42]. To alleviate such textual and language-bias shortcuts, we employ a Solver-Critic Loop, a three-stage iterative debiasing pipeline designed to empirically validate and refine question diďŹiculty: â˘The Blind Solver Agent:An agent attempts to answer the generated question without the video context (V=â ). If this âblindâ agent correctly identifies the answer, the question is flagged as potentially leaked. â˘Diagnostic Agent:A reasoning agent analyzes the Blind Solver Agentâs success to identify the vulnera- bility, such asTone Bias(e.g., the correct answer is the only polite option) orProcess of Elimination(e.g., several distractors contain implausible or mutually in- consistent details, leaving one viable answer without access to the video). It then outputs aRefine Plan to explain how such a shortcut happens and provides suggestions to address the underlying vulnerability. â˘Question Revision Agent:Guided by the Refine Plan, this agent rewrites the question stem, correct answer, or distractors to eliminate language shortcuts. This process iterates until the Blind Solver Agentâs suc- cess rate draws close to a natural random chance or a pre- defined iteration budget is reached. Therefore, multimodal comprehension ability is considered the primary driver for answering the refined questions correctly. Detailed algorithmic pseudocode for this loop is provided on our website [19]. Unless explicitly stated otherwise, we employ Gemini 2.5 Pro [43] for all MLLM-based components of the benchmark construction pipeline, including hierarchical annotation, questionâanswer generation, and the Solverâ Critic refinement loop. D. Quality Control To ensure benchmark reliability, we incorporate human validation at two critical stages of theNARUconstruction pipeline: following the initial question-answer pair gener- ation and after the iterative debiasing refinement stage. We formally recruited40native Japanese experts for the initial validation and28experts for the second-stage evaluation. The annotators verified that each QA item was grounded in the source video, culturally faithful, and appropriate for multiple-choice evaluation. In addition to improving benchmark quality through manual correction and filtering, this two-stage validation process provides an empirical assessment of the effectiveness of the automatic generation and refinement pipeline. Initial QA verification.Across a two-week annota- tion phase, a panel of 40 annotators verified that every question was answerable from the video and provided a single correct answer. They directly revised ambiguous questions, overlapping options, missing correct answers, and incorrect labels. Items that could not be reliably cor- rected were removed. Among 1,500 candidate questions, 178 received an invalid-question flag (not answerable); 161 were repaired, and 17 were removed, resulting in 1,483 candidates progressing for refinement. In addition, annotators authored a correct answer for 107 items and corrected the assigned label for 177 items. Post-refinement verification.Because shortcut- triggered revision (Step 2-2) could alter the question stem, correct answer, or distractors, a second cohort of 28 an- notators evaluated each refined item against its original version and the source video to ensure overall data quality. Over a two-week review stage, they confirmed question validity, verified answer correctness and specificity, and ensured distractors remained plausible, distinct, and un- ambiguously incorrect. Of 1,483 refined items, 949 were JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20267 accepted without a change, 532 were revised, and two were removed. The overlapping revisions included 436 answer corrections, 108 distractor edits, and 35 question rewrites. Following this verification, the final benchmark contains 1,481 items, and the corresponding statistics are reported in Table I. IV. Evaluation We evaluate both open-source and closed-source MLLMs onNARUto characterize their performance and the specific conditions under which long-form, culturally grounded understanding succeeds or breaks down. We begin with a full-benchmark multiple-choice evaluation to provide a controlled comparison across models and expose performance variances across the nine capability categories defined in Sec. I-A. Because long-video under- standing depends heavily on a modelâs temporal capacity, we systematically vary the number of sampled frames to study whether additional visual context leads to superior narrative and cultural comprehension. Subsequently, we notice that multiple-choice options can inadvertently leak contextual hints; therefore, we transformNARUinto an open-ended format to reassess the models. In this setting, models are prompted to produce free-form answers that we score for factual coverage and grounding quality. A. Experimental Setup Models.We evaluate a diverse set of proprietary and open-source MLLMs equipped with long-video un- derstanding capabilities. The proprietary models com- prise Gemini-3-Flash [44], Gemini-3-Pro [45], and Gemini- 2.5-Flash [43], which we access through their native video-input APIs. The open-source models comprise Qwen3.5-9B [46], Qwen3-VL-8B [47], Qwen2.5-VL-7B [2], MiniCPM-o-2.6 [48], and InternVL3.5 [1]. Inference Setting.For MCQ evaluation, each model answers the complete set of four-choice questions in NARU. By default, we adopt the standard input con- figuration recommended for each model family. For the Gemini models, we sample videos uniformly atfps=0.25, ensuring the native API can process the longest videos in our benchmark while maintaining consistent tempo- ral coverage (i.e., 3600 Frames for 4-hour videos). For the open-source models, we apply the maximum uniform sparse-frame sampling permitted by each modelâs respec- tive context window budget. We report overall accuracy as well as fine-grained performance across the nine capability categories: four narrative (N.1âN.4) and five cultural (C.1â C.5). For the remainder of our diagnostic experiments, we applied stratified sampling across categories to obtain 500 questions from the complete 1,481-questionNARUbench- mark to maintain manageable computational overhead. We use this fixed subset for all models in both the frame- budget sweep (Sec. IV-C) and the open-ended generation evaluation (Sec. IV-D). B. Overall Performance Table I reports the primary multiple-choice results across the full benchmark. Overall performance follows a distinct tiering: proprietary models lead significantly, with Gemini-3-Flash achieving the highest accuracy at 76.2%, followed by Gemini-3-Pro (70.0%) and Gemini-2.5-Flash (51.4%). In contrast, open-source models fall into a lower performance regime (29.6â39.8%), led by Qwen3.5-9B. A deeper look at category-level performance reveals a clear capability-dependent split between narrative and cultural reasoning. While Gemini models achieve approxi- mately 11 percentage points higher accuracy on narrative tasks than on cultural ones, open-source models show virtually no difference between the two dimensions. Within narrative categories, tracking sequential structure (N.2) is universally the easiest task. Crucially, however, the pri- mary performance bottleneck shifts as model capability in- creases: weaker models fail primarily at tracking low-level entity continuity (e.g., N.1 Character/Entity Evolution at 20.0% for InternVL3.5 and 23.2% for MiniCPM-o-2.6), whereas stronger models struggle most with high-level abstraction (N.4 Idea/Thematic Development). Thus, the dominant narrative diďŹiculty shifts across the evaluated models, from maintaining entity continuity in most open- source models to synthesizing higher-level thematic devel- opment in the Gemini family. The cultural domain presents an equally nuanced pic- ture. Although open-source models perform best on C.2 (Kuuki wo Yomu), pragmatic reasoning remains challeng- ing across all evaluated systems. In particular, C.3 (Sub- text Interpretation) proves exceptionally diďŹicult even for frontier models: Gemini-3-Flash drops to 57.4% on C.3, despite its strong 68.2% cultural average, making this the sole category where Gemini-3-Pro yields higher accuracy. Furthermore, several open-source models exhibit severe failure cases, with performance on narrative tracking (N.1) and cultural tasks (C.1, C.4) dipping below the 25% random-guessing baseline. Ultimately, these findings point to two distinct mechanisms underlying performance gains: long-range narrative tracking scales directly with context length and general reasoning, whereas cultural inference is more plausibly bottlenecked by the richness and coverage of relevant knowledge in the pre-training data. C. Effect of Temporal Evidence Long-form video understanding depends not only on model capacity but also on how densely the video is represented within the input context window. To isolate this effect, we evaluate all models on the same 500- question subset while varying the number of sampled framesfâ8,16,32,64,128. This controlled sweep measures whether broader temporal coverage helps models recover the distributed evidence required byNARU. Fig. 3 shows the difference in how model families exploit additional temporal evidence. All three Gemini models im- prove substantially as the frame budget increases: Gemini- 3-Pro rises from approximately 64% with 8 frames to JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20268 Level TaskType of EvidenceCode # Narrative (N, 745) Character/Entity EvolutionA character/entity across segmentsN.1 185 Sequential/Topical FlowEvents or topics over timeN.2 187 Plot/Conflict ProgressionA causal thread or conflictN.3 186 Idea/Thematic DevelopmentMotifs, claims, or narrative cuesN.4 187 Cultural (C, 736) Aizuchi(Conversational Mechanics)Backchannels and response timingC.1 143 Kuuki wo Yomu(Situational Awareness)Social atmosphere or implicit normsC.2 147 Subtext InterpretationSurface utterance plus contextC.3 148 Cultural Context RecognitionCulturally specific referencesC.4 149 Sentiment AnalysisVerbal, visual, and social cuesC.5 149 TABLE I: Statistics of NARU . ModelSampling Rate N.1 N.2 N.3 N.4 Narr. Avg C.1 C.2 C.3 C.4 C.5 Cult. Avg Overall Gemini-3-Flash0.25 FPS83.8 90.9 83.9 78.1 84.2 71.3 69.457.473.2 69.8 68.2 76.2 Gemini-3-Pro0.25 FPS74.678.676.366.874.160.168.764.269.867.166.070.0 Gemini-2.5-Flash0.25 FPS51.3 63.1 58.1 50.855.837.8 49.7 48.6 50.3 48.346.951.4 Qwen3.5-9B128 Frames 34.6 48.7 34.9 34.238.141.3 49.0 34.5 40.9 41.641.439.8 Qwen3VL-8B0.25 FPS35.7 46.5 39.8 35.839.532.2 40.1 33.1 38.9 32.935.437.4 Qwen2.5VL-7B128 Frames 25.4 42.2 30.6 26.231.122.4 40.8 25.7 24.2 28.228.229.7 MiniCPM-o-2.6128 Frames 23.2 41.2 27.4 31.630.825.9 37.4 27.7 22.1 28.228.329.6 InternVL3.564 Frames20.0 36.4 25.8 31.028.332.2 41.5 31.8 35.6 32.234.631.5 TABLE I: MCQ accuracy results on NARU. Thebestand the 2nd bestresults are marked. 8163264128 Number of sampled frames (f) 30 40 50 60 70 Overall MCQ accuracy (%) chance Gemini-3-Pro Gemini-3-Flash Gemini-2.5-Flash Qwen3.5-9B Qwen3VL-8B Qwen2.5VL-7B MiniCPM-o-2.6 InternVL3.5 Fig. 3: Overall multiple-choice accuracy as the number of sampled frames increases from 8 to 128; the dotted gray line denotes the 25% chance level. 71% with 128 frames, Gemini-3-Flash from 53% to 64%, and Gemini-2.5-Flash from 39% to 51%. On the other hands, the open-weight models improve less consistently. Qwen3VL-8B exhibits the largest gain in this group, from about 31% to 38%, whereas Qwen2.5VL-7B, MiniCPM- o-2.6, and InternVL3.5 remain close to 30% even at 128 frames. Overall, improvements are larger and more consis- tent among the Gemini models. One possible explanation is a difference ineffective video context: the amount of temporally distributed evidence a model can integrate, rather than its nominal input capacity. Geminiâs end-to- end video pre-training and token-eďŹicient visual encoders may help it convert additional frames into usable evi- dence [43], whereas support for 64â128 frames in open- weight models [2], [48], [1] does not necessarily imply a comparable ability to select and integrate evidence across the sequence. Beyond overall accuracy, the results shown in Fig. 4 highlight a distinct contrast between task domains: in- Gemini-3-Pro Gemini-3-Flash Gemini-2.5-Flash Qwen3.5-9B Qwen3VL-8B Qwen2.5VL-7B MiniCPM-o-2.6 InternVL3.5 â5 0 5 10 15 20 Accuracy gain (p) +8.6 +3.9 +12.3 +10.0 +20.5 +1.8 +6.4 -0.7 +10.0 +2.5 +2.7 +1.8 +6.3 -3.9 +1.4 +0.4 Narrative Î Cultural Î Fig. 4: Changes in narrative and cultural accuracy between 8 and 128 frames, reported in percentage points (p). ModelN.1 N.2 N.3 N.4 C.1 C.2 C.3 C.4 C.5 Avg Gemini-3-Flash0.78 0.66 0.72 0.75 0.69 0.87 0.93 0.85 0.80 0.78 Gemini-3-Pro 0.77 0.610.710.690.650.870.880.800.72 0.75 Gemini-2.5-Flash 0.62 0.50 0.58 0.65 0.56 0.77 0.84 0.68 0.730.66 Qwen3.5-9B 0.54 0.49 0.47 0.50 0.52 0.66 0.65 0.60 0.59 0.56 Qwen3-VL-8B 0.39 0.23 0.34 0.36 0.47 0.64 0.62 0.39 0.53 0.44 Qwen2.5-VL-7B 0.35 0.24 0.27 0.28 0.43 0.47 0.40 0.28 0.40 0.35 MiniCPM-o-2.6 0.21 0.12 0.15 0.16 0.23 0.26 0.30 0.15 0.31 0.21 InternVL3.50.40 0.14 0.27 0.35 0.43 0.49 0.47 0.39 0.51 0.38 TABLE I: Open-ended correctness on NARU. Scores are FActScore recall (0â1): fraction of reference atomic facts covered by the answer. creasing the frame budget impacts narrative understand- ing far more than cultural understanding. For every model, narrative performance gains consistently outpace cultural gains. Narrative accuracy rises by 1.4 to 20.5 percentage points, whereas cultural accuracy fluctuates between a 3.9- point decline and a 10.0-point gain. This disparity indi- cates that narrative errors are largely caused by missing dispersed events, which additional frame sampling directly resolves. Cultural interpretation, however, depends less JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20269 on visual frequency and more on underlying pragmatic reasoning and domain knowledge. Finally, a notable pattern emerges between Gemini-3- Pro and Gemini-3-Flash. Pro leads across all controlled frame budgets, though its margin over Flash shrinks from 11 points at 8 frames to 7 points at 128 frames. However, in the full-benchmark setting at 0.25 FPS (Table I), Flash leads Pro (76.2% vs. 70.0%). This reversal suggests distinct regime-dependent strengths: Pro exhibits superior low-information reasoning under strict frame constraints, whereas Flash exhibits steeper scaling dynamics, capital- izing on high-frequency temporal inputs to yield larger performance gains as context grows denser. D. Open-Ended Evaluation MCQ evaluation provides a comparison across models, but a correct prediction may arise from recognizing the most plausible option or eliminating distractors without reconstructing the answer itself (text-only shortcut). We therefore introduce a complementary open-ended evalu- ation that removes all answer choices and requires each model to generate the requested information directly. We convert the standardized 500-question diagnostic subset into free-form questions and evaluate the same model families used in the multiple-choice experiments. Because correct responses can paraphrase the reference answer, lexical overlap and exact match are inadequate measures of correctness. We instead use GPT-5.5 as an automated judge. Following the idea of FActScore [49], for each question, the judge decomposes the reference answer into self-contained atomic facts and determines whether the generated response covers each fact. We use atomic- fact recall as the primary metric, which is defined as the fraction of reference facts covered by the response. To ensure the quality of the automated judge, we com- pare its results with human judgments. That is, three an- notators independently inspect 50 judged responses (sam- pled from Gemini-3-Flash answers) and indicate whether they agree with the judgesâ decisions. Across the resulting 150 verdicts, the annotators accept the judgeâs scores in 90.0% of cases, with individual acceptance rates ranging from 88.0% to 92.0%. Majority voting accepts 48 of the 50 judgments (96.0%), mean pairwise agreement among annotators is 82.7%, and no judgment is rejected by all three annotators. These results support using the auto- mated judge for aggregate analysis. Result.Removing the choices preserves the perfor- mance hierarchy observed in the multiple-choice evalua- tion: the Gemini family maintains a substantial advantage, achieving 0.66â0.78 atomic-fact recall compared to 0.21â 0.56 across open-source models (Table I). However, we observe that open-ended evaluation yields a distinctly different capability profile than multiple-choice evaluation. The most striking cross-format discrepancy appears in N.2 (Sequential/Topical Flow). While N.2 is the strongest narrative subcategory across all models in the multiple- choice setting (Table I), it degrades into the weakest nar- rative category for seven of the eight models under open- ended generation. A plausible explanation is that multiple- choice options serve as structural scaffolds for temporal organization, enabling models to recognize pre-sequenced event chains. Without these candidate options, models must independently retrieve, synthesize, and chronolog- ically reconstruct events distributed across the video making sequential flow substantially more challenging in open-ended settings. A second notable shift is that the relative diďŹiculty between narrative and cultural understanding reverses across formats. In the multiple-choice evaluation, six of the eight models achieve higher accuracy on narrative questions than on cultural ones. In contrast, every model achieves higher atomic-fact recall on cultural questions under open-ended evaluation. One factor that makes this reversal visible is the graded atomic-fact metric: multiple- choice evaluation applies a binary correctness criterion, whereas atomic-fact recall grants partial credit whenever a response recovers a subset of ground-truth facts. Across all models, 78.9% of cultural responses recover at least one reference fact, compared to 66.3% of narrative responses. Because cultural prompts often outline the underlying interaction before asking for its implicit significance, this contextual grounding helps models generate partially cor- rect observations even when missing the full interpreta- tion. Conversely, narrative questions frequently demand the unprompted reconstruction of distributed events, lead- ing to a higher rate of zero-recall responses. V. Conclusion This paper introducedNARU, a benchmark of 1,481 questions grounded in 155 Japanese long-form videos totalling 146.8 hours and spanning four narrative and five cultural dimensions.NARUcombines a hierarchical annotation-to-QA pipeline with iterative shortcut removal and two-stage verification by 68 native Japanese anno- tators to construct high-quality questions at this tem- poral scale. Evaluation across various model configura- tions shows that both performance and behavior vary substantially across model families, including how models respond to denser frame sampling and how they be- have under open-ended settings. Crucially, our results show that current open-source models lag significantly behind top commercial models in both narrative reasoning and culturally grounded understanding. We hopeNARU serves as a foundational catalyst for developing next- generation MLLMs capable of genuine, high-context video understanding. References [1]W. Wang, Z. Gao, L. Gu, H. Pu, L. Cui, X. Wei, Z. Liu, L. Jing, S. Ye, J. Shaoet al., âInternvl3. 5: Advancing open-source multimodal models in versatility, reasoning, and eďŹiciency,â arXiv preprint arXiv:2508.18265, 2025. [2]S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tanget al., âQwen2. 5-vl technical report,â arXiv preprint arXiv:2502.13923, 2025. JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 202610 [3]B. Zhang, K. Li, Z. Cheng, Z. Hu, Y. Yuan, G. Chen, S. Leng, Y. Jiang, H. Zhang, X. Liet al., âVideollama 3: Frontier multi- modal foundation models for image and video understanding,â arXiv preprint arXiv:2501.13106, 2025. [4]Y. Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhuet al., âVideo understanding with large language models: A survey,âIEEE Transactions on Circuits and Systems for Video Technology, 2025. [5]E. Marrese-Taylor, Y. Hamazono, T. Ishigaki, G. TopiÄ, Y. Miyao, I. Kobayashi, and H. Takamura, âOpen-domain video commentary generation,â inProc. EMNLP. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, p. 7326â7339. [Online]. Available: https://aclanthology.org/2022.emnlp-main.495/ [6]H. Zou, T. Luo, G. Xie, F. Lv, G. Wang, J. Chen, Z. Wang, H. Zhang, H. Zhanget al., âFrom seconds to hours: Reviewing multimodal large language models on comprehensive long video understanding,âarXiv preprint arXiv:2409.18938, 2024. [7]H. Hua, Y. Tang, C. Xu, and J. Luo, âV2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning,â inProc. AAAI, vol. 39, no. 4, 2025, p. 3599â3607. [8]E. Song, W. Chai, G. Wang, Y. Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y. Zhanget al., âMoviechat: From dense token to sparse memory for long video understanding,â inProc. CVPR, 2024, p. 18221â18232. [9]G. J. Faure, J.-F. Yeh, M.-H. Chen, H.-T. Su, S.-H. Lai, and W. H. Hsu, âHermes: temporal-coherent long-form understand- ing with episodes and semantics,â inProc. ICCV, 2025, p. 22911â22921. [10]A. Mukherjee and S. Ghosh, âToward socially aware vision- language models: Evaluating cultural competence through mul- timodal story generation,â inProc. ICCV, 2025, p. 1491â1501. [11]Z. Chen, H. Lin, K. Li, Z. Luo, Y. Deng, and J. Ma, âMemearena: Automating context-aware unbiased evaluation of harmfulness understanding for multimodal large language models,â inProc. EMNLP, 2025, p. 17648â17670. [12]R. Kriz, K. Sanders, D. Etter, K. Murray, C. Carpenter, H. Rec- knor, J. Guallar-Blasco, A. Martin, E. Yang, and B. Van Durme, âMultivent 2.0: A massive multilingual benchmark for event- centric video retrieval,â inProc. CVPR, 2025, p. 24149â24158. [13]H. Wu, D. Li, B. Chen, and J. Li, âLongvideobench: A bench- mark for long-context interleaved video-language understand- ing,âAdvances in Neural Information Processing Systems, vol. 37, p. 28828â28857, 2024. [14]W. Wang, Z. He, W. Hong, Y. Cheng, X. Zhang, J. Qi, M. Ding, X. Gu, S. Huang, B. Xuet al., âLvbench: An extreme long video understanding benchmark,â inProc. ICCV, 2025, p. 22958â 22967. [15]J. Yu, Y. Wu, M. Chu, Z. Ren, Z. Huang, P. Chu, R. Zhang, Y. He, Q. Li, S. Liet al., âVrbench: A benchmark for multi-step reasoning in long narrative videos,â inProc. ICCV, 2025, p. 21655â21666. [16]Z. Wu, Z. Liu, A. Chen, J. Zhang, R. Li, H. Ge, Z. Wang, C. Xiao, and C. Liang, âStoryvideoqa: Scaling deep video un- derstanding with a large-scale, multi-genre and auto-generated dataset,âInternational Journal of Computer Vision, vol. 134, no. 6, p. 308, 2026. [17]B. S. Shafique, A. Vayaniet al., âA culturally-diverse multilin- gual multimodal video benchmark & model,â inProc. EMNLP, 2025, p. 20009â20033. [18]Yano Research Institute, âVideo content business market for fy2025,â Sep. 2025. [Online]. Available: https://w. yanoresearch.com/en/press-release/show/press_id/3919 [19]Website of this Paper, https://ma-labo.github.io/naru/. [20]D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, âVideo question answering via gradually refined attention over appearance and motion,â inACM Multimedia, 2017. [21]J. Xiao, X. Shang, A. Yao, and T.-S. Chua, âNext-qa: Next phase of question-answering to explaining temporal actions,â in Proc. CVPR, 2021, p. 9777â9786. [22]M. Maaz, H. Rasheed, S. Khan, and F. Khan, âVideo-chatgpt: Towards detailed video understanding via large vision and lan- guage models,â inProc. ACL (Volume 1: Long Papers), 2024, p. 12585â12602. [23]K. Li, Y. Wang, Y. Heet al., âMvbench: A comprehensive multi- modal video understanding benchmark,â inProc. CVPR, 2024, p. 22195â22206. [24]J. Zhou, Y. Shuet al., âMlvu: Benchmarking multi-task long video understanding,â inProc. CVPR, 2025, p. 13691â13701. [25]R. Rawal, K. Saifullah, M. FarrĂŠ, R. Basri, D. Jacobs, G. Somepalli, and T. Goldstein, âCinepile: A long video question answering dataset and benchmark,âarXiv preprint arXiv:2405.08813, 2024. [26]C. Fu, Y. Dai, Y. Luo, L. Li, S. Renet al., âVideo-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,â inProc. CVPR, 2025, p. 24108â24118. [27]D. Ma, H. Yuanet al., âScalelong: A multi-timescale benchmark for long video understanding,â inThe Fourteenth International Conference on Learning Representations, 2026. [Online]. Available: https://openreview.net/forum?id=95sD6KKq51 [28]S. Choi, K.-W. On, Y.-J. Heo, A. Seoet al., âDramaqa: Character-centered video story understanding with hierarchical qa,â inProc. AAAI, vol. 35, no. 2, 2021, p. 1166â1174. [29]S. You, B. Yuan, and B.-K. Bao, âScvbench: A benchmark with multi-turn dialogues for story-centric video understanding,â in Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, 2025, p. 2287â2295. [30]H. Ha, J. Ge, B. Feng, K. Ma, and G. Chakraborty, âNarra- tivetrack: Evaluating video language models beyond the frame,â arXiv preprint arXiv:2601.01095, 2026. [31]R. Jain, K. Doshi, B. Uzkent, and G. Kessler, âNarrative aligned long form video question answering,â inProc. CVPR, 2026, p. 8765â8774. [32]A. Zadeh, M. Chan, P. P. Liang, E. Tong, and L.-P. Morency, âSocial-iq: A question answering benchmark for artificial social intelligence,â inProc. CVPR, 2019, p. 8807â8817. [33]X.-Y. Guo, Y.-F. Li, and R. Haf, âDesiq: Towards an unbiased, challenging benchmark for social intelligence understanding,â in Proc. EMNLP, 2023, p. 3169â3180. [34]S. Nayak, K. Jain, R. Awal, S. Reddy, S. Van Steenkiste, L. A. Hendricks, K. StaĹczak, and A. Agrawal, âBenchmarking vision language models for cultural understanding,â inProc. EMNLP, 2024, p. 5769â5790. [35]J. M. Zacks and K. M. Swallow, âEvent segmentation,âCurrent directions in psychological science, vol. 16, no. 2, p. 80â84, 2007. [36]E. T. Hall,Beyond culture. Anchor, 1976. [37]S. Kita and S. Ide, âNodding, aizuchi, and final particles in japanese conversation: How conversation reflects the ideology of communication and social relationships,âJournal of Pragmat- ics, vol. 39, no. 7, p. 1242â1254, 2007. [38]S. K. Maynard, âOn back-channel behavior in japanese and english casual conversation,â 1986. [39]K. Grauman, A. Westburyet al., âEgo4d: Around the world in 3,000 hours of egocentric video,â inProc. CVPR, 2022, p. 18995â19012. [40]J. Yang, S. Liuet al., âEgolife: Towards egocentric life assistant,â inProc. CVPR. IEEE, 2025, p. 28885â28900. [41]A. Lascarides and N. Asher, âSegmented discourse representa- tion theory: Dynamic semantics with discourse structure,â in Computing meaning. Springer, 2007, p. 87â124. [42]M. Asadi, J. W. OâSullivan, F. Cao, T. Nedaee, K. Rajabalifardi, F.-F. Li, E. Adeli, and E. Ashley, âMirage: The illusion of visual understanding,âarXiv preprint arXiv:2603.21687, 2026. [43]Gemini Team, âGemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,â Google DeepMind, Tech. Rep., 2025. [Online]. Available: https://storage.googleapis. com/deepmind-media/gemini/gemini_v2_5_report.pdf [44]Google DeepMind, âGemini 3 flash model card,â 2025. [Online]. Available: https://deepmind.google/models/ model-cards/gemini-3-flash/ [45]â, âGemini 3 pro model card,â 2025. [Online]. Available: https://deepmind.google/models/model-cards/gemini-3-pro/ [46]Qwen Team, âQwen3.5,â Feb. 2026. [Online]. Available: https://qwen.ai/blog?id=qwen3.5 [47]S. Bai, Y. Cai, R. Chenet al., âQwen3-vl technical report,â arXiv preprint arXiv:2511.21631, 2025. [48]OpenBMB, âMinicpm-o-2.6 model card,â https://huggingface. co/openbmb/MiniCPM-o-2_6, 2025, accessed: 2026-07-20. [49]S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi, âFactscore: Fine- grained atomic evaluation of factual precision in long form text generation,â inProc. EMNLP, 2023.