Paper deep dive
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/18/2026, 3:01:24 PM
Summary
The paper introduces StoryTeller, a training-free framework for generating long-form audio descriptions (AD) for blind and low-vision audiences. It addresses the limitation of existing Video-Language Models (VLMs) that treat clips independently by maintaining a verified narrative memory and an identity graph. StoryTeller processes raw video and movie titles, optionally retrieving public metadata to resolve names, but strictly filters facts through semantic and VLM verification to ensure grounding. The system uses reinforcement-decay dynamics for memory salience. The authors also introduce StoryAD-QA, a benchmark evaluating narrative comprehension from generated ADs, demonstrating improved coherence and factual grounding over baselines.
Entities (9)
Relation Signals (7)
StoryTeller â maintains â Identity Graph
confidence 95% · StoryTeller maintains an explicit story state... an identity graph for recurring characters
StoryTeller â maintains â Narrative Memory
confidence 95% · StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes
StoryTeller â targets â Audio Description
confidence 95% · We propose StoryTeller, a training-free framework for story-aware long-form AD.
StoryAD-QA â evaluates â Audio Description
confidence 90% · StoryAD-QA, a question-answering benchmark that tests whether a language model can answer story-context questions using only the generated descriptions.
StoryTeller â supports â Blind and Low-Vision
confidence 90% · Long-form audio description (AD) requires... so that blind and low-vision (BLV) audiences can follow a film.
StoryTeller â uses â IMDb
confidence 88% · retrieve public movie metadata... from IMDb
StoryTeller â outperforms â AutoAD-Zero
confidence 85% · StoryTeller consistently improves narrative coherence... over strong baselines... The baseline example is the output of AutoAD-Zero.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. Instead of relying only on local visual cues, StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes, enabling later descriptions to remain coherent, grounded, and contextually informative. Given only raw video and a movie title, StoryTeller can optionally retrieve public movie metadata to resolve names and story context, while accepting only facts that are supported by the video through semantic filtering and VLM verification. The method requires no subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. To evaluate whether generated AD preserves narrative information, we introduce StoryAD-QA, a question-answering benchmark that tests whether a language model can answer story-context questions using only the generated descriptions. Experiments on standard AD benchmarks and diverse long-form videos show that StoryTeller consistently improves narrative coherence, factual grounding, and story comprehension over strong baselines in automatic, QA-based, and human evaluations.
Tags
Links
- Source: https://arxiv.org/abs/2607.11798v1
- Canonical: https://arxiv.org/abs/2607.11798v1
Trouble viewing inline? Open PDF directly â
Full Text
101,103 characters extracted from source content.
Expand or collapse full text
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description Seung Hyun Hahm, Minh T. Dinh, and SouYoung Jin Dartmouth College, USA Seung.Hyun.Hahm.GR,Minh.T.Dinh.GR,SouYoung.Jin@dartmouth.edu Abstract. Long-form audio description (AD) requires more than de- scribing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audi- ences can follow a film. Modern videoâlanguage models (VLMs) are effec- tive on short clips, but they often treat each moment independently, pro- ducing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. In- stead of relying only on local visual cues, StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes, enabling later descriptions to remain coherent, grounded, and contextually informative. Given only raw video and a movie title, Sto- ryTeller can optionally retrieve public movie metadata to resolve names and story context, while accepting only facts that are supported by the video through semantic filtering and VLM verification. The method re- quires no subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. To evalu- ate whether generated AD preserves narrative information, we introduce StoryAD-QA 1 , a question-answering benchmark that tests whether a language model can answer story-context questions using only the gener- ated descriptions. Experiments on standard AD benchmarks and diverse long-form videos show that StoryTeller consistently improves narrative coherence, factual grounding, and story comprehension over strong base- lines in automatic, QA-based, and human evaluations. Keywords: audio description· long-form video understanding· retrieval-augmented generation· narrative grounding· accessibility 1 Introduction Audio description (AD) enables blind and low-vision (BLV) audiences to access visual media by narrating important visual information between dialogue and other sounds. Effective AD goes beyond describing objects or actions in the current frame: it identifies who is present, explains what has just happened, 1 StoryAD-QA benchmark dataset and evaluation code: https://github.com/SEE-AI-Lab/ECCV2026_StoryTeller_StoryAD_QA arXiv:2607.11798v1 [cs.CV] 13 Jul 2026 2S. H. Hahm et al. Alastor Mad Eye Moody A professor at Hogwarts Mad Eye Moodypresents a spiderto the class. A boy covers his headas a spider crawls on it. Subject: Mad Eye Moody Action: holds Object: a spider Location: room with a chalkboard Memory StoryTeller AD Albus Dumbledoremakes his wand move with a flick of his wrist while conversing with someone. Harry is hit by a teacher, while Ron and Hermione look on. Baseline AD Observed Fact Subject: boy Action: covers Object: hair Location: a desk Previous scene Identity Graph Fig. 1: Story-level audio description requires narrative memory. Existing au- dio description (AD) systems typically generate descriptions using only local visual context or static character banks, which often fails to preserve long-range narrative relationships. Our StoryTeller instead maintains a persistent narrative state consist- ing of an identity graph that tracks characters across scenes and a salience-weighted memory that accumulates narrative facts over time. This enables consistent character references and long-range story reasoning across an entire film. The baseline example is the output of AutoAD-Zero. conveys why a moment matters, and connects the scene to the broader narrative. Professional AD writers are able to provide this narrative grounding because they watch the film and write with the whole story in mind. However, producing high-quality AD remains costly and time-intensive, and copyright restrictions make large publicly distributable AD datasets difficult to obtain. Recent advances in videoâlanguage models (VLMs) [2,19,30,36] and large lan- guage models have motivated research on automatic AD generation. State-of-the- art identity-aware AD systems [8,9,27] are typically trained as task-specific mod- els on curated movie datasets such as LSMDC [28] and MAD [31]. These pipelines incorporate identity-aware components built from face-recognition backbones trained on labeled data. While effective in benchmark settings, this reliance on copyrighted training data, curated movie annotations, and specialized identity modules limits their scalability, reproducibility, and applicability to new films. Training-free pipelines, including AutoAD-Zero [38] and M_Narrator [40], have recently emerged as an alternative to supervised AD generation. These methods remove the need for task-specific model training, but they often still rely on resources prepared before AD generation, including character banks, aligned subtitles, or existing AD transcripts. Moreover, they typically propagate context in shallow or static forms, such as concatenating previous descriptions or reusing fixed identity features. Although these strategies can improve local consistency across nearby clips, they do not explicitly model how narrative in- StoryTeller3 formation evolves over time: which characters, relationships, and events should remain salient, and which details should fade as the film progresses. Human viewers do not understand a film by retaining a complete list of pre- vious frames. Instead, they maintain an evolving representation of the story, where central characters, unresolved conflicts, and recurring relationships re- main salient, while incidental visual details gradually lose importance. This per- spective suggests that long-form AD should preserve identity continuity, causal structure, and evolving character interactions across a film, without relying on external resources such as subtitles, scripts, AD transcripts, character banks, or precomputed face identities. We introduce StoryTeller, a training-free framework for long-form AD that maintains an explicit story state as it processes clips in chronological order. Given a sequence of video clips v 1 ,...,v T , the system keeps two forms of memory: an identity graph for recurring characters and a salience-weighted memory for verified story facts. Each new clip updates this state through three steps: track visible identities, extract candidate facts with optional public movie metadata, and verify those facts against the video before updating memory. The memory then reinforces facts that remain relevant and decays facts that no longer matter. StoryTeller takes as input only the raw video and a movie title. The title may be used to retrieve public IMDb plot summaries, character lists, and short character descriptions. This retrieved text is only a source of possible names or context; it is not copied into the AD, stored directly as memory, or treated as ground truth. A fact enters memory only if it is supported by the current clip and accepted by semantic filtering and VLM verification. Beyond generation, we also address evaluation. Standard metrics such as BLEU [25], CIDEr [34], and SPICE [1] measure lexical overlap between generated descriptions and reference captions, but they do not evaluate whether a listener can reconstruct the narrative from the AD alone. To address this limitation, we introduce StoryAD-QA, a narrative-comprehension question-answering bench- mark. For each scene, we generate questions whose answers depend on both the current clip and preceding narrative context. During evaluation, a language model receives only the generated audio descriptions and must answer these questions without access to the video. This protocol directly measures whether the description preserves characters, events, and causal relationships necessary for story understanding rather than merely describing visual details in isolation. Contributions. â We propose StoryTeller, a fully training-free framework for long-form audio description that uses raw video and public title-keyed metadata without subtitles, AD transcripts, manually curated visual character banks, or task- specific fine-tuning. â We introduce an explicit narrative memory mechanism with reinforcementâ decay dynamics that approximates how human viewers track evolving story salience across scenes. 4S. H. Hahm et al. â We present StoryAD-QA, a narrative-comprehension QA benchmark that evaluates whether generated audio descriptions support story understanding beyond lexical similarity. 2 Related Work Benchmarking Audio Description Generation. Movie understanding has often been studied as question answering over long videos [13, 32, 33, 35, 39]. For example, MovieQA [33] uses questions grounded in plot summaries, scripts, and subtitles, while [35] and [13] annotate characters, interactions, and story structure. MovieChat [32] emphasizes on long-video QA with sparse memory and introduces MovieChat-1K. These datasets are crucial for evaluating video reasoning, but they do not directly ask whether an AD gives a listener enough information to understand the story without seeing the video. Datasets designed specifically for AD instead focus on localized narration. LSMDC [28] and MAD [31] provide large-scale movie captioning resources aligned to short clips, and recent AD generation systems [8â10, 26, 27, 38, 40] build on these datasets with identity-aware modules or curated character banks. Despite their usefulness, evaluation in these works often relies on direct string comparison with professional AD using metrics such as CIDEr [34], SPICE [1], and ROUGE [21], which are highly sensitive to temporal alignment. Addressing this limitation, AutoAD I [8] introduces an evaluation protocol that measures whether generated descriptions are semantically aligned with nearby ground- truth annotations rather than requiring exact sentence matches. Building on this direction, AutoAD I [10] further proposes CRITIC for evaluating character retrieval and LLM-AD-Eval for sentence-level semantic scoring with LLMs. More recently, question-answering evaluation has been explored to assess the usefulness of AD. ADQA [15] evaluates whether generated AD supports visual appreciation and narrative understanding in a certain video segment via questions generated from reference descriptions. Instead, our benchmark focuses on long-form narrative grounding, evaluating whether ADs preserve information required for reasoning over story context that spans multiple scenes. Identity Modeling in Video Narratives. Maintaining consistent character identity across scenes is essential for long-form video understanding and AD. Sev- eral recent AD systems explicitly rely on character banksâprecomputed reposito- ries of character-specific visual features, e.g., face embeddings or reference images used to recognize recurring characters throughout a movie [8â10]. In contrast, StoryTeller requires no character bank or precomputed face identities. Instead, it discovers recurring identities directly from repeated face observations, allowing character identities to emerge dynamically as narrative evidence accumulates. Retrieval-Augmented Generation. Retrieval-augmented generation (RAG) improves grounding by incorporating external evidence during generation [5, 18]. However, naive retrieval can introduce irrelevant or contradictory infor- mation and may amplify hallucinations in multimodal models [12, 29]. Prior work improves retrieval quality through dense retrievers and structured fusion StoryTeller5 mechanisms [7, 14, 16], but these approaches focus primarily on factual knowl- edge grounding rather than narrative coherence. Our approach instead performs narrative-aware retrieval restricted to movie-scoped sources and triggered only when scene observations indicate potentially story-relevant events. Memory for Long-Horizon Reasoning. Memory mechanisms have been widely explored for long-horizon multimodal reasoning. Recent systems such as VideoAgent [4], Optimus-1 [20], and M3-Agent [23] maintain explicit tem- poral memories to support extended video reasoning. Other long-video systems employ episodic memories for VideoQA [37], retrieval-augmented memories for long-video comprehension [24], or global audio-visual character representations for dense video description [11]. These methods primarily target long-video QA, embodied reasoning, video comprehension, or dense video description. In con- trast, StoryTeller is designed for accessibility-oriented audio description, where generated descriptions are intended to complement the original movie audio rather than summarize the entire film. Instead of maintaining a general-purpose memory, StoryTeller stores a lightweight narrative memory of visually verified story facts with evolving salience weights, allowing important characters and events to persist across scenes while transient details gradually fade. 3 StoryTeller Figure 2 illustrates the overall pipeline of StoryTeller. The input is a movie divided into chronological clips v 1 ,...,v T . StoryTeller processes the clips se- quentially while carrying a persistent narrative state that summarizes story in- formation accumulated from previously processed clips. This state enables the system to track recurring characters, extract scene-level events, and preserve long-range narrative context throughout the movie. Before processing clip v t , the system maintains the narrative state S tâ1 = (G tâ1 ,M tâ1 ),(1) where G tâ1 is the identity graph storing visual representations of recurring characters, and M tâ1 is the narrative memory storing verified story facts to- gether with their salience weights. Processing clip v t updates the narrative state according to S t = Ί(v t ,S tâ1 ),(2) where Ί denotes the overall state update operator. For each new clip, StoryTeller updates its narrative state through three mod- ules. First, the identity graph update links visible faces to existing character nodes when there is sufficient visual evidence, allowing the system to maintain consistent references to recurring characters across clips. Second, grounded fact induction proposes structured facts about the current scene, optionally using public movie metadata as hints for names or story context. These candidate facts are then filtered to ensure that they are semantically relevant to the scene 6S. H. Hahm et al. Movie-Scoped Retrieval Scene Representation Movie Metadata Candidate Facts SubjectActionObject verify Movie Context Fact Saliency Weight describe Face Tracklets Assign or Update Identity VLM Reinforce or Decay Saliency Insert Context REJECT ! Narrative Memory âł ! Identity Graph " ! Narrative State ! ! Clip # ! ! " Character Representation Character ID c $ summarize ACCEPT Fig. 2: Overview of StoryTeller. Before processing clip v t , the system maintains the narrative state S tâ1 = (G tâ1 ,M tâ1 ), where G tâ1 is the identity graph and M tâ1 is the narrative memory. Given v t , StoryTeller (1) updates the identity graph to main- tain character continuity, (2) extracts structured factual candidates using the updated identity graph, optional public movie metadata, and the previous narrative memory, and (3) updates the narrative memory using semantic filtering, VLM verification, and reinforcementâdecay salience dynamics. Public movie metadata may suggest character names or plot context, but only visually verified facts are stored in narrative memory. and supported by the video. Third, the salience-based narrative memory updates the importance of stored facts: facts that remain relevant are reinforced, while facts that no longer matter gradually decay. This allows important characters, events, and relationships to persist across scenes without carrying forward every transient detail. Besides the visual clips, the movie title is used to retrieve optional public movie metadata. The retrieved metadata may suggest character names or plot context, but it cannot by itself create a memory entry or directly influence the generated narration. A fact is added to the narrative memory only if it is sup- ported by the current clip and accepted by the semantic filtering and VLM verification steps. 3.1 Identity Graph Update The identity graph G tâ1 stores character representations accumulated from pre- viously processed clips. Its role is to associate recurring appearances of the same person across clips, fostering consistent character references throughout the movie. Unlike conventional character banks in the AD literature [10,38], which store static character representations, our identity graph is updated continuously as new visual observations and verified narrative evidence become available. This allows recurring characters to be tracked consistently across scenes, while novel character identities can be established progressively as sufficient supporting ev- idence accumulates. Tracklets and Embeddings. For each clip v t , faces are detected in each frame and associated into short temporal trackletsT k , where each tracklet represents a single face observed over consecutive frames. A pretrained ArcFace encoder Ï [3] extracts an embedding v t,i = Ï(x t,i ) for each face crop x t,i . The embedding of StoryTeller7 tracklet T k is obtained by averaging the embeddings of all face observations in the tracklet: e k = 1 |T k | X iâT k v t,i . (3) Identity Graph Representation. The identity graph at time t is defined as G t =(c j ,ÎŒ j ) N t j=1 , (4) where c j denotes a character identity andÎŒ j is the mean embedding of all tracklets assigned to that character. The graph therefore maintains a compact representation of characters observed up to time t. Identity Assignment. For each new tracklet embedding e k , we compute its cosine similarity to every existing identity embedding: j â = arg max j cos(e k ,ÎŒ j ).(5) If cos(e k ,ÎŒ j â ) > Ï, the tracklet is assigned to identity c j â . Otherwise, a new identity node is created and the tracklet remains unnamed until sufficient evidence is accumulated to associate it with a character. Identity Reseeding. Characters often appear before their names are revealed. Accordingly, every face tracklet is initially represented as an anonymous iden- tity. The VLM may consult optional public movie metadata when proposing a character name, but the current clip must provide sufficient visual and narrative evidence before the association is accepted. Once verified, the corresponding anonymous identity is assigned the character name, and the identity graph is updated by creating a new named node or updating an existing one. Previously observed or future tracklets that are visually similar can then be associated with the same identity. Thus, the similarity threshold Ï serves only as a candidate matching criterion rather than proof of identity. Tracklets that lack sufficient vi- sual or narrative evidence remain anonymous until additional evidence becomes available. 3.2 Grounded Fact Induction Given the current clip v t , the updated identity graph G t containing recurring character identities, and the narrative memoryM tâ1 summarizing previously ac- cumulated story information, StoryTeller then inducts facts, i.e., extracts struc- tured factual candidates and filters them before they are inserted into memory. The objective is to identify facts that are both visually supported and consistent with the evolving narrative context to achieve a new narrative memory M t . Scene Representation. We first summarize the visual content of the clip. A VLM generates a concise scene summary s t describing the primary event in v t . The summary is then embedded as s t = Δ(s t ), where Δ(·) denotes the text embedding function. 8S. H. Hahm et al. Public Movie Metadata Retrieval. Before fact induction, StoryTeller uses the movie title to retrieve optional public movie metadata, including plot sum- maries, character lists, and short character descriptions from IMDb. Each para- graph is embedded, and the paragraphs most similar to the current scene sum- mary s t are retrieved. The retrieved metadata serves only as auxiliary context: it may suggest character names or plot context, but it is never copied into the narration or directly stored in the narrative memory. All stored facts must be supported by the current clip and accepted by the semantic filtering and VLM verification steps. If little or no metadata is available, retrieval simply returns fewer or no passages. The remainder of the pipeline is unchanged: facts are still extracted and verified, the identity graph and narrative memory are updated from visual evidence, and unnamed characters remain anonymous until sufficient evidence supports an identity assignment. Structured Fact Extraction. Given clip v t , the VLM proposes structured factual candidates conditioned on three sources of context: (i) recurring character identities from the identity graph G t , (i) optional public movie metadata, and (i) relevant narrative context retrieved from the memoryM tâ1 . Each candidate is represented as f = (subject, action, object, context), where the subject may correspond to either a named character or an anony- mous identity. The extracted candidates describe potential narrative updates, including character actions, interactions, object states, and scene context. This stage is intentionally permissive: it generates hypotheses rather than committing facts to memory. Candidate facts, including proposed character names inferred from optional public movie metadata, are accepted only after the semantic filtering and VLM verification. Semantic Filtering. The extracted candidates may include facts that are only weakly related to the current clip or inconsistent with the evolving narrative. Be- fore applying computationally expensive VLM verification, we use a lightweight semantic filter to retain candidates that are either semantically related to the current scene or consistent with previously verified facts. Each candidate fact f is embedded as z f = Δ(f). We compute g f = max cos(z f , s t ), max m cos(z f ,Δ(f m )) , (6) where f m denotes the m-th fact stored in the narrative memory M tâ1 . The first term measures semantic similarity to the current scene summary, while the second measures consistency with previously verified facts. Candidate facts with g f > Ï are forwarded to VLM verification. We use a conservative high- recall threshold for Ï to discard only clearly unrelated candidates while retaining plausible facts for subsequent verification. Verification Stage. Candidates that pass semantic filtering are evaluated by a dedicated VLM verifier using the current clip. The verifier outputs ACCEPT or REJECT. Only accepted facts are inserted into the narrative memory or used to StoryTeller9 update named identities. This stage prevents semantic consistency from being mistaken for visual evidence: a candidate may agree with retrieved metadata or previously verified memory, but it cannot update the identity graph or narrative memory unless the video supports it. 3.3 Salience-Based Narrative Memory After verification, accepted facts are inserted into the narrative memory M t . Each memory entry consists of a validated fact and an associated salience weight that determines its influence on future narration. Formally, the memory at time t is M t =(f m ,w (t) m ),(7) where f m is the m-th validated fact and w (t) m is its salience weight at time t. Accepted facts are inserted into the narrative memory with an initial salience weight w (0) m = 0.25. Relevance Computation. For the current scene summary s t , we compute the semantic similarity between the scene embedding s t and each stored fact embedding z m = Δ(f m ): r (t) m = cos(s t , z m ).(8) ReinforcementâDecay Dynamics. The relevance score r (t) m determines how strongly fact f m should be reinforced. Memory weights evolve through reinforce- ment and decay: at each scene, every fact decays slightly, and facts relevant to the current scene are reinforced in proportion to their relevance: w (t) m = λw (tâ1) m + αr (t) m , λâ (0, 1).(9) Here λ is the decay factor, applied to the previous weight, and α is the reinforce- ment strength. A fact that stays relevant across scenes keeps gaining weight and remains salient, while a fact that stops being relevant (r (t) m = 0) loses a fixed fraction of its weight each scene and gradually fades. Supplementary Sec. S5 visualizes persistent and transient memory dynamics. 4 StoryAD-QA: Evaluating Narrative Comprehension Existing captioning metrics such as CIDEr [34], SPICE [1], ROUGE [21], and BLEU [25] measure lexical or semantic similarity between generated and refer- ence descriptions. While effective for evaluating caption quality at the scene level, they do not assess whether a sequence of audio descriptions preserves the infor- mation needed to understand a story. Narrative comprehension requires tracking characters, relationships, events, and causal developments across scenes, yet ADs are essentially concise and may omit important context. As a result, high cap- tioning scores do not necessarily indicate strong story-level understanding. To 10S. H. Hahm et al. address this limitation, we introduce StoryAD-QA, a multiple-choice benchmark that evaluates whether generated audio descriptions retain sufficient information for downstream narrative reasoning. Rather than comparing generated descrip- tions against reference text, StoryAD-QA measures whether questions about the story can be answered using only the generated AD. Main Video Rationale:In the target scene, the man (who was identified in the context as receiving a haircut) is seen walking out onto the sidewalk, where he pulls a cigarette pack from his pocket, lights a cigarette, and stands there smoking it. A. He hails a passing black vehicle. B. He checks the time on his wristwatch. C. He lights and smokes a cigarette. D. He makes a phone call on a mobile device. E. He adjusts his tie while looking in a window. After leaving the building where he was seen receiving a haircut, what does the man do while standing on the sidewalk? Video src:/gpudata3/minh/MovieQA/60s/3034_IDES_OF_MARCH/ segment_0031.mp4 A. He gets into the driver's seat and drives away. B. The rear window rolls down to reveal a man sitting inside. C. He places a package on the hood of the vehicle. D. A police officer exits the vehicle and asks for his identification. E. He uses the car's side mirror to adjust his glasses. After the man in glasses stops to light a cigarette outside a building, what happens when he approaches the black SUV parked on the street? File:3034_IDES_OF_MARCH/segment_0068.mp4 Index:61/74 Video src:/gpudata3/minh/MovieQA/30s/3034_IDES_OF_MARCH/seg ment_0068.mp4 Track A Track B man_1 man_2 Context Video Fig. 3: Example of questions from two tracks in StoryAD-QA. Grey panels show frames from the preceding context window used in Track B for question gener- ation. The context reveals that the man was receiving a haircut inside the building, enabling the question to reference this earlier event. Blue panels show the main clip, where the man stands outside the building and lights a cigarette, which determines the correct answer. During evaluation, only the AD of the main clip is provided, so the description must preserve or reintroduce the earlier context. 4.1 Benchmark Construction StoryAD-QA is constructed from movie clips derived from the MAD-Eval [31] dataset. In total, the benchmark contains 2,574 questions across two evalua- tion tracks: 1,611 in Track A (segment-only QA) and 963 in Track B (context- conditioned QA). See Supplementary Table S6 for additional dataset statistics. Each example in StoryAD-QA consists of a video segment and, following prior work [15,17], a multiple-choice question with five answer options, exactly one of which is correct. Questions are generated automatically using a vision- language model (VLM) that observes the original video along with short movie- level metadata describing the overall story context. The VLM is instructed to focus on clearly observable events or outcomes and to avoid relying on dialogue, external knowledge, or subjective interpretation. This ensures that answers are grounded in visual evidence rather than speculation. All answer choices are ran- domly shuffled before evaluation to reduce positional bias. During evaluation, the answering model receives only the AD text for the evaluated segment and the answer choices; it receives no video, retrieved IMDb/Wikipedia metadata, movie title, or dialogue transcript. StoryTeller11 4.2 Evaluation Protocol To evaluate whether generated AD preserves narrative information beyond indi- vidual frames, StoryAD-QA uses multiple-choice question answering over contin- uous movie segments. Questions are generated from video segments, but during evaluation an answerer (a Language Model) receives only the generated AD text, together with the question and answer choices, without access to the origi- nal video. Accuracy is the fraction of questions answered correctly from the AD text alone. We report two complementary evaluation tracks (See Table 3). Track A: Segment-only QA. In Track A, questions are generated from a single continuous video segment of 30, 60, 120, or 240 seconds. During evaluation, the answering model receives the generated ADs for that same segment and selects the correct option. This track asks whether the AD captures the key events, entities, and relationships within the observed segment. Track B: Context-conditioned QA. In Track B, each question is constructed from a fixed 30-second main clip together with its preceding context. The ques- tion generator observes the main clip plus c â 30, 60, 90 seconds of video im- mediately before it. During evaluation, however, the answerer receives only the AD generated for the 30-second main clip, along with the question and answer choices; it does not receive the preceding context or the original video. This setting tests whether the main-clip AD preserves or reintroduces contextual in- formation from earlier scenes, such as character identities, ongoing actions, and narrative relationships. All questions retained in StoryAD-QA are manually verified before inclusion. Human annotators check that each question is visually grounded, has plausible distractors, and has a correct answer. Low-quality questions are regenerated and re-verified. Additional details are provided in Supplementary Sections S8.2âS8.3. 5 Experiments We evaluate StoryTeller along three dimensions: (i) captioning quality on MAD- Eval, (i) narrative comprehension with the proposed StoryAD-QA benchmark, and (i) the contribution of individual pipeline components through ablation studies. All experiments are conducted without task-specific training, subtitles, scripts, AD transcripts, character banks, or precomputed face identities. 5.1 Experimental Setup Datasets. We evaluate our method on MAD-Eval [31], which consists of 10 movies from the LSMDC [28] dataset. LSMDC contains 202 movies paired with professionally authored audio descriptions aligned to the corresponding video clips. To evaluate long-range narrative understanding, we further introduce and evaluate on StoryAD-QA, a question-answering benchmark constructed from non-overlapping segments of the MAD-Eval movies. Implementation Details. We use Qwen3-VL-2B-Instruct [19] for scene sum- marization and structured fact extraction. We use Qwen3-VL-8B-Thinking for 12S. H. Hahm et al. fact verification, which outputs an ACCEPT/REJECT decision for each candi- date fact. Text embeddings are computed with Qwen3-VL-Embedding-2B [19], and face embeddings are extracted with ArcFace [3]. Gemini3-Flash [6] is used only for StoryAD-QA question generation and answering. For question gener- ation, Gemini3-Flash receives the video and a short Wikipedia overview. For answering, it receives only the AD text, question, and answer choices. IMDb re- trieval is used only by the AD-generation pipeline and is defined in Section 3.2. Additional implementation details are provided in Supplementary Sec. S4, and the full prompts are provided in Supplementary Sec. S10. Threshold Calibration. Identity assignment uses a fixed cosine similarity threshold of Ï = 0.58 for ArcFace embeddings. The semantic grounding threshold is set to Ï = 0.20, calibrated from grounding-score statistics collected on separate calibration runs to provide a high-recall semantic filter that removes only clearly unrelated candidates before VLM verification. Both thresholds are fixed across all experiments and are not tuned on the evaluation data. Additional calibration statistics and threshold analyses are provided in Supplementary Sec. S4.2. Runtime. On one A100-SXM4-80GB GPU, StoryTeller processes a feature- length movie in approximately 3â4 hours. The semantic filter is approximately 3,000Ă faster than VLM verification, substantially reducing unnecessary verifier calls. Memory Parameters. Narrative memory uses the reinforcementâdecay mech- anism described in Sec.3.3. All hyperparameters are fixed across experiments; complete values are provided in Supplementary Table S4.3. 5.2 MAD-Eval Benchmark Comparison Table 1 compares StoryTeller with prior audio description systems while separat- ing required resources from captioning performance. We additionally evaluate a variant without public movie metadata retrieval in the ablation study (Table 2). We additionally evaluate StoryTeller on MovieChat [32] using 9 videos and 27 questions. Results are reported in Supplementary Sec. S7. 5.3 Ablation Study To evaluate the contribution of each component in StoryTeller, we construct controlled ablation variants on MAD-Eval (Table 2). Each variant removes one module while keeping the remaining pipeline unchanged. All experiments use the Qwen3-VL backbone to isolate the contribution of the proposed components. A1 (VLM-only narration). Direct VLM narration without structured fact extraction, identity tracking, verification, or narrative memory. A2 (No schema). Replaces structured facts with scene summaries while pre- serving the sequential pipeline. A3 (No memory). Removes cross-clip narrative memory while retaining per- clip fact extraction and verification. A4 (No identity). Removes persistent identity tracking, leaving facts in generic form without consistent character grounding. StoryTeller13 Table 1: Comparison on MAD-Eval. Train denotes whether task-specific training is required. Curated Res. includes subtitles, scripts, AD transcripts, aligned captions, character banks, or precomputed face identities. Public Meta. means public movie text retrieved by title, such as IMDb plot summaries, character lists, and short character de- scriptions. VLM indicates the underlying vision-language model used by each method. Our method requires neither additional training nor curated or precomputed movie resources; the default setting uses public movie metadata. Method SetupMetrics ModelTrain Curated Res. Public Meta.VLMCIDEr SPICE ROUGE-L AutoAD-I [9]âĂ-14.34.411.9 AutoAD-I [8]âĂ-19.2-13.4 AutoAD-I [10]âĂ-24.0-- M-Vid [22]ĂâGPT-4V6.16.19.8 M-Narrator [40] ĂâGPT 413.95.213.4 AutoAD-Zero [38] ĂâVideoLLaMA2 22.47.314.4 StoryTellerĂâVideoLLaMA2 19.19.016.0 StoryTellerĂâQwen3-VL21.46.715.3 Table 2: Ablation results on MAD-Eval. A1âA4 remove one internal module while keeping the remaining modules unchanged. A5 disables public IMDb metadata only; identity, schema, and memory remain active. VariantPublic Meta. Identity Schema Memory CIDEr SPICE ROUGE-L StoryTeller (Full)â21.4 6.715.3 A1 (VLM only)Ă Ă Ă11.44.39.6 A2 (No schema)â Ăâ15.44.412.5 A3 (No memory)â Ă18.06.013.2 A4 (No identity)âĂâ17.95.612.3 A5 (No IMDb)Ăâ17.24.812.2 A5 (No IMDb metadata). Disables public movie metadata retrieval while retaining identity tracking, structured facts, verification, and narrative memory. Table 2 shows that each proposed component contributes to StoryTeller, with every ablation lowering MAD-Eval scores. Removing identity grounding (A4) leads to relative drops of 16.4% in CIDEr, 16.4% in SPICE, and 19.6% in ROUGE-L, even though public movie metadata remains available. This suggests that external metadata alone is insufficient for maintaining narrative consistency, and that persistent character resolution provides complementary information for coherent storytelling. A5 further isolates IMDb metadata: disabling it lowers CIDEr/SPICE/ROUGE-L by 19.6%/28.4%/20.3%, showing that public meta- data helps but is not the sole driver of performance. 5.4 StoryAD-QA Narrative Evaluation Table 3 reports results on the StoryAD-QA benchmark. In Track A, which evaluates understanding within a single video segment, StoryTeller consistently outperforms AutoAD-Zero across all segment lengths. The gains remain clear 14S. H. Hahm et al. Table 3: StoryAD-QA QA benchmark (Accuracy). (a) Segment-only QA. (b) Context-conditioned QA with c seconds of prior context plus a 30 s main clip; only the 30 s main-clip ADs are provided during evaluation. Reference AD denotes professional human-written AD text. (a) Track A: Segment-only QA Model30 60 120 240 Reference AD 0.953 0.963 0.995 0.981 AutoAD-Zero 0.840 0.830 0.896 0.917 StoryTeller 0.892 0.928 0.953 0.972 (b) Track B: Context-conditioned QA Model30+30 60+30 90+30 Reference AD 0.828 0.825 0.806 AutoAD-Zero 0.662 0.654 0.673 StoryTeller 0.754 0.716 0.788 for longer segments, suggesting that narrative-aware descriptions better capture story-level information beyond local clip content. Track B evaluates contextual reasoning by generating questions using preced- ing context while providing only the ADs of the 30-second main clip during eval- uation. Across all settings, StoryTeller achieves higher accuracy than AutoAD- Zero, with the largest improvement of +11.5 points in the 90+30 setting. These results indicate that the generated narration more effectively preserves contex- tual cues such as character identities and ongoing events, enabling the answering model to recover information originating from earlier clips. The âReference ADâ row uses the professional AD text supplied with MAD- Eval as the answererâs input. Because professional AD is scene-limited and may omit details needed for some generated questions, its accuracy is not necessarily 100%. Table 4 presents component ablations on StoryAD-QA. Several simplified variants remain competitive on short or local settings, but the complete sys- tem achieves the best performance on the longest and most context-dependent settings: 240s in Track A and 90+30 in Track B. Removing narrative memory (A3), the identity graph (A4), or IMDb metadata (A5) reduces performance in these long-range settings, highlighting the role of each component. Notably, even without IMDb metadata, A5 remains competitive with or above AutoAD-Zero across all settings, indicating that the gains are not solely attributable to public movie information. 5.5 Qualitative Results Figure 4 compares professional reference AD, AutoAD-Zero, and StoryTeller, il- lustrating differences in identity consistency and narrative recall across clips. Additional qualitative examples are provided in Supplementary Fig. S1 and Secs. S2âS3; human evaluation and preference results are reported in Sec. S9. 6 Conclusion We introduced StoryTeller, a training-free framework for long-form audio de- scription that uses narrative memory to capture evolving story context across StoryTeller15 Table 4: Component ablations on StoryAD-QA accuracy. A1: VLM only; A2: no structured schema; A3: no narrative memory; A4: no identity graph; A5: no IMDb metadata. Track ATrack B Variant30s60s120s240s30+30 60+30 90+30 Full0.892 0.928 0.953 0.972 0.7540.716 0.788 A10.9080.9040.9280.9110.747 0.7500.751 A20.9050.9090.9320.931 0.7960.7090.751 A30.8990.9070.9180.8910.7560.7120.779 A40.909 0.9040.9080.9210.794 0.7500.774 A50.8920.9090.9080.9010.7610.7160.770 Fig. 4: Qualitative comparison of StoryTeller, AutoAD-Zero [38], and professional reference AD. (GT). scenes. By maintaining an identity graph and structured narrative memory, the system generates coherent descriptions without task-specific training or curated movie resources; public movie metadata can suggest context, but only video- supported facts enter memory. We also introduced StoryAD-QA, a benchmark for evaluating whether generated AD supports narrative understanding. Exper- iments show that narrative-aware modeling improves contextual grounding and story comprehension in long-form video description. Acknowledgments This work was supported by startup funds provided by Dartmouth College. The authors also acknowledge support from the National Science Foundation under CAREER Award No. 2541968. 16S. H. Hahm et al. References 1. Anderson, P., Fernando, B., Johnson, M., Gould, S.: Spice: Semantic propositional image caption evaluation. In: European Conference on Computer Vision (ECCV). Springer (2016) 2. Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., Zhu, Y., Zhang, W., Luo, Z., Zhao, D., Bing, L.: Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms (2024) 3. Deng, J., Guo, J., Xue, N., Zafeiriou, S.: Arcface: Additive angular margin loss for deep face recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 4690â4699 (2019) 4. Fan, Y., Ma, X., Wu, R., Du, Y., Li, J., Gao, Z., Li, Q.: Videoagent: A memory- augmented multimodal agent for video understanding. In: European Conference on Computer Vision (ECCV). p. 75â92. Springer (2024) 5. Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., Wang, H.: Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 (2023) 6. Google: Gemini 3 Flash Preview [large language model] (2026), https://ai. google.dev/gemini-api/docs/models/gemini-3-flash-preview, accessed: June 29, 2026 7. Guu, K., Lee, K., Tung, Z., Pasupat, P., Chang, M.: Retrieval augmented language model pre-training. In: Proceedings of the International Conference on Machine Learning (ICML). vol. 119, p. 3929â3938. PMLR (2020), https://proceedings. mlr.press/v119/guu20a.html 8. Han, T., Bain, M., Nagrani, A., Varol, G., Xie, W., Zisserman, A.: AutoAD I: The Sequel - who, when, and what in movie audio description. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2023) 9. Han, T., Bain, M., Nagrani, A., Varol, G., Xie, W., Zisserman, A.: AutoAD: Movie description in context. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023) 10. Han, T., Bain, M., Nagrani, A., Varol, G., Xie, W., Zisserman, A.: Autoad i: The prequel - back to the pixels. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 18164â18174 (June 2024) 11. He, Y., Lin, Y., Wu, J., Zhang, H., Zhang, Y., Le, R.: Storyteller: Improving long video description through global audio-visual character identification. arXiv preprint arXiv:2411.07076 (2024). https://doi.org/10.48550/arXiv.2411.07076 12. Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., Liu, T.: A survey on hallucination in large language mod- els: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 1(1) (January 2024). https://doi.org/10.1145/3703155 13. Huang, Q., Xiong, Y., Rao, A., Wang, J., Lin, D.: Movienet: A holistic dataset for movie understanding. In: European Conference on Computer Vision (ECCV) (2020) 14. Izacard, G., Grave, E.: Leveraging passage retrieval with generative models for open domain question answering. In: Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics (EACL). p. 874â880 (2021) 15. Kala, D., Khandelwal, E., Tapaswi, M.: What you see is what you ask: Evaluating audio descriptions. In: Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V. (eds.) Proceedings of the Conference on Empirical Methods in Natural Language StoryTeller17 Processing (EMNLP). p. 23496â23518. Association for Computational Linguis- tics, Suzhou, China (Nov 2025). https://doi.org/10.18653/v1/2025.emnlp- main.1199, https://aclanthology.org/2025.emnlp-main.1199/ 16. Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., Yih, W.t.: Dense passage retrieval for open-domain question answering. In: Pro- ceedings of the Conference on Empirical Methods in Natural Language Process- ing (EMNLP). p. 6769â6781. Association for Computational Linguistics (2020). https://doi.org/10.18653/v1/2020.emnlp-main.550, https://aclanthology. org/2020.emnlp-main.550/ 17. Lei, J., Yu, L., Bansal, M., Berg, T.: Tvqa: Localized, compositional video question answering. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). p. 1369â1379 (2018) 18. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., KĂŒttler, H., Lewis, M., Yih, W.t., RocktĂ€schel, T., Riedel, S., Kiela, D.: Retrieval-augmented generation for knowledge-intensive NLP tasks. In: Advances in Neural Information Processing Systems (NeurIPS). vol. 33, p. 9459â9474 (2020) 19. Li, M., Zhang, Y., Long, D., Chen, K., Song, S., Bai, S., Yang, Z., Xie, P., Yang, A., Liu, D., Zhou, J., Lin, J.: Qwen3-vl-embedding and qwen3-vl-reranker: A unified framework for state-of-the-art multimodal retrieval and ranking. arXiv preprint arXiv:2601.04720 (2026) 20. Li, Z., Xie, Y., Shao, R., Chen, G., Jiang, D., Nie, L.: Optimus-1: Hybrid multi- modal memory empowered agents excel in long-horizon tasks. Advances in Neural Information Processing Systems (NeurIPS) 37, 49881â49913 (2024) 21. Lin, C.Y.: ROUGE: A package for automatic evaluation of summaries. In: Text Summarization Branches Out. p. 74â81. Association for Computational Linguis- tics, Barcelona, Spain (Jul 2004), https://aclanthology.org/W04-1013/ 22. Lin, K., Ahmed, F., Li, L., Lin, C.C., Azarnasab, E., Yang, Z., Wang, J., Liang, L., Liu, Z., Lu, Y., Liu, C., Wang, L.: M-VID: Advancing video understanding with gpt-4v(ision). arXiv preprint arXiv:2310.19773 (2023) 23. Long, L., He, Y., Ye, W., Pan, Y., Lin, Y., Li, H., Zhao, J., Li, W.: Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory. arXiv preprint arXiv:2508.09736 (2025). https://doi.org/10.48550/arXiv.2508.09736 24. Luo, Y., Zheng, X., Li, G., Yin, S., Lin, H., Fu, C., Huang, J., Ji, J., Chao, F., Luo, J., Ji, R.: Video-rag: Visually-aligned retrieval-augmented long video com- prehension. Advances in Neural Information Processing Systems (NeurIPS) 38, 168008â168033 (2026) 25. Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: Bleu: a method for automatic evaluation of machine translation. In: Isabelle, P., Charniak, E., Lin, D. (eds.) Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. p. 311â318. Association for Computational Linguistics, Philadelphia, Pennsylvania, USA (Jul 2002). https://doi.org/10.3115/1073083.1073135, https://aclanthology.org/P02-1040/ 26. Park, J., Ye, J., Lee, S., Ka, H.W., Han, D.: Narrad: Automatic generation of audio descriptions for movies with rich narrative context. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). p. 409â419. IEEE (2025) 27. Raajesh, H., Desanur, N.R., Khan, Z., Tapaswi, M.: Micap: A unified model for identity-aware movie descriptions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024) 18S. H. Hahm et al. 28. Rohrbach, A., Torabi, A., Rohrbach, M., Tandon, N., Pal, C., Larochelle, H., Courville, A., Schiele, B.: Movie description. International Journal of Computer Vision (IJCV) 123(1), 94â120 (2017). https://doi.org/10.1007/s11263-016- 0987-1 29. Shao, H., Qian, S., Xiao, H., Song, G., Zong, Z., Wang, L., Liu, Y., Li, H.: Vi- sual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning. Advances in Neural Information Processing Systems (NeurIPS) 37, 8612â8642 (2024) 30. Singh, A., Fry, A., Perelman, A., Tart, A., Ganesh, A., El-Kishky, A., McLaugh- lin, A., Low, A., Ostrow, A., Ananthram, A., Nathan, A., Luo, A., Helyar, A., Madry, A., Efremov, A., Spyra, A., Baker-Whitcomb, A., Beutel, A., Karpenko, A., Makelov, A., Neitz, A., Wei, A., Barr, A., Kirchmeyer, A., Ivanov, A., Chris- takis, A., Gillespie, A., Tam, A., Bennett, A., Wan, A., Huang, A., Sandjideh, A.M., Yang, A., Kumar, A., Saraiva, A., Vallone, A., Gheorghe, A., Garcia, A.G., Braunstein, A., Liu, A., Schmidt, A., Mereskin, A., Mishchenko, A., Applebaum, A., Rogerson, A., Rajan, A., Wei, A., Kotha, A., Srivastava, A., Agrawal, A., Vi- jayvergiya, A., Tyra, A., Nair, A., Nayak, A., Eggers, B., Ji, B., Hoover, B., Chen, B., Chen, B., Barak, B., Minaiev, B., Hao, B., Baker, B., Lightcap, B., McKinzie, B., Wang, B., Quinn, B., Fioca, B., Hsu, B., Yang, B., Yu, B., Zhang, B., Brenner, B., Zetino, C.R., Raymond, C., Lugaresi, C., Paz, C., Hudson, C., Whitney, C., Li, C., Chen, C., Cole, C., Voss, C., Ding, C., Shen, C., Huang, C., Colby, C., Hallacy, C., Koch, C., Lu, C., Kaplan, C., Kim, C., Minott-Henriques, C., Frey, C., Yu, C., Czarnecki, C., Reid, C., Wei, C., Decareaux, C., Scheau, C., Zhang, C., Forbes, C., Tang, D., Goldberg, D., Roberts, D., Palmie, D., Kappler, D., Levine, D., Wright, D., Leo, D., Lin, D., Robinson, D., Grabb, D., Chen, D., Lim, D., Salama, D., Bhattacharjee, D., Tsipras, D., Li, D., Yu, D., Strouse, D., Williams, D., Hunn, D., Bayes, E., Arbus, E., Akyurek, E., Le, E.Y., Widmann, E., Yani, E., Proehl, E., Sert, E., Cheung, E., Schwartz, E., Han, E., Jiang, E., Mitchell, E., Sigler, E., Wallace, E., Ritter, E., Kavanaugh, E., Mays, E., Nikishin, E., Li, F., Such, F.P., de Avila Belbute Peres, F., Raso, F., Bekerman, F., Tsimpourlas, F., Chantzis, F., Song, F., Zhang, F., Raila, G., McGrath, G., Briggs, G., Yang, G., Parascandolo, G., Chabot, G., Kim, G., Zhao, G., Valiant, G., Leclerc, G., Salman, H., Wang, H., Sheng, H., Jiang, H., Wang, H., Jin, H., Sikchi, H., Schmidt, H., Aspegren, H., Chen, H., Qiu, H., Lightman, H., Covert, I., Kivlichan, I., Silber, I., Sohl, I., Hammoud, I., Clavera, I., Lan, I., Akkaya, I., Kostrikov, I., Kofman, I., Etinger, I., Singal, I., Hehir, J., Huh, J., Pan, J., Wilczynski, J., Pachocki, J., Lee, J., Quinn, J., Kiros, J., Kalra, J., Samaroo, J., Wang, J., Wolfe, J., Chen, J., Wang, J., Harb, J., Han, J., Wang, J., Zhao, J., Chen, J., Yang, J., Tworek, J., Chand, J., Landon, J., Liang, J., Lin, J., Liu, J., Wang, J., Tang, J., Yin, J., Jang, J., Morris, J., Flynn, J., Ferstad, J., Heidecke, J., Fishbein, J., Hallman, J., Grant, J., Chien, J., Gordon, J., Park, J., Liss, J., Kraaijeveld, J., Guay, J., Mo, J., Lawson, J., McGrath, J., Vendrow, J., Jiao, J., Lee, J., Steele, J., Wang, J., Mao, J., Chen, K., Hayashi, K., Xiao, K., Salahi, K., Wu, K., Sekhri, K., Sharma, K., Singhal, K., Li, K., Nguyen, K., Gu-Lemberg, K., King, K., Liu, K., Stone, K., Yu, K., Ying, K., Georgiev, K., Lim, K., Tirumala, K., Miller, K., Ahmad, L., Lv, L., Clare, L., Fauconnet, L., Itow, L., Yang, L., Romaniuk, L., Anise, L., Byron, L., Pathak, L., Maksin, L., Lo, L., Ho, L., Jing, L., Wu, L., Xiong, L., Mamitsuka, L., Yang, L., McCallum, L., Held, L., Bourgeois, L., Engstrom, L., Kuhn, L., Feuvrier, L., Zhang, L., Switzer, L., Kondraciuk, L., Kaiser, L., Joglekar, M., Singh, M., Shah, M., Stratta, M., Williams, M., Chen, M., Sun, M., Cayton, M., Li, M., Zhang, M., Aljubeh, M., StoryTeller19 Nichols, M., Haines, M., Schwarzer, M., Gupta, M., Shah, M., Guan, M.Y., Huang, M., Dong, M., Wang, M., Glaese, M., Carroll, M., Lampe, M., Malek, M., Shar- man, M., Zhang, M., Wang, M., Pokrass, M., Florian, M., Pavlov, M., Wang, M., Chen, M., Wang, M., Feng, M., Bavarian, M., Lin, M., Abdool, M., Rohaninejad, M., Soto, N., Staudacher, N., LaFontaine, N., Marwell, N., Liu, N., Preston, N., Turley, N., Ansman, N., Blades, N., Pancha, N., Mikhaylin, N., Felix, N., Handa, N., Rai, N., Keskar, N., Brown, N., Nachum, O., Boiko, O., Murk, O., Watkins, O., Gleeson, O., Mishkin, P., Lesiewicz, P., Baltescu, P., Belov, P., Zhokhov, P., Pronin, P., Guo, P., Thacker, P., Liu, Q., Yuan, Q., Liu, Q., Dias, R., Puckett, R., Arora, R., Mullapudi, R.T., Gaon, R., Miyara, R., Song, R., Aggarwal, R., Marsan, R., Yemiru, R., Xiong, R., Kshirsagar, R., Nuttall, R., Tsiupa, R., Eldan, R., Wang, R., James, R., Ziv, R., Shu, R., Nigmatullin, R., Jain, S., Talaie, S., Altman, S., Arnesen, S., Toizer, S., Toyer, S., Miserendino, S., Agarwal, S., Yoo, S., Heon, S., Ethersmith, S., Grove, S., Taylor, S., Bubeck, S., Banesiu, S., Amdo, S., Zhao, S., Wu, S., Santurkar, S., Zhao, S., Chaudhuri, S.R., Krishnaswamy, S., Shuaiqi, Xia, Cheng, S., Anadkat, S., Fishman, S.P., Tobin, S., Fu, S., Jain, S., Mei, S., Egoian, S., Kim, S., Golden, S., Mah, S., Lin, S., Imm, S., Sharpe, S., Yadlowsky, S., Choudhry, S., Eum, S., Sanjeev, S., Khan, T., Stramer, T., Wang, T., Xin, T., Gogineni, T., Christianson, T., Sanders, T., Patwardhan, T., Degry, T., Shadwell, T., Fu, T., Gao, T., Garipov, T., Sriskandarajah, T., Sherbakov, T., Korbak, T., Kaftan, T., Hiratsuka, T., Wang, T., Song, T., Zhao, T., Peterson, T., Kharitonov, V., Chernova, V., Kosaraju, V., Kuo, V., Pong, V., Verma, V., Petrov, V., Jiang, W., Zhang, W., Zhou, W., Xie, W., Zhan, W., McCabe, W., DePue, W., Ellsworth, W., Bain, W., Thompson, W., Chen, X., Qi, X., Xiang, X., Shi, X., Dubois, Y., Yu, Y., Khakbaz, Y., Wu, Y., Qian, Y., Lee, Y.T., Chen, Y., Zhang, Y., Xiong, Y., Tian, Y., Cha, Y., Bai, Y., Yang, Y., Yuan, Y., Li, Y., Zhang, Y., Yang, Y., Jin, Y., Jiang, Y., Wang, Y., Wang, Y., Liu, Y., Stubenvoll, Z., Dou, Z., Wu, Z., Wang, Z.: Openai gpt-5 system card. arXiv preprint arXiv:2601.03267 (2025) 31. Soldan, M., Pardo, A., AlcĂĄzar, J.L., Caba, F., Zhao, C., Giancola, S., Ghanem, B.: Mad: A scalable dataset for language grounding in videos from movie audio descriptions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 5026â5035 (June 2022) 32. Song, E., Chai, W., Wang, G., Zhang, Y., Zhou, H., Wu, F., Chi, H., Guo, X., Ye, T., Zhang, Y., Lu, Y., Hwang, J.N., Wang, G.: Moviechat: From dense token to sparse memory for long video understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 18221â 18232 (June 2024) 33. Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: Movieqa: Understanding stories in movies through question-answering. In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR). p. 4631â4640 (2016) 34. Vedantam, R., Zitnick, C.L., Parikh, D.: CIDEr: Consensus-based image descrip- tion evaluation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 4566â4575 (2015) 35. Vicol, P., Tapaswi, M., Castrejon, L., Fidler, S.: Moviegraphs: Towards understand- ing human-centric situations from videos. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2018). https://doi.org/10.1109/ CVPR.2018.00895 20S. H. Hahm et al. 36. Wang, W., Gao, Z., Gu, L., Pu, H., Cui, L., Wei, X., Liu, Z., Jing, L., Ye, S., Shao, J., Wang, Z., Chen, Z., Zhang, H., Yang, G., Wang, H., Wei, Q., Yin, J., Li, W., Cui, E., Chen, G., Ding, Z., Tian, C., Wu, Z., Xie, J., Li, Z., Yang, B., Duan, Y., Wang, X., Hou, Z., Hao, H., Zhang, T., Li, S., Zhao, X., Duan, H., Deng, N., Fu, B., He, Y., Wang, Y., He, C., Shi, B., He, J., Xiong, Y., Lv, H., Wu, L., Shao, W., Zhang, K., Deng, H., Qi, B., Ge, J., Guo, Q., Zhang, W., Zhang, S., Cao, M., Lin, J., Tang, K., Gao, J., Huang, H., Gu, Y., Lyu, C., Tang, H., Wang, R., Lv, H., Ouyang, W., Wang, L., Dou, M., Zhu, X., Lu, T., Lin, D., Dai, J., Su, W., Zhou, B., Chen, K., Qiao, Y., Wang, W., Luo, G.: Internvl3. 5: Advancing open- source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265 (2025) 37. Wang, Y., Zhang, L., Liu, J., Yan, J., Zhang, Z., Zheng, J., Ma, A., Ling, R., Yang, X., Wu, D., Chen, X., Li, X.: Video-em: Event-centric episodic memory for long-form video understanding. arXiv preprint arXiv:2508.09486 (2025). https: //doi.org/10.48550/arXiv.2508.09486 38. Xie, J., Han, T., Bain, M., Nagrani, A., Varol, G., Xie, W., Zisserman, A.: AutoAD- Zero: A training-free framework for zero-shot audio description. In: Proceedings of the Asian Conference on Computer Vision (ACCV). p. 2265â2281 (2024) 39. Yue, Z., Zhang, Q., Hu, A., Zhang, L., Wang, Z., Jin, Q.: Movie101: A new movie understanding benchmark. In: Proceedings of the Association for Computational Linguistics (ACL). p. 4669â4684 (2023) 40. Zhang, C., Lin, K., Yang, Z., Wang, J., Li, L., Lin, C.C., Liu, Z., Wang, L.: Mm- narrator: Narrating long-form videos with multimodal in-context learning. In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion (CVPR). p. 13647â13657 (2024) StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description âSupplementary Materialâ Seung Hyun Hahm, Minh T. Dinh, and SouYoung Jin Dartmouth College S1 Overview This supplementary document provides additional details and experimental re- sults supporting the main paper. In particular, we include: â Additional qualitative comparisons between StoryTeller, AutoAD-Zero, and professional reference AD. â Failure cases of baseline methods and discussion of how StoryTeller improves narrative consistency. â Analysis of narrative memory dynamics, including visualization of memory weights over time. â A backbone comparison explaining how the VideoLLaMA2 and Qwen3-VL configurations differ across MAD-Eval and StoryAD-QA results. â Further details and validation of the StoryAD-QA benchmark. â Human evaluation of generated audio descriptions. â Implementation details and hyperparameters used in the proposed frame- work. â Full prompt templates used in the StoryTeller pipeline. These materials aim to improve transparency and reproducibility of the pro- posed framework. S2 Additional Qualitative Results This section supports the qualitative analysis in Section 5.5 of the main paper by providing more examples beyond the main-paper figure. Figure S1 presents additional qualitative comparisons between StoryTeller, AutoAD-Zero, and professional reference AD across four movies: Harry Pot- ter and the Goblet of Fire, The Ides of March, Roommate, and Legion. Each example shows three representative frames together with the generated audio descriptions. Across these examples, StoryTeller generally produces descriptions that in- clude more concrete scene context visible in the frames, while AutoAD-Zero often focuses on shorter action-oriented phrases or more generic descriptions. The differences can be observed directly in the examples. 2S. H. Hahm et al. Fig. S1: Additional qualitative comparisons between StoryTeller, AutoAD-Zero, and professional reference AD across four movies. StoryTeller captures broader narrative context and scene semantics, while AutoAD-Zero often focuses on isolated actions or produces generic descriptions. In Harry Potter and the Goblet of Fire, the frames show Harry fighting with the chained dragon. StoryTeller describes Harry leaping over a rocky cliff while being pursued by the dragon, capturing the spatial interaction between the char- acter and the environment. AutoAD-Zero instead focuses on Harry casting a spell with his wand. Both descriptions refer to the same moment, but they emphasize different aspects of the scene. The example from The Ides of March illustrates a scene with visible campaign signage and buildings in the background. StoryTeller describes the police vehicle and the political signs visible in the frames. In contrast, AutoAD-Zero produces a more generic description about officers walking toward a building, which does not explicitly mention the visual text present in the scene. In the Roommate example, the first frame shows a character lying in bed, followed by outdoor campus scenes. StoryTeller describes both the indoor and outdoor settings shown across the frames. AutoAD-Zero produces a shorter de- scription focused mainly on the sleeping character and nearby people walking. Finally, the Legion example shows a car driving along a desert road while a dust storm forms in the distance. StoryTeller describes the car moving toward the storm and also mentions a small insect visible on the car vent in the final frame. AutoAD-Zero instead produces a more general description of the car driving and a person looking toward a dust cloud. These examples highlight how the generated descriptions differ in content emphasis across methods. StoryTeller tends to include environmental details visible in the frames, while AutoAD-Zero often produces shorter descriptions centered on basic actions. StoryTeller3 S3 Comparative Error Analysis Fig. S2: Qualitative comparison of audio descriptions. Examples comparing StoryTeller, AutoAD-Zero, and professional reference AD across several movies. Text highlighted in blue corresponds to information that is visually grounded in the frames and appears in both StoryTeller and the professional reference AD. Text highlighted in red marks content that is less consistent with the visual evidence in the scene. These examples illustrate common challenges in automatic audio description, including grounding textual cues, interpreting scene entities, and maintaining character identity across scenes. Figure S2 presents representative qualitative examples comparing Story- Teller, AutoAD-Zero [4], and professional reference AD across several movies. These examples highlight common challenges in automatic audio description, including difficulties in grounding textual cues, identifying scene entities, and maintaining character identity across scenes. In the figure, text highlighted in blue corresponds to information that is con- sistent with the visual frames and appears in both StoryTeller and the profes- sional reference AD. Text highlighted in red marks content that is less consistent with the visual evidence in the scene. Missing textual or symbolic cues. In the The Ides of March example, the frame clearly shows a note containing the message âMeet me in the stairwell at noon.â StoryTeller and the professional reference AD both capture this textual information, while the baseline description instead focuses on a generic action (âHe walksâ), overlooking the written content visible in the scene. Object grounding ambiguity. In the Harry Potter and the Goblet of Fire example, the scene shows a figure approaching through fog behind the tents. StoryTeller and the professional reference AD describe the approaching figure, whereas the baseline description refers to a âred circle,â which does not clearly correspond to any object visible in the scene. Scene interpretation differences. The Battle: Los Angeles example illus- trates a case where the baseline description interprets the television display as people interacting with computer screens. In contrast, StoryTeller and the pro- fessional reference AD identify the meteor shower shown on the display. Character identity ambiguity. In the final example, the scene shows a char- acter writing a note. StoryTeller correctly identifies the character as Jason and describes the action accordingly, while the baseline description refers to the char- acter as âshe,â which loses the identity information present in the narrative con- text. 4S. H. Hahm et al. These examples highlight common challenges in grounding audio descrip- tions to visual evidence and maintaining narrative consistency across scenes. By incorporating structured narrative facts and an evolving memory mechanism, StoryTeller is able to produce descriptions that better preserve scene context and character identity. S4 Implementation Details S4.1 Model Configuration We use the following models in the StoryTeller pipeline: â VLM for scene summarization and fact extraction: Qwen3-VL-2B- Instruct â VLM for fact verification: Qwen3-VL-8B-Thinking â Text embedding model: Qwen3-VL-Embedding-2B â VLM for StoryAD-QA generation and answering: Gemini3-Flash for initial question generation and answering; Gemini 3.5 Flash is used only to regenerate questions that fail the first manual verification pass. â Face recognition encoder: ArcFace [1] Within each reported configuration, the same generation backbone is used consistently unless a table explicitly compares backbones. Scene-level observa- tions are first summarized into structured facts, which are then verified and stored in the narrative memory. During generation, StoryTeller retrieves the most relevant verified facts and identity-consistent context to produce the fi- nal audio description. The AD pipeline may also retrieve IMDb plot summaries, character lists, and short character descriptions using the movie title. These texts can suggest possible names or plot context, but are not copied into narration or treated as professional reference AD; accepted facts must be supported by the clip. Except for Gemini, all models are executed locally on a cluster comprising 8 NVIDIA A100-SXM4-80GB GPUs and 8 NVIDIA RTX 6000 Ada Generation GPUs. On one allocated A100-SXM4-80GB GPU, generating audio descriptions for a single film requires approximately 3â4 hours, corresponding to roughly 650 clips per movie. Across the 10 full movie runs, the pipeline processes 6,520 clips and 13,974 candidate facts; verification takes 3.68± 1.73 seconds per fact block, uses 473K tokens in total, and accepts 94.7% of candidate facts. The coarse semantic gate is approximately 3,000Ă faster than VLM verification. Generating StoryAD-QA questions for the 10-movie evaluation set takes approximately 10â 15 minutes per track. In practice, generating the complete StoryAD-QA benchmark (2,574 ques- tions across 10 movies) incurs an API cost of under $20, while evaluating the complete benchmark costs less than $5. These values are approximate and may vary depending on API pricing and the modelâs internal reasoning-token usage. StoryTeller5 Table S1: Hyperparameters used in StoryTeller. ParameterValue Identity threshold Ï0.58 Grounding threshold Ï0.20 Reinforcement strength α 0.35 Decay factor λ0.92 Initial memory weight w (0) m 0.25 Table S2: Effect of different semantic grounding thresholds. âVerifier calls savedâ is measured relative to Ï = 0.20. ÏCalls savedAccepted facts 0.200.0%100.0% 0.401.4%98.7% 0.456.1%94.3% 0.5019.3%81.6% 0.5543.2%57.9% For StoryAD-QA question generation only, Gemini3-Flash receives the video and a short Wikipedia movie overview. During question answering, it receives only the generated AD text, question, and answer choicesânever the source video, movie title, IMDb or Wikipedia text, subtitles, or other retrieved text. S4.2 Hyperparameters Table S1 summarizes the hyperparameters used in StoryTeller. These values are fixed globally across all experiments and are not tuned on the evaluation datasets. Grounding Threshold (Ï). The semantic grounding threshold determines which candidate facts are forwarded to the computationally expensive VLM verification stage. To calibrate this threshold, we executed the complete Story- Teller pipeline on a separate set of movies and recorded the grounding score of every extracted candidate fact, producing approximately 2,700 candidate facts. Among candidate facts ultimately accepted by the verifier, the minimum observed similarity was approximately 0.21, the 5th percentile was 0.36, and the median was 0.55. Based on these statistics, we selected Ï = 0.20, slightly below the minimum observed similarity, so that the semantic filter functions as a high- recall gate. Its purpose is not to decide factual correctness, but to remove clearly unrelated candidates before VLM verification. Table S2 reports the trade-off between computational cost and recall under alternative threshold values. Higher thresholds reduce the number of verifier calls but also discard visually supported facts. We therefore adopt the conservative setting Ï = 0.20 to prioritize narrative recall over computational savings. 6S. H. Hahm et al. Table S3: Hyperparameter robustness on the test split. Performance remains sta- ble under moderate variation. StoryAD-QA results are reported as accuracy (%) for Track A. The best performance for each metric is bolded. Setting CIDEr StoryAD-QA Accuracy (%) 30s 60s 120s 240s Default21.489.2 92.895.397.2 α -20%20.289.4 94.4 96.297.2 α +20%21.090.5 94.4 97.2 98.2 λ -5%21.989.7 94.295.397.2 λ +5%21.890.1 94.094.896.3 Identity Threshold (Ï). Identity assignment uses a fixed cosine similarity threshold of Ï = 0.58 for ArcFace embeddings. This value was selected empiri- cally to balance missed associations and incorrect identity merges, and is kept fixed across all experiments. Hyperparameter Robustness. Although StoryTeller introduces several global hyperparameters (identity threshold Ï, grounding threshold Ï, reinforcement strength α, and decay rate λ), these are fixed once and not tuned on evaluation data. To assess robustness, we vary α within±20% and λ within±5% while keeping the remaining parameters fixed. Performance on the test split remains stable, with StoryAD-QA accuracy varying by less than 3 percentage points and CIDEr by less than 6 percentage points. Results are summarized in Table S3. S4.3 Memory Overhead of Identity Reseeding Identity reseeding maintains a small backlog of unresolved (anonymous) track- lets until sufficient evidence is available for identity assignment. Each unre- solved tracklet stores a single 512-dimensional ArcFace embedding (approxi- mately 2 KB). Across all movies in our experiments, retaining this anonymous backlog required less than 3.5 MB per movie, indicating that delayed identity assignment introduces negligible memory overhead. S4.4 Limitations and Discussion StoryTeller maintains consistent character references using the identity graph described in Section 3.1 of the main paper which associates face tracklets across clips through clustering and similarity-based assignment. In rare cases, clustering errors may temporarily associate a tracklet with the wrong identity. When this occurs, the generated description may mention an incorrect character name for the corresponding scene. Because StoryTeller maintains a persistent narrative state, such errors can often be corrected when stronger visual or narrative evidence becomes available. StoryTeller7 In particular, the reseeding mechanism described in Section 3.1 of the main paper allows identities to be updated once additional observations clarify the character assignment. However, if incorrect associations persist across several clips, the resulting descriptions may temporarily reduce narrative accuracy. Interestingly, this behavior also highlights the importance of maintaining an explicit narrative memory rather than relying solely on frame-level captioning. By storing and updating structured narrative facts over time, the system can revise earlier assumptions and maintain consistent story context as additional evidence appears. Future work may further improve robustness by integrating stronger identity verification or multimodal reasoning mechanisms. S5 Analysis of Narrative Memory Dynamics To better understand how StoryTeller maintains narrative context across long video sequences, we analyze the evolution of memory weights associated with verified narrative facts. Each fact f m stored in the narrative memory is associated with a salience weight w (t) m that reflects its current narrative importance. When later scenes provide visual or semantic evidence supporting f m , the memory module reinforces its salience by increasing w (t) m . Conversely, if f m is not observed again, its salience gradually decreases through decay. Figure S3 visualizes representative trajectories from three movies: Signs, Bat- tle: Los Angeles, and Harry Potter and the Goblet of Fire. For each movie we plot two facts: a persistent narrative cue (solid line) and a transient observation (dashed line). These trajectories illustrate how StoryTeller selectively retains story-relevant information while allowing short-lived visual details to fade from memory. S5.1 Memory Weight Evolution Figure S3 highlights two characteristic patterns of the memory update mech- anism. Solid curves correspond to persistent narrative facts whose weights increase when similar visual evidence appears in later scenes. Dashed curves cor- respond to transient observations that occur only once and therefore decay when they are not reinforced. Star markers indicate scene indices selected for the qualitative examples shown in Fig. S4. Signs. In Signs, the persistent fact corresponds to the interaction between the children and the family dog near the farmhouse porch. Because similar config- urations recur across several scenes, the memory module repeatedly reinforces this fact, resulting in a steadily increasing weight. In contrast, a transient ob- servation involving Bo grilling food outside appears only briefly and is never revisited, causing the associated weight to decay rapidly. 8S. H. Hahm et al. (a) Signs(b) Battle: Los Angeles(c) Goblet of Fire Fig. S3: Memory-weight trajectories w (t) m for representative facts f m extracted from Signs, Battle: Los Angeles, and Harry Potter and the Goblet of Fire. Solid curves de- note persistent narrative cues that reappear across multiple scenes and therefore receive repeated reinforcement by the memory update mechanism. Dashed curves de- note transient observations that occur only once and gradually decay when they are not reinforced. Star markers indicate the scene indices used in the qualitative case stud- ies shown in Fig. S4, where the corresponding video frames and StoryTeller-generated audio descriptions are displayed. Battle: Los Angeles. For Battle: Los Angeles, the persistent fact captures the squad advancing through a dark, smoke-filled environment while protecting civil- ians. Multiple neighboring scenes depict variations of this tactical situation, lead- ing to repeated reinforcement and a sustained memory weight. By comparison, a brief close-up of a wounded soldier appears only once during the moment and quickly fades from memory. Harry Potter and the Goblet of Fire. The third example contrasts a brief for- est insert with a repeated corridor sequence. The persistent fact corresponds to Harry being escorted through a stone corridor by Moody. Each return to the hall- way configuration reinforces the stored fact and increases its weight. The forest insert, however, functions only as a short atmospheric transition and therefore decays quickly. S5.2 Case Study While the weight trajectories reveal how narrative facts evolve in memory, the qualitative examples in Fig. S4 illustrate how this salience affects the generated audio descriptions. Each panel presents representative frames together with the narration gener- ated by StoryTeller. Words highlighted in red correspond to persistent narrative cues whose memory weights remain high across multiple scenes. Because these facts remain salient, StoryTeller continues referencing the same narrative context even when the camera framing shifts or visual details change. For example, in Signs, the generated description consistently refers to the interaction between the children and the dog despite changes in viewpoint. Sim- ilarly, in Battle: Los Angeles, the narration preserves the context of soldiers ad- vancing through smoke-filled interiors while protecting civilians. In Harry Potter StoryTeller9 Fig. S4: Representative frames corresponding to the facts highlighted in Fig. S3, to- gether with StoryTeller-generated audio descriptions. Persistent narrative cues (high- lighted in red) remain relevant across multiple scenes and are reinforced by the memory mechanism. Transient observations (highlighted in blue) appear only briefly and quickly decay. and the Goblet of Fire, repeated corridor shots reinforce the interaction between Harry and Moody as they move through the dim stone corridor. In contrast, words highlighted in blue correspond to transient visual details that appear only briefly. Because these observations are not reinforced by later scenes, their memory weights decay and they do not influence subsequent de- scriptions. These examples demonstrate the intended behavior of the StoryTeller mem- ory mechanism: facts tied to ongoing narrative context accumulate weight and remain retrievable across scenes, while isolated observations gradually fade. This selective retention enables the system to maintain story-level continuity without overloading the memory with short-lived details. 10S. H. Hahm et al. S6 Backbone Model Comparison We compare StoryTeller with two generation backbones to separate architec- tural effects from backbone choice. The VideoLLaMA2 row corresponds to the StoryAD-QA setting reported in the main narrative-evaluation table, while the Qwen3-VL row corresponds to the default implementation used for the Qwen3- VL MAD-Eval result in the main comparison table. Table S4: Comparison of different VLM backbones. StoryAD-QA results are reported as accuracy (%) for Track A (blue shade) and Track B (orange shade). BackboneCIDErStoryAD-QA Accuracy (%) 30s60s120s240s30+3060+3090+30 Qwen3-VL21.489.292.895.397.275.471.678.8 VideoLLaMA2 19.185.990.292.995.472.770.972.4 S7 Additional Evaluation on MovieChat We additionally evaluate StoryTeller on a subset of the MovieChat-1K test set [3], containing 9 videos and 27 questions. Table S5 compares methods under the input modality appropriate to each setting. MovieChat follows the stan- dard MovieChat setup and accordingly receives the original source video (visual modality) as input. In contrast, AutoAD-Zero and StoryTeller are AD-generation systems, so we evaluate them by replacing the video with their generated AD text and asking the same LLM answerer to answer the MovieChat questions from text alone. This measures how much long-video QA information is preserved in the generated descriptions. For a fair no-curated-resource comparison between AD systems, AutoAD- Zero is evaluated without its character bank, and StoryTeller is evaluated with- out public IMDb metadata. We also report StoryTeller with IMDb metadata to measure the additional benefit of public movie context. As shown in Table S5, StoryTeller outperforms AutoAD-Zero under the same generated-AD input set- ting, remains strong without IMDb metadata, and improves further when public metadata is available. StoryAD-QA remains our primary narrative benchmark because it provides denser coverage of story events, with approximately one question per 30s compared with three questions per movie in this MovieChat subset. S8 StoryAD-QA Benchmark Overview and Validation This section provides additional details about the construction and validation of the StoryAD-QA benchmark introduced in Section 4 of the main paper. We StoryTeller11 Table S5: Additional evaluation on MovieChat The subset contains 9 videos and 27 questions. The subset contains 9 videos and 27 questions. MovieChat â uses the original source video, while AutoAD-Zero and StoryTeller use generated AD text as input to the same LLM answerer. MethodInputAcc.Score MovieChatSource video0.6233.230 Human reference captionShort caption0.0370.185 AutoAD-Zero (no character bank) Generated AD 0.7043.593 StoryTeller (no IMDb)Generated AD 0.7413.778 StoryTeller (with IMDb)Generated AD 0.778 3.852 Table S6: Distribution of the 2,574 questions in StoryAD-QA across evaluation tracks and segment durations. Track A uses only the target segment, while Track B provides a 30, 60, or 90-second preceding context window plus a fixed 30-second target segment. TrackSegment Duration # Questions Track A (segment-only QA)30s860 60s430 120s212 240s109 Track A Total1,611 Track B (context-conditioned QA) 30s context + 30s449 60s context + 30s294 90s context + 30s220 Track B Total963 summarize the distribution of question types and segment durations, describe the manual verification protocol used for all retained questions, and report validation rates and annotation consistency. S8.1 Benchmark Overview and Statistics Table S6 summarizes the distribution of the final retained StoryAD-QA questions across the two evaluation tracks. Track A evaluates models using only the target segment, while Track B requires additional narrative context from preceding segments. To capture events unfolding across multiple shots, we construct segments by grouping consecutive clips into temporal windows of varying duration. In Track A, questions are associated with segments ranging from 30 to 240 sec- onds. In Track B, each question is paired with a 30-second target segment and an additional context window preceding it, encouraging models to reason over narrative dependencies across segments. 12S. H. Hahm et al. Fig. S5: Annotation interface used for human evaluation. Annotators watch the video segment, review the multiple-choice question and answer options, and record judgments for several verification criteria. Answer Randomization. To prevent positional bias from the model generating the questions, the answer options are randomly shuffled before being provided to the language model during evaluation. API Usage. Although the API safety filters were set to the least restrictive available configuration, a small number of AD inputs were still blocked by the service and could not be processed. These instances were treated as missing values and excluded from aggregate results. S8.2 Manual Verification All questions retained in StoryAD-QA are manually verified before inclusion. Each question is reviewed through a custom Gradio annotation interface, which presents annotators with the corresponding video segment together with the generated question, answer options, the labeled correct answer, and the accom- panying rationale produced during question generation. Figure S5 demonstrates the interface from where annotators watch the corresponding video segment and the corresponding generated questions and submit their evaluation of several quality metrics designed to assess the integrity of the questionâanswer pairs. The review covers both the question and its answer set: annotators check that the question is visually grounded and unambiguous, that the labeled answer is uniquely supported by the video evidence, and that the distractor options are plausible but incorrect. Questions that fail verification are discarded or regener- ated and then checked again before they can enter the final benchmark. Verification criteria We use Correct answer to assess whether the labeled answer is correct and uniquely supported by the visual evidence in the clip. In addition, we adopt two quality metrics from [2]. Visually grounded evaluates whether the correct answer can be determined solely from the visual content of the clip, without relying on dialogue, audio, or external knowledge. Plausible StoryTeller13 Table S7: Human verification results on all questions of the initial StoryAD-QA bench- mark questions. Rates report the proportion of questions satisfying each criterion. SettingGroundedPlausible DistractorsCorrect Ans. Track A 30 s0.8900.8570.752 60 s 0.9010.8750.795 120 s0.8920.8870.814 240 s 0.9630.9360.833 Track B 60 s 0.8380.8270.744 90 s0.8220.7840.703 120 s 0.8860.8310.699 Unsure3345 distractors assesses whether the incorrect answer options constitute realistic alternatives that could reasonably be confused with the correct answer given the scene. In addition to the binary Yes/No labels, annotators could also select Unsure when the judgment was genuinely ambiguous. The verification statistics for the full benchmark are presented in Table Tab. S7. S8.3 Question Regeneration Questions that received a negative rating in any evaluation criterion during the initial human verification were discarded and regenerated using the same video clip and prompt, but with a stronger VLM, namely Gemini 3.5 Flash. In total, we regenerated 773 questions, of which 618 were retained after filtering out clips that the VLM determined lacked sufficient information to support a meaningful question. All regenerated questions then underwent a second round of manual verification by a human annotator. Every question in this pass was accepted and included in the final benchmark. S9 Human Evaluation of Audio Descriptions To evaluate the perceptual quality of generated audio descriptions, we conducted a human study comparing StoryTeller with AutoAD-Zero. S9.1 Study Setup We randomly sampled 50 clips from the evaluation set, each with an average du- ration of approximately 20 seconds. For each clip, three annotators watched the video and compared two candidate audio descriptions: one generated by Story- Teller and one generated by AutoAD-Zero. The descriptions were presented with 14S. H. Hahm et al. Fig. S6: Human evaluation interface used in the study. Annotators watch a video clip and compare two candidate audio descriptions (Option A and Option B), whose order is randomized per trial. They select whether one description is better, whether the two are comparable (Tie), or whether both descriptions are inadequate. anonymized labels (Option A and Option B), and their order was randomized for each trial to avoid positional bias. Because each clip spans approximately 20 seconds, the candidate descriptions are formed by concatenating multiple con- secutive AD sentences generated for that clip, resulting in short multi-sentence descriptions. Figure S6 illustrates the evaluation interface. Annotators could replay the clip and optionally view the reference audio narration (ground truth) for additional context, but no system identifiers or auxiliary metadata were shown in order to maintain a blind evaluation. For each clip, annotators selected one of four options: â Option A better: description A provides a better audio description. â Option B better: description B provides a better audio description. â Tie: both descriptions are of comparable quality. â Both bad: neither description adequately describes the scene. Annotators were instructed to judge descriptions based on two criteria: (i) how accurately the description reflects the visible content of the clip, and (i) whether the narration conveys the scene clearly and coherently for a blind or low-vision audience. The study involved three human evaluators, each rating all 50 clips. S9.2 Results Table S8 summarizes the human preference results. Across the 50 evaluated clips, annotators preferred StoryTeller substantially more often than AutoAD- Zero, with StoryTeller receiving 76.0% of the preferences compared to 5.8% for StoryTeller15 Table S8: Human preference comparison between StoryTeller and AutoAD-Zero. MethodPreference (%) StoryTeller76.0 AutoAD-Zero5.8 Tie9.1 Both bad9.1 the baseline while 9.1% of the evaluations resulted ina tie and 9.1% judged both descriptions inadequate. In many cases, annotators selected StoryTeller when its descriptions captured scene context or environmental cues visible in the frames, whereas AutoAD-Zero descriptions tended to focus on shorter or more generic actions. S10 Prompt Templates This section lists the prompt templates used in the StoryTeller pipeline, orga- nized according to the modules described in Sec. 3 of the main paper: scene summarization, structured fact extraction, fact verification, and narration gen- eration. The prompts guide the model to distinguish observed evidence from auxiliary cues such as public movie metadata. In particular, the fact-extraction schema records whether each candidate fact is suggested by visual, audio, or metadata input. Metadata may help propose candidates, such as possible char- acter names, but these candidates are passed through a separate verification step before being added to memory or used in narration. S10.1 Scene Summarization Prompt Stage 1: Scene Summarization. This prompt generates a concise summary of the visual content of a single video clip. The goal of this stage is to produce a short, grounded description that captures the main observable event without introducing speculative interpretation. The summary serves as the initial repre- sentation of the scene and is later expanded into structured narrative facts. Instruction. You create concise, evidence-grounded clip summaries for audio de- scription. Rules: present tense; one sentence; less than 25 words; describe only what is clearly visible or audible; no speculation, no spoilers; if identity is unclear, use neutral labels. Prompt. Summarize the clip in one concise sentence, describing only what is clearly visible or audible. S10.2 Structured Fact Extraction Prompt Stage 2: Structured Fact Extraction. Given the scene summary and video clip, this prompt extracts structured facts describing observable actions, entities, 16S. H. Hahm et al. and locations. The output follows a strict JSON schema in order to produce atomic, verifiable facts that can be independently checked. These structured facts form the basis of the narrative memory used by StoryTeller. Instruction. You extract structured, evidence-grounded facts from a video clip for audio description. Return ONLY valid JSON exactly matching the schema. No markdown. No extra keys. No inference. Keep facts atomic (one action each) and limit to 6 to 10 items. Use neutral labels if identity is unclear. Fill the JSON schema below using only what is clearly visible or audible in the clip. Schema: "scene_summary": "one sentence visual summary", "time_window": "start-end seconds or m:s-m:s", "facts": [ "subject": "who performed the action", "action": "verb phrase", "object": "what/who was acted upon; use empty string if not needed", "location": "where in the scene", "time": "timestamp or short range", "evidence": "visual|audio|metadata", "confidence": "high|medium|low", "notes": "optional clarifications", "cluster_id": "if the action involves a TRACKLET_GUIDE entry, copy its tracklet_id (e.g., trk_0003)", "character_id": "when 100% certain, supply the canonical character_id from metadata" ], "uncertain_observations": ["optional list describing ambiguities"] TRACKLET_GUIDE (cluster IDs visible in this clip): tracklet_hint Spoiler-safe metadata snippets that may match this moment: context_snippets Metadata hints for this movie (use only if they match the video): metadata_hint S10.3 Fact Verification Prompt Stage 3: Fact Verification. To prevent hallucinated or weakly grounded ob- servations from entering the memory, each extracted fact is passed through a verification step. The model acts as a strict fact checker and determines whether the proposed fact is clearly supported by the visual evidence in the clip. Only accepted facts are retained for downstream narration. StoryTeller17 Instruction. You are a strict video-grounded fact checker. Accept only if the clip clearly supports the fact. If ambiguous, inferred, off-screen, or not shown: reject. Prompt. You are verifying whether the FACT is truly visible or supported in this video clip. Use the video as the primary source. On the first line respond exactly with either ACCEPT or REJECT, then provide a one-sentence justification citing what you observed. FACT: fact_text S10.4 Final Input Prompt Stage 4: Final Narration Generation. The final prompt produces the audio description used for evaluation. It conditions on the verified fact set and optional narrative memory context from previous clips. The prompt instructs the model to generate a concise, visually grounded narration that reflects the scene while avoiding references to cameras or recording artifacts. Instruction. You are an in-scene narrator describing events to blind or low-vision listeners. Deliver the narration as a storyteller within the world of the film. Never mention cameras, screens, or phrases such as âin this video,â âon screen,â âtoward the camera,â or âlooks at the camera.â If a draft would mention recording equipment, rewrite it so the line instead describes how characters relate to one another or to their environment. Rely on grounded details from the FACT_BLOCK, keep the chronology clear, and favor concise active sentences. When a FACT_BLOCK entry includes a character_id, refer to that person by the provided first name only. Do not invent or guess names. Use the provided first name when the character is first mentioned; afterwards, pronouns such as he, she, or they may be used. Highlight motion, intent, and key sounds. Omit filler and self-referential phrasing. Keep the final description under 20 words. Use MEMORY_CONTEXT only when it directly supports continuity without speculation. Clip duration: duration:.2fs Output format: Return JSON in the form "summarised_AD": "... ". Examples. For a 0.8s clip: "summarised_AD": "She looks at Riker." For a 1.4s clip: "summarised_AD": "Paul looks at his wife lovingly." For a 2.6s clip: "summarised_AD": "Returning to the room, Sarah peers into the darkness." For a 3.5s clip: "summarised_AD": "Stephen looks at Sara as she walks to the door." S10.5 StoryAD-QA Question Generation Prompt Track A (segment-only QA). The Track A prompt asks the VLM to gener- ate a multiple-choice question about the main visual event within a single clip. For question generation only, the VLM also receives a short Wikipedia movie overview as background. The prompt explicitly restricts the model to informa- tion that is directly visible, prohibits the use of dialogue, external knowledge, or 18S. H. Hahm et al. speculative reasoning, and requires a short rationale grounded in visual evidence. If the clip does not contain a clear self-contained event, the model is instructed to output SKIP. You are given: - A video clip from a movie. - A movie overview for background context only. Overview: movie_overview IMPORTANT: This clip may still be incomplete relative to the full movie. Do NOT assume anything outside what is visible in this clip. Do NOT use prior knowledge of the movie, characters, or story. Use only what is directly shown. TASK: Your task is to generate ONE 5-option multiple-choice question about the main visual event or progression shown in this clip. If the clip does NOT contain a clear, self-contained action, reaction, reveal, or causeâeffect moment that is fully shown, output: SKIP The question MUST reflect the main development of the clip, not a small detail from a brief moment. Focus on the most important or climactic visible event. Do NOT generate a question if: - The moment is routine or transitional - The answer depends on events outside this clip - It requires guessing thoughts, intentions, or emotions - It relies on what happens before or after the clip If uncertain, output: SKIP Requirements: - 1â2 sentence question - Answerable using only this clip - No âWhyâ questions - No references to before or after the clip - No psychological interpretation - Focus on clearly visible actions or outcomes All incorrect options must be plausible. Output format: Question: <question> A. B. C. D. E. StoryTeller19 Correct Answer: <letter> Rationale: <1â2 sentences describing only visible evidence> Track B (context-conditioned QA). The Track B prompt introduces an explicit separation between a context window and a target scene. The context may introduce entities or situations, but the correct answer must be visually ver- ifiable in the target scene. The prompt further discourages references to times- tamps, dialogue, or outside knowledge to ensure questions remain grounded in observable visual information. You are given a movie video clip and a brief movie overview for background only. Overview: movie_overview The first context_len seconds of the clip are CONTEXT. The remaining part of the clip is the TARGET SCENE. The context may introduce people, objects, or locations that help interpret the target scene. Generate ONE multiple-choice question (5 options A-E) that requires using the context to understand what is happening in the target scene. Key principle: The answer must be visible in the TARGET SCENE, but the CONTEXT should help identify or interpret the object, person, or situation refer- enced in the question. Preferred patterns: - An object or person appears in the context, and the question asks about its state, action, or location in the target scene. - Something in the target scene is ambiguous unless the viewer remembers the context. - The context introduces an item or situation whose consequence or change becomes visible in the target scene. Rules: - The question must be answerable using only what is visibly shown in the video. - The correct answer must be visually verifiable in the TARGET SCENE (after the first context_len seconds). - The context should help identify what the question refers to, but should not contain the answer itself. - Do not rely on character names, dialogue, subtitles, or outside knowledge. - Focus on observable actions, objects, interactions, or spatial relationships. - Avoid trivial decorative details that do not affect the scene. - Do not ask âWhyâ questions or infer emotions, intentions, or story meaning. - Do not mention numeric timestamps, seconds, or "the first context_len seconds". - Do not refer to the video itself (no phrases like âin this clipâ or âon screenâ). 20S. H. Hahm et al. The question MUST reflect the main development of the clip, not a small detail from a brief moment. Focus on the most important or climactic visible event. Do NOT generate a question if: - The moment is routine or transitional - The answer depends on events outside this clip - It requires guessing thoughts, intentions, or emotions - It relies on what happens before or after the clip If there is no clear visual moment after the context that connects to something introduced earlier, output: SKIP Output format: Question: <question> A. <option> B. <option> C. <option> D. <option> E. <option> Correct Answer: <letter> Rationale: <1â2 sentences describing only the visible evidence in the target scene that confirms the answer> S10.6 StoryAD-QA Question Answering Prompt During evaluation, we provide the language model with only textual inputs: the question and the corresponding audio descriptions for the scene. The model is instructed to select an answer using only the information contained in the AD, without the use of external knowledge, prior familiarity with the movie, or assumptions about events not described in the AD text. In addition, the model is required to produce a short rationale citing explicit evidence from the AD, encouraging answers that are grounded in the provided narration. To reduce positional bias, the answer options are randomly shuffled before being presented to the language model during evaluation. You are given: - Audio Descriptions (AD) for a movie scene. - A multiple-choice question about that scene. StoryTeller21 Question: question Audio Descriptions (AD): ad_text IMPORTANT: Use only information explicitly stated in the AD text. Do NOT use prior knowledge of the movie. Do NOT assume events outside the AD. Do NOT guess charactersâ thoughts or intentions unless the AD explicitly states them. TASK: Select the best answer (A, B, C, D, or E) based only on the AD. Output format: Answer: <letter> Rationale: <1â2 sentences citing only explicit AD evidence> References 1. Deng, J., Guo, J., Xue, N., Zafeiriou, S.: Arcface: Additive angular margin loss for deep face recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 4690â4699 (2019) 2. Kala, D., Khandelwal, E., Tapaswi, M.: What you see is what you ask: Evaluating audio descriptions. In: Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V. (eds.) Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). p. 23496â23518. Association for Computational Linguistics, Suzhou, China (Nov 2025). https://doi.org/10.18653/v1/2025.emnlp-main. 1199, https://aclanthology.org/2025.emnlp-main.1199/ 3. Song, E., Chai, W., Wang, G., Zhang, Y., Zhou, H., Wu, F., Chi, H., Guo, X., Ye, T., Zhang, Y., Lu, Y., Hwang, J.N., Wang, G.: Moviechat: From dense token to sparse memory for long video understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 18221â18232 (June 2024) 4. Xie, J., Han, T., Bain, M., Nagrani, A., Varol, G., Xie, W., Zisserman, A.: AutoAD- Zero: A training-free framework for zero-shot audio description. In: Proceedings of the Asian Conference on Computer Vision (ACCV). p. 2265â2281 (2024)