Paper deep dive
Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding
Quanwei Tang, Dong Zhang, Shoushan Li, Guodong Zhou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/26/2026, 5:27:18 AM
Summary
The paper introduces LongAudioQA, a dataset for long-form audio meeting understanding (LAMU), and proposes the Graph-based Retrieval-Generation Agent (GRGA) model. GRGA addresses acoustic information loss and context fragmentation by modeling heterogeneous audio features into a multi-dimensional graph and using agent planning for retrieval and answer generation. It outperforms existing Speech LLMs and RAG baselines on factual, inferential, temporal, summarization, and acoustic-aware tasks.
Entities (15)
Relation Signals (13)
Dong Zhang → affiliatedwith → Soochow University
confidence 95% · Dong Zhang Affiliation: School of Computer Science & Technology, NLP Lab, Soochow University
Quanwei Tang → affiliatedwith → Soochow University
confidence 95% · Quanwei Tang Affiliation: School of Computer Science & Technology, NLP Lab, Soochow University
GRGA → solves → Acoustic Missing
confidence 95% · GRGA models heterogeneous audio features... to address these issues [acoustic information loss]
GRGA → solves → Context Fragmentation
confidence 95% · GRGA... leverages agent planning for retrieval... to address... poor long-term context memory
GRGA → uses → LongAudioQA
confidence 95% · we construct the LongAudioQA dataset and propose the GRGA model
LongAudioQA → derivedfrom → DailyTalk
confidence 90% · We mainly construct the speech question-answer pairs from the following raw dataset: ... DailyTalk
LongAudioQA → derivedfrom → AMI Meeting
confidence 90% · We mainly construct the speech question-answer pairs from the following raw dataset: ... AMI Meeting
LongAudioQA → derivedfrom → AliMeeting
confidence 90% · We mainly construct the speech question-answer pairs from the following raw dataset: AliMeeting
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation.
Tags
Links
- Source: https://arxiv.org/abs/2608.24048v1
- Canonical: https://arxiv.org/abs/2608.24048v1
Trouble viewing inline? Open PDF directly →
Full Text
130,971 characters extracted from source content.
Expand or collapse full text
Don’t Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding Quanwei Tang Affiliation: School of Computer Science & Technology, NLP Lab, Soochow University, China Dong Zhang Affiliation: School of Computer Science & Technology, NLP Lab, Soochow University, China Shoushan Li Affiliation: School of Computer Science & Technology, NLP Lab, Soochow University, China Guodong Zhou Affiliation: School of Computer Science & Technology, NLP Lab, Soochow University, China Affiliation: Jiangsu Key Lab of Language Computing, Suzhou. dzhang@suda.edu.cn Abstract While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation. GitHub for data and code. 1 Introduction **footnotetext: Corresponding Author: Dong Zhang Long-form audio meeting understanding (LAMU) has attracted significant attention in speech processing. However, previous studies have only focused on transcription recognition for long-form multi-party meetings Yu et al. (2022a); Yu et al. (2022b); Jain et al. (2024). Consequently, there is a lack of dedicated question answering (QA) datasets for long audio meetings, despite the critical importance of this task. To address this gap, we construct the LongAudioQA dataset for LAMU. Distinct from short-form conversation QA, it is designed to capture three core dimensions of long-form audio meeting content: complex semantics, multi-speaker interactions, and quite long timestamps. Figure 1: Existing Speech LLMs and our GRGA for long-form audio QA. By indexing conversations across Semantic, Speaker, and Timestamp dimensions, our model enables precise reasoning. Although existing Speech LLMs KimiTeam et al. (2025); Xu et al. (2025a) have demonstrated strong performance across various tasks, they prioritize textual context by ASR over speech context You et al. (2022); Lin et al. (2024). Specifically, these models typically map the input speech to the textual context, which inevitably leads to the loss of valuable acoustic information. For example, regarding the query “sudden loud voice at the 19-minute mark” in Figure 1, the inability to access voice volume prevents natural understanding and appropriate response (as the result of Text LLM in the Figure). We denote this phenomenon as Acoustic Missing. Additionally, prior QA on speech conversations has been exclusively centered on short-form audio clips or dialogues (<< 30s) Johnson et al. (2024); Zhao et al. (2025); Hu et al. (2026). When applied to extremely long audio meetings, these methods fail to capture long-term dependencies. Let’s return to the query in Figure 1, we not only need to locate the content of the 19-minute timestamp, but also to trace the answer back to the description of 1:30 (as our GRGA). Without an explicit design for long-range semantic understanding, even RAG approaches Tang et al. (2025); Li et al. (2025); Wang et al. (2024) fail to identify the correct rationale (as Text RAG). We designate this phenomenon as Context Fragmentation. To address these challenges, we propose the Graph-based Retrieval-Generation Agent (GRGA) model. Specifically, we first model heterogeneous features from the audio, including not only acoustic information (e.g., voice and tone) but also many speaker attributes (e.g., role: project manager, gender: male), into a unified multi-dimensional graph structure. Then, we leverage agent planning to retrieve question-relevant clues from this multi-dimensional graph and generate the final answer. In summary, we contribute: ∙ We design and construct the novel dataset LongAudioQA for LAMU. ∙ We propose GRGA to handle both acoustic missing and context fragmentation. ∙ We conduct both automatic and human evaluation on three datasets of our LongAudioQA. Category Core Competency Subtasks & Example Queries Factual Explicit Retrieval ∙ Entity Retrieval: “Who mentioned the project code ‘Alpha’?” ∙ Attribute Association: “What is the budget proposed by the Manager?” ∙ Keyword Locating: Matching specific text spans. Inferential Multi-hop Reasoning ∙ Causal Inference: Linking a problem raised early with the final decision. ∙ Coreference Resolution: Identifying what “that plan” refers to. ∙ Stance Analysis: Tracking how a speaker’s attitude evolves. Temporal Time Awareness ∙ Absolute Localization: “What topic was discussed at the 30-min mark?” ∙ Relative Sequencing: “What was discussed after ‘market research’?” ∙ Frequency Stats: Counting term occurrences in a window. Summarization Aggregation ∙ Topic Summarization: Synthesizing consensus on “backend architecture”. ∙ Speaker Profiling: Summarizing a participant’s main contributions. Acoustic-Aware Multi-modal Alignment ∙ Emotion & Intensity: “Who seemed most agitated when discussing the budget?” ∙ Cross-modal Localization & Causal: “What is the reason for the person’s sudden loud voice at the 15-minute mark?” Table 1: Taxonomy of Meeting Questions in our LongAudioQA. Categories are distinguished by background colors: Factual, Inferential, Temporal, Summarization, and Acoustic-Aware. The Acoustic-Aware category uniquely requires grounding textual semantics with paralinguistic acoustic signals. 2 Dataset Construction Existing speech QA benchmarks Zhifei et al. (2025) predominantly focus on short-context span extraction or simple intent classification. They often treat the dialogue as a flat sequence of text, neglecting the intricate graph-like structure of meeting interactions (e.g., speaker turns, cross-references, and temporal dynamics). Consequently, current models are rarely tested on their ability to perform multi-hop reasoning or temporal grounding over long-form recordings. For instance, answering a question like “How did the speaker’s attitude change after the 30-minute discussion?” requires a model to not only localize information but also aggregate evidence across distinct time slices and model the causal dependencies between nodes. Therefore, we propose our LongAudioQA. 2.1 Data Selection and Collection We mainly construct the speech question-answer pairs from the following raw dataset: AliMeeting. AliMeeting Yu et al. (2022a); Yu et al. (2022b) is a large-scale Mandarin speech dataset tailored for multi-speaker meeting ASR and SD. A distinguishing feature of this corpus is the simultaneous recording of overlapping far-field audio and individual near-field references. This setup is designed to address the “who said what when” problem, testing the model’s ability to handle speaker overlap and diarization in complex settings. AMI Meeting. The AMI Conference Corpus Jain et al. (2024) consists of meeting audio recordings captured using far-field microphones, primarily capturing interactions among non-native English speakers. It is characterized by acoustic complexity, presenting significant challenges commonly encountered in real conference settings, such as background noise and reverberation issues. DailyTalk. In contrast to the aforementioned conference-centric corpora, DailyTalk Lee et al. (2022) focuses on high-quality open-domain dyadic conversations. This corpus comprises 2,541 dialogues. DailyTalk provides clean acoustic environment with an emphasis on conversational fluency and prosodic features. We concatenated 20 clips to create inputs of intermediate length (<<10 minutes). While shorter than full meetings, this duration serves to ensure our method performs well on longer audio segments while maintaining effectiveness on shorter ones. 2.2 Question Definition and Data Annotation To bridge the gap, we constructed a novel dataset designed to expand the boundaries of long-form meeting understanding. Unlike existing research that relies solely on factual retrieval, our dataset introduces a diversified question taxonomy, encompassing factual, inferential, temporal, and acoustic-aware as shown in Table 1). This design enables us to systematically evaluate models’ ability to transition from explicit pattern matching to advanced semantic reasoning within complex interaction graphs. Following methodologies for automated data annotation using LLMs Lian et al. (2025a); Lian et al. (2025b), we employ a hybrid strategy combining large language model-driven automated annotation with human-assisted verification. Specifically, this approach leverages large language models to generate questions from specific local segments of raw conference data based on predefined query types and examples. The models then search conference content for answers and extract detailed evidence. This process is iterative: each batch of generated question-answer pairs undergoes human verification to ensure quality before advancing to the next iteration. While the generation process leverages an “oracle” mechanism to access localized evidence, the core challenge of this study lies in retrieving answers from global, long-context conference data without prior knowledge of relevant segments. This evaluates the model’s ability to overcome context fragmentation. 2.3 Dataset Quality Control Given that automated generation inevitably introduces noise, we enforce a rigorous verification protocol to ensure data reliability. Expert annotators inspect candidate samples to discard hallucinated questions (where answers are absent), rewrite ambiguous references, and validate logic depth (hop count) to ensure complexity balance. Specifically, we require that the generated timestamp evidence yields an Intersection-over-Union (IoU) of >0.9>0.9 with the ground truth. This process results in a high-quality dataset with a human inter-annotator agreement rate of κ=0.91κ=0.91 (Cohen’s Kappa). Finally, the duration and sample statistics of our LongAudioQA, along with the question distribution, are summarized in Table 2 and Figure 2. Dataset Dur (h) Avg. Dur (s) Avg. Turns # Q&A # Dial. AliMeeting 14.91 1,935 815 3,013 28 AMI 18.24 1,930 757 3,243 34 DailyTalk 21.59 610 186 10,200 128 Table 2: Statistical information among DailyTalk, AMI, and AliMeeting datasets. Dur: Duration, Avg: Average. Figure 2: Question Distribution Across Corpora 3 Methodology Figure 3: The overall architecture of our graph-based retrieval-generation agent GRGA. The framework models long-form audio as a Multi-dimensional Graph (bottom-left) to capture semantic, temporal, and speaker dependencies. Our GRGA Planning Process consists of: Query Decomposition, Planning, Execution, Synthesis, and Reflection. The core challenge in long-form audio QA lies in handling composite queries that involve semantic, temporal, and speaker relationships. Existing methods often rely on “one-shot” retrieval, which fails to capture the intricate dependencies in meeting data. To address this, we propose a novel Graph-based Retrieval-Generation Agent (GRGA) that mimics the cognitive process of a human expert: it does not merely search, but rather plans and reflects. Cognitive Inspiration. To design an effective agent, we draw inspiration from the cognitive strategies human experts employ to navigate complex meetings. We observe that resolving composite queries typically adheres to a “Search-Reason-Verify” cognitive loop, which adapts dynamically to the query type. For instance, in factual tasks (e.g., verifying budget details), humans execute a targeted keyword scan followed by contextual verification—a mechanism we emulate through retrieval tools. Conversely, inferential tasks (e.g., discerning the rationale behind a rejection) necessitates maintaining working memory while traversing causal chains. Similarly, humans naturally construct a mental timeline to address temporal queries, whereas summarization involves synthesizing fragmented details into high-level concepts. Consequently, we architect our agent to replicate these cognitive behaviors by incorporating a Planner for logical orchestration, a Synthesizer module for holistic information aggregation, and a Self-Reflector for answer verification. 3.1 Problem Formulation We formulate the task as a Partially Observable Markov Decision Process (POMDP) Lauri et al. (2023a), defined by the tuple ℳ=⟨,,,ℛ,Ω,,γ⟩M= ,A,T,R, ,O,γ . The pre-constructed multi-dimensional graph G serves as the environment with which the agent interacts. Belief State (btb_t) & Policy (π): The belief state btb_t, represented by the query Q and interaction history HtH_t, serves as the sufficient statistic for decision-making. We leverage a pre-trained LLM as the policy π(at∣bt)π(a_t b_t). Implemented via the Query Decomposer and Execution Planner, the policy utilizes in-context reasoning to determine the next optimal action without parameter updates. Action (A) & Observation (Ω ): The action space A comprises topological retrieval operations. The Tool Executor executes ata_t on G, yielding a partial observation ot∈Ωo_t∈ (e.g., specific acoustic cues or semantic nodes), which updates btb_t. Reward (ℛR): The Self-Reflector acts as a verifier, providing a heuristic reward rtr_t. It evaluates whether oto_t aligns with the reasoning chain, prompting the planner to prune incorrect paths or terminate generation. Since the model is training-free, our objective is to approximate the optimal trajectory τ∗=b0,a0,…,aTτ^*=\b_0,a_0,…,a_T\ that maximizes the reliability of the final answer. The process terminates when the Answer Synthesizer determines that sufficient evidence is gathered (rt>thresholdr_t>threshold) or the maximum step limit is reached. Figure 3 illustrates our framework. 3.2 Audio-to-Graph Indexing 3.2.1 Acoustic Semantic Alignment To initialize the graph nodes, we transform the continuous audio stream into discrete, semantically meaningful units with precise temporal boundaries. VAD and Transcription. We employ FSMN-VAD to segment audio A into clips, filtering silence. These clips are transcribed by a high-fidelity ASR model to obtain text T. We apply CTC-based Forced Alignment: for each token wiw_i, we perform constrained Viterbi decoding to align text with acoustic features, yielding exact boundaries [tstart,tend][t_start,t_end]. Speaker Diarization. Identifying “who spoke” is insufficient; understanding “who they are” (e.g., role, stance) is key to reasoning. We propose a multimodal profiling mechanism. We extract speaker embeddings using ERes2NetV2 and perform incremental clustering. If the cosine similarity between a new embedding and existing clusters is below a threshold η=0.8η=0.8, a new speaker ID is created. Semantic Role Generation. A raw ID (e.g., Speaker_0) lacks semantic context. We construct a Speaker Profile kP_k for each unique speaker SkS_k. We aggregate all utterances belonging to SkS_k and prompt an LLM to distill attributes: k=LLM(Concat(vi.text∣vi.spk=Sk))P_k=LLM(Concat(\v_i.text v_i.spk=S_k\)) (1) The output is a structured profile, e.g., Role: Project Manager, Gender: Male, Stance: Conservative. This profile is stored as a node attribute. 3.2.2 Multi-dimensional Graph Modeling Beyond the raw meeting transcripts, we also incorporate the acoustic features (e.g., speaker identity, start/end timestamps, and corresponding speech) and explicitly model the dialogue flow as a heterogeneous graph =(,ℰ)G=(V,E). Nodes (V): Each utterance is treated as a fundamental node viv_i, enriched with attributes including text transcripts, speaker identity, start/end timestamps, and corresponding speech, etc. Edges (ℰE): To support the diverse reasoning tasks, we construct multiple edge types: (1) Temporal Edges (etempe_temp) connecting adjacent utterances (vi→vi+1v_i→ v_i+1); (2) Reply-To Edges (ereplye_reply): Captures conversational turns. We establish a directed edge vi→vjv_i→ v_j if vjv_j occurs within a dynamic window Δt<5s t<5s after viv_i and speakers are different; (3) SameSpeaker Edges (espke_spk) Connects nodes vi,vjv_i,v_j where spk(vi)=spk(vj)spk(v_i)=spk(v_j). This enables the agent to aggregate scattered opinions of a specific person; (4) Entity Edges (ente_ent) derived from entity keyword co-occurrence or discourse relations. (5) Semantic Edges (eseme_sem) To solve context forgetting, we employ an LLM to detect coreference chains. If vjv_j contains pronouns (e.g., “that idea”) referring to an entity in viv_i, we create a semantic link vj→semviv_j semv_i. 3.3 GRGA Planning Process Belief Initialization: Query Decomposition. Given a user query Q, the belief state is initialized by projecting the unstructured query into a structured constraint space C. We define the decomposition function fdec:→f_dec:Q as: =fdec(Q)=ce,c,ct,cmC=f_dec(Q)=\c_e,c_c,c_t,c_m\ (2) where ce,c,ct,cmc_e,c_c,c_t,c_m denote constraints on entities, concepts, time, and metadata, respectively. The initial belief is set as b0=H0=Q,b_0=H_0=\Q,C\. Policy Network: Execution Planning. The Execution Planner approximates the policy πθ _θ parameterized by an In-Context Learning (ICL ) Dong et al. (2024) guided LLM. At step t, it generates a multi-step plan tP_t conditioned on the current belief btb_t: t∼πθ(⋅∣bt)P_t _θ(· b_t) (3) The plan tP_t is a sequence of atomic operations derived from the Action Space A (see Table 11): t=[op1,op2,…,opk],opi∈P_t=[op_1,op_2,…,op_k], op_i (4) This formulation enables long-horizon planning, allowing the agent to chain logical operations (e.g., Search → Filter) before interacting with the environment. Environment Interaction: Tool Execution. The Execution Engine functions as the transition interface. It executes the plan tP_t on the graph G to obtain an observation oto_t: ot=Exec(t,)=s1,s2,…,sno_t=Exec(P_t,G)=\s_1,s_2,…,s_n\ (5) where each segment si=(texti,timei,spki)s_i=(text_i,time_i,spk_i). Upon receiving oto_t, the agent updates its belief state via a deterministic transition function ψ: bt+1=ψ(bt,t,ot)=bt⊕t,otb_t+1=ψ(b_t,P_t,o_t)=b_t \P_t,o_t\ (6) Reward Estimation: Synthesis & Reflection. To guide the reasoning trajectory refinement, we employ a two-stage mechanism as reward function. 1) Answer Synthesis. The synthesizer generates a candidate answer A^t A_t and citations CitetCite_t based on the accumulated evidence in bt+1b_t+1: (A^t,Citet)=fsyn(Q,bt+1)( A_t,Cite_t)=f_syn(Q,b_t+1) (7) 2) Reflection as Sparse Reward. The Reflector evaluates the logical entailment between the evidence E⊂bt+1E⊂ b_t+1 and the answer A^t A_t, assigning a verification score sver∈[0,1]s_ver∈[0,1]: sver=fref(Q,A^t,E)s_ver=f_ref(Q, A_t,E) (8) The reward rtr_t is defined as a threshold function: rt=1if sver≥τ(Success, Terminate)−βif sver<τ(Failure, Re-plan)r_t= cases1&if s_ver≥τ (Success, Terminate)\\ -β&if s_ver<τ (Failure, Re-plan) cases (9) If rt<0r_t<0, the negative feedback fbtfb_t (critique) is injected into the belief state: bt+1←bt+1∪fbtb_t+1← b_t+1∪\fb_t\, prompting the policy π to generate a corrective plan t+1P_t+1 in the next iteration. 4 Experimentation Method AliMeeting AMI Meeting DailyTalk Fact. Infer. Temp. Summ. Acou. Fact. Infer. Temp. Summ. Acou. Fact. Infer. Temp. Summ. Acou. Speech as Context Qwen3-Omni 29.31 27.33 15.81 19.06 16.50 20.96 26.38 15.12 16.67 12.15 34.26 38.43 2.43 2.11 9.39 Audio Flamingo 3 35.70 35.23 13.19 18.09 32.89 16.04 27.94 12.62 10.24 17.06 65.13 61.31 3.83 2.44 40.09 MiMo-Audio 54.50 54.50 16.92 39.91 25.15 52.75 64.74 35.34 38.30 17.17 81.57 78.65 5.58 3.46 31.44 Transcription as Context Qwen3-Omni 47.66 63.53 38.34 32.65 11.70 53.55 46.66 37.32 43.63 14.13 80.11 77.14 62.71 15.71 18.27 Audio Flamingo 3 32.12 45.16 28.23 21.06 9.69 48.52 40.21 22.12 29.16 12.36 76.43 75.26 58.32 13.42 15.85 MiMo-Audio 48.02 64.62 38.29 29.43 15.72 53.62 44.10 37.14 41.49 17.94 80.67 76.16 62.46 15.35 17.52 Both Speech and Transcription as Context with RAG BGE-M3 43.24 29.68 31.56 15.43 18.45 32.56 24.56 23.78 25.36 25.34 46.72 36.31 32.46 34.79 26.21 CLASP 25.37 19.54 18.92 14.76 16.32 24.12 20.31 22.76 11.60 14.79 30.58 33.42 27.34 13.52 11.81 GRGA(ours) 57.31 66.45 39.48 44.29 39.91 59.46 65.29 38.56 48.46 35.68 85.25 79.06 63.58 61.10 52.32 Table 3: Main results on accuracy (%). We compare our proposed method against End-to-End Speech LLM and standard RAG baselines across three datasets. Bold indicates the best performance. Fact.: Factual, Infer.: Inferential, Temp.: Temporal, Summ.: Summarization, Acou.: Acoustic reasoning. To validate the effectiveness of our GRGA, we conduct extensive experiments on our LongAudioQA dataset. 4.1 Baselines and Implementation First, we compare three recent most competitive multi-modal LLMs: Qwen3-Omni Xu et al. (2025a), Audio Flamingo 3 Goel et al. (2025), and MiMo-Audio Xiaomi (2025), which are considered as SOTAs for speech or spoken text QA. Therefore, we implement them in two forms: 1) use whole meeting Speech as Context and 2) use whole meeting Transcription as Context. Second, we investigate two tailored RAG approaches for audio QA: BGE-M3 Chen et al. (2024a) retrieving meeting text through a query, then merging the retrieved text and corresponding audio into the model (TextRAG) and CLASP Abootorabi and Asgari (2025), which are also considered as the SOTAs for RAG-based audio QA, retrieving meeting audio through a query, then merging the retrieved audio and corresponding text into the model (AudioRAG). They cast both meeting Speech and Transcription as Context. Notably, both of them adopt the same LLMs as ours to implement. More details of our GRGA and above baselines can refer to Appendix G. 4.2 Evaluation Metric To rigorously assess the reasoning capabilities of our model, we report Semantic Accuracy across all datasets. Traditional n-gram metrics (e.g., BLEU, Exact Match) often fail to capture the true validity of generated responses, as they penalize correct answers that differ in lexical surface forms from the ground truth. Drawing upon recent methodologies in complex reasoning evaluation Mishra et al. (2025); Shui et al. (2023), we move beyond rigid string matching and establish a robust, automated evaluation pipeline that prioritizes semantic equivalence and logical soundness. Specifically, we employ a high-capability LLM (LLM-as-a-Judge) to approximate human-level judgment. Formally, let =(qi,ai)i=1ND=\(q_i,a_i)\_i=1^N denote the evaluation dataset, where qiq_i is the query and aia_i is the ground-truth answer. Let a^i a_i represent the model-generated response. We define a semantic judgment function J, parameterized by an external expert model (e.g., GPT-OSS-120B OpenAI et al. (2025)): si=(qi,ai,a^i)∈0,1,s_i=J(q_i,a_i, a_i)∈\0,1\, (10) where si=1s_i=1 if and only if J determines that a^i a_i entails the same semantic information as aia_i, and 00 otherwise. The judge is prompted to disregard stylistic differences and focus solely on factual consistency and the correctness of reasoning. Finally, the Semantic Accuracy is computed as the expectation of correct judgments: Accuracy=1N∑i=1Nsi×100%.Accuracy= 1N _i=1^Ns_i× 100\%. (11) This approach ensures a fair comparison by validating whether the model successfully retrieves and reasons over the core knowledge, regardless of its output phrasing. 4.3 Main Results Table 3 shows the performance comparison of all baselines of our GRGA on our proposed three datasets for Long-form speech meeting understanding. From this table, we can see: Overcoming Context Limitations. End-to-End Speech LLMs (e.g., Qwen3-Omni) degrade sharply on long-form datasets, with AudioFlamingo3 dropping from ∼65% 65\% on DailyTalk to ∼16% 16\% on AMI. This confirms that fixed context windows hinder reasoning in long-form audio. Conversely, our GRGA maintains robust performance, outperforming the strongest baseline (MiMo-Audio) on AMI, validating the scalability of our graph-based retrieval beyond context limits. Planning vs. Naive Retrieval. Comparisons with RAG baselines highlight the failure of “one-shot” retrieval. Text RAG, while effective for factual questions, struggles significantly with inferential questions (24.6%24.6\% vs. ours 65.3%65.3\% on AMI). This proves that vector similarity alone cannot capture multi-hop dependencies. Our Query Planner bridges this gap by decomposing queries into logical chains. Additionally, the poor performance of Audio RAG (∼20% 20\%) indicates that raw acoustic retrieval is too noisy, justifying our use of a structured graph intermediate. Dataset WER (↓ ) DER (↓ ) Collar = 0 s Collar = 0.25 s DailyTalk 0.85 6.42 3.41 AliMeeting-far 20.38 16.80 13.10 AMI-sdm 18.84 17.60 14.90 Table 4: Error analysis on benchmark datasets. We report Word Error Rate (WER) and Diarization Error Rate (DER) under strict (0s0\,s) and standard (0.25s0.25\,s) collar (tolerance ) settings. All metrics are reported in percentage (%), and ↓ denotes that lower values are better. 4.4 Analysis and Discussion Figure 4: Human evaluation results. Our method shows significant improvements, particularly in Groundedness across complex meeting scenarios. Human Evaluation. As illustrated in Figure 4, our method consistently outperforms Qwen3-Omni across all datasets and metrics. Notably, the most substantial gap appears in Groundedness (e.g., +1.27 on AllMeeting). While the baseline generates fluent text (high Coherence), it frequently lacks evidentiary support in multi-party dialogues. Our superior grounding directly translates to higher Correctness, confirming that precise temporal anchoring effectively reduces factual errors. Upstream Performance Analysis. Table 4 highlights the noise disparity across datasets. DailyTalk serves as a clean baseline with minimal errors (WER 0.85%0.85\%). Conversely, AMI and AliMeeting present high noise rates, with WERs ∼20% 20\% and DERs >13%>13\%. Despite these significant upstream errors in transcription and diarization, our GRGA maintains high QA accuracy (Table 3), demonstrating the robustness of our iterative planning and reflection mechanisms against noisy inputs. Figure 5: Citation Precision-Recall Trade-off. Results are reported in % with a ±2± 2s tolerance. Evidence Recall & Precision Analysis. To assess the model’s ability to locate supporting evidence, we report the Precision (P) and Recall (R) of generated citations (timestamps) against ground truth spans (with a ±2± 2s tolerance) in Figure 5 and Table 7. Our GRGA significantly outperforms baselines across all datasets. Specifically, compared to standard Text RAG, our method improves Precision by +19.3%+19.3\% on AMI and +6.7%+6.7\% on AliMeeting. This indicates that our Query Planner effectively filters out irrelevant noise, while the Reflection mechanism ensures that retrieved segments are strictly relevant to the answer, minimizing hallucinated citations common in naive retrieval approaches. Ablation Settings Overall AMI Avg. Δ Fact. Infer. Temp. Summ. Acou. Ours (Full) 49.49 - 59.46 65.29 38.56 48.46 35.68 Validation of Query Decomposer w/o Query Planner 46.39 -3.10 57.63 59.35 36.79 45.73 32.43 Validation of Query Planner w/o Query Planner 44.68 -4.80 56.12 56.35 36.41 43.18 31.37 Validation of Reflection w/o Reflection 39.64 -9.84 51.32 49.52 31.10 39.83 26.45 Validation of Action Space w/o Tool: Filtering 43.69 -5.79 57.81 62.27 35.36 41.45 28.58 w/o Tool: Semantic Search 16.28 -33.21 28.96 16.38 12.21 13.69 10.15 w/o Tool: Graph Traversal 38.21 -11.27 45.57 38.62 32.31 41.25 33.31 w/o Tool: Audio Access 34.71 -14.78 47.52 49.12 33.13 31.71 12.06 Table 5: Ablation study on the AMI dataset. We report the accuracy (%) drop when specific components are removed. Δ : Performance degradation compared to the full framework. Ablation Study To validate the necessity of each component in our GRGA, we conduct an ablation study on the AMI dataset as a typical example in Table 5. From this table, we can see that removing any module results in varying degrees of performance degradation for our GRGA. Especially, the results of w/o Semantic Search demonstrate that our approach can indeed solve the semantic understanding problem in long-form speech meeting scenario. The removal of Graph Traversal indicates the importance of our designed multi-dimensional graph. More analysis can be found in Appendix 6. Step Analysis can be found in Appendix C. Case Study can be found in Appendix D. Noise Sensitivity Analysis can be found in Appendix E. 5 Related Work Speech and Meeting Question Answering. Early research in speech QA primarily focused on extracting answer spans from short, single-speaker speech segments, such as Spoken SQuAD Li et al. (2018) and HeySQuAD Wu et al. (2023). They normally focus on answering ranking Hu et al. (2026), instead of generative QA. While recent datasets like AMI Jain et al. (2024) and AliMeeting Yu et al. (2022a) provide rich multi-speaker meeting resources, they are predominantly used for ASR and diarization tasks rather than complex reasoning. Existing QA models on these datasets often treat transcripts as flat text sequences, neglecting the intricate temporal and interpersonal dependencies inherent in meetings. In contrast, our work targets long-form meeting comprehension, requiring models to navigate graph-structured dialogues involving multiple speakers and temporal dynamics. Large Speech Language Models. Recent advancements in multi-modal LLMs have enabled direct processing of audio inputs. However, current Speech LLMs face significant limitations in context length. For instance, Kimi-Audio KimiTeam et al. (2025) is optimized for short clips (<30<30s) and suffers from truncation issues with longer inputs. Even state-of-the-art models like AudioFlamingo3 Goel et al. (2025) are typically constrained to context of approximately 1010 minutes. This bottleneck renders them unsuitable for long-form meeting analysis in the absence of external retrieval mechanisms. Retrieval-Augmented Generation (RAG). RAG has emerged as a standard paradigm for grounding LLMs in external knowledge Lewis et al. (2020). Traditional “Retrieve-then-Generate” approaches rely on dense vector similarity to retrieve relevant contexts in a single pass. However, these methods struggle with multi-hop reasoning, where the evidence is fragmented or requires logical deduction steps not captured by semantic similarity alone Liu et al. (2025); Tang et al. (2025). Recent works have explored recursive retrieval or chain-of-thought prompting to address these limitations. Our framework advances this paradigm by formalizing retrieval as an iterative planning process, specifically tailored to handle the noisy and unstructured nature of ASR transcripts. LLM Agents and Tool Using. The emergence of LLMs has catalyzed the development of autonomous agents capable of using tools to solve complex tasks, exemplified by frameworks like ReAct Yao et al. (2023) and Toolformer Schick et al. (2023). While these agents demonstrate proficiency in open-domain tasks (e.g., web browsing, math) Zhang et al. (2025), their application to structured audio understanding remains underexplored. We bridge this gap by defining a specialized Action Space for meeting analysis (e.g., time_range_search, hybrid_search) and incorporating a Reflection mechanism modeled as a POMDP Lauri et al. (2023b), enabling the agent to self-correct in partially observable audio environments. 6 Conclusion We present GRGA, an agentic framework that addresses the acoustic missing and context forgetting issues in long-form speech understanding. By structuring audio into a multimodal heterogeneous graph and formulating QA as a POMDP, we enable an agent to explicitly plan, navigate, and reason over complex interactions. Extensive experiments on our proposed LongAudioQA benchmarks demonstrate that our GRGA significantly outperforms both end-to-end Speech-LLMs and RAG-based SOTAs. Acknowledgements This work was supported by Jiangsu Province Frontier Program Project (BF2025036), and Hong Kong RGC grant GRF #15611021. They are also with Jiangsu Key Lab of Language Computing, Suzhou. Limitations Our work has three main limitations. (1) Error Propagation: Since graph construction relies on upstream ASR and diarization, severe acoustic noise or speaker overlap may introduce artifacts into the reasoning graph. (2) Inference Latency: The iterative agentic workflow (planning and reflection) incurs higher computational cost than single-turn RAG, limiting real-time deployment. (3) Domain Specificity: Our evaluation focuses on structured meetings; generalizing to unstructured domains like movies or vlogs remains future work. Ethics Statement Dataset Sourcing and Compliance. Our dataset is constructed based on three existing public corpora: AliMeeting Yu et al. (2022a), AMI Meeting Corpus Jain et al. (2024), and DailyTalk Lee et al. (2022). We strictly adhere to the original licensing terms of these datasets (e.g., Creative Commons BY-NC-SA 4.0 and Apache License 2.0). We have reviewed the source data to ensure that no Personally Identifiable Information (PII) beyond what was originally consented to by the participants is exposed. For the AMI corpus, we utilize the specific split intended for academic research involving simulated scenarios, minimizing privacy risks associated with real-world private meetings. Human Annotation and Fair Compensation. We employ highly qualified graduate students as annotators. We ensured that all participants were compensated at a rate significantly above the local minimum wage (approximately $5 per hour), respecting fair labor standards. We also implemented strict protocols to protect annotators from exposure to any potentially harmful or offensive content, although the source datasets are generally free of such material. Mitigation of LLM Bias and Hallucination. We acknowledge that utilizing Large Language Models (LLMs) for data generation may introduce inherent biases or hallucinations. To mitigate this, our human-in-the-loop pipeline strictly enforces factual consistency checks. We explicitly instructed annotators to discard questions that reinforce stereotypes or rely on hallucinated events not present in the audio. The high inter-annotator agreement (κ=0.95κ=0.95) suggests that our verification process effectively filters out low-quality or biased generations. Potential Societal Impact. The technology proposed in this paper aims to enhance productivity and accessibility by structuring long-form audio. However, we recognize the potential dual-use risk regarding unauthorized surveillance or privacy intrusion in workplace settings. We strongly advocate that the deployment of such meeting analysis tools must be accompanied by transparent consent from all recorded parties and robust data encryption measures. This dataset is released solely for research purposes to advance the field of interpretable audio understanding. References Abootorabi and Asgari (2025) M. M. Abootorabi and E. Asgari CLASP: contrastive language-speech pretraining for multilingual multimodal information retrieval. In Advances in Information Retrieval: 47th European Conference on Information Retrieval, ECIR 2025, Lucca, Italy, April 6–10, 2025, Proceedings, Part IV, Berlin, Heidelberg, p. 10–20. External Links: ISBN 978-3-031-88716-1, Link, Document Cited by: §G.1, §H.2, §4.1. Chen et al. (2024a) J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu BGE m3-embedding: multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. External Links: 2402.03216, Link Cited by: §G.1, §H.1, §4.1. Chen et al. (2024b) Y. Chen, S. Zheng, H. Wang, L. Cheng, Q. Chen, S. Zhang, and J. Li ERes2NetV2: boosting short-duration speaker verification performance with computational efficiency. External Links: 2406.02167, Link Cited by: Appendix H. Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, L. Marris, S. Petulla, C. Gaffney, A. Aharoni, N. Lintz, T. C. Pais, H. Jacobsson, I. Szpektor, N. Jiang, K. Haridasan, A. Omran, N. Saunshi, D. Bahri, G. Mishra, E. Chu, T. Boyd, B. Hekman, A. Parisi, C. Zhang, K. Kawintiranon, T. Bedrax-Weiss, O. Wang, Y. Xu, O. Purkiss, U. Mendlovic, I. Deutel, N. Nguyen, A. Langley, F. Korn, L. Rossazza, A. Ramé, S. Waghmare, H. Miller, N. Byrd, A. Sheshan, R. Hadsell, S. Bhardwaj, P. Janus, T. Rissa, D. Horgan, A. Abdagic, L. Belenki, J. Allingham, A. Singh, T. Guidroz, S. Srinivasan, H. Schmit, K. Chiafullo, A. Elisseeff, N. Jha, P. Kolhar, L. Berrada, F. Ding, X. Si, S. B. Mallick, F. Och, S. Erell, E. Ni, T. Latkar, S. Yang, P. Sirkovic, Z. Feng, R. Leland, R. Hornung, G. Wu, C. Blundell, H. Alvari, P. Huang, C. Yip, S. Deur, L. Liu, G. Surita, P. Duque, D. Damen, J. Jia, A. Guez, M. Mircea, A. Sinha, A. Magni, P. Stradomski, T. Marian, V. Galić, W. Chen, H. Husain, A. Singhal, D. Grewe, F. Aubet, S. Song, L. Blanco, L. Rechis, L. Ho, R. Munoz, K. Zheng, J. Hamrick, K. Mather, H. Taitelbaum, E. Rutherford, Y. Lei, K. Chen, A. Shukla, E. Moreira, E. Doi, B. Isik, N. Shabat, D. Rogozińska, K. Kolipaka, J. Chang, E. Vušak, S. Venkatachary, S. Noghabi, T. Bharti, Y. Jun, A. Zaks, S. Green, J. Challagundla, W. Wong, M. Mohammad, D. Hirsch, Y. Cheng, I. Naim, L. Proleev, D. Vincent, A. Singh, M. Krikun, D. Krishnan, Z. Ghahramani, A. Atias, R. Aggarwal, C. Kirov, D. Vytiniotis, C. Koh, A. Chronopoulou, P. Dogra, V. Ion, G. Tyen, J. Lee, F. Weissenberger, T. Strohman, A. Balakrishna, J. Rae, M. Velic, R. de Liedekerke, O. Elyada, W. Yuan, C. Liu, L. Shani, S. Kishchenko, B. Alessio, Y. Li, R. Song, S. Kwei, O. Jankowski, A. Pappu, Y. Namiki, Y. Ma, N. Tripuraneni, C. Cherry, M. Ikonomidis, Y. Ling, C. Ji, B. Westberg, A. Wright, D. Yu, D. Parkinson, S. Ramaswamy, J. Connor, S. H. Yeganeh, S. Grover, G. Kenwright, L. Litchev, C. Apps, A. Tomala, F. Halim, A. Castro-Ros, Z. Li, A. Boral, P. Sho, M. Yarom, E. Malmi, D. Klinghoffer, R. Lin, A. Ansell, P. K. S, S. Zhao, S. Zuo, A. Santoro, H. Cheng, S. Demmessie, Y. Liu, N. Brichtova, A. Culp, N. Braun, D. Graur, W. Ng, N. Mehta, A. Phillips, P. Sundberg, V. Godbole, F. Liu, Y. Katariya, D. Rim, M. Seyedhosseini, S. Ammirati, J. Valfridsson, M. Malihi, T. Knight, A. Toor, T. Lampe, A. Ittycheriah, L. Chiang, C. Yeung, A. Fréchette, J. Rao, H. Wang, H. Srivastava, R. Zhang, R. Rhodes, A. Brand, D. Weesner, I. Figotin, F. Gimeno, R. Fellinger, P. Marcenac, J. Leal, E. Marcus, V. Cotruta, R. Cabrera, S. Luo, D. Garrette, V. Axelrod, S. Baltateanu, D. Barker, D. Chen, H. Toma, B. Ingram, J. Riesa, C. Kulkarni, Y. Zhang, H. Liu, C. Wang, M. Polacek, W. Wu, K. Hui, A. N. Reyes, Y. Su, M. Barnes, I. Malhi, A. Siddiqui, Q. Feng, M. Damaschin, D. Pighin, A. Steiner, S. Yang, R. S. Boppana, S. Ivanov, A. Kandoor, A. Shah, A. Mujika, D. Huang, C. A. Choquette-Choo, M. Patel, T. Yu, T. Creswell, Jerry, Liu, C. Barros, Y. Razeghi, A. Roy, P. Culliton, B. Xiong, J. Pan, T. Strohmann, T. Powell, B. Seal, D. DeCarlo, P. Shyam, K. Katircioglu, X. Wang, C. Hardin, I. Odisho, J. Broder, O. Chang, A. Nair, A. Shtefan, M. O’Brien, M. Agarwal, S. Potluri, S. Goyal, A. Jhindal, S. Thakur, Y. Stuken, J. Lyon, K. Toutanova, F. Feng, A. Wu, B. Horn, A. Wang, A. Cullum, G. Taubman, D. Shrivastava, C. Shi, H. Tomlinson, R. Patel, T. Tu, A. M. Oflazer, F. Pongetti, M. Yang, A. A. Taïga, V. Perot, N. W. Pierse, F. Han, Y. Drori, I. Iturrate, A. Chakrabarti, L. Yeung, D. Dopson, Y. Chen, A. Kulshreshtha, T. Guo, P. Pham, T. Schuster, J. Chen, A. Polozov, J. Xing, H. Zhou, P. Kacham, D. Kukliansky, A. Miech, S. Yaroshenko, E. Chi, S. Douglas, H. Fei, M. Blondel, P. Myla, L. Madmoni, X. Wu, D. Keysers, K. Kjems, I. Albuquerque, L. Yu, J. D’sa, M. Plantan, V. Ionescu, J. S. Elias, A. Gupta, M. R. Vuyyuru, F. Alcober, T. Zhou, K. Ji, F. Hartmann, S. Puttagunta, H. Song, E. Amid, A. Stefanoiu, A. Lee, P. Pucciarelli, E. Wang, A. Raul, S. Petrov, I. Tian, V. Anklin, N. Nti, V. Gomes, M. Schumacher, G. Vesom, A. Panagopoulos, K. Bousmalis, D. Andor, J. Jacob, Y. Zhang, B. Rosgen, M. Kecman, M. Tung, A. Belias, N. Goodman, P. Covington, B. Wieder, N. Saxena, E. Davoodi, M. Huang, S. Maddineni, V. Roulet, F. Campbell-Ajala, P. G. Sessa, Xintian, Wu, G. Lai, P. Collins, A. Haig, V. Sakenas, X. Xu, M. Giustina, L. E. Shafey, P. Charoenpanit, S. Garg, J. Ainslie, B. Severson, M. G. Arenas, S. Pathak, S. Rajayogam, J. Feng, M. Bakker, S. Li, N. Wichers, J. Rogers, X. Geng, Y. Li, R. Jagerman, C. Jia, N. Olmert, D. Sharon, M. Mauger, S. Mariserla, H. Ma, M. Mohabey, K. Kim, A. Andreev, S. Pollom, J. Love, V. Jain, P. Agrawal, Y. Schroecker, A. Fortin, M. Warmuth, J. Liu, A. Leach, I. Blok, G. P. Girirajan, R. Aharoni, B. Uria, A. Sozanschi, D. Goldberg, L. Ionita, M. T. Ribeiro, M. Zlocha, V. Birodkar, S. Lachgar, L. Yuan, H. Choudhury, M. Ginsberg, F. Zheng, G. Dibb, E. Graves, S. Lokhande, G. Rasskin, G. Muraru, C. Quick, S. Tata, P. Sermanet, A. Chawla, I. Karo, Y. Wang, S. Zhang, O. Keller, A. Dragan, G. Su, I. Chou, X. Liu, Y. Tao, S. Prabhakara, M. Wilson, R. Liu, S. Wang, G. Evans, D. Du, A. Castaño, G. Prasad, M. E. Mahdy, S. Gerlach, M. Reid, J. Kahn, A. Zait, T. S. Pillai, T. Ulrich, G. Wang, J. Wassenberg, E. Farkash, K. Yalasangi, C. Wang, M. Bauza, S. Bucher, T. Liu, J. Yan, G. Leung, V. Sindhwani, P. Barnes, A. Singh, I. Jurin, J. Chang, N. K. Bhumihar, S. Eiger, G. Citovsky, B. Withbroe, Z. Li, S. Xue, N. D. Santo, G. Stoyanov, Y. Raimond, S. Zheng, Y. Gao, V. Listík, S. Kwasiborski, R. Saputro, A. Ozturel, G. Mallya, K. Majmundar, R. West, P. Caron, J. Wei, L. Castrejon, S. Vikram, D. Ramachandran, N. Dhawan, J. Park, S. Smoot, G. van den Driessche, Y. Blau, C. Malik, W. Liang, R. Hirsch, C. N. dos Santos, E. Weinstein, A. van den Oord, S. Lall, N. FitzGerald, Z. Jiang, X. Yang, D. Webster, A. Elqursh, A. Pope, G. Rotival, D. Raposo, W. Zhu, J. Dean, S. Alabed, D. Tran, A. Gupta, Z. Gleicher, J. Austin, E. Rosseel, M. Umekar, D. Das, Y. Sun, K. Chen, K. Misiunas, X. Zhou, Y. Di, A. Loo, J. Newlan, B. Li, V. Ramasesh, Y. Xu, A. Chen, S. Gandhe, R. Soricut, N. Gupta, S. Hu, S. El-Sayed, X. Garcia, I. Brusilovsky, P. Chen, A. Bolt, L. Huang, A. Gurney, Z. Zhang, A. Pritzel, J. Wilkiewicz, B. Seybold, B. K. Shamanna, F. Fischer, J. Dean, K. Gill, R. Mcilroy, A. Bhowmick, J. Selier, A. Yang, D. Cheng, V. Magay, J. Tan, D. Varma, C. Walder, T. Kocisky, R. Nakashima, P. Natsev, M. Kwong, I. Gog, C. Zhang, S. Dieleman, T. Jimma, A. Ryabtsev, S. Brahma, D. Steiner, D. Du, A. Žužul, M. Žanić, M. Raghavachari, W. Gierke, Z. Zheng, D. Petrova, Y. Dauphin, Y. Liu, I. Kessler, S. Hand, C. Duvarney, S. Kim, H. Lee, L. Hussenot, J. Hui, J. Smith, D. Jain, J. Xia, G. S. Tomar, K. Amiri, D. Phan, F. Fuchs, T. Weyand, N. Tomasev, A. Cordell, X. Liu, J. Mallinson, P. Joshi, A. Crawford, A. Suggala, S. Chien, N. Fernando, M. Sanchez-Vargas, D. Williams, P. Crone, X. Luo, I. Karpov, J. Shan, T. Thurk, R. Strudel, P. Voigtlaender, P. Patil, T. Dozat, A. Khodaei, S. Singla, P. Ambroszczyk, Q. Wu, Y. Chang, B. Roark, C. Hegde, T. Ding, A. Filos, Z. Wu, A. S. Pinto, S. Liu, S. Khanna, A. Pandey, S. Mcloughlin, Q. Li, S. Haves, A. Zhou, E. Buchatskaya, I. Leal, P. de Boursac, N. Akazawa, N. Anderson, T. Chen, K. Somandepalli, C. Liang, S. Goenka, S. Winkler, A. Grushetsky, Y. Ding, J. Smith, F. Ye, J. Pont-Tuset, E. Li, R. Li, T. Golany, D. Wegner, T. Jiang, O. Barak, Y. Shangguan, E. Vértes, R. Wong, J. Bornschein, A. Tudor, M. Bevilacqua, T. Schaul, A. S. Rawat, Y. Zhao, K. Axiotis, L. Meng, C. McLean, J. Lai, J. Beattie, N. Kushman, Y. Liu, B. Kutzman, F. Lang, J. Ye, P. Netrapalli, P. Mishra, M. Khan, M. Goel, R. Willoughby, D. Tian, H. Zhuang, J. Chen, Z. Tsai, T. Kementsietsidis, A. Khare, J. Keeling, K. Xu, N. Waters, F. Altché, A. Popat, B. Mittal, D. Saxton, D. E. Badawy, M. Mathieu, Z. Zheng, H. Zhou, N. Ranka, R. Shin, Q. Duan, T. Salimans, I. Mihailescu, U. Shaham, M. Chang, Y. Assael, N. Dikkala, M. Izzard, V. Cohen-Addad, C. Graves, V. Feinberg, G. Chung, D. Strouse, D. Karmon, S. Sharifzadeh, Z. Ashwood, K. Pham, J. Blanton, A. Vasiloff, J. Barber, M. Geller, A. Zhou, F. Zubach, T. Huang, L. Zhang, H. Gupta, M. Young, J. Proskurnia, R. Votel, V. Gabeur, G. Barcik, A. Tripathi, H. Yu, G. Yan, B. Changpinyo, F. Pavetić, A. Coyle, Y. Fujii, J. G. Mendez, T. Zhou, H. Rajamani, B. Hechtman, E. Cao, D. Juan, Y. Tan, V. Dalibard, Y. Du, N. Clay, K. Yao, W. Jia, D. Vijaykumar, Y. Zhou, X. Bai, W. Hung, S. Pecht, G. Todorov, N. Khadke, P. Gupta, P. Lahoti, A. Autef, K. Duddu, J. Lee-Thorp, A. Bykovsky, T. Misiunas, S. Flennerhag, S. Thangaraj, J. McGiffin, Z. Nado, M. Kunesch, A. Noever, A. Hertz, M. Liang, V. Stone, E. Palmer, S. Daruki, A. Pramanik, S. Põder, A. Kyker, M. Khan, E. Sluzhaev, M. Ritter, A. Ruderman, W. Zhou, C. Nagpal, K. Vodrahalli, G. Necula, P. Barham, E. Pavlick, J. Hartford, I. Shafran, L. Zhao, M. Mikuła, T. Eccles, H. Shimokawa, K. Garg, L. Vilnis, H. Chen, I. Shumailov, K. Lee, A. Abdelhamed, M. Xie, V. Cohen, E. Hlavnova, D. Malkin, C. Sitawarin, J. Lottes, P. Coquinot, T. Yu, S. Kumar, J. Zhang, A. Mahendru, Z. Ahmed, J. Martens, T. Chen, A. Boag, D. Peng, C. Devin, A. Klimovskiy, M. Phuong, D. Vainstein, J. Xie, B. Ramabhadran, N. Howard, X. Yu, G. Goswami, J. Cui, S. Shleifer, M. Pinto, C. Yeh, M. Yang, S. Javanmardi, D. Ethier, C. Lee, J. Orbay, S. Kotecha, C. Bromberg, P. Shaw, J. Thornton, A. G. Rosenthal, S. Gu, M. Thomas, I. Gemp, A. Ayyar, A. Ushio, A. Selvan, J. Wee, C. Liu, M. Majzoubi, W. Yu, J. Abernethy, T. Liechty, R. Pan, H. Nguyen, Qiong, Hu, S. Perrin, A. Arora, E. Pitler, W. Wang, K. Shivakumar, F. Prost, B. Limonchik, J. Wang, Y. Gao, T. Cour, S. Buch, H. Gui, M. Ivanova, P. Neubeck, K. Chan, L. Kim, H. Chen, N. Goyal, D. Chung, L. Liu, Y. Su, A. Petrushkina, J. Shen, A. Joulin, Y. Xu, S. X. Lin, Y. Kulizhskaya, C. Chelba, S. Vasudevan, E. Collins, V. Bashlovkina, T. Lu, D. Fritz, J. Park, Y. Zhou, C. Su, R. Tanburn, M. Sushkov, M. Rasquinha, J. Li, J. Prendki, Y. Li, P. LV, S. Sharma, H. Fitoussi, H. Huang, A. Dai, P. Dao, M. Burrows, H. Prior, D. Qin, G. Pundak, L. L. Sjoesund, A. Khurshudov, Z. Zhu, A. Webson, E. Kemp, T. Tan, S. Agrawal, S. Sargsyan, L. Cheng, J. Stephan, T. Kwiatkowski, D. Reid, A. Byravan, A. H. Michaely, N. Heess, L. Zhou, S. Goenka, V. Carpenter, A. Levskaya, B. Wang, R. Roberts, R. Leblond, S. Chikkerur, S. Ginzburg, M. Chang, R. Riachi, Chuqiao, Xu, Z. Borsos, M. Pliskin, J. Pawar, M. Lustman, H. Kirkwood, A. Anand, A. Chaudhary, N. Kalb, K. Milan, S. Augenstein, A. Goldie, L. Prince, K. Raman, Y. Sun, V. Xia, A. Cohen, Z. Huo, J. Camp, S. Ellis, L. Zilka, D. V. Torres, L. Patel, S. Arora, B. Chan, J. Adler, K. Ayoub, J. Liang, F. Jamil, J. Jiang, S. Baumgartner, H. Sun, Y. Karov, Y. Akulov, H. Zheng, I. Cai, C. Fantacci, J. Rubin, A. R. Acha, M. Wang, N. D’Souza, R. Sathyanarayana, S. Dai, S. Rowe, A. Simanovsky, O. Goldman, Y. Kuang, X. Pan, A. Rosenberg, T. Rojas-Esponda, P. Dutta, A. Zeng, I. Jurenka, G. Farquhar, Y. Bansal, S. Iqbal, B. Roelofs, G. Joung, P. Beak, C. Ryu, R. Poplin, Y. Wu, J. Alayrac, S. Buthpitiya, O. Ronneberger, C. Habtegebriel, W. Li, P. Cavallaro, A. Wei, G. Bensky, T. Denk, H. Ganapathy, J. Stanway, P. Joshi, F. Bertolini, J. Lo, O. Ma, Z. Charles, G. Sampemane, H. Sahni, X. Chen, H. Askham, D. Gaddy, P. Young, J. Tan, M. Eyal, A. Bražinskas, L. Zhong, Z. Wu, M. Epstein, K. Bailey, A. Hard, K. Lee, S. Goldshtein, A. Ruiz, M. Badawi, M. Lochbrunner, J. Kearns, A. Brown, F. Pardo, T. Weber, H. Yang, P. Jiang, B. Akin, Z. Fu, M. Wainwright, C. Zou, M. Gaba, P. Manzagol, W. Kan, Y. Song, K. Zainullina, R. Lin, J. Ko, S. Deshmukh, A. Jindal, J. Svensson, D. Tyam, H. Zhao, C. Kaeser-Chen, S. Baird, P. Moradi, J. Hall, Q. Guo, V. Tsang, B. Liang, F. Pereira, S. Ganesh, I. Korotkov, J. Adamek, S. Thiagarajan, V. Tran, C. Chen, C. Tar, S. Jain, I. Dasgupta, T. Bilal, D. Reitter, K. Zhao, G. Vezzani, Y. Gehman, P. Mehta, L. Beltrone, X. Dotiwalla, S. Guadarrama, Z. Abbas, S. Karp, P. Georgiev, C. Ferng, M. Brockschmidt, L. Peng, C. Hirnschall, V. Verma, Y. Bi, Y. Xiao, A. Dabush, K. Xu, P. Wallis, R. Parker, Q. Wang, Y. Xu, I. Safarli, D. Tewari, Y. Zhang, S. Kim, A. Gesmundo, M. Thomas, S. Levi, A. Chowdhury, K. Rao, P. Garst, S. Conway-Rahman, H. Ran, K. McKinney, Z. Xiao, W. Yu, R. Agrawal, A. Stjerngren, C. Ionescu, J. Chen, V. Sharma, J. Chiu, F. Liu, K. Franko, C. Sanford, X. Cai, P. Michel, S. Ganapathy, J. Labanowski, Z. Garrett, B. Vargas, S. Sun, B. Gale, T. Buschmann, G. Desjardins, N. Ghelani, P. Jain, M. Verma, C. Asawaroengchai, J. Eisenschlos, J. Harlalka, H. Kazawa, D. Metzler, J. Howland, Y. Jian, J. Ades, V. Shah, T. Gangwani, S. Lee, R. Ring, S. M. Hernandez, D. Reich, A. Sinha, A. Sathe, J. Kovac, A. Gill, A. Kannan, A. D’olimpio, M. Sevenich, J. Whang, B. Kim, K. C. Sim, J. Chen, J. Zhang, S. Lall, Y. Matias, B. Jia, A. Friesen, S. Nasso, A. Thapliyal, B. Perozzi, T. Yu, A. Shekhawat, S. Huda, P. Grabowski, E. Wang, A. Sreevatsa, H. Dib, M. Hassen, P. Schuh, V. Milutinovic, C. Welty, M. Quinn, A. Shah, B. Wang, G. Barth-Maron, J. Frye, N. Axelsson, T. Zhu, Y. Ma, I. Giannoumis, H. Sedghi, C. Ye, Y. Luan, K. Aydin, B. Chandra, V. Sampathkumar, R. Huang, V. Lavrenko, A. Eleryan, Z. Hong, S. Hansen, S. M. Carthy, B. Samanta, D. Ćevid, X. Wang, F. Li, M. Voznesensky, M. Hoffman, A. Terzis, V. Sehwag, G. Fidel, L. He, M. Cai, Y. He, A. Feng, M. Nikoltchev, S. Phatale, J. Chase, R. Lawton, M. Zhang, T. Ouyang, M. Tragut, M. H. Manshadi, A. Narayanan, J. Shen, X. Gao, T. Bolukbasi, N. Roy, X. Li, D. Golovin, L. Panait, Z. Qin, G. Han, T. Anthony, S. Kudugunta, V. Patraucean, A. Ray, X. Chen, X. Yang, T. Bhatia, P. Talluri, A. Morris, A. Ražnatović, B. Brownfield, J. An, S. Peng, P. Kane, C. Zheng, N. Duduta, J. Kessinger, J. Noraky, S. Liu, K. Rong, P. Veličković, K. Rush, A. Goldin, F. Wei, S. M. R. Garlapati, C. Pantofaru, O. Kwon, J. Ni, E. Noland, J. D. Trapani, F. Beaufays, A. G. Roy, Y. Chow, A. Turker, G. Cideron, L. Mei, J. Clark, Q. Dou, M. Bošnjak, R. Leith, Y. Du, A. Yazdanbakhsh, M. Nasr, C. Kwak, S. S. Sheth, A. Kaskasoli, A. Anand, B. Lakshminarayanan, S. Jerome, D. Bieber, C. Chu, A. Senges, T. Shen, M. Sridhar, N. Ndebele, B. Beyret, S. Mohamed, M. Chen, M. Freitag, J. Guo, L. Liu, P. Roit, H. Chen, S. Yan, T. Stone, J. Co-Reyes, J. Cole, S. Scellato, S. Azizi, H. Hashemi, A. Jin, A. Iyer, M. Valentine, A. György, A. Ahuja, D. H. Diaz, C. Lee, N. Clement, W. Kong, D. Garmon, I. Watts, K. Bhatia, K. Gupta, M. Miecnikowski, H. Vallet, A. Taly, E. Loper, S. Joshi, J. Atwood, J. Chick, M. Collier, F. Iliopoulos, R. Trostle, B. Gunel, R. Leal-Cavazos, A. M. Hrafnkelsson, M. Guzman, X. Ju, A. Forbes, J. Emond, K. Chauhan, B. Caine, L. Xiao, W. Zeng, A. Moufarek, D. Murphy, M. Meng, N. Gupta, F. Riedel, A. Das, E. Lawal, S. Narayan, T. Sosea, J. Swirhun, L. Friso, B. Neyshabur, J. Lu, S. Girgin, M. Wunder, E. Yvinec, A. Pyne, V. Carbune, S. Rijhwani, Y. Guo, T. Doshi, A. Briukhov, M. Bain, A. Hitron, X. Wang, A. Gupta, K. Chen, C. Du, W. Zhang, D. Shah, A. Akula, M. Dylla, A. Kachra, W. Kuo, T. Zou, L. Wang, L. Xu, J. Zhu, J. Snyder, S. Menon, O. Firat, I. Mordatch, Y. Yuan, N. Ponomareva, R. Blevins, L. Moore, W. Wang, P. Chen, M. Scholz, A. Dwornik, J. Lin, S. Li, D. Antognini, T. I, X. Song, M. Miller, U. Kalra, A. Raveret, O. Akerlund, F. Wu, A. Nystrom, N. Godbole, T. Liu, H. DeBalsi, J. Zhao, B. Liu, A. Caciularu, L. Lax, U. Khandelwal, V. Langston, E. Bailey, S. Lattanzi, Y. Wang, N. Kovelamudi, S. Mondal, G. Guruganesh, N. Hua, O. Roval, P. Wesołowski, R. Ingale, J. Halcrow, T. Sohn, C. Angermueller, B. Raad, E. Stickgold, E. Lu, A. Kosik, J. Xie, T. Lillicrap, A. Huang, L. L. Zhang, D. Paulus, C. Farabet, A. Wertheim, B. Wang, R. Joshi, C. Ko, Y. Wu, S. Agrawal, L. Lin, X. Sheng, P. Sung, T. Breland-King, C. Butterfield, S. Gawde, S. Singh, Q. Zhang, R. Apte, S. Shetty, A. Hutter, T. Li, E. Salesky, F. Lebron, J. Kanerva, M. Paganini, A. Nguyen, R. Vallu, J. Peter, S. Velury, D. Kao, J. Hoover, A. Bortsova, C. Bishop, S. Jakobovits, A. Agostini, A. Agarwal, C. Liu, C. Kwong, S. Tavakkol, I. Bica, A. Greve, A. GP, J. Marcus, L. Hou, T. Duerig, R. Moroshko, D. Lacey, A. Davis, J. Amelot, G. Wang, F. Kim, T. Strinopoulos, H. Wan, C. L. Lan, S. Krishnan, H. Tang, P. Humphreys, J. Bai, I. H. Shtacher, D. Machado, C. Pang, K. Burke, D. Liu, R. Aravamudhan, Y. Song, E. Hirst, A. Singh, B. Jou, L. Bai, F. Piccinno, C. K. Fu, R. Alazard, B. Meiri, D. Winter, C. Chen, M. Zhang, J. Heitkaemper, J. Lambert, J. Lee, A. Frömmgen, S. Rogulenko, P. Nair, P. Niemczyk, A. Bulyenov, B. Xu, H. Shemtov, M. Zadimoghaddam, S. Toropov, M. Wirth, H. Dai, S. Gollapudi, D. Zheng, A. Kurakin, C. Lee, K. Bullard, N. Serrano, I. Balazevic, Y. Li, J. Schalkwyk, M. Murphy, M. Zhang, K. Sequeira, R. Datta, N. Agrawal, C. Sutton, N. Attaluri, M. Chiang, W. Farhan, G. Thornton, K. Lin, T. Choma, H. Nguyen, K. Dasgupta, D. Robinson, I. Comşa, M. Riley, A. Pillai, B. Mustafa, B. Golan, A. Zandieh, J. Lespiau, B. Porter, D. Ross, S. Rajayogam, M. Agarwal, S. Venugopalan, B. Shahriari, Q. Yan, H. Xu, T. Tobin, P. Dubov, H. Shi, A. Recasens, A. Kovsharov, S. Borgeaud, L. Dery, S. Vasanth, E. Gribovskaya, L. Qiu, M. Mahdieh, W. Skut, E. Nielsen, C. Zheng, A. Yu, C. G. Bostock, S. Gupta, A. Archer, C. Rawles, E. Davies, A. Svyatkovskiy, T. Tsai, Y. Halpern, C. Reisswig, B. Wydrowski, B. Chang, J. Puigcerver, M. H. Taege, J. Li, E. Schnider, X. Li, D. Dena, Y. Xu, U. Telang, T. Shi, H. Zen, K. Kastner, Y. Ko, N. Subramaniam, A. Kumar, P. Blois, Z. Dai, J. Wieting, Y. Lu, Y. Zeldes, T. Xie, A. Hauth, A. Ţifrea, Y. Li, S. El-Husseini, D. Abolafia, H. Zhou, W. Ding, S. Ghalebikesabi, C. Guía, A. Maksai, Á. Weisz, S. Arik, N. Sukhanov, A. Świetlik, X. Jia, L. Yu, W. Wang, M. Brand, D. Bloxwich, S. Kirmani, Z. Chen, A. Go, P. Sprechmann, N. Kannen, A. Carin, P. Sandhu, I. Edkins, L. Nooteboom, J. Gupta, L. Maggiore, J. Azizi, Y. Pritch, P. Yin, M. Gupta, D. Tarlow, D. Smith, D. Ivanov, M. Babaeizadeh, A. Goel, S. Kambala, G. Chu, M. Kastelic, M. Liu, H. Soltau, A. Stone, S. Agrawal, M. Kim, K. Soparkar, S. Tadepalli, O. Bunyan, R. Soh, A. Kannan, D. Kim, B. J. Chen, A. Halumi, S. Roy, Y. Wang, O. Sercinoglu, G. Gibson, S. Bhatnagar, M. Sano, D. von Dincklage, Q. Ren, B. Mitrevski, M. Olšák, J. She, C. Doersch, Jilei, Wang, B. Liu, Q. Tan, T. Yakar, T. Warkentin, A. Ramirez, C. Lebsack, J. Dillon, R. Mathews, T. Cobley, Z. Wu, Z. Chen, J. Simon, S. Nath, T. Sainath, A. Bendebury, R. Julian, B. Mankalale, D. Ćurko, P. Zacchello, A. R. Brown, K. Sodhia, H. Howard, S. Caelles, A. Gupta, G. Evans, A. Bulanova, L. Katzen, R. Goldenberg, A. Tsitsulin, J. Stanton, B. Schillings, V. Kovalev, C. Fry, R. Shah, K. Lin, S. Upadhyay, C. Li, S. Radpour, M. Maggioni, J. Xiong, L. Haas, J. Brennan, A. Kamath, N. Savinov, A. Nagrani, T. Yacovone, R. Kappedal, K. Andriopoulos, L. Lao, Y. Li, G. Rozhdestvenskiy, K. Hashimoto, A. Audibert, S. Austin, D. Rodriguez, A. Ruoss, G. Honke, D. Karkhanis, X. Xiong, Q. Wei, J. Huang, Z. Leng, V. Premachandran, S. Bileschi, G. Evangelopoulos, T. Mensink, J. Pavagadhi, D. Teplyashin, P. Chang, L. Xue, G. Tanzer, S. Goldman, K. Patel, S. Li, J. Wiesner, I. Zheng, I. Stewart-Binks, J. Han, Z. Li, L. Luo, K. Lenc, M. Lučić, F. Xue, R. Mullins, A. Guseynov, C. Chang, I. Galatzer-Levy, A. Zhang, G. Bingham, G. Hu, A. Hartman, Y. Ma, J. Griffith, A. Irpan, C. Radebaugh, S. Yue, L. Fan, V. Ungureanu, C. Sorokin, H. Teufel, P. Li, R. Anil, D. Paparas, T. Wang, C. Lin, H. Peng, M. Shum, G. Petrovic, D. Brady, R. Nguyen, K. Macherey, Z. Li, H. Singh, M. Yenugula, M. Iinuma, X. Chen, K. Kopparapu, A. Stern, S. Dave, C. Thekkath, F. Perot, A. Kumar, F. Li, Y. Xiao, M. Bilotti, M. H. Bateni, I. Noble, L. Lee, A. Vázquez-Reina, J. Salazar, X. Yang, B. Wang, E. Gruzewska, A. Rao, S. Raghuram, Z. Xu, E. Ben-David, J. Mei, S. Dalmia, Z. Zhang, Y. Liu, G. Bansal, H. Pankov, S. Schwarcz, A. Burns, C. Chan, S. Sanghai, R. Liang, E. Liang, A. He, A. Stuart, A. Narayanan, Y. Zhu, C. Frank, B. Fatemi, A. Sabne, O. Lang, I. Bhattacharya, S. Settle, M. Wang, B. McMahan, A. Tacchetti, L. B. Soares, M. Hadian, S. Cabi, T. Chung, N. Putikhin, G. Li, J. Chen, A. Tarango, H. Michalewski, M. Kazemi, H. Masoom, H. Sheftel, R. Shivanna, A. Vadali, R. Comanescu, D. Reid, J. Moore, A. Neelakantan, M. Sander, J. Herzig, A. Rosenberg, M. Dehghani, J. Choi, M. Fink, R. Hayes, E. Ge, S. Weng, C. Ho, J. Karro, K. Krishna, L. N. Thiet, A. Skerry-Ryan, D. Eppens, M. Andreetto, N. Sarma, S. Bonacina, B. K. Ayan, M. Nawhal, Z. Shan, M. Dusenberry, S. Thakoor, S. Gubbi, D. D. Nguyen, R. Tsarfaty, S. Albanie, J. Mitrović, M. Gandhi, B. Chen, A. Epasto, G. Stephanov, Y. Jin, S. Gehman, A. Amini, J. Weber, F. Behbahani, S. Xu, M. Allamanis, X. Chen, M. Ott, C. Sha, M. Jastrzebski, H. Qi, D. Greene, X. Wu, A. Toki, D. Vlasic, J. Shapiro, R. Kotikalapudi, Z. Shen, T. Saeki, S. Xie, A. Cassirer, S. Bharadwaj, T. Kiyono, S. Bhojanapalli, E. Rosenfeld, S. Ritter, J. Mao, J. G. Oliveira, Z. Egyed, B. Bandemer, E. Parisotto, K. Kinoshita, J. Pluto, P. Maniatis, S. Li, Y. Guo, G. Ghiasi, J. Tarbouriech, S. Chatterjee, J. Jin, Katrina, Xu, J. Palomaki, S. Arnold, M. Sewak, F. Piccinini, M. Sharma, B. Albrecht, S. Purser-haskell, A. Vaswani, C. Chen, M. Wisniewski, Q. Cao, J. Aslanides, N. M. Phu, M. Sieb, L. Agubuzu, A. Zheng, D. Sohn, M. Selvi, A. Andreassen, K. Subudhi, P. Eruvbetine, O. Woodman, T. Mery, S. Krause, X. Ren, X. Ma, J. Luo, D. Chen, W. Fan, H. Griffiths, C. Schuler, A. Li, S. Zhang, J. Sarr, S. Luo, R. Patana, M. Watson, D. Naboulsi, M. Collins, S. Sidhwani, E. Hoogeboom, S. Silver, E. Caveness, X. Zhao, M. Rodriguez, M. Deines, L. Bai, P. Griffin, M. Tagliasacchi, E. Xue, S. R. Babbula, B. Pang, N. Ding, G. Shen, E. Peake, R. Crocker, S. S. Raghvendra, D. Swisher, W. Han, R. Singh, L. Wu, V. Pchelin, T. Munkhdalai, D. Alon, G. Bacon, E. Robles, J. Bulian, M. Johnson, G. Powell, F. T. Ferreira, Y. Li, F. Benzing, M. Velimirović, H. Soyer, W. Kong, Tony, Nguyên, Z. Yang, J. Liu, J. van Amersfoort, D. Gillick, B. Sun, N. Rauschmayr, K. Zhang, S. Zhan, T. Zhou, A. Frolov, C. Yang, D. Vnukov, L. Rouillard, H. Li, A. Mandhane, N. Fallen, R. Venkataraman, C. H. Hu, J. Brennan, J. Lee, J. Chang, M. Sundermeyer, Z. Pan, R. Ke, S. Tong, A. Fabrikant, W. Bono, J. Gu, R. Foley, Y. Mao, M. Delakis, D. Bhaswar, R. Frostig, N. Li, A. Zipori, C. Hope, O. Kozlova, S. Mishra, J. Djolonga, C. Schiff, M. A. Merey, E. Briakou, P. Morgan, A. Wan, A. Hassidim, R. Skerry-Ryan, K. Sengupta, M. Jasarevic, P. Kallakuri, P. Kunkle, H. Brennan, T. Lieber, H. Mansoor, J. Walker, B. Zhang, A. Xie, G. Žužić, A. Chukwuka, A. Druinsky, D. Cho, R. Yao, F. Naeem, S. Butt, E. Kim, Z. Jia, M. Jordan, A. Lelkes, M. Kurzeja, S. Wang, J. Zhao, A. Over, A. Chakladar, M. Prasetya, N. Jha, S. Ganapathy, Y. Cong, P. Shroff, C. Saroufim, S. Miryoosefi, M. Hammad, T. Nasir, W. Xi, Y. Gao, Y. Maeng, B. Hora, C. Cheng, P. Haghani, Y. Lewenberg, C. Lu, M. Matysiak, N. Raisinghani, H. Wang, L. Baugher, R. Sukthankar, M. Giang, J. Schultz, N. Fiedel, M. Chen, C. Lee, T. Dey, H. Zheng, S. Paul, C. Smith, A. Ly, Y. Wang, R. Bansal, B. Perz, S. Ricco, S. Blank, V. Keshava, D. Sharma, M. Chow, K. Lad, K. Jalan, S. Osindero, C. Swanson, J. Scott, A. Ilić, X. Li, S. R. Jonnalagadda, A. S. Soudagar, Y. Xiong, B. Batsaikhan, D. Jarrett, N. Kumar, M. Shah, M. Lawlor, A. Waters, M. Graham, R. May, S. Ramos, S. Lefdal, Z. Cankara, N. Cano, B. O’Donoghue, J. Borovik, F. Liu, J. Grimstad, M. Alnahlawi, K. Tsihlas, T. Hudson, N. Grigorev, Y. Jia, T. Huang, T. P. Igwe, S. Lebedev, X. Tang, I. Krivokon, F. Garcia, M. Tan, E. Jia, P. Stys, S. Vashishth, Y. Liang, B. Venkatraman, C. Gu, A. Kementsietsidis, C. Zhu, J. Jung, Y. Bai, M. J. Hosseini, F. Ahmed, A. Gupta, X. Yuan, S. Ashraf, S. Nigam, G. Vasudevan, P. Awasthi, A. M. Gilady, Z. Mariet, R. Eskander, H. Li, H. Hu, G. Garrido, P. Schlattner, G. Zhang, R. Saxena, P. Dević, K. Muralidharan, A. Murthy, Y. Zhou, M. Choi, A. Wongpanich, Z. Wang, P. Shah, Y. Xu, Y. Huang, S. Spencer, A. Chen, J. Cohan, J. Wang, J. Tompson, J. Wu, R. Haroun, H. Li, B. Huergo, F. Yang, T. Yin, J. Wendt, M. Bendersky, R. Chaabouni, J. Snaider, J. Ferret, A. Jindal, T. Thompson, A. Xue, W. Bishop, S. M. Phal, A. Sharma, Y. Sung, P. Radhakrishnan, M. Shomrat, R. Ingle, R. Vij, J. Gilmer, M. D. Istin, S. Sobell, Y. Lu, E. Nottage, D. Sadigh, J. Willcock, T. Zhang, S. Xu, S. Brown, K. Lee, G. Wang, Y. Zhu, Y. Tay, C. Kim, A. Gutierrez, A. Sharma, Y. Xian, S. Seo, C. Cui, E. Pochernina, C. Baetu, K. Jastrzębski, M. Ly, M. Elhawaty, D. Suh, E. Sezener, P. Wang, N. Yuen, G. Tucker, J. Cai, Z. Yang, C. Wang, A. Muzio, H. Qian, J. Yoo, D. Lockhart, K. R. McKee, M. Guo, M. Mehrotra, A. Mendonça, S. V. Mehta, S. Ben, C. Tekur, J. Mu, M. Zhu, V. Krakovna, H. Lee, A. Maschinot, S. Cevey, H. Choe, A. Bai, H. Srinivasan, D. Gasaway, N. Young, P. Siegler, D. Holtmann-Rice, V. Piratla, K. Baumli, R. Yogev, A. Hofer, H. van Hasselt, S. Grant, Y. Chervonyi, D. Silver, A. Hogue, A. Agarwal, K. Wang, P. Singh, F. Flynn, J. Lipschultz, R. David, L. Bellot, Y. Yang, L. Le, F. Graziano, K. Olszewska, K. Hui, A. Maurya, N. Parotsidis, W. Chen, T. Oguntebi, J. Kelley, A. Baddepudi, J. Mauerer, G. Shaw, A. Siegman, L. Yang, S. Shetty, S. Roy, Y. Song, W. Stokowiec, R. Burnell, O. Savant, R. Busa-Fekete, J. Miao, S. Ghosh, L. MacDermed, P. Lippe, M. Dektiarev, Z. Behrman, F. Mentzer, K. Nguyen, M. Wei, S. Verma, C. Knutsen, S. Dasari, Z. Yan, P. Mitrichev, X. Wang, V. Shejwalkar, J. Austin, S. Sunkara, N. Potti, Y. Virin, C. Wright, G. Liu, O. Riva, E. Pot, G. Kochanski, Q. Le, G. Balasubramaniam, A. Dhar, Y. Liao, A. Bloniarz, D. Shukla, E. Cole, J. Lee, S. Zhang, S. Kafle, S. Vashishtha, P. Mahmoudieh, G. Chen, R. Hoffmann, P. Srinivasan, A. D. Lago, Y. B. Shalom, Z. Wang, M. Elabd, A. Sharma, J. Oh, S. Kothawade, M. Le, M. Monteiro, S. Yang, K. Alarakyia, R. Geirhos, D. Mincu, H. Garnes, H. Kobayashi, S. Mariooryad, K. Krasowiak, Zhixin, Lai, S. Mourad, M. Wang, F. Bu, O. Aharoni, G. Chen, A. Goyal, V. Zubov, A. Bapna, E. Dabir, N. Kothari, K. Lamerigts, N. D. Cao, J. Shar, C. Yew, N. Kulkarni, D. Mahaarachchi, M. Joshi, Z. Zhu, J. Lichtarge, Y. Zhou, H. Muckenhirn, V. Selo, O. Vinyals, P. Chen, A. Brohan, V. Mehta, S. Cogan, R. Wang, T. Geri, W. Ko, W. Chen, F. Viola, K. Shivam, L. Wang, M. C. Elish, R. A. Popa, S. Pereira, J. Liu, R. Koster, D. Kim, G. Zhang, S. Ebrahimi, P. Talukdar, Y. Zheng, P. Poklukar, A. Mikhalap, D. Johnson, A. Vijayakumar, M. Omernick, M. Dibb, A. Dubey, Q. Hu, A. Suman, V. Aggarwal, I. Kornakov, F. Xia, W. Lowe, A. Kolganov, T. Xiao, V. Nikolaev, S. Hemingray, B. Li, J. Iljazi, M. Rybiński, B. Sandhu, P. Lu, T. Luong, R. Jenatton, V. Govindaraj, Hui, Li, G. Dulac-Arnold, W. Park, H. Wang, A. Modi, J. Pouget-Abadie, K. Greller, R. Gupta, R. Berry, P. Ramachandran, J. Xie, L. McCafferty, J. Wang, K. Gupta, H. Lim, B. Bratanič, A. Brock, I. Akolzin, J. Sproch, D. Karliner, D. Kim, A. Goedeckemeyer, N. Shazeer, C. Schmid, D. Calandriello, P. Bhatia, K. Choromanski, C. Montgomery, D. Dua, A. Ramalho, H. King, Y. Gao, L. Nguyen, D. Lindner, D. Pitta, O. Johnson, K. Salama, D. Ardila, M. Han, E. Farnese, S. Odoom, Z. Wang, X. Ding, N. Rink, R. Smith, H. T. Lehri, E. Cohen, N. Vats, T. He, P. Gopavarapu, A. Paszke, M. Patel, W. V. Gansbeke, L. Loher, L. Castro, M. Voitovich, T. von Glehn, N. George, S. Niklaus, Z. Eaton-Rosen, N. Rakićević, E. Jue, S. Perel, C. Zhang, Y. Bahat, A. Pouget, Z. Xing, F. Huot, A. Shenoy, T. Bos, V. Coriou, B. Richter, N. Noy, Y. Wang, S. Ontanon, S. Qin, G. Makarchuk, D. Hassabis, Z. Li, M. Sharma, K. Venkatesan, I. Kemaev, R. Daniel, S. Huang, S. Shah, O. Ponce, Warren, Chen, M. Faruqui, J. Wu, S. Andačić, S. Payrits, D. McDuff, T. Hume, Y. Cao, M. Tessler, Q. Wang, Y. Wang, I. Rendulic, E. Agustsson, M. Johnson, T. Lando, A. Howard, S. G. S. Padmanabhan, M. Daswani, A. Banino, M. Kilgore, J. Heek, Z. Ji, A. Caceres, C. Li, N. Kassner, A. Vlaskin, Z. Liu, A. Grills, Y. Hou, R. Sukkerd, G. Cheon, N. Shetty, L. Markeeva, P. Stanczyk, T. Iyer, Y. Gong, S. Gao, K. Gopalakrishnan, T. Blyth, M. Reynolds, A. Bhoopchand, M. Bilenko, D. Gharibian, V. Zayats, A. Faust, A. Singh, M. Ma, H. Jiao, S. Vijayanarasimhan, L. Aroyo, V. Yadav, S. Chakera, A. Kakarla, V. Meshram, K. Gregor, G. Botea, E. Senter, D. Jia, G. Kovacs, N. Sharma, S. Baur, K. Kang, Y. He, L. Zhuo, M. Kostelac, I. Laish, S. Peng, L. O’Bryan, D. Kasenberg, G. R. Rao, E. Leurent, B. Zhang, S. Stevens, A. Salazar, Y. Zhang, I. Lobov, J. Walker, A. Porter, M. Redshaw, H. Ke, A. Rao, A. Lee, H. Lam, M. Moffitt, J. Kim, S. Qiao, T. Koo, R. Dadashi, X. Song, M. Sundararajan, P. Xu, C. Kawamoto, Y. Zhong, C. Barbu, A. Reddy, M. Verzetti, L. Li, G. Papamakarios, H. Klimczak-Plucińska, M. Cassin, K. Kavukcuoglu, R. Swavely, A. Vaucher, J. Zhao, R. Hemsley, M. Tschannen, H. Ge, G. Menghani, Y. Yu, N. Ha, W. He, X. Wu, M. Song, R. Sterneck, S. Zinke, D. A. Calian, A. Marsden, A. C. Ruiz, M. Hessel, A. Gueta, B. Lee, B. Farris, M. Gupta, Y. Li, M. Saleh, V. Misra, K. Xiao, P. Mendolicchio, G. Buttimore, V. Krayvanova, N. Nayakanti, M. Wiethoff, Y. Pande, A. Mirhoseini, N. Lao, J. Liu, Y. Hua, A. Chen, Y. Malkov, D. Kalashnikov, S. Gupta, K. Audhkhasi, Y. Zhai, S. Kopalle, P. Jain, E. Ofek, C. Meyer, K. Baatarsukh, H. Strejček, J. Qian, J. Freedman, R. Figueira, M. Sokolik, O. Bachem, R. Lin, D. Kharrat, C. Hidey, P. Xu, D. Duan, Y. Li, M. Ersoy, R. Everett, K. Cen, R. Santamaria-Fernandez, A. Taubenfeld, I. Mackinnon, L. Deng, P. Zablotskaia, S. Viswanadha, S. Goel, D. Yates, Y. Deng, P. Choy, M. Chen, A. Sinha, A. Mossin, Y. Wang, A. Szlam, S. Hao, P. K. Rubenstein, M. Toksoz-Exley, M. Aperghis, Y. Zhong, J. Ahn, M. Isard, O. Lacombe, F. Luisier, C. Anastasiou, Y. Kalley, U. Prabhu, E. Dunleavy, S. Bijwadia, J. Mao-Jones, K. Chen, R. Pasumarthi, E. Wood, A. Dostmohamed, N. Hurley, J. Simsa, A. Parrish, M. Pajarskas, M. Harvey, O. Skopek, Y. Kochinski, J. Rey, V. Rieser, D. Zhou, S. J. Lee, T. Acharya, G. Li, J. Jiang, X. Zhang, B. Gipson, E. Mahintorabi, M. Gelmi, N. Khajehnouri, A. Yeh, K. Lee, L. Matthey, L. Baker, T. Pham, H. Fu, A. Pak, P. Gupta, C. Vasconcelos, A. Sadovsky, B. Walker, S. Hsiao, P. Zochbauer, A. Marzoca, N. Velan, J. Zeng, G. Baechler, D. Driess, D. Jain, Y. Huang, L. Tao, J. Maggs, N. Levine, J. Schneider, E. Gemzer, S. Petit, S. Han, Z. Fisher, D. Zelle, C. Biles, E. Ie, A. Fadeeva, C. Liu, J. V. Franco, A. Collister, H. Zhang, R. Wang, R. Zhao, L. Kieliger, K. Shuster, R. Zhu, B. Gong, L. Chan, R. Sun, S. Basu, R. Zimmermann, J. Hayes, A. Bapna, J. Snoek, W. Yang, P. Datta, J. A. Abdallah, K. Kilgour, L. Li, S. Mah, Y. Jun, M. Rivière, A. Karmarkar, T. Spalink, T. Huang, L. Gonzalez, D. Tran, A. Nowak, J. Palowitch, M. Chadwick, E. Talius, H. Mehta, T. Sellam, P. Fränken, M. Nicosia, K. He, A. Kini, D. Amos, S. Basu, H. Jobe, E. Shaw, Q. Xu, C. Evans, D. Ikeda, C. Yan, L. Jin, L. Wang, S. Yadav, I. Labzovsky, R. Sampath, A. Ma, C. Schumann, A. Siddhant, R. Shah, J. Youssef, R. Agarwal, N. Dabney, A. Tonioni, M. Ambar, J. Li, I. Guyon, B. Li, D. Soergel, B. Fang, G. Karadzhov, C. Udrescu, T. Trinh, V. Raunak, S. Noury, D. Guo, S. Gupta, M. Finkelstein, D. Petek, L. Liang, G. Billock, P. Sun, D. Wood, Y. Song, X. Yu, T. Matejovicova, R. Cohen, K. Andra, D. D’Ambrosio, Z. Deng, V. Nallatamby, E. Songhori, R. Dangovski, A. Lampinen, P. Botadra, A. Hillier, J. Cao, N. Baddi, A. Kuncoro, T. Yoshino, A. Bhagatwala, M. Ranzato, R. Schaeffer, T. Liu, S. Ye, O. Sarvana, J. Nham, C. Kuang, I. Gao, J. Baek, S. Mittal, A. Wahid, A. Gergely, B. Ni, J. Feldman, C. Muir, P. Lamblin, W. Macherey, E. Dyer, L. Kilpatrick, V. Campos, M. Bhutani, S. Fort, Y. Ahmad, A. Severyn, K. Chatziprimou, O. Ferludin, M. Dimarco, A. Kusupati, J. Heyward, D. Bahir, K. Villela, K. Millican, D. Marcus, S. Bahargam, C. Unlu, N. Roth, Z. Wei, S. Gopal, D. Ghoshal, E. Lee, S. Lin, J. Lees, D. Lee, A. Hosseini, C. Fan, S. Neel, M. Wu, Y. Altun, H. Cai, E. Piqueras, J. Woodward, A. Bissacco, S. Haykal, M. Bordbar, P. Sundaram, S. Hodkinson, D. Toyama, G. Polovets, A. Myers, A. Sinha, T. Levinboim, K. Krishnakumar, R. Chhaparia, T. Sholokhova, N. B. Gundavarapu, G. Jawahar, H. Qureshi, J. Hu, N. Momchev, M. Rahtz, R. Wu, A. P. S, K. Dhamdhere, M. Guo, U. Gupta, A. Eslami, M. Schain, M. Blokzijl, D. Welling, D. Orr, L. Bolelli, N. Perez-Nieves, M. Sirotenko, A. Prasad, A. Kar, B. D. B. Pigem, T. Terzi, G. Weisz, D. Ghosh, A. Mavalankar, D. Madeka, K. Daugaard, H. Adam, V. Shah, D. Berman, M. Tran, S. Baker, E. Andrejczuk, G. Chole, G. Raboshchuk, M. Mirzazadeh, T. Kagohara, S. Wu, C. Schallhart, B. Orlando, C. Wang, A. Rrustemi, H. Xiong, H. Liu, A. Vezer, N. Ramsden, S. Chang, S. Mudgal, Y. Li, N. Vieillard, Y. Hoshen, F. Ahmad, A. Slone, A. Hua, N. Potikha, M. Rossini, J. Stritar, S. Prakash, Z. Wang, X. Dong, A. Nazari, E. Nehoran, K. Tekelioglu, Y. Li, K. Badola, T. Funkhouser, Y. Li, V. Yerram, R. Ganeshan, D. Formoso, K. Langner, T. Shi, H. Li, Y. Yamamori, A. Panda, A. Saade, A. S. Scarpati, C. Breaux, C. Carey, Z. Zhou, C. Hsieh, S. Bridgers, A. Butryna, N. Gupta, V. Tulsyan, S. Woo, E. Eltyshev, W. Grathwohl, C. Parks, S. Benjamin, R. Panigrahy, S. Dodhia, D. D. Freitas, C. Sauer, W. Song, F. Alet, J. Tolins, C. Paduraru, X. Zhou, B. Albert, Z. Zhang, L. Shu, M. Bansal, S. Nguyen, A. Globerson, O. Xiao, J. Manyika, T. Hennigan, R. Rong, J. Matak, A. Bakalov, A. Sharma, D. Sinopalnikov, A. Pierson, S. Roller, G. Brown, M. Gao, T. Fukuzawa, A. Ghafouri, K. Vassigh, I. Barr, Z. Wang, A. Korsun, R. Jayaram, L. Ren, T. Zaman, S. Khan, Y. Lunts, D. Deutsch, D. Uthus, N. Katz, M. Samsikova, A. Khalifa, N. Sethi, J. Sun, L. Tang, U. Alon, X. Luo, D. Yu, A. Nayyar, B. Petrini, W. Truong, V. Hellendoorn, N. Chinaev, C. Alberti, W. Wang, J. Hu, V. Mirrokni, A. Balashankar, A. Aharon, A. Mehta, A. Iscen, J. Kready, L. Manning, A. Mohananey, Y. Chen, A. Tripathi, A. Wu, I. Petrovski, D. Hwang, M. Baeuml, S. Chandrakaladharan, Y. Liu, R. Coaguila, M. Chen, S. Ma, P. Tafti, S. Tatineni, T. Spitz, J. Ye, P. Vicol, M. Rosca, A. Puigdomènech, Z. Yahav, S. Ghemawat, H. Lin, P. Kirk, Z. Nabulsi, S. Brin, B. Bohnet, K. Caluwaerts, A. S. Veerubhotla, D. Zheng, Z. Dai, P. Petrov, Y. Xu, R. Mehran, Z. Xu, L. Zintgraf, J. Choi, S. A. Hombaiah, R. Thoppilan, S. Reddi, L. Lew, L. Li, K. Webster, K. Sawhney, L. Lamprou, S. Shakeri, M. Lunayach, J. Chen, S. Bagri, A. Salcianu, Y. Chen, Y. Donchev, C. Magister, S. Nørly, V. Rodrigues, T. Izo, H. Noga, J. Zou, T. Köppe, W. Zhou, K. Lee, X. Long, D. Eisenbud, A. Chen, C. Schenck, C. M. To, P. Zhong, E. Taropa, M. Truong, O. Levy, D. Martins, Z. Zhang, C. Semturs, K. Zhang, A. Yakubovich, P. Moreno, L. McConnaughey, D. Lu, S. Redmond, L. Weerts, Y. Bitton, T. Refice, N. Lacasse, A. Conmy, C. Tallec, J. Odell, H. Forbes-Pollard, A. Socala, J. Hoech, P. Kohli, A. Walton, R. Wang, M. Sazanovich, K. Zhu, A. Kapishnikov, R. Galt, M. Denton, B. Murdoch, C. Sikora, K. Mohamed, W. Wei, U. First, T. McConnell, L. C. Cobo, J. Qin, T. Avrahami, D. Balle, Y. Watanabe, A. Louis, A. Kraft, S. Ariafar, Y. Gu, E. Rives, C. Yoon, A. Rusu, J. Cobon-Kerr, C. Hahn, J. Luo, Yuvein, Zhu, N. Ahuja, R. Benenson, R. L. Kaufman, H. Yu, L. Hightower, J. Zhang, D. Ni, L. A. Hendricks, G. Wang, G. Yona, L. Jain, P. Barrio, S. Bhupatiraju, S. Velusamy, A. Dafoe, S. Riedel, T. Thomas, Z. Yuan, M. Bellaiche, S. Panthaplackel, K. Kloboves, S. Jauhari, C. Akbulut, T. Davchev, E. Gladchenko, D. Madras, A. Chuklin, T. Hill, Q. Yuan, M. Madhavan, L. Leonhard, D. Scandinaro, Q. Chen, N. Niu, A. Douillard, B. Damoc, Y. Onoe, F. Pedregosa, F. Bertsch, C. Leichner, J. Pagadora, J. Malmaud, S. Ponda, A. Twigg, O. Duzhyi, J. Shen, M. Wang, R. Garg, J. Chen, U. Evci, J. Lee, L. Liu, K. Kojima, M. Yamaguchi, A. Rajendran, A. Piergiovanni, V. K. Rajendran, M. Fornoni, G. Ibagon, H. Ragan, S. M. Khan, J. Blitzer, A. Bunner, G. Sun, T. Kosakai, S. Lundberg, N. Elue, K. Guu, S. Park, J. Park, A. Narayanaswamy, C. Wu, J. Mudigonda, T. Cohn, H. Mu, R. Kumar, L. Graesser, Y. Zhang, R. Killam, V. Zhuang, M. Giménez, W. A. Jishi, R. Ley-Wild, A. Zhai, K. Osawa, D. Cedillo, J. Liu, M. Upadhyay, M. Sieniek, R. Sharma, T. Paine, A. Angelova, S. Addepalli, C. Parada, K. Majumder, A. Lamp, S. Kumar, X. Deng, A. Myaskovsky, T. Sabolić, J. Dudek, S. York, F. de Chaumont Quitry, J. Nie, D. Cattle, A. Gunjan, B. Piot, W. Khawaja, S. Bang, S. Wang, S. Khodadadeh, R. R, P. Rawlani, R. Powell, K. Lee, J. Griesser, G. Oh, C. Magalhaes, Y. Li, S. Tokumine, H. N. Vogel, D. Hsu, A. BC, D. Jindal, M. Cohen, Z. Yang, J. Yuan, D. de Cesare, T. Bruguier, J. Xu, M. Roy, A. Jacovi, D. Belov, R. Arya, P. Meadowlark, S. Cohen-Ganor, W. Ye, P. Morris-Suzuki, P. Banzal, G. Song, P. Ponnuramu, F. Zhang, G. Scrivener, S. Zaiem, A. R. Rochman, K. Han, B. Ghazi, K. Lee, S. Drath, D. Suo, A. Girgis, P. Shenoy, D. Nguyen, D. Eck, S. Gupta, L. Yan, J. Carreira, A. Gulati, R. Sang, D. Mirylenka, E. Cooney, E. Chou, M. Ling, C. Fan, B. Coleman, G. Tubone, R. Kumar, J. Baldridge, F. Hernandez-Campos, A. Lazaridou, J. Besley, I. Yona, N. Bulut, Q. Wellens, A. Pierigiovanni, J. George, R. Green, P. Han, C. Tao, G. Clark, C. You, A. Abdolmaleki, J. Fu, T. Chen, A. Chaugule, A. Chandorkar, A. Rahman, W. Thompson, P. Koanantakool, M. Bernico, J. Ren, A. Vlasov, S. Vassilvitskii, M. Kula, Y. Liang, D. Kim, Y. Huang, C. Ye, D. Lepikhin, and W. Helmholz Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. External Links: 2507.06261, Link Cited by: §G.1. Dong et al. (2024) Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, Z. Wu, B. Chang, X. Sun, L. Li, and Z. Sui A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 1107–1128. External Links: Link, Document Cited by: §3.3. Goel et al. (2025) A. Goel, S. Ghosh, J. Kim, S. Kumar, Z. Kong, S. Lee, C. H. Yang, R. Duraiswami, D. Manocha, R. Valle, and B. Catanzaro Audio flamingo 3: advancing audio intelligence with fully open large audio language models. In NeurIPS 2025 : Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §G.1, §4.1, §5. Hu et al. (2026) J. Hu, Z. Li, B. Qi, L. Guoming, and P. Wang End-to-end contrastive language-speech pretraining model for long-form spoken question answering. In AAAI 2026, Cited by: §1, §5. Jain et al. (2024) A. Jain, F. Cunha, M. J. Bunsen, J. S. Cañas, L. Pasi, N. Pinoy, F. Helsing, J. Russo, M. Botham, M. Sabourin, J. Fréchette, A. Anctil, Y. Lopez, E. Navarro, F. P. Pimentel, A. C. Zamora, J. A. R. Silva, J. Gagnon, T. August, K. Bjerge, A. G. Segura, M. Bélisle, Y. Basset, K. P. McFarland, D. Roy, T. T. Høye, M. Larrivée, and D. Rolnick Insect identification in the wild: the ami dataset. External Links: 2406.12452, Link Cited by: §1, §2.1, §5, Dataset Sourcing and Compliance.. Johnson et al. (2024) A. Johnson, P. Plantinga, P. Sun, S. Gadiyaram, A. Girma, and A. Emami Efficient SQA from long audio contexts: A policy-driven approach. In 25th Annual Conference of the International Speech Communication Association, Interspeech 2024, Kos, Greece, September 1-5, 2024, I. Lapidot and S. Gannot (Eds.), Cited by: §1. KimiTeam et al. (2025) KimiTeam, D. Ding, Z. Ju, Y. Leng, S. Liu, T. Liu, Z. Shang, K. Shen, W. Song, X. Tan, H. Tang, Z. Wang, C. Wei, Y. Xin, X. Xu, J. Yu, Y. Zhang, X. Zhou, Y. Charles, J. Chen, Y. Chen, Y. Du, W. He, Z. Hu, G. Lai, Q. Li, Y. Liu, W. Sun, J. Wang, Y. Wang, Y. Wu, Y. Wu, D. Yang, H. Yang, Y. Yang, Z. Yang, A. Yin, R. Yuan, Y. Zhang, and Z. Zhou Kimi-audio technical report. External Links: 2504.18425, Link Cited by: §1, §5. Lauri et al. (2023a) M. Lauri, D. Hsu, and J. Pajarinen Partially observable markov decision processes in robotics: a survey. IEEE Transactions on Robotics 39 (1), p. 21–40. External Links: Document Cited by: §3.1. Lauri et al. (2023b) M. Lauri, D. Hsu, and J. Pajarinen Partially observable markov decision processes in robotics: a survey. IEEE Transactions on Robotics 39 (1), p. 21–40. External Links: ISSN 1941-0468, Link, Document Cited by: §5. Lee et al. (2022) K. Lee, K. Park, and D. Kim DailyTalk: spoken dialogue dataset for conversational text-to-speech. External Links: 2207.01063 Cited by: §2.1, Dataset Sourcing and Compliance.. Lewis et al. (2020) P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33, p. 9459–9474. External Links: Link Cited by: §5. Li et al. (2018) C. Li, S. Wu, C. Liu, and H. Lee Spoken squad: a study of mitigating the impact of speech recognition errors on listening comprehension. External Links: 1804.00320, Link Cited by: §5. Li et al. (2025) Q. Li, T. Xiao, Z. Li, P. Wang, M. Shen, and H. Zhao Dialogue-rag: enhancing retrieval for llms via node-linking utterance rewriting. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 24423–24438. Cited by: §1. Lian et al. (2025a) Z. Lian, H. Chen, L. Chen, H. Sun, L. Sun, Y. Ren, Z. Cheng, B. Liu, R. Liu, X. Peng, J. Yi, and J. Tao AffectGPT: A new dataset, model, and benchmark for emotion understanding with multimodal large language models. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, External Links: Link Cited by: §2.2. Lian et al. (2025b) Z. Lian, H. Sun, L. Sun, H. Chen, L. Chen, H. Gu, Z. Wen, S. Chen, S. Zhang, H. Yao, B. Liu, R. Liu, S. Liang, Y. Li, J. Yi, and J. Tao OV-MER: towards open-vocabulary multimodal emotion recognition. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, External Links: Link Cited by: §2.2. Lin et al. (2024) C. Lin, G. Lin, Y. Chuang, W. Wu, S. Li, A. Mohamed, H. Lee, and L. Lee SpeechDPR: end-to-end spoken passage retrieval for open-domain spoken question answering. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2024, Seoul, Republic of Korea, April 14-19, 2024, p. 12476–12480. External Links: Link, Document Cited by: §1. Liu et al. (2025) H. Liu, Z. Wang, X. Chen, Z. Li, F. Xiong, Q. Yu, and W. Zhang HopRAG: multi-hop reasoning for logic-aware retrieval-augmented generation. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 1897–1913. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §5. Mishra et al. (2025) V. Mishra, B. Pathiraja, M. Parmar, S. Chidananda, J. Srinivasa, G. Liu, A. Payani, and C. Baral Investigating the shortcomings of LLMs in step-by-step legal reasoning. In Findings of the Association for Computational Linguistics: NAACL 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 7795–7826. External Links: Link, Document, ISBN 979-8-89176-195-7 Cited by: §G.2, §4.2. OpenAI et al. (2025) OpenAI, :, S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, B. Barak, A. Bennett, T. Bertao, N. Brett, E. Brevdo, G. Brockman, S. Bubeck, C. Chang, K. Chen, M. Chen, E. Cheung, A. Clark, D. Cook, M. Dukhan, C. Dvorak, K. Fives, V. Fomenko, T. Garipov, K. Georgiev, M. Glaese, T. Gogineni, A. Goucher, L. Gross, K. G. Guzman, J. Hallman, J. Hehir, J. Heidecke, A. Helyar, H. Hu, R. Huet, J. Huh, S. Jain, Z. Johnson, C. Koch, I. Kofman, D. Kundel, J. Kwon, V. Kyrylov, E. Y. Le, G. Leclerc, J. P. Lennon, S. Lessans, M. Lezcano-Casado, Y. Li, Z. Li, J. Lin, J. Liss, Lily, Liu, J. Liu, K. Lu, C. Lu, Z. Martinovic, L. McCallum, J. McGrath, S. McKinney, A. McLaughlin, S. Mei, S. Mostovoy, T. Mu, G. Myles, A. Neitz, A. Nichol, J. Pachocki, A. Paino, D. Palmie, A. Pantuliano, G. Parascandolo, J. Park, L. Pathak, C. Paz, L. Peran, D. Pimenov, M. Pokrass, E. Proehl, H. Qiu, G. Raila, F. Raso, H. Ren, K. Richardson, D. Robinson, B. Rotsted, H. Salman, S. Sanjeev, M. Schwarzer, D. Sculley, H. Sikchi, K. Simon, K. Singhal, Y. Song, D. Stuckey, Z. Sun, P. Tillet, S. Toizer, F. Tsimpourlas, N. Vyas, E. Wallace, X. Wang, M. Wang, O. Watkins, K. Weil, A. Wendling, K. Whinnery, C. Whitney, H. Wong, L. Yang, Y. Yang, M. Yasunaga, K. Ying, W. Zaremba, W. Zhan, C. Zhang, B. Zhang, E. Zhang, and S. Zhao Gpt-oss-120b. External Links: 2508.10925, Link Cited by: §G.2, §4.2. Radford et al. (2022) A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever Robust speech recognition via large-scale weak supervision. External Links: 2212.04356, Link Cited by: Appendix E. Schick et al. (2023) T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom Toolformer: language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 68539–68551. External Links: Link Cited by: §5. Shui et al. (2023) R. Shui, Y. Cao, X. Wang, and T. Chua A comprehensive evaluation of large language models on legal judgment prediction. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 7337–7348. External Links: Link, Document Cited by: §G.2, §4.2. Tang et al. (2025) Q. Tang, S. Y. M. Lee, J. Wu, D. Zhang, S. Li, E. Cambria, and G. Zhou A comprehensive graph framework for question answering with mode-seeking preference alignment. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 21504–21523. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1, §5. Wang et al. (2024) M. Wang, I. Shafran, H. Soltau, W. Han, Y. Cao, D. Yu, and L. E. Shafey Retrieval augmented end-to-end spoken dialog models. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2024, Seoul, Republic of Korea, April 14-19, 2024, p. 12056–12060. External Links: Link, Document Cited by: §1. Wu et al. (2023) Y. Wu, S. Rallabandi, R. Srinivasamurthy, P. P. Dakle, A. Gon, and P. Raghavan HeySQuAD: a spoken question answering dataset. External Links: 2304.13689 Cited by: §5. Xiaomi (2025) L. Xiaomi MiMo-audio: audio language models are few-shot learners. External Links: Link Cited by: §G.1, §4.1. Xu et al. (2025a) J. Xu, Z. Guo, H. Hu, Y. Chu, X. Wang, J. He, Y. Wang, X. Shi, T. He, X. Zhu, Y. Lv, Y. Wang, D. Guo, H. Wang, L. Ma, P. Zhang, X. Zhang, H. Hao, Z. Guo, B. Yang, B. Zhang, Z. Ma, X. Wei, S. Bai, K. Chen, X. Liu, P. Wang, M. Yang, D. Liu, X. Ren, B. Zheng, R. Men, F. Zhou, B. Yu, J. Yang, L. Yu, J. Zhou, and J. Lin Qwen3-omni technical report. arXiv preprint arXiv:2509.17765. Cited by: §G.1, Appendix H, §1, §4.1. Xu et al. (2025b) K. Xu, F. Xie, X. Tang, and Y. Hu FireRedASR: open-source industrial-grade mandarin speech recognition models from encoder-decoder to llm integration. External Links: 2501.14350, Link Cited by: Appendix H. Yao et al. (2023) S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao ReAct: synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: §5. You et al. (2022) C. You, N. Chen, F. Liu, S. Ge, X. Wu, and Y. Zou End-to-end spoken conversational question answering: task, dataset and model. In Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, United States, p. 1219–1232. External Links: Link, Document Cited by: §1. Yu et al. (2022a) F. Yu, S. Zhang, Y. Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Ma, X. Xu, and H. Bu M2MeT: the ICASSP 2022 multi-channel multi-party meeting transcription challenge. In Proc. ICASSP, Cited by: §1, §2.1, §5, Dataset Sourcing and Compliance.. Yu et al. (2022b) F. Yu, S. Zhang, P. Guo, Y. Fu, Z. Du, S. Zheng, W. Huang, L. Xie, Z. Tan, D. Wang, Y. Qian, K. A. Lee, Z. Yan, B. Ma, X. Xu, and H. Bu Summary on the ICASSP 2022 multi-channel multi-party meeting transcription grand challenge. In Proc. ICASSP, Cited by: §1, §2.1. Zhang et al. (2025) W. Zhang, X. Li, K. Dong, Y. Wang, P. Jia, X. Li, Y. Zhang, D. Xu, Z. Du, H. Guo, R. Tang, and X. Zhao Process vs. outcome reward: which is better for agentic RAG reinforcement learning. NIPS 2025. Cited by: §5. Zhao et al. (2025) Z. Zhao, Y. Jiang, H. Liu, Y. Wang, and Y. Wang LibriSQA: A novel dataset and framework for spoken question answering with large language models. IEEE Trans. Artif. Intell. 6 (11), p. 2884–2895. External Links: Link, Document Cited by: §1. Zhifei et al. (2025) X. Zhifei, M. Lin, Z. Liu, P. Wu, S. Yan, and C. Miao Audio-reasoner: improving reasoning capability in large audio language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 23840–23862. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2. Appendix A More Analysis in Ablation Study Ablation Settings Overall AMI Avg. Δ Fact. Infer. Temp. Summ. Acou. Ours (Full) 49.49 - 59.46 65.29 38.56 48.46 35.68 Validation of Query Decomposer w/o Query Planner 46.39 -3.10 57.63 59.35 36.79 45.73 32.43 Validation of Query Planner w/o Query Planner 44.68 -4.80 56.12 56.35 36.41 43.18 31.37 Validation of Reflection w/o Reflection 39.64 -9.84 51.32 49.52 31.10 39.83 26.45 Validation of Action Space w/o Tool: Filtering 43.69 -5.79 57.81 62.27 35.36 41.45 28.58 w/o Tool: Semantic Search 16.28 -33.21 28.96 16.38 12.21 13.69 10.15 w/o Tool: Graph Traversal 38.21 -11.27 45.57 38.62 32.31 41.25 33.31 w/o Tool: Audio Access 34.71 -14.78 47.52 49.12 33.13 31.71 12.06 Table 6: Ablation study on the AMI dataset. We report the accuracy (%) drop when specific components are removed. Δ : Performance degradation compared to the full framework. To validate the necessity of each component in AudioGraph, we conduct an ablation study on the AMI dataset (Table 6). Impact of Cognitive Modules. Removing the Query Planner leads to a significant drop of 4.80%, particularly in inferential tasks (-8.94%). This confirms that complex queries (e.g., multi-hop reasoning) cannot be solved by single-step retrieval; explicit planning is essential for decomposing intents. Most notably, the removal of the Reflection mechanism causes a sharp decline of 9.84%. Without the “verify” loop, the agent is prone to hallucination, accepting the first retrieved chunk even if it is irrelevant. This validates our hypothesis that a POMDP-style feedback loop is critical for robustness. Impact of Graph Tools. The Graph Traversal tool proves indispensable (Δ=−11.27% =-11.27\%). When disabled, the agent degrades to a flat-text searcher, failing to aggregate scattered information via speaker edges (espke_spk) or temporal edges (etempe_temp). Furthermore, removing Audio Access severely impacts acoustic-aware questions (Accuracy 35.68% → 12.06%), demonstrating that text transcripts alone are insufficient for capturing paralinguistic cues like emotion or speaker overlap. Finally, Semantic Search serves as the foundational entry point; its removal collapses the system (Δ=−33.21% =-33.21\%), as the agent loses the ability to locate initial evidence nodes. Appendix B Evidence Precision Recall Analysis Method DailyTalk AMI AliMeeting P R P R P R Qwen3-Omni 32.6 24.6 25.7 27.3 29.3 28.2 + Text RAG 67.3 69.1 56.9 65.2 72.4 68.3 + Audio RAG 61.6 63.6 45.2 42.7 57.3 49.6 + Ours 84.2 77.2 76.2 69.7 79.1 74.1 Table 7: Citation accuracy comparison. Results are reported in % with a ±2± 2s tolerance. P: Precision, R: Recall. Table 7 compares the evidence citation accuracy across three datasets, demonstrating that our method consistently outperforms both the vanilla model and RAG baselines. Appendix C Step Analysis Figure 6: Average reasoning steps across different question types. The agent exhibits adaptive computation: solving simple Factual queries requires minimal steps, while complex Inferential and Temporal queries trigger deeper reasoning chains. The stacked colors illustrate the distribution of tool usage, showing heavy reliance on Graph Traversal for multi-hop reasoning. Complexity-Aware Reasoning. The results demonstrate a clear correlation between question difficulty and planning depth. For Factual queries (e.g., “What is the budget?”), the agent adopts a “shortcut” strategy, typically resolving the intent in just 1.8 steps using primarily Semantic Search. This indicates high efficiency. In contrast, Inferential and Temporal queries trigger significantly longer trajectories (avg. >>4 steps). This confirms that the agent is actively performing multi-hop reasoning—iteratively traversing etempe_temp or espke_spk edges to aggregate scattered evidence, rather than relying on a single retrieval pass. Tool Usage Distribution. The stacked breakdown reveals task-specific tool preferences. Summarization tasks show a dominant usage of Filter tools (Green) to isolate specific time ranges or speakers. Crucially, Acoustic-aware questions exhibit a high frequency of Audio Access (Red) calls in the final steps, validating that the agent correctly learns to “listen” to the raw audio only when textual transcripts are insufficient (e.g., verifying emotion). Figure 7: Noise Sensitivity Analysis. QA Accuracy comparison under varying upstream error rates. The rightmost panel highlights the widening performance gap (Δ Accuracy), showing that GRGA’s advantage grows in noisier environments. Appendix D Human Evaluation. Dataset Method Human Ratings (1-5 Scale) Avg. IAA Corr. Grd. Coh. AMI Qwen3-Omni 1.79 2.12 3.76 2.56 0.78 Ours 2.25 3.25 4.35 3.28 AllMeeting Qwen3-Omni 1.93 2.24 4.32 2.83 0.75 Ours 2.36 3.51 4.37 3.41 DailyTalk Qwen3-Omni 3.43 3.72 4.12 3.76 0.82 Ours 4.11 4.23 4.37 4.24 Overall Average Qwen3-Omni 2.38 2.69 4.07 - - Average Ours 2.91 3.66 4.36 Improvement (%) +0.52 +0.97 +0.30 Table 8: Human evaluation results across different datasets. Models are assessed on three dimensions: Correctness (Corr.), Groundedness (Grd.), and Coherence (Coh.) using a 1-5 scale. We also report average scores (Avg.) and Inter-Annotator Agreement (IAA). To validate real-world performance, we conducted a blind evaluation on 150 queries randomly sampled across five question types from AMI, AliMeeting, and DailyTalk. Three graduate student annotators assessed the responses (IAA=0.78 on avg.). Results in Table 8 shows: The “Fluency vs. Factualit” Gap. While the baseline (Qwen3-Omni) maintains high Coherence (4.07), its low Groundedness score (2.69) indicates a tendency to generate fluent but hallucinated content. In contrast, our method achieves a massive +0.97 improvement in Groundedness. This confirms that our approach does not merely generate text but accurately anchors answers to precise timestamps and speakers, effectively mitigating hallucinations. Grounding Drives Correctness. There is a clear positive correlation between groundedness and correctness. By explicitly locating evidence, our method achieves a +0.52 gain in Correctness. This demonstrates that superior localization capabilities directly translate to more factually accurate answers, particularly for reasoning-heavy questions. Robustness in Complex Scenarios. The performance gap is most pronounced on challenging datasets like AMI and AliMeeting (noisy, multi-party interactions) compared to the cleaner DailyTalk. Our method maintains distinct advantages in these “hard” settings (e.g., +1.13 Groundedness on AMI), proving its robustness where standard omni-models struggle with speaker attribution and temporal reasoning. Stage Module / Action Execution Details & Internal State User Query: “What is the team’s general attitude to the project deadline and workload?” Phase 1 Intent Analysis Entities: [‘team’, ‘project deadline’, ‘workload’] Concept: [‘general attitude’] Plan Execution 1. Hybrid Search: query=“team project deadline workload” (→ 10 hits) 2. Traverse Relations: Expand context (→ 64 nodes found). Draft Answer “The team generally views the deadline as tight… members feeling urgency.” (Generic summary, lacks specific details) Reflection (Critic) Assessment: ✗Confidence: 2/5 Reasoning: “Supported is False. Need reverse confirmation for ‘tight’ and ‘work-intensive’ claims.” ↪ Decision: Trigger Re-Plan. Phase 2 Re-Plan Strategy New Goal: Validate sentiment strength (e.g., stress, specific complaints). Tool: hybrid_search(query=‘project deadline focused effort’, alpha=0.6) Evidence Retrieval The new search strategy retrieves highly specific emotional utterances: ∙ utt180: “…only got six days to do the topic segmentation” ∙ utt225: “shall we change the deadlines?” ∙ utt349: “that’s a hassle too it’s a lot of work” ∙ utt553: “i wanna do it superficially” Final Response “The team’s attitude indicates the project deadline is perceived as tight… One member states they only have six days [utt180]. Another asks ‘shall we change the deadlines?’ [utt225]. Another calls it ‘a hassle’ [utt349]… expresses a preference to handle it ‘superficially’ [utt553].” Final Reflection Assessment: ✓Confidence: 5/5 Status: Supported. Table 9: Case Study on Self-Correction. Initially, the model retrieved broad context but generated a low-confidence (2/5), generic answer. The Reflection module detected the lack of specific evidence and triggered a re-planning step focusing on “focused effort.” This allowed the model to retrieve concrete complaints (e.g., “six days,” “hassle,” “superficially”), resulting in a highly grounded (5/5) final response. Appendix E Noise Sensitivity Analysis To evaluate the robustness of our proposed GRGA against upstream errors, we conducted a sensitivity analysis by simulating varying levels of Word Error Rate (WER) and Diarization Error Rate (DER). We compare our model with the baseline Text RAG system across a 3×33× 3 grid of noise conditions. We define three noise levels for both ASR and Diarization components: Low (Oracle): We use Ground Truth (GT) transcripts and speaker labels to simulate an ideal scenario. Mid (Standard): We utilize the outputs from our standard pipeline (as described in Section H), representing a realistic deployment scenario (DER ≈ 17.6%, WER ≈ 18.8%). High (Simulated Noise): For High WER, we utilize a lower-performance lightweight ASR model (openai/whisper-small Radford et al. (2022)) to generate transcripts with high error rates (WER ≈ 33.6%). For High DER, we simulated noise by randomly shuffling a subset of speaker labels within a meeting while preserving timestamp boundaries. This represents a worst-case scenario where speaker identity information is highly unreliable (DER ≈ 52.4%). Figure 7 visualizes the impact of noise, revealing distinct behaviors: Graceful Degradation. Text RAG suffers catastrophic drops as noise increases, falling from 54.47%54.47\% to 24.67%24.67\% in the High/High setting. This underscores the fragility of standard RAG when explicit text or speaker cues are noisy. Conversely, GRGA demonstrates graceful degradation, retaining 39.54%39.54\% accuracy even under the harshest conditions. The “Noise Buffer” Effect. Notably, the performance gap between GRGA and the baseline widens from +5.37%+5.37\% (Low/Low) to +14.87%+14.87\% (High/High). This indicates that our graph structure acts as an effective noise buffer. By leveraging topological connectivity and temporal constraints, the agent can compensate for corrupted signals, reducing the system’s reliance on perfect transcription and diarization. Appendix F Case Study Table 9 presents a case study illustrating our method’s self-correction capability, where the agent initially generates a low-confidence response but, triggered by the reflection module, re-plans to retrieve specific evidence (e.g., concrete complaints about deadlines), ultimately producing a highly grounded and accurate answer. Appendix G Experimental Setting Details G.1 Baselines Qwen3-Omni. Qwen3-Omni Xu et al. (2025a) is a state-of-the-art (SOTA) end-to-end omni-modal model that processes interleaved inputs (e.g., text and audio) and generates text responses. It supports speech inputs of up to 40 minutes. The model adopts a Mixture-of-Experts (MoE) architecture with 30B parameters, of which 3B are activated per inference. Notably, it achieves performance comparable to top-tier proprietary models (e.g., Gemini 2.5 Pro Comanici et al. (2025)) across a wide range of audio understanding benchmarks. AudioFlamingo3. We select Audio Flamingo 3 Goel et al. (2025) as a SOTA audio-language model. Built upon Whisper encoder and 7B LLM, it employs an “on-demand thinking” mechanism to facilitate chain-of-thought reasoning. Crucially, its architecture processes inputs via 30-second windows with a strict context cap of 10 minutes. This baseline serves to quantify the performance degradation of context-constrained models when applied to long meeting scenarios. MiMo-Audio. MiMo-Audio Xiaomi (2025) is a SOTA end-to-end audio language model. It combines a 1.2B audio tokenizer with a 7B LLM. The model supports a 128k context window and downsamples audio tokens to 6.25 Hz, enabling efficient processing of long sequences. Crucially, its instruction-tuning stage incorporates “thinking” mechanisms, allowing chain-of-thought reasoning directly over audio inputs. Comparing against MiMo-Audio helps assess whether our graph-based structural reasoning provides advantages beyond large-scale end-to-end pretraining. Text RAG Baseline. We implement a standard dense retrieval baseline using BGE-M3 Chen et al. (2024a) embeddings. Meeting transcripts are segmented into utterances with prepended metadata (e.g., [time] SPK: text). Given a query q, we retrieve the top-k (k=10,λ=0.25k=10,λ=0.25) most relevant segments based on cosine similarity, get text result and corresponding audio clips. Crucially, to maintain discourse coherence, retrieved segments are temporally reordered before being fed into the Speech-LLM. This baseline represents the conventional approach to handling long meetings without graph-based structural reasoning. Audio RAG Baseline. We construct a cross-modal retrieval baseline using CLASP Abootorabi and Asgari (2025), which aligns audio and text in a shared semantic space. The long audio is segmented into clips based on speaker diarization timestamps. Given a text query q, we retrieve the top-k most relevant audio clips via cosine similarity between the query embedding and CLASP-encoded audio embeddings. These retrieved raw audio clips, along with their speaker metadata, are then fed directly into the Speech-LLM (same backbone as ours) to generate the response. This baseline evaluates the efficacy of retrieving raw acoustic features versus our proposed graph-based semantic navigation. G.2 Evaluation Metric: Semantic Accuracy To rigorously assess the reasoning capabilities of our model, we report Semantic Accuracy across all datasets. Traditional n-gram metrics (e.g., BLEU, Exact Match) often fail to capture the true validity of generated responses, as they penalize correct answers that differ in lexical surface forms from the ground truth. Drawing upon recent methodologies in complex reasoning evaluation Mishra et al. (2025); Shui et al. (2023), we move beyond rigid string matching and establish a robust, automated evaluation pipeline that prioritizes semantic equivalence and logical soundness. Specifically, we employ a Strong LLM (LLM-as-a-Judge) to approximate human-level judgment. Formally, let =(qi,ai)i=1ND=\(q_i,a_i)\_i=1^N denote the evaluation dataset, where qiq_i is the query and aia_i is the ground-truth answer. Let a^i a_i represent the model-generated response. We define a semantic indicator function (⋅)I(·) parameterized by an external judge J (e.g., GPT-OSS-120B OpenAI et al. (2025)): si=(qi,ai,a^i)∈0,1,s_i=J(q_i,a_i, a_i)∈\0,1\, (12) where si=1s_i=1 if and only if J determines that a^i a_i entails the same semantic information as aia_i, and 00 otherwise. The judge is prompted to ignore stylistic differences and focus strictly on factual consistency and reasoning correctness. Finally, the Semantic Accuracy is computed as the expectation of correct judgments: Accuracy=1N∑i=1Nsi×100%.Accuracy= 1N _i=1^Ns_i× 100\%. (13) This approach ensures a fair comparison by validating whether the model successfully retrieves and reasons over the core knowledge, regardless of its output phrasing. Appendix H Implementation Details All the methods that did not explicitly state the model used Qwen3-Omni Xu et al. (2025a). Graph Construction Setup. For the initial audio processing, we utilize the FireRedASR-AED Xu et al. (2025b) model for Automatic Speech Recognition (ASR) to transcribe the raw audio meeting data. To obtain precise word-level timestamps (i.e., tstart_start and tendt_end), we employ the Montreal Forced Aligner (MFA). Speaker Verification: We use a pre-trained ERes2NetV2 Chen et al. (2024b) model to extract speaker embeddings. The cosine similarity threshold for speaker clustering is set to η=0.8η=0.8 based on validation set performance. Edge Construction: The temporal window for establishing “Reply-To” edges is set to Δt=5 t=5 seconds. Attribute Extraction: We employ Qwen3-Omni (via vllm API) to extract speaker profiles and resolve coreference chains during the node enrichment phase. Agent Planning Configuration The core planning mechanism of GRGA depends on specific hyper-parameters governing the exploration-exploitation trade-off: Backbone Model: The Policy Network πθ _θ is parameterized by Qwen3-Omni, operated with a temperature of 00 to ensure deterministic reasoning paths. Reward Function: As defined in Eq. (9), the self-reflection verification threshold is set to τ=4τ=4 (1-5 Scale). The penalty factor for failed verification is set to β=0.5β=0.5. Search Constraints: The maximum depth for the reasoning tree is limited to Kmax=5K_max=5 steps. If the agent fails to reach a verified conclusion within these steps, the process terminates and returns the current best candidate. All experiments were conducted on a server cluster equipped with 8 × Ascend 910B (64GB) GPUs. H.1 Text RAG Baseline. To benchmark our method against a standard text-based retrieval pipeline, we implement a dense Retrieval-Augmented Generation (RAG) baseline over meeting transcripts. Context construction. We segment the meeting transcript into individual utterances and treat each utterance as one retrieval unit. Each unit is formatted with temporal and speaker metadata: ui=[starti - endi] SPEAKER_ID: texti.u_i= [start_i - end_i ] SPEAKER\_ID: text_i. Dense retrieval. We use BGE-M3 Chen et al. (2024a) as the embedding backbone. Given a query q, we compute embeddings for the query and each utterance: q=fBGE(q)∈ℝ1024,i=fBGE(ui)∈ℝ1024.e_q=f_BGE(q) ^1024, _i=f_BGE(u_i) ^1024. We apply L2 normalization to enable cosine-based scoring: ^=∥2,si=cos(^q,^i)=^q⊤^i. e= e _2, s_i= ( e_q, e_i)= e_q e_i. We rank all utterances by sis_i and retrieve the top-k segments (k=10k=10) that satisfy a minimum similarity threshold λ=0.25λ=0.25: sim=TopK(ui∣si≥λ,k).C_sim=TopK (\u_i s_i≥λ\,\,k ). Temporal reordering. To preserve discourse coherence, we reorder the retrieved utterances by their start time before forming the final context: =SortByTime(sim).C=SortByTime(C_sim). Generation. We concatenate the reordered segments as context c=Concat()c=Concat(C) and prompt the generator to answer strictly based on the retrieved context. We use the same backbone LLM as in our main method for a fair comparison: y=LLM(q,c).y=LLM(q,c). H.2 Audio RAG Baseline. Given a text query q and a long audio recording A, we build a retrieval-augmented generation (RAG) baseline based on CLASP Abootorabi and Asgari (2025), a multilingual audio-text representation model that maps raw speech into a shared 768-dimensional semantic space. Audio segmenttation. We segment the long audio A into K shorter clips akk=1K\a_k\_k=1^K, by speaker diarization. Embedding computation. We compute the query embedding and the audio clip embeddings using CLASP: q _q =ftext(q)∈ℝ768, =f_text(q) ^768, (14) k _k =faudio(ak)∈ℝ768,k=1,…,K, =f_audio(a_k) ^768, k=1,…,K, (15) where faudio(⋅)f_audio(·) encodes speech (e.g., via HuBERT and spectrogram encoders with a fusion module), and ftext(⋅)f_text(·) encodes text into the same representation space. Retrieval. We rank audio clips by cosine similarity and select the most relevant clip (or top-M clips): sk s_k =cos(q,k)=q⊤k‖q‖2‖k‖2, = (e_q,e_k)= e_q e_k\|e_q\|_2\,\|e_k\|_2, (16) k⋆ k =argmaxk∈1,…,Ksk. = _k∈\1,…,K\s_k. (17) Optionally, we retrieve =TopM(skk=1K)K=TopM(\s_k\_k=1^K) for multi-evidence prompting. Generation. We feed the query and retrieved evidence to a Speech LLM to generate the final response: y y =LLM(q,akk∈), =LLM (q,\a_k\_k ), (18) where the evidence is provided, contains retrieval speaker diarization and corresponding audio clips. Appendix I Multi-dimensional Meeting Database Construction Algorithm Complexity. Let M be the number of VAD clips, and let FmF_m be the number of acoustic frames in clip m. Let NwN_w be the total number of word tokens after forced alignment, NsN_s the total number of diarization segments, and NuN_u the number of utterances. ASR inference and CTC forced alignment are linear in the input length, costing (∑m=1MFm)O\! ( _m=1^MF_m ) time. Speaker diarization requires embedding extraction over frames plus clustering over segments; in practice this is (∑m=1MFm)O\! ( _m=1^MF_m ) for feature extraction, and an additional clustering cost by AHC is (Ns2)O(N_s^2). Speaker assignment via overlap matching costs (Nw+Ns)O(N_w+N_s) using the two-pointer implementation in Alg. 1 (a naive all-pairs matching would be (NwNs)O(N_wN_s)). Building utterances from word streams is (Nw)O(N_w), and adding temporal edges in the meeting graph is (Nu)O(N_u). The space complexity is dominated by storing word- and utterance-level annotations, i.e., (Nw+Nu)O(N_w+N_u), plus the storage overhead of the text and vector indices for utterance retrieval. Algorithm 1 Multi-dimensional Meeting Database Construction Input: Long meeting audio X; session id sidsid; hyper-parameters θ=max_clip_len,vad_pad,βsil,αovlp,Kθ=\ max\_clip\_len, vad\_pad, _sil, _ovlp,K\ Output: Structured database D; meeting graph G=(V,E)G=(V,E); indices ℐI 1 X←Preprocess(X)X← Preprocess(X) ; // resample/mono/normalize 2 ←FSMN_VAD(X,vad_pad,max_clip_len)C← FSMN\_VAD(X, vad\_pad, max\_clip\_len) ; // clips with absolute spans 3 Initialize tables Clips, Words, SpkSegs, Utterances; 4 Initialize graph G=(V,E)G=(V,E); 5 foreach clip c∈c do // process each VAD clip 6 Clips ← Clips ∪(sid,c.id,c.ts,c.te,c.wav)∪\(sid,c.id,c.t_s,c.t_e,c.wav)\; 7 (T,)←ASR(c.wav)(T,P)← ASR(c.wav) ; // transcript and CTC posteriors 8 ←CTC_ForcedAlign(T,)W← CTC\_ForcedAlign(T,P); ; // =(wi,ai,bi,pi)W=\(w_i,a_i,b_i,p_i)\ word-level timestamps 9 ←Diarize(c.wav)S← Diarize(c.wav) ; // =(spkj,sj,ej,embj,confj)S=\(spk_j,s_j,e_j,emb_j,conf_j)\ 10 ←ClusterAndSmooth()S← ClusterAndSmooth(S) ; // merge/smooth speaker segments 11 SpkSegs ← SpkSegs ∪(sid,c.id,)∪\(sid,c.id,S)\; 12 ~←AssignSpeakerToWordsFast(,,αovlp) W← AssignSpeakerToWordsFast(W,S, _ovlp); 13 Words ← Words ∪(sid,c.id,~)∪\(sid,c.id, W)\; 14 ←BuildUtterances(~,βsil)U← BuildUtterances( W, _sil) ; // group into utterances 15 foreach utterance u∈u do 16 Utterances ← Utterances ∪(sid,u.id,u.spk,u.ts,u.te,u.text,u.audio_ref,u.conf)∪\(sid,u.id,u.spk,u.t_s,u.t_e,u.text,u.audio\_ref,u.conf)\; 17 V←V∪u.idV← V∪\u.id\ ; // utterance nodes in G 18 E←E∪AddTemporalEdges(Utterances)E← E∪ AddTemporalEdges( Utterances) ; // time adjacency 19 E←E∪AddSpeakerEdges(Utterances)E← E∪ AddSpeakerEdges( Utterances) ; // same-speaker links (optional) 20 E←E∪AddEntityOrTopicEdges(Utterances)E← E∪ AddEntityOrTopicEdges( Utterances) ; // optional NER/topic links 21 ℐbm25←BuildBM25Index(Utterances.text)I_bm25← BuildBM25Index( Utterances.text); 22 ℐvec←BuildVectorIndex(Utterances,K)I_vec← BuildVectorIndex( Utterances,K); 23 ℐ←ℐbm25,ℐvecI←\I_bm25,I_vec\; 24 ←Clips,Words,SpkSegs,Utterances,ℐD←\ Clips, Words, SpkSegs, Utterances,I\; 25 return ,G,ℐD,G,I; Appendix J Retrieval-Generation Agent Algorithm Algorithm 2 GRGA with Formal Notation Input: Q∈Q , =(V,E,X)G=(V,E,X) Output: A^∈ A with Cite⊆V×TCite V× T // Decomposition: fdec:→f_dec:Q 1 =fdec(Q)=ce,c,ct,cmC=f_dec(Q)=\c_e,c_c,c_t,c_m\ 2 b0=H0=Q,b_0=H_0=\Q,C\; k←0k← 0 // POMDP Loop 3 while k<Kmaxk<K_max do // Policy: πθ:ℬ→Δ(∗) _θ:B→ (A^*) 4 k∼πθ(⋅∣bk)P_k _θ(· b_k) where bk=b0∪Hkb_k=b_0∪ H_k 5 if k=⊥P_k= then 6 return ⊥ // Transition: T:×→Δ()T:S×A→ (S) 7 ok=⨁op∈kExec(op,)o_k= _op _kExec(op,G) where ok∈o_k // Belief Update: ψ:ℬ×→ℬψ:B×A×O 8 bk+1=ψ(bk,k,ok)=bk∪k,okb_k+1=ψ(b_k,P_k,o_k)=b_k∪\P_k,o_k\ // Synthesis: fsyn:×ℬ→×f_syn:Q×B ×C 9 (A^k,Citek)=fsyn(Q,bk+1)( A_k,Cite_k)=f_syn(Q,b_k+1) // Reward: R:×→ℝR:S×A 10 sver=fref(Q,A^k,ok)∈[0,1]s_ver=f_ref(Q, A_k,o_k)∈[0,1] 11 rk=1if sver≥τ−βotherwiser_k= cases1&if s_ver≥τ\\ -β&otherwise cases 12 if rk>0r_k>0 then 13 return (A^k,Citek)( A_k,Cite_k) 14 else 15 Hk+1=Hk∪(k,ok,A^k,rk)H_k+1=H_k∪\(P_k,o_k, A_k,r_k)\; k←k+1k← k+1 16 return ⊥ Tool Time Complexity Description Filter O(|V|+|E|)O(|V|+|E|) Metadata filter Search O(|V|⋅dembed+klogk)O(|V|· d_embed+k k) Vector retrieval + sorting GraphTraversal O(d¯⋅depth)O( d·depth) BFS (d¯ d: avg degree) Temporal O(1)O(1) Indexed temporal query AudioAccess O(1)O(1) Get audio clip Table 10: Time Complexity of Graph Operations J.1 Complexity Analysis We provide a comprehensive theoretical analysis of the time and space complexity of the proposed GRGA algorithm. We examine each component’s computational cost and derive the overall complexity bounds. J.1.1 Time Complexity Analysis The total time complexity of GRGA is determined by the iterative POMDP loop and its constituent operations: Ttotal=Tdec+∑k=0Kmax(Tplan(k)+Texec(k)+Tsyn(k)+Tref(k))T_total=T_dec+ _k=0^K_ (T_plan^(k)+T_exec^(k)+T_syn^(k)+T_ref^(k) ) (19) where KmaxK_ is the maximum number of iterations. Query Decomposition. The decomposition function fdec:→f_dec:Q involves a single LLM inference: Tdec=O(|Q|⋅dmodel)T_dec=O(|Q|· d_model) (20) where |Q||Q| denotes the query length in tokens and dmodeld_model is the model dimension. This is a one-time operation outside the main loop. Execution Planning. At iteration k, the policy network πθ _θ generates a plan conditioned on the belief state bkb_k: Tplan(k)=O(|bk|⋅dmodel+Lplan⋅||)T_plan^(k)=O(|b_k|· d_model+L_plan·|A|) (21) where: • |bk|=O(|Q|+||+k⋅|oavg|)|b_k|=O(|Q|+|C|+k·|o_avg|) grows linearly with iteration count • LplanL_plan is the average plan length • |||A| is the action space size Tool Execution. The execution engine applies LplanL_plan operations on the meeting graph =(V,E)G=(V,E). The complexity depends on the tool type: The worst-case execution time is: Texec(k)=O(Lplan⋅max|V|⋅dembed,|V|log|V|)T_exec^(k)=O(L_plan· \|V|· d_embed,|V| |V|\) (22) Answer Synthesis. The synthesizer fsynf_syn processes accumulated evidence: Tsyn(k)=O(|⋃i=0koi|⋅dmodel+Lanswer⋅dmodel)T_syn^(k)=O ( | _i=0^ko_i |· d_model+L_answer· d_model ) (23) where LanswerL_answer is the generated answer length. Note that |⋃i=0koi|=O(k⋅|oavg|) | _i=0^ko_i |=O(k·|o_avg|) grows with iterations. Reflection. The reflector freff_ref evaluates logical entailment: Tref(k)=O((|Q|+|A^k|+|ok|)⋅dmodel)T_ref^(k)=O ((|Q|+| A_k|+|o_k|)· d_model ) (24) Aggregate Complexity Ttotal=O(Kmax⋅[(|Q|+k⋅|oavg|)⋅dmodel+Lplan⋅|V|⋅dembed]) splitT_total=O (K_ · [&(|Q|+k·|o_avg|)· d_model\\ &+L_plan·|V|· d_embed ] ) split (25) Under typical conditions where |V|≫|Q||V| |Q| and dembed≈dmodeld_embed≈ d_model, this simplifies to: Ttotal=O(Kmax⋅|V|⋅dmodel)T_total=O(K_ ·|V|· d_model) (26) In practice, Kmax∈[1,5]K_ ∈[1,5] and early stopping (when sver≥τs_ver≥τ) significantly reduces the average iteration count K¯<Kmax K<K_ . J.2 Space Complexity Analysis The space requirements consist of three main components: Stotal=SgraphS_total=S_graph (27) Graph Storage. The meeting graph requires storage for structure and embeddings: Sgraph=O(|V|+|E|+|V|⋅dembed)S_graph=O(|V|+|E|+|V|· d_embed) (28) For a typical 30 minutes meeting: • |V|≈900|V|≈ 900 nodes • dembed=1024d_embed=1024 (BGE-M3) • Storage: ∼ 2MB (structure) + ∼ 3.5MB (embeddings) Appendix K Tools Details The table 11 shows the definitions of the atomic tools in our Action Space (A). Category Tool Signature Description & Purpose Retrieval keyword_search(query) BM25-based search for precise entity/term lookup. semantic_search(query) Dense vector retrieval for abstract semantic concepts. hybrid_search(query, α=0.6α=0.6) Weighted combination of keyword and semantic scores to optimize recall. Filtering filte_time_range(tstart,tendt_start,t_end) Retrieves utterances strictly within a specified time window. filte_speaker(nodes, spk_id) Filters a set of candidate nodes by speaker identity. Traversal traverse_relations(nodes, depth=k) Walks along graph edges (e.g., Next, Reply-To) to trace dialogue threads for multi-hop reasoning. Audio audio_segment(tstart,tendt_start,t_end) Retrieves raw waveform data to ground text in acoustic signals (e.g., for emotion detection). Table 11: The definitions of the atomic tools in our Action Space (A). The Query Planner invokes these tools to interact with the meeting graph and retrieve evidence. Appendix L QA Examples Header. The box title Example # (QA_TYPE, Key=#) indicates the example index, its question type (e.g., Factual), and a unique identifier Key used to retrieve the corresponding record in the released JSON files. Evidence section (lower part). Below the dashed line, we display the verbatim evidence excerpt(s) for the referenced utterance IDs. Each excerpt is formatted as: Evidence [utt_id]: [t_start -- t_end] SPEAKER_#: transcript. Here, [t_start -- t_end] denotes the absolute timestamps (in seconds) within the meeting, and SPEAKER_# is the diarization-based speaker label. This layout makes the supervision explicitly grounded: readers can directly verify that the answer is entailed by the cited utterance(s), and systems can be evaluated on both answer correctness and evidence retrieval. Example 1 (Factual, Key=1) Question: What specific technical difficulty is Speaker 4 facing regarding XML files? Answer: Speaker 4 has written code to read and remove the XML, but is struggling to store the data (e.g., into a vector) or display it on the screen using Java. Evidence (utterance IDs): [513] Rationale: Speaker 4 explicitly states their current coding roadblock regarding Java vectors and display. Evidence [513]: [1083.43 - 1100.41] SPEAKER_4: i’ve been trying to g write something to read the x. m. l. and get rid of it and i can get rid of it but i’m having trouble putting it anywhere else so it will come up on the screen for the moment i haven’t managed to put it into a vector or whatever in java to play with it. Example 2 (Inferential, Key=26) Question: How did the group decide to resolve the issue of having too many pop-up windows? Answer: The group discussed using tabs (similar to Mozilla) or toggle buttons within a single window to switch between views like the full transcription and the summary, rather than opening separate windows. Evidence (utterance IDs): [265, 266, 269, 273, 304, 314] Rationale: Multiple speakers contribute to the idea of using tabs/toggles to manage content within one window frame to avoid clutter. Evidence [265]: [657.56 - 663.43] SPEAKER_2: so maybe you can just like choose the s same window for transcription and summary. Evidence [266] : [661.99 - 664.73] SPEAKER_1: hmm so like have a tab there. Evidence [269] : [665.32 - 669.42] SPEAKER_2: yeah yeah tabs are nice. Evidence [273] : [669.91 - 671.05] SPEAKER_4: mozilla style. Evidence [304] : [697.23 - 703.29] SPEAKER_2: yeah uh change the contents of the same window like from transcription to summary. Evidence [314] : [719.03 - 725.56] SPEAKER_2: no no it could be like transcription summary like two buttons and you just press on which ever you want. Example 3 ( Summarization, Key=53) Question: What are the key interface design decisions made during this meeting? Answer: The team decided to: (1) use buttons instead of right-click menus for speaker characterisation, (2) use tabs or buttons to toggle between transcription and summary in the same window rather than separate windows, and (3) implement right-click menus for topics with options to view all meetings containing that topic. Evidence (utterance IDs): [76, 84, 88, 158, 167, 266, 269, 314] Rationale: These utterances capture the major UI/UX consensus points reached through discussion. Evidence [76]: [207.74 - 220.21] SPEAKER_2: i guess a button button makes a bit more sense ’cause otherwise you don’t really know that oh what if i right click now what happens then it’s like more if it’s visual. Evidence [84]: [225.08 - 232.04] SPEAKER_4: it’s more idiot proof isn’t it it’s got a button. Evidence [88]: [229.16 - 231.24] SPEAKER_3: that’s true yeah it’s more intuitive really isn’t it. Evidence [158]: [399.55 - 402.78] SPEAKER_3: so we could do that in a similar way do it right click as well. Evidence [167]: [411.28 - 424.13] SPEAKER_3: so we have basically two options of of browsing the meetings is by either um searching and opening individual observations and when then we have the interlinking by right click basically. Evidence [266]: [661.99 - 664.73] SPEAKER_1: hmm so like have a tab there. Evidence [269]: [665.32 - 669.42] SPEAKER_2: yeah yeah tabs are nice. Evidence [314]: [719.03 - 725.56] SPEAKER_2: no no it could be like transcription summary like two buttons and you just press on which ever you want. Example 4 (Temporal, Key=60) Question: At what point does the team shift from discussing GUI to discussing the interim prototype? Answer: Around 810–824 seconds (approximately 13–14 minutes into the meeting). Evidence (utterance IDs): [361, 363, 365] Rationale: The topic pivot occurs when Speaker 3 asks what prototype they should aim for. Evidence [361]: [816.58 - 822.38] SPEAKER_3: and finally the prototype he spoke about what kind of prototype could we produce. Evidence [363]: [824.27 - 828.76] SPEAKER_3: because i’m i’m just you know i go into the lab and i say right what am i gonna change today. Evidence [365]: [828.76 - 836.48] SPEAKER_3: you know and it kind of just it just develops i’m not aiming for anything do we wanna aim for something. Example 5 (Acoustic, Key=80) Question: Why did Speaker 3 express sadness when saying “well i hope so” at approximately 24 seconds? Answer: Speaker 3 sounded sad when expressing hope that others had done some work, likely in response to Speaker 4’s question “has anybody done anything” and Speaker 1’s admission “not a lot no”, suggesting disappointment about the team’s progress. Evidence (utterance IDs): [13, 11, 12] Rationale: This combines the sad emotion label from node 13 with the contextual text from surrounding nodes to explain the cause. Evidence [11]: [ 19.07 - 23.69] SPEAKER_4: has anybody done anything. Evidence [12]: [ 21.47 - 23.29] SPEAKER_1: not a lot no. Evidence [13]: [ 24.17 - 25.22] SPEAKER_3: well i hope so. Appendix M Manual Evaluation Tool Figure 8: Data Display Figure 9: Manual Evaluation Appendix N Data Construction Prompt Prompt for Acoustic-Aware QA Generation # Role You are an expert AI assistant specializing in multimodal speech dataset construction. Your task is to generate high-quality Acoustic-Aware QA pairs based on the provided audio and speech transcription and metadata. # Input Data format An audio clip. A list of Nodes. Each Node contains: ‘turn_id‘: [‘timestamp start‘-‘timestamp end‘] Speaker ‘speaker‘: ‘text‘ (Emotion: ‘emotion‘) (Volume: ‘Volume‘) # Task Definition: Acoustic-Aware QA You must generate questions that cannot be answered by reading the text alone. The answer must require checking the Acoustic Features (specifically the ‘emotion‘ and ‘timestamp‘ fields). ## Critical Requirement 1: Focus on High Arousal Emotions You must prioritize questions regarding strong emotions such as Anger, Happiness, Excitement, or Sadness. ### How to handle "Neutral" data: If the input data contains only "Neutral" labels: • Do: Ask Verification Questions checking for the presence of strong emotions. – Good Example: "Did Speaker 3 sound angry when asking about the annual meeting?" – Good Answer: "No. According to the audio data, Speaker 3 maintained a neutral tone." • Don’t: Ask passive descriptive questions like "What was the emotion?". ## Critical Requirement 2: NO Confidence Scores • Strictly Forbidden: Do NOT mention the numeric confidence scores (e.g., ‘0.99‘, ‘1.0‘) in the Question or the Answer. • Simply state the emotion label as a fact (e.g., "The speaker was angry"). ## Allowed Question Patterns 1. Emotion-Content Attribution (Why is he angry?): "Why did spk3 sound angry at the 15th minute?" (Combine Emotion label + Text analysis). 2. Specific Time Identification: "Who sounded the most excited around 10 seconds?" 3. Emotion Verification (For Neutral Data): "When Speaker 1 said [Text], did they express happiness?" "Did the speaker sound furious or annoyed during the discussion?" ## NOT Allowed Question Patterns • Wrong Questions: Q: Why did Speaker 3 sound sad when saying ’i think the search works as well’ around 45 seconds? A: Speaker 3 did not sound sad at that moment. # Output Format (Strict JSON) Output a valid JSON List. Ensure ‘evidence_nodes‘ is a list of integers. Using ’ instead of " for citation ‘json [ "question": "The specific ’question’ string.", "answer": "The brief ’answer’ derived from the text.", "evidence_nodes": [10, 12], // Must be a list of turn_ids "reasoning": "Brief explanation of the logic." ] ‘ # Input: input_data # Output: Prompt for Graph-based QA Generation # Role You are an expert data annotator specialized in Graph-based Audio Understanding and Conversational AI. Your goal is to construct a massive, high-quality Graph RAG dataset from raw ASR transcripts. # Input Data Structure A list of Nodes. Attributes: ‘turn_id‘: [‘timestamp start‘-‘timestamp end‘] Speaker ‘speaker‘: ‘text‘ # Mission Your mission is to exhaustively mine the provided transcript for every possible piece of information and convert it into a Question-Answer (QA) pair. Do not stop at just one question per category. Aim to generate as many valid, distinct QA pairs as the data supports (e.g., 10-20 pairs for a medium-length segment). # Critical Guidelines (The "Don’t"s) 1. Filter Noise: Do NOT generate Factual questions about "YEAH", "OKAY", "M-HMM" (backchannels). • Bad: "Who said ’Yeah’?" • Good (Inferential): Use "Yeah" nodes only as evidence for "Agreement" or "Consensus". 2. Avoid Hallucination: Every answer must be strictly supported by the ‘evidence_nodes‘. 3. No Generic Questions: Avoid vague questions like "What happened?". Be specific: "What software bug did Speaker A mention?" # Question Generation Strategy (How to generate MANY questions) To maximize the number of QAs, apply these specific strategies: ## Strategy A: Entity-Centric Mining (Factual) Scan for every entity in the text. Generate a question for each. • Entities: Project names, specific numbers, technical terms, people’s names, locations, file paths. • Trigger: "I see a phone number ’04555’." -> QA: "What is the phone number mentioned?" ## Strategy B: Logic & Linkage (Inferential) Look for connections between distant or adjacent nodes. • Cause & Effect: Node X proposes something -> Node Y rejects it. -> QA: "Why was the proposal in Node X rejected?" • Clarification: Node A asks a question -> Node B answers it. -> QA: "How did Speaker B respond to A’s inquiry about [Topic]?" • Sentiment: Look for emotional words or emphatic language. -> QA: "Which speaker expressed frustration about the file system?" ## Strategy C: Time Anchors (Temporal) • Absolute: "What topic was introduced exactly at the 10-minute mark?" • Relative: "What was discussed immediately before the discussion about the budget?" • Duration: "How long did the debate about ’UI Design’ last?" (Calculate from start/end timestamps of the cluster). ## Strategy D: Scene Understanding (Summarization) • Topic Segmentation: Identify where the topic shifts. -> QA: "Summarize the main points discussed regarding [Topic Name]." • Speaker Role: "Based on the segment, what seems to be the role of Speaker 0?" (e.g., Manager, Technical Lead). # Output Format (Strict JSON) Output a valid JSON List. Ensure ‘evidence_nodes‘ is a list of integers. ‘json [ "type": "factual | inferential | temporal | summarization", "question": "The specific question string.", "answer": "The precise answer derived from the text.", "evidence_nodes": [10, 12], // Must be a list of turn_ids "reasoning": "Brief explanation of the logic." ] ‘ # Input: input_data # Output: