Paper deep dive
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/19/2026, 5:11:24 AM
Summary
This survey reviews multi-turn conversational AI, analyzing the evolution from text-only dialogue to multimodal and omni-modal systems. It highlights that while modality support has advanced, systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, and cultural alignment. The paper organizes literature around datasets, models, training, and evaluation, identifying significant gaps in session-level competence and resource diversity.
Entities (11)
Relation Signals (8)
Multi-turn Conversational AI â requires â Persistent memory
confidence 95% ¡ This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory...
Multi-turn Conversational AI â requires â Cross-turn grounding
confidence 95% ¡ ...ground responses across modalities, tools, and external knowledge...
Current Systems â struggleswith â Persistent memory
confidence 95% ¡ current systems still struggle with persistent memory...
Current Systems â struggleswith â Cultural alignment
confidence 95% ¡ current systems still struggle with ... cultural alignment.
AudioLLMs â extends â Text-only dialogue
confidence 90% ¡ AudioLLMs extend this setting to spoken interaction by processing speech with text
Omni-modal systems â extends â AudioLLMs
confidence 90% ¡ Omni-modal models extend AudioLLMs by jointly handling text, speech, vision, and sometimes video.
This Survey â follows â PRISMA-ScR
confidence 90% ¡ we followed PRISMA-ScR framework
Evaluation â uses â LLM-as-Judge
confidence 90% ¡ Benchmarks increasingly use LLM-as-judge and mixed approaches
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (this https URL)
Tags
Links
- Source: https://arxiv.org/abs/2608.17605v1
- Canonical: https://arxiv.org/abs/2608.17605v1
Trouble viewing inline? Open PDF directly â
Full Text
166,652 characters extracted from source content.
Expand or collapse full text
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury Qatar Computing Research Institute, Qatar syeda.faiza.ahmed@gmail.com, shchowdhury@hbku.edu.qa Abstract Conversational AI is moving beyond isolated text prompts toward sustained, multimodal in- teraction. In real conversations, users clar- ify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to main- tain and update memory, ground responses across modalities, tools, and external knowl- edge, and adapt across languages and cultures. This study reviewsmulti-turn conversational AIacross text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni- modal systems, and tool-augmented agents. We organize the literature arounddatasets andbenchmarks,modeling paradigms,train- ing strategies,evaluation setups, andcross- cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent in- teraction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full- duplex interaction, robust evaluation, and cul- tural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. 1 1 Introduction Conversational AI is becoming a primary interface through which users interact with large language models (LLMs). These interactions often extend beyond single prompts. Users clarify goals, re- vise requests, ask follow-up questions, switch top- ics, introduce new evidence, and return to earlier points in the conversation ( Kwan et al.,2024;Bai et al.,2024;Reddy et al.,2019). We usesession- level multi-turn interactionto refer to a complete 1 multiturn-conversational-resources Figure 1: Three-axis view of conversational AI cov- ering modality complexity, interaction depth, and cultural-linguistic diversity. conversation with more than one exchange, where later turns may depend on earlier requests, re- sponses, evidence, actions, or state. A capable model must therefore do more than answer the cur- rent query. It must preserve context, resolve cross- turn references, update earlier assumptions, and maintain coherence across the session. Multi-turn dialogue is thus a distinct modeling and evaluation problem, not merely a longer version of single- turn question answering ( Deshpande et al.,2025). This problem has become more important as conversational AI expands beyond text. Early dia- logue systems focused on text-based state tracking, response generation, persona consistency, and task completion ( Wu et al.,2020;Zhang et al.,2020b; Roller et al.,2021). AudioLLMs extend this set- ting to spoken interaction by processing speech with text (Zhang et al.,2023;Chu et al.,2024; Tang et al.,2024). Full-duplex systems support more flexible turn-taking, including real-time in- terruption and overlapping speech ( DĂŠfossez et al., 2024). More recent omni-modal models jointly handle text, speech, vision, and often video within unified architectures (Hurst et al.,2024;Xu et al., 2025b;Fu et al.,2026;Chen et al.,2025c;Tong et al. ,2025). Figure2illustrates this progression. Each stage broadens what systems can perceive and produce, while making memory, grounding, timing, and alignment across turns more difficult. 1 arXiv:2608.17605v1 [cs.CL] 18 Aug 2026 Recent evidence shows that even frontier mod- els still struggle with sustained interaction. Mod- els can underuse relevant information even when it remains inside the context window (Liu et al., 2024a). They also degrade when a task is dis- tributed across several turns rather than stated as a complete single-turn instruction (Laban et al.,2026). MultiChallenge further shows that frontier models remain below human-level reli- ability on realistic multi-turn instruction follow- ing (Deshpande et al.,2025). These failures be- come more complex in spoken, multimodal, and tool-augmented settings. Speech recognition or acoustic-token errors can propagate across turns, visual grounding can decay as the dialogue moves away from the original image or video, and tool- using agents must preserve external state in ad- dition to dialogue history. Evaluation setups that score turn-level responses can therefore miss fail- ures that only emerge at the session level ( Kwan et al. ,2024;Graf et al.,2026). Related Surveys.Prior surveys provide useful coverage of dialogue systems, multi-turn LLM ca- pabilities, agents, and multimodal LLMs, how- ever, they study these areas largely in isola- tion. Dialogue-system surveys review the tran- sition from modular or retrieval-based systems to LLM-based dialogue, however, remain largely text-centered ( Wang et al.,2023;Yi et al.,2025). Surveys on multi-turn interaction discuss instruc- tion following, memory, and consistency, however, do not cover spoken or multimodal interaction in depth ( Zhang et al.,2025b;Li et al.,2025f). Work on LLM agents focuses mainly on tool use and agent evaluation ( Guan et al.,2026), while sur- veys of multimodal LLMs often treat dialogue as a downstream application rather than the central unit of analysis (Zhang et al.,2024b). Cultural and di- alectal variation is also largely absent from these discussions. In contrast, this survey treatssus- tained interactionas theprimary unit of analysis across modalities, systems, resources, and evalua- tion methods. Table 1summarizes how our scope differs from the closest related surveys. Appendix Table11provides a more detailed comparison of their conceptual framing, modality coverage, and key distinctions. Survey Scope and Selection.We organize our coverage along three dimensions:interaction depth(multi-turn sessions where later responses depend on earlier context),modality complexity SurveyMT Speech Vision Tri Cultural Evalâ (YI ET AL.,2025)377773 (WANG ET AL.,2023)377777 (ZHANG ET AL.,2025B)377773 (LI ET AL.,2025F)3777â˘3 (GUAN ET AL.,2026)377773 (ZHANG ET AL.,2024B)7â˘3777 This survey333333 Table 1: Comparison with closely related surveys.MT = multi-turn focus;Speech= spoken dialogue/Audi- oLLMs;Vision= image/video;Tri= joint textâspeechâ vision;Cultural= cross-lingual, dialectal, or cultur- ally grounded coverage;Evalâ= multi-turn evaluation gaps.â˘denotes partial coverage. (text through speech to omni-modal systems), and cultural and linguistic diversity(multilingual, di- alectal, and culturally grounded settings). Fig- ure1illustrates this joint space. We include work that advances dialogue across turns or de- velops thedata,models,training methods, and evaluationneeded for sustained interaction, span- ning topics such astext dialogue, spoken dia- logue, AudioLLMs, vision-language and video di- alogue, omni-modal systems, conversational re- trieval, tool use, and culturally grounded re- sources. Details keywords are provided in Ap- pendix Sec. A.4. To ensure a systematic, and reproducible review of the literature, we followed PRISMA-ScR framework ( Page et al., 2021). To maximize coverage, we searched Google Scholar and Semantic Scholar and ob- tained papers from IEEE Xplore, ACM Digital Li- brary, Elsevier, DBLP, the ACL Anthology and arXiv. Papers include publications from â ACL, NeurIPS, ICLR, ICML, Interspeech, SIGDIAL, CVPR, ICCV, AAAI and related venues. Our ini- tial search identifiedâź4K papers. After applying the PRISMA screening, we selected 200 papers for detailed review. Contributions.This survey makes following con- tributions. â˘We provide a unified survey of multi-turn con- versational AI spanning text, speech, vision, video, agentic, and culturally grounding. â˘We organize datasets, benchmarks, models, training strategies, and evaluation methods around session-level interaction rather than isolated responses. â˘We identify open challenges in memory, co- herence, paralinguistic understanding, multi- modal grounding, real-time interaction, and cultural alignment. 2 Findings.This survey highlights several gaps in the current multi-turn conversational AI. â˘Capability gap.Longer context windows, stronger base models, and broader modality support do not guarantee session-level com- petence. Systems still struggle with mem- ory, grounding, assumption revision, tool state, spoken timing, and cross-cultural adaptation. â˘Resource gap.Existing datasets remain dom- inated by English with (over 80%), text-only, and text-image interaction. Speech, video, omni-modal, and culturally grounded multi- turn resources remain limited. â˘Evaluation gap.Benchmarks increasingly use LLM-as-judge and mixed approaches, but session-level metrics, human validation, and reproducible scoring remain underdeveloped. â˘Integration gap.Most resources test one dimension at a time. Few combine long- horizon memory, multilinguality, speech, vi- sual grounding, tool use, safety, and cultural alignment in the same evaluation setting. 2 Multi-turn Dialogue We define aconversationas an interaction episode between a user and a system, consisting of ordered turns. Amulti-turn conversationis a ses- sion withT>1exchanges, where later turns may depend on earlier requests, responses, evidence, actions, or state. This dependency makes the full session the natural unit of analysis rather than a single response. We define a session,D, as D=(u t ,y t ,c t ,m t ,z t ) T t=1 , y t =f θ (u t ,c t ,m t ,z t ). Hereu t is the user input at turnt,y t is the sys- tem response,c t is the accumulated dialogue con- text from earlier turns,m t is the multimodal con- text, andz t is optional external state. The multi- modal context may include speech, images, video, or other non-text signals, and may be empty in text- only settings. The external state may include mem- ory, retrieved evidence, tool outputs, task state, or user profile information. The functionf θ de- notes the conversational model. This formulation makes the multi-turn setting explicit because each response can depend on the current input, the pre- vious dialogue, modality-specific context, and ex- ternal state. Modality and Context.The user inputu t and system responsey t may be expressed astext, speech,images,video, ortheir combination. The multimodal contextm t captures modality specific information such as acoustic cues, visual content, or video events. The dialogue contextc t captures what has already happened in the session, while the external statez t captures information that may persist or change. Since these signals carry dif- ferent information, the model must decide what to retain, update, or ignore at each turn. Dialogue Types.Dialogue systems serve dif- ferent interaction functions. We considertask oriented dialoguefor goal completion via state tracking (Wu et al.,2020;Eric et al.,2020), open-domain dialoguefor coherent conversation across topics (Zhang et al.,2020b;Roller et al., 2021),knowledge grounded dialoguefor re- sponses grounded in documents, databases, or re- trieval (Varshney et al.,2022),social and em- pathetic dialoguefor socially appropriate re- sponses (Zhang et al.,2024d),persona grounded dialoguefor consistency with a user or character profile ( Gosling et al.,2023;Oh et al.,2023),agen- tic dialoguefor planning and tool use (Yao et al., 2022;Lu et al.,2024), andmultimodal grounded dialoguefor reasoning over text, speech, images, video, or their combination ( Xue et al.,2025). These types often overlap, so they serve as ana- lytical lenses not strict dataset boundaries. Evaluation Focus.A multi-turn system is mainly evaluated at theturnandsessionlevels. Turn level evaluation asks whether the current re- sponse is correct and useful. Session level evalu- ation asks whether the system remains consistent, grounded, and useful as earlier turns shape later ones. Some settings add further requirements. Per- sistent memory extends the problem beyond one session, tool using systems add external actions and state changes, and spoken systems may re- quire full duplex evaluation for interruptions, over- lapping speech, and response timing. Our focus is interaction within a multi-turn session, while cross session memory and agentic interaction are cov- ered as related extensions when they affect session level behavior. 3 Datasets and Benchmarks Datasets and benchmarks for multi-turn dialogue differ not only in modality, but also in what they assume a system must preserve across turns. Text resources emphasize context tracking, instruction 3 retention, persona consistency, retrieval, and so- cial intelligence. Spoken resources add acous- tic cues, turn-taking, prosody, speaker variation, ASR errors, and full-duplex timing. Multimodal and video resources add visual, audio-visual, and temporal grounding. Cultural and linguistic re- sources expose whether these abilities transfer be- yond English-centered settings. We organize datasets by modality in Table2 and benchmarks by evaluation focus in Table3. Overall, they cover text-only, spoken, multimodal, video, and culturally grounded interaction. Some findings are as follows. â˘English and text remain dominant.Outside the cultural and cross-lingual block, 43 of 52 dataset entries are English-only. The bench- mark table shows the same pattern, with 32 of 53 benchmarks focusing on text-only dataset. â˘Modality coverage is expanding, however, varies across modalities.Table 2lists 18 text- image resources, however, only 8 spoken and 6 video-oriented datasets. Many multimodal resources still use static images rather than streaming, temporal, or spoken interaction. â˘Task coverage is broadening.Though QA and instruction following dominate, newer re- sources add memory, personalization, tool use, retrieval, social interaction, and emotional in- telligence. Long-horizon and multi-session di- alogue remain rare. â˘Cultural coverage is growing, however, of- ten single-turn.Resources such as CVQA, Dallah, MMA-ASIA, OASIS, and Shawarma Chats improve cultural coverage. However, 4 of the 5 entries in this block are single-turn and do not test sustained interaction. â˘Evaluation needs stronger calibration. LLM-as-judge and mixed scoring now dom- inate many benchmarks. These setups scale well, but require stronger human evaluation and clearer reproducibility. Safety and robust- ness evaluation also remains text-centered. 4 Modeling Paradigm Modeling paradigms have shifted from modu- lar dialogue pipelines to general-purpose LLMs, AudioLLMs, omni-modal systems, and tool- augmented agents, as shown in Figure 2. Ap- pendix Table5summarizes the details. This evolu- tion has broadened the interaction interface, how- ever, it has not solved the central multi-turn chal- Multi-Turn Complexity Classical & Pre-LLM (Text) 2018--2021 Instruction- Tuned LLMs (Text) 2022--2023 AudioLLMs (Text + Speech) 2023--2024 Omni-Modal (Text + Speech + Vision) 2024--2026 DialoGPT TOD-BERT InstructGPT Vicuna SpeechGPT Qwen2-Audio Gemini-2.5 Qwen2.5-Omni Evolution of Conversational Paradigms Context Window Limits Cross-Turn Grounding Decay Temporal Alignment Errors Real-Time Duplex Interruptions Figure 2: Evolution of conversational architectures. Multi-turn challenges escalate as systems evolve from simple text to complex omni-modal agents. lenge. Most systems still treat dialogue history as a flat context, while only a smaller body of work explicitly models memory, state updates, selective grounding, and efficient long-session reasoning. Classical systemsseparated language understand- ing, dialogue state tracking, policy learning, and response generation, often using rule-based, statis- tical, or neural sequence-to-sequence components. Transformer-eradialogue models began to con- solidate these functions through pretraining, multi- task learning, knowledge injection, and explicit memory mechanisms. Instruction-tuned LLMsthen shifted dialogue modeling away from task-specific managers to- ward general-purpose models trained or adapted for instruction following, open-ended interaction, and structured task behavior through prompting or function-like abstractions. The next stage added speech, vision, and real-time interaction. AudioLLMsextend conversational AI from text- only interaction to spoken dialogue by modeling speech alongside language. Early systems rep- resent speech as discrete tokens within an LLM framework, enabling speech input and output to be handled as part of the same sequence model- ing problem. Later models broaden this direction toward chat-ready audio understanding, generic hearing, and interleaved speech-text modeling. A second line targets real-time spoken inter- action, supporting streaming generation, control- lable voice, emotion, timbre, and full-duplex dia- logue with low latency. Textless spoken dialogue models further show that dialogue can be modeled directly from raw audio, preserving turn-taking and paralinguistic cues often lost in transcripts. Overall, AudioLLMs move dialogue systems to- ward speech-native interaction. Yet most still han- dle multi-turn context through the underlying lan- 4 ReferenceMod. TaskTurnsSize Lang. Split Curation Text-only P ERSONA C HAT (Zhang et al.,2018) TPS 6 â 20 10,981/164K u en TrH COQA (Reddy et al.,2019)TQA6â208,399 enTr+Ev H MULTIWOZ2.1 (Eric et al.,2020)TTOD/DST6â20.10,438 enTr+Ev H PIPPA (Gosling et al.,2023)TPS>2025,940/>1M u enTrH+S ULTRACHAT(Ding et al.,2023)TInstâ¤51.5M enTrS MT-BENCH(Zheng et al.,2023)TInstâ¤580 q enEvH WILDCHAT(Zhao et al.,2024b)TODâ¤51M ML(68) Tr+Ev H MT-BENCH-101 (Bai et al.,2024)TInstâ¤51,388 enEvH+S MT-EVAL(Kwan et al.,2024)TInst6â20168 enEvH+S PRODIGY(Occhipinti et al.,2024)TPS420,850 enTr+Ev H+S LONGMEMEVAL(Wu et al.,2025b)TMemMS500 q enEvH+S MULTICHALLENGE(Deshpande et al.,2025)TInstâ¤10 / 5 avg273 enEvH+S LMSYS-CHAT-1M (Zheng et al.,2024)TODâ¤51M ML(154) Tr+Ev H+S Ď-BENCH(Yao et al.,2025)TTODâ165tasks enEvS+H PERSONAMEM(Jiang et al.,2025)TPers/MemMS180+ histories /âź6K q enEvS+H CONSISTENTCHAT(Chen et al.,2025a)TInst6--2015K / 224K u enTrS DOCTALK(Lee et al.,2025a)TKG/QA>20730K enTrH+S TOOLWOZ (Lattimer et al.,2025)TTOD+Toolâ7,849 scenarios enTr+Ev H+S SOTOPIA (Zhou et al.,2024a)TSocial/PS6--20450 tasks enEvH+S DIALSIM(Kim et al.,2024b)TMem>2018.99K turns / 1.31K sess. enEvS+H Spoken (audio / speech) SPOKENWOZ (Si et al.,2023)T+S TOD/DST>205.7K / 249 h enTr+Ev H DEEPDIALOGUE(Koudounas et al.,2025)T+S OD/EI3â1040,150 enTrS + H AUDIOMULTICHALLENGE( Gosai et al.,2025)SInst6â20452 enEvH+S MENASPEECHBANK(Ali et al.,2026)T+S TOD/Persâ¤5417K ar*Tr+Ev H+S C3 (Ma et al.,2025)T+S QA6--201,079 en+zh EvH+S ASK-QA (Chen et al.,2025d)T+S QAâ¤57,830 enTr+Ev S MULTI-BENCH(Deng et al.,2025)SOD+EI6--201.5K en+zh EvH+S MSIB (Tong et al.,2025)T+S OD6â20244 enEvS+H Multimodal VISDIAL(Das et al.,2017)T+V QA6â20âź123K / 1.23M q enTr+Ev H MMDIALOG(Feng et al.,2023a)T+V ODâ¤51.08M enTrH IMAD (Moskvoretskii et al.,2024)T+V OD6â204864 enTrH INFOVISDIAL(Wen et al.,2023)T+V KG6â2050K enTrS DIALOGCC (Lee et al.,2024)T+V OD6â2083K enTrH+S L O C O M O ( Maharana et al.,2024) T+V Mem > 20 10 / 1,986 q en EvH+S TMDIALOG(Lei et al.,2025)T+V Inst6â2067.9K / 329 enTr+Ev H+S MMDU-45K( Liu et al.,2024d)T+V Inst6â2045K / 410K q enTrS+H CONVBENCH(Liu et al.,2024b)T+V Instâ¤5577 enEvH+S CB-300K (Tian et al.,2025b)T+V QAâ¤5340k/717K q enTr+Ev H+S MMDIAG(Liu et al.,2025a)T+V QAâ¤5639K q enTr+Ev H+S MULTIVERSE(Lee et al.,2025b)T+V Instâ¤5647 enEvH+S MMRC ( Xue et al.,2025)T+V Mem+QA6â205,120 / 28K q enEvH ALIGNMMBENCH(Wu et al.,2025d)T+V Instâ4,978 q zhEvH+S MEM-GALLERY(Bei et al.,2026)T+V MemMS240 sess. / 1,711 q enEvH+S MMMB (Tong et al.,2025)T+V Mem6â20300 enEvS+H DIALOGBEN( Huang et al.,2025b)T+V T2Iâ¤59,957 en+zh EvS MMMT-IF (Epstein et al.,2024)T+V Inst1--2071 enEvH+S Video (text + video / streaming) MT-VIDEO-BENCH(Pan et al.,2025)T+Vid QA6â201,000 / 5,887 q enEvH+S OMNIMMI ( Wang et al.,2025c)T+S+Vid QAâ¤51,121 v /2,290 q enEvH+S SCVBENCH( You et al.,2025)T+Vid QA6â20925 v / 7,280 q enEvH+S COGSTREAM(Zhao et al.,2026b)T+Vid QAâ1,088 v / 59,032 q enTr+Ev H+S IVCR-200K (Han et al.,2024)T+Vid Ret6--2012,516 v / 201,631 q en+zh Tr+Ev S+H SVBENCH(Yang et al.,2025)T+Vid QAâ¤549,979 q / 1,353 v zhTr+Ev H+S Cultural ( â single-turn) CVQA â (Mogrovejo et al.,2024)T+V QA110,374 q ML(31) EvH DALLAH â (Alwajih et al.,2024)T+V QA120Ă6 q ar(6D) EvH MMA-ASIA â ( Weihua et al.,2025)T+S+V QA127,000 q ML(10) EvH+S OASIS â (Alam et al.,2025)T+S+V QA114.8M q en+ar* Tr+Ev H+S SHAWARMACHATS(Zeinalipour et al.,2025)TOD+KG6--2030K ar*(3D) Tr+Ev H+S Table 2: Multi-turn dialoguedatasetsorganised by modality.Mod.: T = text, S = speech/audio, V = image/vision, Vid = video.Task: TOD = task-oriented dialogue, OD = open-domain, QA = question answering, PS = persona / roleplay, Pers. = personalization / user profiling, KG = knowledge-grounded, Inst = instruction-following bench- mark, T2I = text-to-image generation / editing, EI = emotional intelligence, Ret = retrieval, Mem = long-term mem- ory.Turns: approximate range per session (â¤5/6-20/>20/ MS = multi-session).Size: # dialogues unless marked: u = utterances; q = QA pairs / instances; h = hours; v = videos.Split: Tr = training data, Ev = evaluation data, Tr+Ev = both.Curation: H = human; S = synthetic / LLM-generated;Lang.: ML(*) = Multilingual; Number in parenthesis indicate number of languages. D = Dialects. Datasets marked â aresingle-turn; they are included solely as cultural / cross-lingual baselines. guage model, with limited explicit modeling of long-horizon spoken memory, interruptions, and cross-turn acoustic grounding. Omni-modalmodels extend AudioLLMs by jointly handling text, speech, vision, and some- times video. Recent systems explore several de- sign choices, including streaming duplex inter- action, low-latency speech and vision alignment, shared multimodal tokenization, modality-specific encoders, and separate modules for reasoning and speech generation. These designs move conversa- tional AI from speech-text interaction toward uni- fied perception and generation across modalities. Audio-visual dialogueforms a more specialized 5 BenchmarkMod.TurnsEval General multi-turn instruction-following & consistency MT-BENCH(Zheng et al.,2023)T2J MT-BENCH-101 (Bai et al.,2024)T3J MT-EVAL(Kwan et al.,2024)T7M MULTICHALLENGE(Deshpande et al.,2025)T5avg /â¤10M IHEVAL(Zhang et al.,2025e)T2R TURNWISE (Graf et al.,2026)T2â8J TOD-PROCBENCH(Ghazarian et al.,2025)TâR EVOLIF (Jia et al.,2025)TâM PARROT-BENCH(Sun et al.,2024)T8J Ď-BENCH(Yao et al.,2025)TâM STRUCTFLOWBENCH(Li et al.,2025b)T4.14 avg J PERSONAMEM(Jiang et al.,2025)Tâ¤60 sessions Ac CORAL (Cheng et al.,2025)T8.26 avg M TOOLSANDBOX(Lu et al.,2025)T13.9 avg rM TURNBENCH-MS (Zhang et al.,2025d)T>20M MINT (Wang et al.,2024)Tâ¤5R SOTOPIA (Zhou et al.,2024a)Tâ¤20M AGENTBOARD(Ma et al.,2024)T3--25M MEMORYAGENTBENCH(Hu et al.,2025)TâM Multilingual & cross-lingual M2LINGUAL(Maheshwary et al.,2025)T2J CMT-EVAL(Tian et al.,2025a)TmultiJ ALIGNMMBENCH(Wu et al.,2025d)T+Vâ avgM MT-BENCH-HI(Kamath et al.,2025)T2J CUDIALOG(Cao et al.,2024)T8sent.(5+3)R INDOTOD (Kautsar et al.,2023)T2.63/4.06avg R Multimodal (image + text) CONVBENCH(Liu et al.,2024b)T+V3J MMDU (Liu et al.,2024d)T+V15avg /27max J MMCR (Yan et al.,2025a)T+V4/8J MMRC ( Xue et al.,2025)T+V$15.2avg M MULTIVERSE(Lee et al.,2025b)T+V3.91avg J MEM-GALLERY(Bei et al.,2026)T+VMSM MMMT-IF ( Epstein et al.,2024)T+V1--20M MMMB ( Tong et al.,2025)T+Vâ¤15J Spoken & video AUDIOMULTICHALLENGE(Gosai et al.,2025)S3â8M MT-VIDEO-BENCH(Pan et al.,2025)T+Vid5â8J OMNIMMI ( Wang et al.,2025c)T+S+Vid stream /1â3M SCVBENCH(You et al.,2025)T+Vid6.16avg R COGSTREAM(Zhao et al.,2026b)T+Vidstream;5.02J AVHBENCH(Sung-Bin et al.,2025)T+S+VâM MTALK-BENCH( Du et al.,2025b)S2â3M FD-BENCH(Peng et al.,2025)Sâ¤5M MULTI-BENCH(Deng et al.,2025)S8 avgM SVBENCH(Yang et al.,2025)T+Vidstream;4.29J MSIB (Tong et al.,2025)S2--10M Robustness, fairness & safety FB-BENCH(Li et al.,2025d)T2M FAIRMT-BENCH( Fan et al.,2025)TmultiM SYCON-BENCH(Hong et al.,2025)TmultiR CURSE OFMULTI-MODALITIES(Leng et al.,2026) T+S+VâM LOST INMULTI-TURN(Laban et al.,2026)TmultiM X-TEAMING(Rahman et al.,2025)Tavg 5.10 M CRESCENDO( Russinovich et al.,2025)T1--7R DIAHALU(Chen et al.,2024b)T6.91 avg H+J SAFEDIALBENCH(Cao et al.,2026)T3--10M Table 3: Multi-turn evaluation benchmarks grouped by modality and evaluation focus.Mod.denotes modal- ity, with T for text, S for speech, V for vision, and Vid for video.Turnsreports the average number of turns per instance or session when available, and otherwise gives a qualitative range.Evaldenotes the evaluation protocol, with R for rule-based scoring, H for human evaluation, J for LLM-as-judge, M for mixed evalua- tion, and Ac for accuracy. Benchmarks that evaluate only single-turn ability are excluded. direction in which visual cues support speaker tracking, turn-taking, and grounding rather than serving only as general perceptual input. AV- Dialog uses lip-centered visual features to iden- tify the target speaker, predict turn transitions, and generate responses under noise and compet- ing speech (Chen et al.,2026). MAViD instead combines multimodal understanding with synchro- nized audio-video response generation through a ConductorâCreator architecture, though it is not evaluated evaluated for sustained multi-turn interaction (Pang et al.,2025). Despite this progress, most omni-modal systems still treat dia- logue history as a flat sequence. Multi-turn-aware methods address grounded memory, context man- agement, and long-session inference cost, while long-horizon grounding, reference tracking, and session-level coherence remain open challenges. Agenticdialogue systems extend multi-turn mod- eling by framing conversation as a loop of plan- ning, tool use, observation, and revision. Re- Act interleaves reasoning and actions (Yao et al., 2022), Reflexion uses feedback and episodic mem- ory to improve later attempts (Shinn et al.,2023), and MetaGPT enables role-based collaboration through structured workflows ( Hong et al.,2024). Recent systems apply these ideas to tool use, function calling, web navigation, recommenda- tion, video reasoning, and live multimodal inter- action ( Wu et al.,2024;Lu et al.,2024). These set- tings introduce three key challenges. First, agents must track external states and dependencies that may not appear in the dialogue history; ToolSand- box evaluates such stateful and implicit dependen- cies. Second, tool calls can produce undesirable or irreversible effects, motivating evaluation of both task completion and harmful side effects. Third, failed actions require replanning and backtrack- ing. Reflexion and SCoRe support self-correction, but do not provide transactional rollback for al- ready executed external actions. Overall, model- ing has progressed from modular state tracking to unified, multimodal, and agentic systems, while long-horizon grounding and cross-turn memory re- main key open challenges. Appendix A.2provides additional model-level details. 5 Training Strategies Multi-turn dialogue training must account for de- pendencies across turns, delayed rewards, chang- ing user intent, and context-sensitive response quality. We group existing training approaches into five families: (i) supervised fine-tuning, (i) re- inforcement learning and preference optimization, (i)multi-task learning,(iv)synthetic data gener- ation, and(v)conversational retrieval-augmented 6 training. In Appendix Table5, we summarize modeling paradigms, while Table6presents the main training strategies. Table7further compares these different families in terms of their mecha- nisms, suitable settings, strengths, limitations, and modality coverage. Supervised fine-tuning.Supervised fine-tuning requires conversations that reflect real multi-turn phenomena such as follow-up questions, anaphora, ellipsis, topic shifts, and safety escalation. Ultra- Chat and WildChat provide large-scale synthetic and real multi-turn chat data, while Parrot ex- plicitly targets referential phenomena in follow-up turns. Domain-specific resources such as Aquila- Med, Qilin-Med, and Zhongjing adapt SFT to multi-turn clinical dialogue, and XGuard-Train ex- tends SFT to multi-turn adversarial safety training. Reinforcement learning and preference opti- mization.Multi-turn reinforcement learning shifts the learning signal from isolated responses to full interactions. InstructGPT establishes the standard RLHF pipeline, while ArCHer and MT- RLHF adapt optimization to longer conversations and conversation-level preferences. Preference methods such as DMPO, Multi-turn DPO/KTO, Parrot, and SDPO optimize preferences at re- sponse, segment, or trajectory level. Other work trains action selection, tool use, self-correction, proactive interaction, and long-horizon sparse- reward behavior across turns. Clinical and tutor- ing settings further combine SFT, RLHF, DPO, ex- pert feedback, and multi-agent RL for specialized multi-turn interaction. Multi-task learning.Multi-task learning im- proves transfer by sharing representations across dialogue objectives. PPTOD jointly trains re- sponse generation, dialogue state tracking, and policy prediction, while TOD-BERT learns dialogue-aware representations across task- oriented dialogue corpora. DAMSEL shows that auxiliary objectives are useful for conversational QA when labeled speech data is limited. Synthetic data generation.Synthetic generation is widely used to scale multi-turn data and con- trol dialogue structure. UltraChat, MMDU-45K, MMDiag, and TMDialog generate instruction- following, multimodal, note-taking, and context- modeling dialogues. Other pipelines target Arabic dialogue, emotional speech, tool use, coherent in- tent trajectories, multi-topic information seeking, reviewer-guided regeneration, task-oriented boot- strapping, and multi-turn API/tool-use simulation. Conversational RAG.Conversational RAG trains models to retrieve, filter, and use evidence across turns. PK-ICR and commonsense retrieval ap- proaches ground responses in persona, external knowledge, and knowledge graphs. ChatQA, ChatQA-2, HAConvDR, UniConv, IterCQR, and related work improve conversational retrieval, query reformulation, history selection, and joint retrieval-generation training. KEDiT and CORAL further explore efficient evidence integration and citation-aware conversational RAG training. To translate this comparison into practice, in Table8, we map common deployment goals to the training strategy. In AppendixA.2, we provide additional details on training strategies. 6 Evaluation of Multi-turn Dialogue Evaluating multi-turn dialogue requires moving beyond single-response metrics. A response may be fluent and factually correct on its own; how- ever, it can still fail at the session level if it ig- nores earlier instructions, loses user preferences, misuses retrieved evidence, mishandles tool state, or breaks grounding across modalities. Evaluation therefore needs to measure both local response quality and cross-turn consistency, memory, state tracking, grounding, interaction success, and ro- bustness across the full dialogue. Existing evaluation methods fall into five broad groups.(i)Surface-form and task-oriented met- rics, such as overlap scores, semantic similarity, dialogue state tracking, and task success, remain useful, though they provide limited insight into session-level coherence.(i)Session-level met- rics assess instruction retention, memory recall, constraint satisfaction, feedback integration, and dialogue-level hallucination.(i)Agentic and retrieval-grounded evaluation measures tool use, stateful execution, evidence retrieval, and source attribution.(iv)Speech-native and full-duplex evaluation adds speech quality, paralinguistic cues, interruption handling, latency, and response tim- ing. (v) LLM-as-a-judge and human evaluation support open-ended assessment; however, their re- liability depends on rubric quality and may be affected by judge bias, rater inconsistency, and small ranking differences. In Table 10, we com- pare these five families by what they measure, the systems they best suit, and their main limi- tations, showing that no single family fully cap- tures session-level competence. Appendix A.3 7 provides a detailed discussion of these approaches, and Table9summarizes representative metrics and frameworks. Evaluation gaps.Current setups make multi-turn failures more visible, however, several gaps re- main.(i)Strong single-turn performance does not reliably transfer to session-level interaction. (i)Agentic dialogue lacks unified evaluation for hidden state, tool-side effects, source attribution, and repeated-trial reliability.(i)Spoken and full- duplex systems still lack integrated session-level evaluation that jointly measures semantic correct- ness, audio quality, timing, interruptions, and par- alinguistic behaviour.(iv)Multi-turn evaluation remains focused in English and high-resource set- tings, leaving multilingual, dialectal, and cultur- ally grounded evaluation split across separate re- sources. These gaps show that current evalua- tion still measures many components of dialogue competence separately, rather than testing reliable session-level behaviour across languages, modali- ties, tools, and time. 7 Challenges, and Future Directions The studied literature shows progress in datasets, modeling, training, and evaluation, however, sus- tained multi-turn interaction remains unresolved. Stronger single-turn models alone do not achieve multi-turn competence. Systems must preserve context, revise assumptions, ground responses, coordinate tools, and maintain quality across long, multimodal, and culturally diverse sessions. Mechanisms for memory, cross-turn grounding, speech-native interaction, robust revision, and session-level evaluation remain limited. Persistent memory and state management. Current long-context models can process more di- alogue history, however, they still struggle to re- trieve, update, and apply the right information across turns. This problem grows in multi-session settings, where preferences, facts, and task states change over time. Future work should move be- yond flat context windows toward explicit mem- ory systems that separate dialogue state, and task- specific working memory. These systems should support updates, forgetting, conflict resolution, and transparent retrieval across sessions. Cross-turn grounding.Grounding weakens when responses depend on earlier turns, external evidence, visual content, acoustic cues, or tool outputs. Full dialogue history can add noise in retrieval-augmented dialogue, while multimodal references often decay across turns. Future work should develop selective grounding methods that identify the relevant dialogue history, evidence, and modality streams for each turn. This requires structured dialogue state rather than treating the full history as a flat input sequence. Speech-native and full-duplex interaction. Spoken dialogue adds timing, prosody, interrup- tions, repairs, noise, and paralinguistic cues that transcript-only pipelines often miss. Full-duplex systems make this harder because users and systems may speak at the same time. Future work should treat speech as an interaction medium, not only as an input modality. This requires training and evaluation for interruption handling, response timing, spoken clarification, voice-grounded tool use, and paralinguistic understanding. Robustness and revision.Multi-turn systems often commit to early assumptions and fail to re- vise them after later corrections. They are also vulnerable to adversarial escalation, where indi- vidually benign turns accumulate into unsafe or incorrect outcomes. Future work should improve session-level robustness by enabling models to de- tect uncertainty, recover from mistakes, and main- tain safety throughout the interaction. Evaluation mismatch.Current evaluation meth- ods capture only parts of multi-turn behavior. Surface-form metrics, task-oriented metrics, agen- tic benchmarks, speech-native evaluation settings, and LLM-as-judge methods remain difficult to compare. No unified framework yet measures session-level competence across memory, ground- ing, tool use, speech, safety, and user satisfac- tion. Future evaluation should move from turn- level scoring to session-level assessment, with ex- plicit measures of state consistency, evidence use, revision, timing, and repeated-trial reliability. Cultural and linguistic coverage. Most re- sources and evaluations still focus on English and high-resource settings. Dialects, code-switching, cultural pragmatics, and region-specific knowl- edge shape how users express intent and judge ap- propriate responses. Future benchmarks should move beyond isolated language-specific datasets toward comparable multilingual and culturally grounded evaluation suites. This is especially im- 8 portant for multimodal settings, where language, culture, accent, and visual context interact. Beyond turn-by-turn exchange.Most systems still follow a simple user-query and assistant- response loop. This framing does not fully capture proactive suggestions, mid-utterance clarification, streaming perception, emotion-aware interaction, or mixed-initiative collaboration. Future systems need interaction models that treat dialogue as con- tinuous, stateful, and adaptive rather than as a se- quence of independent turns. 8 Conclusion Multi-turn conversational AI is shifting from text- only dialogue toward sustained interaction across speech, vision, video, tools, and culturally diverse contexts. This study shows that modality support has advanced faster than session-level competence. Current systems can increasingly perceive, speak, and act, but still struggle to maintain memory, re- vise assumptions, preserve grounding, coordinate external state, handle spoken timing, and adapt across languages and cultures. We argue that fu- ture work should evaluate conversational systems not only by the quality of isolated responses, but by their ability to sustain coherent, grounded, safe, and culturally appropriate interaction across turns. Limitations This study focuses on recent work on multi-turn conversational AI across data, models, training, and evaluation. While we aim to cover major direc- tions, the literature is growing quickly, and some new or concurrent resources may not be included. We cover diverse systems into broad categories, which may hide finer differences in architecture, training data, deployment setting, or evaluation se- tups. Societal/Broader Impact This study aims to clarify the capabilities, gaps, and risks of multi-turn conversational AI. By orga- nizing prior work across datasets, models, training, and evaluation, it can help researchers design sys- tems that better support long-horizon interaction, multilingual access, speech-based interfaces, mul- timodal assistance, and culturally grounded com- munication. Ethical Considerations This study does not introduce new datasets or run experiments with human participants. References Abdelrahman Abdallah, Mahmoud Kasem, Mah- moud Abdalla, Mohamed Mahmoud, Mohamed Elkasaby, Yasser Elbendary, and Adam Jatowt. 2024. Arabicaqa: A comprehensive dataset for arabic question answering. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2049--2059. Marwa Abdulhai, Isadora White, Charlie Victor Snell, Charles Sun, Joey Hong, Yuexiang Zhai, Kelvin Xu, and Sergey Levine. 2025.LMRL gym: Benchmarks for multi-turn reinforcement learning with language models . InForty-second International Conference on Machine Learning. Inclusion AI, Biao Gong, Cheng Zou, Chuanyang Zheng, Chunluan Zhou, Canxiang Yan, Chunx- iang Jin, Chunjie Shen, Dandan Zheng, Fudong Wang, and 1 others. 2025. Ming-omni: A uni- fied multimodal model for perception and gen- eration.arXiv preprint arXiv:2506.09344. Firoj Alam, Ali Ezzat Shahroor, Md Arid Hasan, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Mohamed Bayan Kmainasi, Shammur Absar Chowdhury, Basel Mousi, Fahim Dalvi, Nadir Durrani, and 1 others. 2025. Everydaymmqa: A multilingual and multimodal framework for culturally grounded spoken visual qa.arXiv preprint arXiv:2510.06371. Zien Sheikh Ali, Hunzalah Hassan Bhatti, Ra- bindra Nath Nandi, Shammur Absar Chowd- hury, and Firoj Alam. 2026. MENASpeech- Bank: A reference voice bank with persona- conditioned multi-turn conversations for audi- ollms.arXiv preprint arXiv:2602.07036. Fakhraddin Alwajih, Gagan Bhatia, and Muham- mad Abdul-Mageed. 2024.Dallah: A dialect- aware multimodal large language model for Arabic . InProceedings of the Second Ara- bic Natural Language Processing Conference, pages 320--336, Bangkok, Thailand. Associa- tion for Computational Linguistics. 9 Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jia- heng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, and Wanli Ouyang. 2024.MT-bench-101: A fine-grained bench- mark for evaluating large language models in multi-turn dialogues. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 7421--7454, Bangkok, Thailand. Association for Computational Linguistics. Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2020.PLATO: Pre-trained dia- logue generation model with discrete latent vari- able. InProceedings of the 58th Annual Meet- ing of the Association for Computational Lin- guistics, pages 85--96, Online. Association for Computational Linguistics. Yuanchen Bei, Tianxin Wei, Xuying Ning, Yan- jun Zhao, Zhining Liu, Xiao Lin, Yada Zhu, Hendrik Hamann, Jingrui He, and Hanghang Tong. 2026. Mem-gallery: Benchmarking mul- timodal long-term conversational memory for mllm agents.arXiv preprint arXiv:2601.03515. Luel Hagos Beyene, Vivek Verma, Min Ma, Je- sujoba O Alabi, Fabian David Schmidt, Joyce Nakatumba-Nabende, and David Ifeoluwa Ade- lani. 2025. msteb: Massively multilingual eval- uation of llms on speech and text tasks.arXiv preprint arXiv:2506.08400. Aydar Bulatov, Yury Kuratov, and Mikhail Burt- sev. 2022. Recurrent memory transformer.Ad- vances in Neural Information Processing Sys- tems, 35:11079--11091. Yash Butala, Siddhant Garg, Pratyay Banerjee, and Amita Misra. 2024.ProMISe: A proac- tive multi-turn dialogue dataset for information- seeking intent resolution. InFindings of the As- sociation for Computational Linguistics: EACL 2024, pages 1774--1789, St. Julianâs, Malta. As- sociation for Computational Linguistics. Hongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng, Zhixin Bai, Zhe Cao, Meng Fang, Fan Feng, Jiaheng Liu, Boyan Wang, Tianpei Yang, Jing Huo, Yang Gao, Fanyu Meng, Xi Yang, Chao Deng, and Junlan Feng. 2026. Safedial- bench: A fine-grained safety evaluation bench- mark for large language models in multi-turn di- alogues with diverse jailbreak attacks. InThe Fourteenth International Conference on Learn- ing Representations. Yong Cao, Min Chen, and Daniel Hershcovich. 2024.Bridging cultural nuances in dialogue agents through cultural value surveys. InFind- ings of the Association for Computational Lin- guistics: EACL 2024, pages 929--945, St. Ju- lianâs, Malta. Association for Computational Linguistics. Haonan Chen, Zhicheng Dou, Kelong Mao, Jiong- nan Liu, and Ziliang Zhao. 2024a. General- izing conversational dense retrieval via LLM- cognition data augmentation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2700--2718, Bangkok, Thailand. Association for Computational Linguistics. Jiawei Chen, Xinyan Guan, Qianhao Yuan, Mo Guozhao, Weixiang Zhou, Yaojie Lu, Hongyu Lin, Ben He, Le Sun, and Xianpei Han. 2025a. Consistentchat: Building skeleton- guided consistent multi-turn dialogues for large language models from scratch. InProceed- ings of the 2025 Conference on Empirical Meth- ods in Natural Language Processing, pages 8426--8452. Junjie Chen, Yao Hu, Junjie Li, Kangyue Li, Kun Liu, Wenpeng Li, Xu Li, Ziyuan Li, Feiyu Shen, Xu Tang, and 1 others. 2025b. Fireredchat: A pluggable, full-duplex voice interaction system with cascaded and semi-cascaded implementa- tions.arXiv preprint arXiv:2509.06502. Kai Chen, Yunhao Gou, Runhui Huang, Zhili Liu, Daxin Tan, Jing Xu, Chunwei Wang, Yi Zhu, Yihan Zeng, Kuo Yang, and 1 others. 2025c. Emova: Empowering language models to see, hear and speak with vivid emotions. InProceed- ings of the Computer Vision and Pattern Recog- nition Conference , pages 5455--5466. Kedi Chen, Qin Chen, Jie Zhou, He Yishen, and Liang He. 2024b.DiaHalu: A dialogue-level hallucination evaluation benchmark for large language models . InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 9057--9079, Miami, Florida, USA. Asso- ciation for Computational Linguistics. 10 Maximillian Chen, Ruoxi Sun, and Sercan O Arik. 2025d.Data-centric improvements for enhanc- ing multi-modal understanding in spoken con- versation modeling. InFindings of the Associa- tion for Computational Linguistics: ACL 2025, pages 1366--1387, Vienna, Austria. Association for Computational Linguistics. Maximillian Chen, Ruoxi Sun, Tomas Pfister, and Sercan O Arik. 2025e.Learning to clarify: Multi-turn conversations with action-based con- trastive self-training. InThe Thirteenth Interna- tional Conference on Learning Representations. Nuo Chen, Hongguang Li, Jianhui Chang, Juhua Huang, Baoyuan Wang, and Jia Li. 2025f. Com- press to impress: Unleashing the potential of compressive memory in real-world long-term conversations. InProceedings of the 31st Inter- national Conference on Computational Linguis- tics, pages 755--773. Qian Chen, Yafeng Chen, Yanni Chen, Mengzhe Chen, Yingda Chen, Chong Deng, Zhihao Du, Ruize Gao, Changfeng Gao, Zhifu Gao, Yabin Li, Xiang Lv, Jiaqing Liu, Haoneng Luo, Bin Ma, Chongjia Ni, Xian Shi, Jialong Tang, Hui Wang, and 17 others. 2025g. Minmo: A multi- modal large language model for seamless voice interaction.Preprint, arXiv:2501.06282. Tuochao Chen, Bandhav Veluri, Hongyu Gong, and Shyamnath Gollakota. 2026. Av-dialog: Spoken dialogue models with audio-visual in- put. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 42208--42225. Wenxi Chen, Ziyang Ma, Ruiqi Yan, Yuzhe Liang, Xiquan Li, Ruiyang Xu, Zhikang Niu, Yanqiao Zhu, Yifan Yang, Zhanxun Liu, and 1 others. 2025h. Slam-omni: Timbre-controllable voice interaction system with single-stage training. In Findings of the Association for Computational Linguistics: ACL 2025, pages 2262--2282. Zhipeng Chen, Kun Zhou, Beichen Zhang, Zheng Gong, Xin Zhao, and Ji-Rong Wen. 2023.Chat- CoT: Tool-augmented chain-of-thought reason- ing on chat-based large language models . InFindings of the Association for Compu- tational Linguistics: EMNLP 2023, pages 14777--14790, Singapore. Association for Com- putational Linguistics. Yiruo Cheng, Kelong Mao, Ziliang Zhao, Guant- ing Dong, Hongjin Qian, Yongkang Wu, Tet- suya Sakai, Ji-Rong Wen, and Zhicheng Dou. 2025.CORAL: Benchmarking multi-turn con- versational retrieval-augmented generation. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 1308--1330, Albuquerque, New Mexico. Association for Computational Linguistics. Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. 2024.Videollama 2: Advanc- ing spatial-temporal modeling and audio un- derstanding in video-llms.arXiv preprint arXiv:2406.07476. Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, and 1 others. 2024. Qwen2-audio technical report.arXiv preprint arXiv:2407.10759. Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023. Qwen-audio: Advanc- ing universal audio understanding via unified large-scale audio-language models .Preprint, arXiv:2311.07919. Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdi- nov. 2019. Transformer-XL: Attentive language models beyond a fixed-length context. InPro- ceedings of the 57th Annual Meeting of the As- sociation for Computational Linguistics, pages 2978--2988, Florence, Italy. Association for Computational Linguistics. Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, JosĂŠ MF Moura, Devi Parikh, and Dhruv Batra. 2017. Visual Dia- log. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3268--3276. Alexandre DĂŠfossez, Laurent MazarĂŠ, Manu Orsini, AmĂŠlie Royer, Patrick PĂŠrez, HervĂŠ JĂŠgou, Edouard Grave, and Neil Zeghidour. 2024. Moshi: a speech-text foundation 11 model for real-time dialogue.arXiv preprint arXiv:2410.00037. Yayue Deng, Guoqiang Hu, Haiyang Sun, Xi- angyu Zhang, Haoyang Zhang, Fei Tian, Xuerui Yang, Gang Yu, and Eng Siong Chng. 2025. Multi-bench: A multi-turn interactive bench- mark for assessing emotional intelligence abil- ity of spoken dialogue models.Preprint, arXiv:2511.00850. Kaustubh Deshpande, Ved Sirdeshmukh, Jo- hannes Baptist Mols, Lifeng Jin, Ed-Yeremai Hernandez-Cardona, Dean Lee, Jeremy Kritz, Willow E. Primack, Summer Yue, and Chen Xing. 2025. MultiChallenge: A realistic multi- turn conversation evaluation benchmark chal- lenging to frontier LLMs. InFindings of the As- sociation for Computational Linguistics: ACL 2025, pages 18632--18702, Vienna, Austria. As- sociation for Computational Linguistics. Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023.Enhancing chat lan- guage models by scaling high-quality instruc- tional conversations . InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3029--3051, Sin- gapore. Association for Computational Linguis- tics. Hanwen Du, Bo Peng, and Xia Ning. 2025a. Sapi- ent: Mastering multi-turn conversational recom- mendation with strategic planning and monte carlo tree search. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 2629--2648. Yuhao Du, Qianwei Huang, Guo Zhu, Zhanchen Dai, Shunian Chen, Qiming Zhu, Le Pan, Ming- hao Chen, Yuhao Zhang, Li Zhou, and 1 oth- ers. 2025b. Mtalk-bench: Evaluating speech-to- speech models in multi-turn dialogues via arena- style and rubrics protocols.arXiv preprint arXiv:2508.18240. Haodong Duan, Jueqi Wei, Chonghua Wang, Hongwei Liu, Yixiao Fang, Songyang Zhang, Dahua Lin, and Kai Chen. 2024. BotChat: Evaluating LLMsâ capabilities of having multi- turn dialogues. InFindings of the Association for Computational Linguistics: NAACL 2024, pages 3184--3200, Mexico City, Mexico. Asso- ciation for Computational Linguistics. Erik Ekstedt and Gabriel Skantze. 2020.TurnGPT: a transformer-based language model for predict- ing turn-taking in spoken dialog. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 2981--2990, Online. Asso- ciation for Computational Linguistics. Abdellah El Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah Alsayadi, Nadia Ghezaiel Hammouda, Alshima Alkhazimi, Aya Hamod, Al-Ya s Al- Ghafri, Wesam El-Sayed, Asila Al sharji, Mo- hamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, and 28 others. 2026. Alexan- dria: A Multi-Domain Dialectal Arabic Ma- chine Translation Dataset for Culturally Inclu- sive and Linguistically Diverse LLMs. InPro- ceedings of the 64th Annual Meeting of the As- sociation for Computational Linguistics, Cali- fornia, United States. Association for Compu- tational Linguistics. Elliot L. Epstein, Kaisheng Yao, Jing Li, Xinyi Bai, and Hamid Palangi. 2024.Mmmt- if: A challenging multimodal multi-turn in- struction following benchmark .Preprint, arXiv:2409.18216. Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar, Anuj Goyal, Peter Ku, and Dilek Hakkani-Tur. 2020. MultiWOZ 2.1: A con- solidated multi-domain dialogue dataset with state corrections and state tracking baselines. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 422--428, Marseille, France. European Language Re- sources Association. Zhiting Fan, Ruizhe Chen, Tianxiang Hu, and Zuozhu Liu. 2025. Fairmt-bench: Benchmark- ing fairness for multi-turn dialogue in conver- sational llms. InInternational Conference on Learning Representations, volume 2025, pages 190--218. Qingkai Fang, Yan Zhou, Shoutao Guo, Shaolei Zhang, and Yang Feng. 2025. LLaMA-Omni 12 2: Llm-based real-time spoken chatbot with au- toregressive streaming speech synthesis. InPro- ceedings of the 63rd Annual Meeting of the As- sociation for Computational Linguistics. Jiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, and Qingwei Lin. 2023a.MMDialog: A large- scale multi-turn dialogue dataset towards multi- modal open-domain conversation. InProceed- ings of the 61st Annual Meeting of the Asso- ciation for Computational Linguistics (Volume 1: Long Papers), pages 7348--7363, Toronto, Canada. Association for Computational Lin- guistics. Yichun Feng, Jiawei Wang, Lu Zhou, Zhen Lei, and Yixue Li. 2026. Doctoragent-rl: A multi-agent collaborative reinforcement learn- ing system for multi-turn clinical dialogue. In ICASSP 2026-2026 IEEE International Confer- ence on Acoustics, Speech and Signal Process- ing (ICASSP), pages 16952--16956. IEEE. Yujie Feng, Zexin Lu, Bo Liu, Liming Zhan, and Xiao-Ming Wu. 2023b.Towards LLM-driven dialogue state tracking . InProceedings of the 2023 Conference on Empirical Methods in Nat- ural Language Processing, pages 739--755, Sin- gapore. Association for Computational Linguis- tics. Chaoyou Fu, Haojia Lin, Xiong Wang, YiFan Zhang, Yunhang Shen, Xiaoyu Liu, Haoyu Cao, Zuwei Long, Heting Gao, Ke Li, Long MA, Xi- awu Zheng, Rongrong Ji, Xing Sun, Caifeng Shan, and Ran He. 2026. VITA-1.5: Towards GPT-4o level real-time vision and speech inter- action . InThe Thirty-ninth Annual Conference on Neural Information Processing Systems. Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2024.GPTScore: Evaluate as you desire . InProceedings of the 2024 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Technologies (Volume 1: Long Pa- pers), pages 6556--6576, Mexico City, Mexico. Association for Computational Linguistics. Kuofeng Gao, Shu-Tao Xia, Ke Xu, Philip Torr, and Jindong Gu. 2025a.Benchmarking open- ended audio dialogue understanding for large audio-language models. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 4763--4784, Vienna, Austria. Asso- ciation for Computational Linguistics. Zhaolin Gao, Wenhao Zhan, Jonathan Daniel Chang, Gokul Swamy, KiantĂŠ Brantley, Ja- son D. Lee, and Wen Sun. 2025b.Regressing the relative future: Efficient policy optimization for multi-turn rlhf. InThe Thirteenth Interna- tional Conference on Learning Representations (ICLR). Milica GaĹĄi Ě c, Catherine Breslin, Matthew Hen- derson, Dongho Kim, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, and Steve Young. 2013.POMDP-based dialogue manager adap- tation to extended domains . InProceedings of the SIGDIAL 2013 Conference, pages 214--222, Metz, France. Association for Computational Linguistics. Sarik Ghazarian, Abhinav Gullapalli, Swair Shah, Anurag Beniwal, Nanyun Peng, Narayanan Sadagopan, and Zhou Yu. 2025. Tod-procbench: Benchmarking complex instruction-following in task-oriented dialogues. InNeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models. Advait Gosai, Tyler Vuong, Utkarsh Tyagi, Steven Li, Wenjia You, Miheer Bavare, Arda Uçar, Zhongwang Fang, Brian Jang, Bing Liu, and Yunzhong He. 2025. Audio multichallenge: A multi-turn evaluation of spoken dialogue sys- tems on natural human interaction.Preprint, arXiv:2512.14865. Tear Gosling, Alpin Dale, and Yinhe Zheng. 2023. Pippa: A partially synthetic conversational dataset.arXiv preprint arXiv:2308.05884. Victoria Graf, Valentina Pyatkin, Nouha Dziri, Nathan Lambert, and Hannaneh Hajishirzi. 2026. Turnwise: The gap between single-and multi-turn language model capabilities.arXiv preprint arXiv:2603.16759 . Shengyue Guan, Jindong Wang, Jiang Bian, Bin Zhu, Jian-Guang Lou, and Haoyi Xiong. 2026. Evaluating llm-based agents for multi-turn con- versations: A survey.ACM Transactions on In- telligent Systems and Technology. 13 Qingpei Guo, Kaiyou Song, Zipeng Feng, Ziping Ma, Qinglong Zhang, Sirui Gao, Xuzheng Yu, Yunxiao Sun, Tai-Wei Chang, Jingdong Chen, and 1 others. 2025. M2-omni: Advancing omni-mllm for comprehensive modality support with competitive performance.arXiv preprint arXiv:2502.18778. Ning Han, Yawen Zeng, Shaohua Long, Chengqing Li, Sijie Yang, Dun Tan, Zemin Liu, Jianfeng Dong, and Jingjing Chen. 2024. IVCR-200k: A large-scale benchmark for interactive video corpus retrieval. Seungmin Han, Haeun Kwon, Ji jun Park, and Taeyang Yoon. 2025.Contextuallvlm-agent: A holistic framework for multi-turn visually- grounded dialogue and complex instruction fol- lowing.Preprint, arXiv:2508.15164. Jiseung Hong, Grace Byun, Seungone Kim, and Kai Shu. 2025.Measuring sycophancy of language models in multi-turn dialogues. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 2239--2259, Suzhou, China. Association for Computational Linguistics. Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xi- awu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and JĂźrgen Schmidhuber. 2024. MetaGPT: Meta programming for a multi-agent collaborative framework. InThe Twelfth Inter- national Conference on Learning Representa- tions. Yuanzhe Hu, Yu Wang, and Julian McAuley. 2025. Evaluating memory in LLM agents via incre- mental multi-turn interactions. InICML 2025 Workshop on Long-Context Foundation Models. Kai Huang, Hao Zou, Bochen Wang, Ye Xi, Zhen Xie, and Hao Wang. 2025a. Aircache: Acti- vating inter-modal relevancy kv cache compres- sion for efficient large vision-language model inference. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 23958--23967. Minbin Huang, Yanxin Long, Xinchi Deng, Rui- hang Chu, Jiangfeng Xiong, Xiaodan Liang, Hong Cheng, Qinglin Lu, and Wei Liu. 2025b. DialogGen: Multi-modal interactive dialogue system with multi-turn text-image generation. InFindings of the Association for Compu- tational Linguistics: NAACL 2025, pages 411--426, Albuquerque, New Mexico. Associ- ation for Computational Linguistics. Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card.arXiv preprint arXiv:2410.21276. Yunah Jang, Kang-il Lee, Hyunkyung Bae, Hwan- hee Lee, and Kyomin Jung. 2024.IterCQR: It- erative conversational query reformulation with retrieval guidance. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguis- tics: Human Language Technologies (Volume 1: Long Papers), pages 8121--8138, Mexico City, Mexico. Association for Computational Linguistics. Qi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang, Shibo Wang, Dun Pei, Xiangyang Zhu, and Guangtao Zhai. 2025. One battle after another: Probing llmsâ limits on multi-turn instruction following with a benchmark evolving frame- work.arXiv preprint arXiv:2511.03508. Bowen Jiang, Zhuoqun Hao, Young Min Cho, Bryan Li, Yuan Yuan, Sihao Chen, Lyle Un- gar, Camillo Jose Taylor, and Dan Roth. 2025. Know me, respond to me: Benchmarking LLMs for dynamic user profiling and personalized re- sponses at scale. InSecond Conference on Lan- guage Modeling. Sunghee Jung, Donghun Lee, Shinbok Lee, Gaeun Seo, Daniel Lee, Byeongil Ko, Junrae Cho, Ki- hyun Kim, EungGyun Kim, and Myeongcheol Shin. 2025.DiaTool-DPO: Multi-turn di- rect preference optimization for tool-augmented large language models . InProceedings of the 26th Annual Meeting of the Special Inter- est Group on Discourse and Dialogue, pages 397--416, Avignon, France. Association for Computational Linguistics. Anusha Kamath, Kanishk Singla, Rakesh Paul, Raviraj Bhuminand Joshi, Utkarsh Vaidya, San- jay Singh Chauhan, and Niranjan Wartikar. 2025.Benchmarking Hindi LLMs: A new suite 14 of datasets and a comparative analysis. InPro- ceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardiza- tion for Human-Centric AI in Indian Languages (BHASHA 2025), pages 52--68, Mumbai, India. Association for Computational Linguistics. Eleftherios Kapelonis, Efthymios Georgiou, and Alexandros Potamianos. 2022.A multi-task bert model for schema-guided dialogue state tracking. InProc. Interspeech 2022, pages 2733--2737. Yannis Katsis, Sara Rosenthal, Kshitij Fadnis, Chulaka Gunasekara, Young-Suk Lee, Lucian Popa, Vraj Shah, Huaiyu Zhu, Danish Contrac- tor, and Marina Danilevsky. 2025.mt RAG: A multi-turn conversational benchmark for eval- uating retrieval-augmented generation systems. Transactions of the Association for Computa- tional Linguistics, 13:784--808. Muhammad Kautsar, Rahmah Nurdini, Samuel Cahyawijaya, Genta Winata, and Ayu Purwari- anti. 2023.IndoToD: A multi-domain Indone- sian benchmark for end-to-end task-oriented di- alogue systems . InProceedings of the First Workshop in South East Asian Language Pro- cessing, pages 85--99, Nusa Dua, Bali, Indone- sia. Association for Computational Linguistics. Jang-Hyun Kim, Junyoung Yeom, Sangdoo Yun, and Hyun Oh Song. 2024a.Compressed con- text memory for online language model interac- tion. InThe Twelfth International Conference on Learning Representations. Jiho Kim, Woosog Chay, Hyeonji Hwang, Daeun Kyung, Hyunseung Chung, Eunbyeol Cho, Yohan Jo, and Edward Choi. 2024b. Dialsim: A real-time simulator for evaluating long-term multi-party dialogue understanding of conversa- tional agents. Aobo Kong, Wentao Ma, Shiwan Zhao, Yong- bin Li, Yuchuan Wu, Ke Wang, Xiaoqian Liu, Qicheng Li, Yong Qin, and Fei Huang. 2025. Sdpo: Segment-level direct preference opti- mization for social agents. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 12409--12423. Alkis Koudounas, Moreno La Quatra, and Elena Baralis. 2025.Deepdialogue: A multi- turn emotionally-rich spoken dialogue dataset. Preprint, arXiv:2505.19978. Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, JD Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Pre- cup, Feryal Behbahani, and Aleksandra Faust. 2025. Training language models to self-correct via reinforcement learning. InProceedings of the International Conference on Learning Rep- resentations (ICLR). Oral presentation. Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024.MT- eval: A multi-turn capabilities evaluation bench- mark for large language models . InProceed- ings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing, pages 20153--20177, Miami, Florida, USA. Associa- tion for Computational Linguistics. Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. 2026.LLMs get lost in multi-turn conversation . InThe Fourteenth In- ternational Conference on Learning Represen- tations. Barrett Martin Lattimer, Varun Prashant Gangal, Ryan McDonald, and Yi Yang. 2025.Sparse rewards can self-train dialogue agents . InFind- ings of the Association for Computational Lin- guistics: ACL 2025, pages 25395--25413, Vi- enna, Austria. Association for Computational Linguistics. Jing Yang Lee, Hamed Bonab, Nasser Zal- mout, Ming Zeng, Sanket Lokegaonkar, Colin Lockard, Binxuan Huang, Ritesh Sarkhel, and Haodong Wang. 2025a. DocTalk: Scalable graph-based dialogue synthesis for enhancing LLM conversational capabilities. InProceed- ings of the 26th Annual Meeting of the Spe- cial Interest Group on Discourse and Dialogue, pages 658--677, Avignon, France. Association for Computational Linguistics. Young-Jun Lee, Byungsoo Ko, Han-Gyu Kim, Jonghwan Hyeon, and Ho-Jin Choi. 2024. Di- alogcc: An automated pipeline for creating 15 high-quality multi-modal dialogue dataset. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 1938--1963. Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang, Yechan Hwang, Byungsoo Ko, Han-Gyu Kim, Dongyu Yao, Xuankun Rong, Eojin Joo, Seung- Ho Han, and 1 others. 2025b. Multiverse: A multi-turn conversation benchmark for evaluat- ing large vision and language models. InPro- ceedings of the IEEE/CVF International Confer- ence on Computer Vision, pages 708--719. Yiming Lei, Zhizheng Yang, Zeming Liu, Haitao Leng, Shaoguo Liu, Tingting Gao, Qingjie Liu, and Yunhong Wang. 2025. Contextqformer: A new context modeling method for multi- turn multi-modal conversations.arXiv preprint arXiv:2505.23121. Sicong Leng, Yun Xing, Zesen Cheng, Yang Zhou, Hang Zhang, Xin Li, Deli Zhao, Shi- jian Lu, Chunyan Miao, and Lidong Bing. 2026. The curse of multi-modalities: Evaluating hal- lucinations of large multimodal models across language, visual, and audio . InThe Thirty- ninth Annual Conference on Neural Informa- tion Processing Systems Datasets and Bench- marks Track. Patrick Lewis, Barlas Oguz, Ruty Rinott, Se- bastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. InProceedings of the 58th Annual Meeting of the Association for Compu- tational Linguistics, pages 7315--7330, Online. Association for Computational Linguistics. Haoyang Li, Zhanchao Xu, Yiming Li, Xuejia Chen, Darian Li, Anxin Tian, Qingfa Xiao, Cheng Deng, Jun Wang, Qing Li, and 1 others. 2025a. Loopserve: An adaptive dual-phase llm inference acceleration system for multi-turn di- alogues.arXiv preprint arXiv:2507.13681. Jinnan Li, Jinzhe Li, Yue Wang, Yi Chang, and Yuan Wu. 2025b. Structflowbench: A struc- tured flow benchmark for multi-turn instruction following. InFindings of the Association for Computational Linguistics: ACL 2025, pages 9322--9341. Kunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang, Chengfei Lv, Shengyu Zhang, Fan Wu, and Fei Wu. 2025c.MadaKV: Adap- tive modality-perception KV cache eviction for efficient multimodal long-context inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 13306--13318, Vi- enna, Austria. Association for Computational Linguistics. Yadong Li, Haoze Sun, Mingan Lin, Tianpeng Li, Guosheng Dong, Tao Zhang, Bowen Ding, Wei Song, Zhenglin Cheng, Yuqi Huo, Song Chen, Xu Li, Da Pan, Shusen Zhang, Xin Wu, Zheng Liang, Jun Liu, Tao Zhang, Keer Lu, and 8 others. 2024a.Baichuan-omni technical report. Preprint, arXiv:2410.08565. Yinghao Aaron Li, Xilin Jiang, Jordan Daref- sky, Ge Zhu, and Nima Mesgarani. 2024b. Style-talker: Finetuning audio language model and style-based text-to-speech model for fast spoken dialogue generation.Preprint, arXiv:2408.11849. Youquan Li, Miao Zheng, Fan Yang, Guosheng Dong, Bin Cui, Weipeng Chen, Zenan Zhou, and Wentao Zhang. 2025d. Fb-bench: A fine- grained multi-task benchmark for evaluating llms responsiveness to human feedback. In Proceedings of the 2025 Conference on Empir- ical Methods in Natural Language Processing, pages 9282--9302. Yubo Li, Yidi Miao, Xueying Ding, Ramayya Kr- ishnan, and Rema Padman. 2025e. Firm or fickle? evaluating large language models consis- tency in sequential interactions. InFindings of the Association for Computational Linguistics: ACL 2025, pages 6679--6700. Yubo Li, Xiaobin Shen, Xinyu Yao, Xueying Ding, Yidi Miao, Ramayya Krishnan, and Rema Pad- man. 2025f. Beyond single-turn: A survey on multi-turn interactions with large language mod- els.arXiv preprint arXiv:2504.04717. Zekun Li, Zhiyu Chen, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Xin Luna Dong, Adithya Sagar, Xifeng Yan, and Paul A Crook. 2024c. Large language models as zero- shot dialogue state tracker through function calling. InProceedings of the 62nd Annual 16 Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8688--8704. Borui Liao, Yulong Xu, Jiao Ou, Kaiyuan Yang, Weihua Jian, Pengfei Wan, and Di Zhang. 2025. Flexduo: A pluggable system for enabling full- duplex capabilities in speech dialogue systems. arXiv preprint arXiv:2502.13472. Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. 2025. Wildbench: Benchmarking llms with challenging tasks from real users in the wild. InInternational Conference on Learning Repre- sentations, volume 2025, pages 47852--47870. Chin-Yew Lin. 2004. Rouge: A package for auto- matic evaluation of summaries. InText summa- rization branches out, pages 74--81. Guan-Ting Lin, Chen Chen, Zhehuai Chen, and Hung-yi Lee. 2026. Full-duplex-bench-v3: Benchmarking tool use for full-duplex voice agents under real-world disfluency.arXiv preprint arXiv:2604.04847. Weizhe Lin, Bo-Hsiang Tseng, and Bill Byrne. 2021a.Knowledge-aware graph-enhanced GPT-2 for dialogue state tracking . InProceed- ings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 7871--7881, Online and Punta Cana, Domini- can Republic. Association for Computational Linguistics. Weizhe Lin, Bo-Hsiang Tseng, and Bill Byrne. 2021b. Knowledge-aware graph-enhanced gpt-2 for dialogue state tracking. InPro- ceedings of the 2021 conference on empirical methods in natural language processing, pages 7871--7881. Jiazheng Liu, Sipeng Zheng, BĂśrje F. Karlsson, and Zongqing Lu. 2025a.Taking notes brings focus? towards multi-turn multimodal dia- logue learning . InProceedings of the 2025 Conference on Empirical Methods in Natu- ral Language Processing, pages 33303--33324, Suzhou, China. Association for Computational Linguistics. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024a.Lost in the middle: How language models use long contexts.Trans- actions of the Association for Computational Linguistics, 12:157--173. Shuo Liu, Kaining Ying, Hao Zhang, Yue Yang, Yuqi Lin, Tianle Zhang, Chuanhao Li, Yu Qiao, Ping Luo, Wenqi Shao, and Kaipeng Zhang. 2024b.Convbench: A multi-turn conversa- tion evaluation benchmark with hierarchical ab- lation capability for large vision-language mod- els. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Ye Liu, Kevin Qinghong Lin, Chang Wen Chen, and Mike Zheng Shou. 2026.Videomind: A chain-of-loRA agent for temporal-grounded video reasoning . InThe Fourteenth Interna- tional Conference on Learning Representations. Zihan Liu, Wei Ping, Rajarshi Roy, Peng Xu, Chankyu Lee, Mohammad Shoeybi, and Bryan Catanzaro. 2024c. ChatQA: Surpassing GPT-4 on conversational QA and RAG. InThe Thirty- eighth Annual Conference on Neural Informa- tion Processing Systems. Ziyu Liu, Tao Chu, Yuhang Zang, Xilin Wei, Xi- aoyi Dong, Pan Zhang, Zijian Liang, Yuanjun Xiong, Yu Qiao, Dahua Lin, and Jiaqi Wang. 2024d. MMDU: A multi-turn multi-image di- alog understanding benchmark and instruction- tuning dataset for LVLMs. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Zuyan Liu, Yuhao Dong, Jiahui Wang, Ziwei Liu, Winston Hu, Jiwen Lu, and Yongming Rao. 2025b.Ola: Pushing the frontiers of omni-modal language model.Preprint, arXiv:2502.04328. Jiarui Lu, Thomas Holleis, Yizhe Zhang, Bern- hard Aumayer, Feng Nan, Haoping Bai, Shuang Ma, Shen Ma, Mengyu Li, Guoli Yin, Zirui Wang, and Ruoming Pang. 2025. ToolSand- box: A stateful, conversational, interactive eval- uation benchmark for LLM tool use capabili- ties . InFindings of the Association for Com- putational Linguistics: NAACL 2025, pages 1160--1183, Albuquerque, New Mexico. Asso- ciation for Computational Linguistics. 17 Jingyu Lu, Yuhan Wang, Jianming Luo, Yifu Chen, Tianle Liang, Shengpeng Ji, Ziyue Jiang, Xiaoda Yang, Yu Zhang, Xize Cheng, and 1 oth- ers. 2026. A survey of full-duplex spoken dia- logue systems: Architectural hierarchy, interac- tion ontology, and decision state machine.arXiv preprint arXiv:2606.19453. Xing Han Lu, Zden Ë ek Kasner, and Siva Reddy. 2024.WebLINX: Real-world website naviga- tion with multi-turn dialogue. InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Ma- chine Learning Research, pages 33007--33056. PMLR. Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. 2024.Agent- board: An analytical evaluation board of multi- turn LLM agents . InThe Thirty-eight Confer- ence on Neural Information Processing Systems Datasets and Benchmarks Track. Chengqian Ma, Wei Tao, and Steven Y. Guo. 2025. C3: A bilingual benchmark for spoken dialogue models exploring challenges in com- plex conversations . InProceedings of the 2025 Conference on Empirical Methods in Natu- ral Language Processing, pages 22778--22796, Suzhou, China. Association for Computational Linguistics. Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. 2024. Evaluating very long- term conversational memory of llm agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13851--13870. Rishabh Maheshwary, Vikas Yadav, Hoang H Nguyen, Khyati Mahajan, and Sathwik Tejaswi Madhusudhan. 2025. M2Lingual: Enhanc- ing multilingual, multi-turn instruction align- ment in large language models. InProceed- ings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 9676--9713, Albuquerque, New Mexico. Asso- ciation for Computational Linguistics. Potsawee Manakul, Guangzhi Sun, Warit Sirichot- edumrong, Kasima Tharnpipitchai, and Kunat Pipatanakul. 2025.Enhancing low-resource lan- guage and instruction following capabilities of audio language models. InInterspeech 2025. Ahmed Mahmoud Misbah, Mohamed Farouk, and Mustafa AbdulAzim. 2026. Fine-tuning arabic large language models for improved multi-turn dialogue: A blueprint for synthetic data generation and benchmarking.Plos one, 21(2):e0341905. Fengran Mo, Yifan Gao, Chuan Meng, Xin Liu, Zhuofeng Wu, Kelong Mao, Zhengyang Wang, Pei Chen, Zheng Li, Xian Li, Bing Yin, and Meng Jiang. 2025.UniConv: Unifying retrieval and response generation for large language mod- els in conversations . InProceedings of the 63rd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 6936--6949, Vienna, Austria. Association for Computational Linguistics. Fengran Mo, Chen Qu, Kelong Mao, Tianyu Zhu, Zhan Su, Kaiyu Huang, and Jian-Yun Nie. 2024. History-aware conversational dense retrieval. In Findings of the Association for Computational Linguistics: ACL 2024, pages 13366--13378, Bangkok, Thailand. Association for Computa- tional Linguistics. David Orlando Romero Mogrovejo, Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago GĂłn- gora, Aishik Mandal, Sukannya Purkayastha, Jesus-German Ortiz-Barajas, Emilio Villa Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Xin Yong, Zheng Wei Lim, Paula MĂłnica Silva, Jocelyn Dunstan, MĂŠlanie Jouitteau, David LE MEUR, Joan Nwatu, Gan- zorig Batnasan, and 57 others. 2024. CVQA: Culturally-diverse multilingual visual question answering benchmark . InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. V Moskvoretskii, A Frolov, and D Kuznetsov. 2024. Imad: Image-augmented multi-modal dialogue.Journal of Mathematical Sciences, 285(1):72--87. Basel Mousi, Fahim Dalvi, Shammur Chowdhury, Firoj Alam, and Nadir Durrani. 2026.Once 18 correct, still wrong: Counterfactual halluci- nation in multilingual vision-language models. Preprint, arXiv:2602.05437. Tarek Naous, Christian Hokayem, and Hazem Hajj. 2020. Empathy-driven arabic conversa- tional chatbot. InProceedings of the Fifth Arabic Natural Language Processing Workshop, pages 58--68. Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, BenoĂŽt Sagot, Ab- delrahman Mohamed, and Emmanuel Dupoux. 2023. Generative spoken dialogue language modeling.Transactions of the Association for Computational Linguistics, 11:250--266. Tu Anh Nguyen, Benjamin Muller, Bokai Yu, Marta R. Costa-jussa, Maha Elbayad, Sravya Popuri, Christophe Ropers, Paul-Ambroise Duquenne, Robin Algayres, Ruslan Mavlyu- tov, Itai Gat, Mary Williamson, Gabriel Syn- naeve, Juan Pino, BenoĂŽt Sagot, and Emmanuel Dupoux. 2025. SpiRit-LM: Interleaved spoken and written language model.Transactions of the Association for Computational Linguistics, 13:30--52. Jean De Dieu Nyandwi, Yueqi Song, Simran Khanuja, and Graham Neubig. 2025. Ground- ing multilingual multimodal llms with cultural knowledge. InProceedings of the 2025 Con- ference on Empirical Methods in Natural Lan- guage Processing, pages 24198--24242. Daniela Occhipinti, Serra Sinem Tekiro Ě glu, and Marco Guerini. 2024. Prodigy: a profile-based dialogue generation dataset. InFindings of the Association for Computational Linguistics: NAACL 2024, pages 3500--3514. Minsik Oh, Joosung Lee, Jiwei Li, and Guoyin Wang. 2023.PK-ICR: Persona-knowledge in- teractive multi-context retrieval for grounded dialogue . InProceedings of the 2023 Con- ference on Empirical Methods in Natural Lan- guage Processing, pages 16383--16395, Singa- pore. Association for Computational Linguis- tics. Olabiyi Oluwatobi and Erik Mueller. 2020.DL- GNet: A transformer-based model for dialogue response generation. InProceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, pages 54--62, Online. Asso- ciation for Computational Linguistics. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with hu- man feedback.Advances in neural information processing systems, 35:27730--27744. Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoff- mann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Bren- nan, and 1 others. 2021. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews.BMJ, 372. Yaning Pan, Zekun Wang, Qianqian Xie, Yongqian Wen, Yuanxing Zhang, Guohui Zhang, Haoxuan Hu, Zhiyu Pan, Yibing Huang, Zhidong Gan, Yonghong Lin, An Ping, Tianhao Peng, and Jiaheng Liu. 2025. Mt-video-bench: A holistic video understanding benchmark for evaluating multimodal llms in multi-turn dialogues .Preprint, arXiv:2510.17722. Youxin Pang, Jiajun Liu, Lingfeng Tan, Yong Zhang, Feng Gao, Xiang Deng, Zhuoliang Kang, Xiaoming Wei, and Yebin Liu. 2025. Mavid: A multimodal framework for audio- visual dialogue understanding and generation. arXiv preprint arXiv:2512.03034. Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002.Bleu: a method for automatic evaluation of machine transla- tion. InProceedings of the 40th Annual Meet- ing of the Association for Computational Lin- guistics, pages 311--318, Philadelphia, Pennsyl- vania, USA. Association for Computational Lin- guistics. Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kran- thi Gv, Xuzheng He, Haowen Hou, Prze- myslaw Kazienko, Jan Kocon, Jiaming Kong, BartĹomiej Koptyra, and 13 others. 2023. RWKV: Reinventing RNNs for the transformer 19 era. InFindings of the Association for Com- putational Linguistics: EMNLP 2023, pages 14048--14077, Singapore. Association for Com- putational Linguistics. Yizhou Peng, Yi-Wen Chao, Dianwen Ng, Yukun Ma, Chongjia Ni, Bin Ma, and Eng Siong Chng. 2025.FD-Bench: A Full-Duplex Benchmark- ing Pipeline Designed for Full Duplex Spoken Dialogue Systems. InInterspeech 2025, pages 176--180. Akshara Prabhakar, Zuxin Liu, Ming Zhu, Jian- guo Zhang, Tulika Manoj Awalgaonkar, Shiyu Wang, Zhiwei Liu, Haolin Chen, Thai Quoc Hoang, Juan Carlos Niebles, Shelby Heinecke, Weiran Yao, Huan Wang, Silvio Savarese, and Caiming Xiong. 2026.APIGen-MT: Agentic pipeline for multi-turn data generation via sim- ulated agent-human interplay. InThe Thirty- ninth Annual Conference on Neural Informa- tion Processing Systems Datasets and Bench- marks Track. Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, Hamid Palangi, Kai-Wei Chang, Yejin Choi, and Saadia Gabriel. 2025. X-teaming: Multi- turn jailbreaks and defenses with adaptive multi- agents.arXiv preprint arXiv:2504.13203. Siva Reddy, Danqi Chen, and Christopher D. Man- ning. 2019.CoQA: A conversational ques- tion answering challenge.Transactions of the Association for Computational Linguistics, 7:249--266. Stephen Roller and 1 others. 2021.Recipes for building an open-domain chatbot. InProceed- ings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, pages 300--325. AmĂŠlie Royer, Moritz BĂśhle, Gabriel de Marmiesse, Laurent MazarĂŠ, Neil Zeghi- dour, Alexandre DĂŠfossez, and Patrick PĂŠrez. 2025. Vision-speech models: Teaching speech models to converse about images.Preprint, arXiv:2503.15633. Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalan Bor- sos, Felix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, and 1 others. 2023. Audiopalm: A large lan- guage model that can speak and listen (2023). arXiv preprint arXiv:2306.12925. Mark Russinovich, Ahmed Salem, and Ronen El- dan. 2025. Great, now write an article about that: The crescendoMulti-TurnLLMjail- break attack. In34th USENIX Security Sympo- sium (USENIX Security 25), pages 2421--2440. Alexander Scarlatos, Naiming Liu, Jaewook Lee, Richard Baraniuk, and Andrew Lan. 2025. Training llm-based tutors to improve student learning outcomes in dialogues. InInterna- tional Conference on Artificial Intelligence in Education, pages 251--266. Springer. Iulian Serban, Alessandro Sordoni, Yoshua Ben- gio, Aaron Courville, and Joelle Pineau. 2016. Building end-to-end dialogue systems using generative hierarchical neural network models. InProceedings of the AAAI conference on arti- ficial intelligence, volume 30. Lior Shani, Aviv Rosenberg, Asaf Cassel, Oran Lang, Daniele Calandriello, Avital Zipori, Hila Noga, Orgad Keller, Bilal Piot, Idan Szpektor, Avinatan Hassidim, Yossi Matias, and Remi Munos. 2024. Multi-turn reinforcement learn- ing with preference human feedback . InThe Thirty-eighth Annual Conference on Neural In- formation Processing Systems. Ryan Shea and Zhou Yu. 2023. Building per- sona consistent dialogue agents with offline re- inforcement learning. InProceedings of the 2023 Conference on Empirical Methods in Nat- ural Language Processing, pages 1778--1795. Wentao Shi, Mengqi Yuan, Junkang Wu, Qifan Wang, and Fuli Feng. 2024. Direct multi-turn preference optimization for language agents. In Proceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing, pages 2312--2324. Jeonghoon Shim, Gyuhyeon Seo, Cheongsu Lim, and Yohan Jo. 2025. Tooldial: Multi-turn di- alogue generation method for tool-augmented language models. InThe Thirteenth Interna- tional Conference on Learning Representations. Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023. 20 Reflexion: language agents with verbal rein- forcement learning. InThirty-seventh Confer- ence on Neural Information Processing Sys- tems. Shuzheng Si, Wentao Ma, Haoyu Gao, Yuchuan Wu, Ting-En Lin, Yinpei Dai, Hangyu Li, Rui Yan, Fei Huang, and Yongbin Li. 2023. Spo- kenwoz: A large-scale speech-text benchmark for spoken task-oriented dialogue agents.Ad- vances in Neural Information Processing Sys- tems, 36:39088--39118. Yixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta, Deng Cai, Yi-An Lai, and Yi Zhang. 2022.Multi-task pre-training for plug-and-play task-oriented dialogue system. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4661--4676, Dublin, Ireland. As- sociation for Computational Linguistics. Yuchong Sun, Che Liu, Kun Zhou, Jinwen Huang, Ruihua Song, Xin Zhao, Fuzheng Zhang, Di Zhang, and Kun Gai. 2024.Parrot: Enhanc- ing multi-turn instruction following for large language models . InProceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 9729--9750, Bangkok, Thailand. Associ- ation for Computational Linguistics. Kim Sung-Bin, Oh Hyun-Bin, JungMok Lee, Arda Senocak, Joon Son Chung, and Tae-Hyun Oh. 2025.AVHBench: A cross-modal hallu- cination benchmark for audio-visual large lan- guage models . InThe Thirteenth International Conference on Learning Representations. Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neu- ral networks.Advances in neural information processing systems, 27. Changli Tang, Wenyi Yu, Guangzhi Sun, Xi- anzhao Chen, Tian Tan, Wei Li, Lu Lu, Ze- jun MA, and Chao Zhang. 2024.SALMONN: Towards generic hearing abilities for large lan- guage models . InThe Twelfth International Conference on Learning Representations. Yuqi Tang, Kehua Feng, Yunfeng Wang, Zhiwen Chen, Chengfei Lv, Gang Yu, Qiang Zhang, Keyan Ding, and Huajun Chen. 2025. Learn- ing an efficient multi-turn dialogue evaluator from multiple llm judges.arXiv preprint arXiv:2508.00454. Kimi Team. 2024.Kimi-audio technical report. Preprint, arXiv:2409.06523. Siyu Tian, Kaijie Mo, Yupei Wang, and Renfen Hu. 2025a.CMT-eval: A novel Chinese multi- turn dialogue evaluation dataset addressing real- world conversational challenges. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 18279--18303, Suzhou, China. Association for Computational Linguis- tics. Yunjie Tian, Tianren Ma, Lingxi Xie, and Qixiang Ye. 2025b. Chatterbox: Multimodal referring and grounding with chain-of-questions. InPro- ceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 7401--7409. Wenwen Tong, Hewei Guo, Dongchuan Ran, Jiangnan Chen, Jiefan Lu, Kaibin Wang, Ke- qiang Li, Xiaoxu Zhu, Jiakui Li, Kehan Li, Xueheng Li, Lumin Li, Chenxu Guo, Jiasheng Zhou, Jiandong Chen, Xianye Wu, Jiahao Wang, Silei Wu, Lei Chen, and 7 others. 2025. Interactiveomni: A unified omni-modal model for audio-visual multi-turn dialogue .Preprint, arXiv:2510.13747. Dennis Ulmer, Elman Mansimov, Kaixiang Lin, Lijia Sun, Xibin Gao, and Yi Zhang. 2024. Bootstrapping LLM-based task-oriented dia- logue agents via self-talk. InFindings of the As- sociation for Computational Linguistics: ACL 2024, pages 9500--9522, Bangkok, Thailand. Association for Computational Linguistics. Norawit Urailertprasert, Peerat Limkonchotiwat, Supasorn Suwajanakorn, and Sarana Nutanong. 2024.SEA-VQA: Southeast Asian cultural context dataset for visual question answer- ing . InProceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR), pages 173--185, Bangkok, Thailand. Association for Computational Linguistics. Deeksha Varshney, Akshara Prabhakar, and Asif Ekbal. 2022.Commonsense and named en- tity aware knowledge grounded dialogue gener- ation . InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan- guage Technologies, pages 1322--1335, Seattle, 21 United States. Association for Computational Linguistics. Chengyao Wang, Zhisheng Zhong, Bohao Peng, Senqiao Yang, Yuqi Liu, Haokun Gui, Bin Xia, Jingyao Li, Bei Yu, and Jiaya Jia. 2025a. Mgm-omni: Scaling omni llms to person- alized long-horizon speech.arXiv preprint arXiv:2509.25131. Haoyu Wang, Yuxin Chen, Liang Luo, Buyun Zhang, Ellie Dingqiao Wen, and Pan Li. 2026. Implicit turn-wise policy optimization for proactive user-llm interaction.arXiv preprint arXiv:2603.23550. Hongru Wang, Lingzhi Wang, Yiming Du, Liang Chen, Jingyan Zhou, Yufei Wang, and Kam-Fai Wong. 2023. A survey of the evolution of lan- guage model-based dialogue systems.arXiv e- prints, pages arXivâ2311. Xingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng, and Heng Ji. 2024. MINT: Evaluating LLMs in multi-turn interac- tion with tools and language feedback. InThe Twelfth International Conference on Learning Representations. Xiong Wang, Yangze Li, Chaoyou Fu, Yike Zhang, Yunhang Shen, Lei Xie, Ke Li, Xing Sun, and Long MA. 2025b.Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM . InForty-second International Conference on Machine Learning. Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu, Dongyan Zhao, and Zilong Zheng. 2025c. Omnimmi: A comprehensive multi-modal inter- action benchmark in streaming video contexts . In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18925--18935. Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Wang, Ning Shi, Siyu Li, Yizhi Li, Hao- ran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, and Wenhao Huang. 2025d. MIO: A foundation model on multi- modal tokens. InProceedings of the 2025 Con- ference on Empirical Methods in Natural Lan- guage Processing, pages 5077--5099, Suzhou, China. Association for Computational Linguis- tics. Zheng Weihua, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, and 16 others. 2025. MMA-ASIA: A multilingual and multimodal alignment framework for culturally grounded evaluation. arXiv preprint arXiv:2510.08608. Bingbing Wen, Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Bill Howe, and Lijuan Wang. 2023.Infovisdial: An informative vi- sual dialogue dataset by bridging large multi- modal and language models.arXiv preprint arXiv:2312.13503. Boyong Wu, Chao Yan, Chen Hu, Cheng Yi, Chengli Feng, Fei Tian, Feiyu Shen, Gang Yu, Haoyang Zhang, Jingbei Li, Mingrui Chen, Peng Liu, Wang You, Xiangyu Tony Zhang, Xingyuan Li, Xuerui Yang, Yayue Deng, Yechang Huang, Yuxin Li, and 90 others. 2025a. Step-audio 2 technical report.Preprint, arXiv:2507.16632. Chien-Sheng Wu, Steven CH Hoi, Richard Socher, and Caiming Xiong. 2020. Tod-bert: Pre- trained natural language understanding for task- oriented dialogue. InProceedings of the 2020 conference on empirical methods in natural lan- guage processing (EMNLP), pages 917--929. Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. 2025b. Longmemeval: Benchmarking chat assistants on long-term interactive memory. InThe Thir- teenth International Conference on Learning Representations. Jiangxu Wu, Cong Wang, TianHuang Su, Lin Haozhi, JunYang JunYang, Zhangchao Zhangchao, Binqiang Pan, SongpanYang SongpanYang, Mingpeng Mingpeng, Kai Shi, and 1 others. 2025c. Instruct: A review-driven multi-turn conversations generation method for large language models.Findings of the Association for Computational Linguistics: ACL 2025, pages 16578--16595. Qingyang Wu and Zhou Yu. 2024.Stateful memory-augmented transformers for efficient 22 dialogue modeling. InFindings of the Asso- ciation for Computational Linguistics: EACL 2024, pages 853--867, St. Julianâs, Malta. As- sociation for Computational Linguistics. Qinzhuo Wu, Wei Liu, Jian Luan, and Bin Wang. 2024.ToolPlanner: A tool augmented LLM for multi granularity instructions with path plan- ning and feedback. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 18315--18339, Mi- ami, Florida, USA. Association for Computa- tional Linguistics. Yuhang Wu, Wenmeng Yu, Yean Cheng, Yan Wang, Xiaohan Zhang, Jiazheng Xu, Ming Ding, and Yuxiao Dong. 2025d.AlignMM- Bench: Evaluating Chinese multimodal align- ment in large vision-language models. InPro- ceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 6541--6558, Vi- enna, Austria. Association for Computational Linguistics. Yuan Xie, Tianshui Chen, Zheng Ge, and Lionel Ni. 2025.Video-mtr: Reinforced multi-turn rea- soning for long video understanding .Preprint, arXiv:2508.20478. Zhifei Xie and Changqiao Wu. 2024.Mini- omni2: Towards open-source gpt-4o with vi- sion, speech and duplex capabilities .Preprint, arXiv:2410.11190. Xiezhifei. 2024.Mini-Omni: Language models can hear, talk while thinking in streaming . In Submitted to Tsinghua University Course: Ad- vanced Machine Learning. Under review. Wei Xiong, Chengshuai Shi, Jiaming Shen, Aviv Rosenberg, Zhen Qin, Daniele Calandriello, Misha Khalman, Rishabh Joshi, Bilal Piot, Mo- hammad Saleh, Chi Jin, Tong Zhang, and Tianqi Liu. 2025. Building math agents with multi-turn iterative preference learning. InInternational Conference on Learning Representations, vol- ume 2025, pages 25901--25932. Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, and Shafiq Joty. 2025a. Does context matter? contextualjudgebench for evaluating llm-based judges in contextual settings. InPro- ceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 9541--9564. Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. 2025b. Qwen2.5-omni technical report.Preprint, arXiv:2503.20215. Peng Xu, Wei Ping, Xianchao Wu, Chejian Xu, Zi- han Liu, Mohammad Shoeybi, and Bryan Catan- zaro. 2025c.ChatQA 2: Bridging the gap to proprietary LLMs in long context and RAG ca- pabilities . InThe Thirteenth International Con- ference on Learning Representations. Haochen Xue, Feilong Tang, Ming Hu, Yexin Liu, Qidong Huang, Yulong Li, Chengzhi Liu, Zhongxing Xu, Chong Zhang, Chun-Mei Feng, and 1 others. 2025. Mmrc: A large-scale bench- mark for understanding multimodal large lan- guage model in real-world conversation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 22477--22503. Dawei Yan, Yang Li, Qing-Guo Chen, Weihua Luo, Peng Wang, Haokui Zhang, and Chunhua Shen. 2025a. Mmcr: Advancing visual lan- guage model in multimodal multi-turn contex- tual reasoning .Preprint, arXiv:2503.18533. Ruiqi Yan, Xiquan Li, Wenxi Chen, Zhikang Niu, Chen Yang, Ziyang Ma, Kai Yu, and Xie Chen. 2025b.URO-bench: Towards comprehensive evaluation for end-to-end spoken dialogue mod- els . InFindings of the Association for Com- putational Linguistics: EMNLP 2025, pages 17211--17242, Suzhou, China. Association for Computational Linguistics. Songhua Yang, Hanjie Zhao, Senbin Zhu, Guangyu Zhou, Hongfei Xu, Yuxiang Jia, and Hongying Zan. 2024. Zhongjing: Enhancing the chinese medical capabilities of large lan- guage model through expert feedback and real- world multi-turn dialogue. InProceedings of the AAAI conference on artificial intelligence, vol- ume 38, pages 19368--19376. Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue, Shengsheng Qian, Jiahong Wu, Fan Yang, Weiming Dong, and Changsheng Xu. 2025. 23 SVBench: A benchmark with temporal multi- turn dialogues for streaming video understand- ing. InThe Thirteenth International Conference on Learning Representations. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik R. Narasimhan. 2025.Ď-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. InProceedings of the Thirteenth International Conference on Learn- ing Representations. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models.arXiv preprint arXiv:2210.03629. Qichen Ye, Junling Liu, Dading Chong, Peilin Zhou, Yining Hua, Fenglin Liu, Meng Cao, Ziming Wang, Xuxin Cheng, Zhu Lei, and 1 others. 2024a. Qilin-med: Multi-stage knowl- edge injection advanced medical large language model.arXiv preprint arXiv:2310.09089. Qilang Ye, Zitong Yu, Rui Shao, Xinyu Xie, Philip Torr, and Xiaochun Cao. 2024b. Cat: En- hancing multimodal large language model to an- swer questions in dynamic audio-visual scenar- ios. InEuropean Conference on Computer Vi- sion, pages 146--164. Springer. Zihao Yi, Jiarui Ouyang, Zhe Xu, Yuwen Liu, Tianhao Liao, Haohao Luo, and Ying Shen. 2025. A survey on recent advances in llm-based multi-turn dialogue systems.ACM Computing Surveys, 58(6):1--38. Sisi You, Bowen Yuan, and Bing-Kun Bao. 2025. Scvbench: A benchmark with multi-turn dia- logues for story-centric video understanding. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJ- CAI-25, pages 2287--2295. International Joint Conferences on Artificial Intelligence Organiza- tion. Main Track. Wenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Guangzhi Sun, Lu Lu, Yuxuan Wang, and Chao Zhang. 2026. Salmonn-omni: A standalone speech llm with- out codec injection for full-duplex conversation. Advances in Neural Information Processing Sys- tems, 38:24617--24643. Kamyar Zeinalipour, Mohamed Zaky Saad, Oumaima Attafi, Marco Maggini, and Marco Gori. 2025.Shawarma chats: A benchmark ex- act dialogue & evaluation platter in Egyptian, Maghrebi & modern standard ArabicâA triple- dialect feast for hungry language models. In Proceedings of The Third Arabic Natural Lan- guage Processing Conference, pages 472--524, Suzhou, China. Association for Computational Linguistics. Aohan Zeng, Zhengxiao Du, Mingdao Liu, Ke- dong Wang, Shengmin Jiang, Lei Zhao, Yux- iao Dong, and Jie Tang. 2024. GLM-4- Voice: Towards intelligent and human-like end-to-end spoken chatbot.arXiv preprint arXiv:2412.02612. Bo Zhang, Hui Ma, Dailin Li, Jian Ding, Jian Wang, Bo Xu, and HongFei Lin. 2025a. Efficient tuning of large language models for knowledge-grounded dialogue generation. Transactions of the Association for Computa- tional Linguistics, 13:1007--1031. Chen Zhang, Xinyi Dai, Yaxiong Wu, Qu Yang, Yasheng Wang, Ruiming Tang, and Yong Liu. 2025b. A survey on multi-turn interaction capa- bilities of large language models.arXiv preprint arXiv:2501.09959. Chen Zhang, Luis Fernando DâHaro, Yiming Chen, Malu Zhang, and Haizhou Li. 2024a. A comprehensive analysis of the effectiveness of large language models as automatic dialogue evaluators. InProceedings of the AAAI Con- ference on Artificial Intelligence, volume 38, pages 19515--19524. Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023.SpeechGPT: Empowering large lan- guage models with intrinsic cross-modal con- versational abilities . InFindings of the Associ- ation for Computational Linguistics: EMNLP 2023, pages 15757--15773, Singapore. Associa- tion for Computational Linguistics. Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenx- ing Li, Dan Su, Chenhui Chu, and Dong Yu. 2024b. Mm-llms: Recent advances in mul- timodal large language models.Findings of the Association for Computational Linguistics: ACL 2024, pages 12401--12430. 24 Michael JQ Zhang and Eunsol Choi. 2025. Clarify when necessary: Resolving ambiguity through interaction with lms. InFindings of the Asso- ciation for Computational Linguistics: NAACL 2025, pages 5526--5543. Pan Zhang, Xiaoyi Dong, Yuhang Cao, Yuhang Zang, Rui Qian, Xilin Wei, Lin Chen, Yifei Li, Junbo Niu, Shuangrui Ding, Qipeng Guo, Haodong Duan, Xin Chen, Han Lv, Zheng Nie, Min Zhang, Bin Wang, Wen- wei Zhang, Xinyue Zhang, and 10 others. 2024c.Internlm-xcomposer2.5-omnilive: A comprehensive multimodal system for long- term streaming video and audio interactions. Preprint, arXiv:2412.09596. Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018.Personalizing dialogue agents: I have a dog, do you have pets too? InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 2204--2213, Melbourne, Australia. Association for Computational Linguistics. Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou, and Yang Feng. 2025c. Stream-omni: Si- multaneous multimodal interactions with large language-vision-speech model.arXiv preprint arXiv:2506.13642. Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020a.Bertscore: Evaluating text generation with bert . InInterna- tional Conference on Learning Representations. Xiaoyu Zhang, Ruobing Xie, Yougang Lyu, Xin Xin, Pengjie Ren, Mingfei Liang, Bo Zhang, Zhanhui Kang, Maarten de Ri- jke, and Zhaochun Ren. 2024d. Towards empathetic conversational recommender systems. InProceedings of the 18th ACM Conference on Recommender Systems, pages 84--93. Yiran Zhang, Mo Wang, Xiaoyang Li, Kaix- uan Ren, Chencheng Zhu, and Usman Naseem. 2025d.TurnBench-MS: A benchmark for eval- uating multi-turn, multi-step reasoning in large language models . InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 19892--19924, Suzhou, China. Associa- tion for Computational Linguistics. Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020b.DI- ALOGPT : Large-scale generative pre-training for conversational response generation. InPro- ceedings of the 58th Annual Meeting of the As- sociation for Computational Linguistics: Sys- tem Demonstrations, pages 270--278, Online. Association for Computational Linguistics. Zhihan Zhang, Shiyang Li, Zixuan Zhang, Xin Liu, Haoming Jiang, Xianfeng Tang, Yifan Gao, Zheng Li, Haodong Wang, Zhaoxuan Tan, and 1 others. 2025e. Iheval: Evaluating language models on following the instruction hierarchy. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Asso- ciation for Computational Linguistics: Human Language Technologies (Volume 1: Long Pa- pers), pages 8374--8398. Jiahao Zhao, Yunjia Li, Wei Li, and Kazuyoshi Yoshii. 2026a.Abc-eval: Benchmarking large language models on symbolic music under- standing and instruction following . InICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 16072--16076. Lulu Zhao, Weihao Zeng, Xiaofeng Shi, Hua Zhou, Donglin Hao, and Yonghua Lin. 2024a. Aqulia-med llm: pioneering full-process open- source medical language models.arXiv preprint arXiv:2406.12182. Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024b. Wildchat: 1m chatGPT interaction logs in the wild . InThe Twelfth International Conference on Learning Representations. Zicheng Zhao, Kangyu Wang, Shijie Li, Rui Qian, Weiyao Lin, and Huabin Liu. 2026b. Cogstream: Context-guided streaming video question answering. InProceedings of the AAAI Conference on Artificial Intelligence, vol- ume 40, pages 13332--13341. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yong- hao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. 2024. LMSYS-chat-1m: A large-scale real- world LLM conversation dataset. InThe Twelfth 25 International Conference on Learning Repre- sentations. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging llm-as-a- judge with mt-bench and chatbot arena. Ad- vances in Neural Information Processing Sys- tems, 36:46595--46623. Xuhui Zhou, Hao Zhu, Leena Mathur, Ruo- hong Zhang, Haofei Yu, Zhengyang Qi, Louis- Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. 2024a.SO- TOPIA: Interactive evaluation for social intelli- gence in language agents. InThe Twelfth In- ternational Conference on Learning Represen- tations. Yifei Zhou, Song Jiang, Yuandong Tian, Jason Weston, Sergey Levine, Sainbayar Sukhbaatar, and Xian Li. 2025. Sweet-rl: Training multi- turn llm agents on collaborative reasoning tasks. Preprint, arXiv:2503.15478. Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, and Aviral Kumar. 2024b.ArCHer: Training language model agents via hierarchi- cal multi-turn RL . InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learn- ing Research, pages 62178--62209. PMLR. Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2025.JudgeLM: Fine-tuned large lan- guage models are scalable judges. InThe Thir- teenth International Conference on Learning Representations. A Appendix A.1 Additional Resources A.1.1 Text-only Multi-turn Resources Foundational benchmarks. CoQA ( Reddy et al.,2019) contains conversational question answering grounded in source passages and emphasizes dependencies across previous turns. MultiWOZ 2.1 (Eric et al.,2020) provides multi- domain task-oriented dialogues with corrected state annotations and is widely used for dialogue state tracking, inform rate, and task success. Per- sonaChat ( Zhang et al.,2018) pairs open-domain conversations with explicit persona profiles, mak- ing it a standard resource for persona-consistency research. Capability evaluation.MT-Bench-101 (Bai et al.,2024) introduces a hierarchical taxonomy for fine-grained multi-turn evaluation. MT- Eval (Kwan et al.,2024) evaluates recollection, expansion, refinement, and follow-up, showing that models can degrade when a task is distributed across turns. MultiChallenge (Deshpande et al., 2025) targets realistic human-LLM conversations and evaluates instruction retention, inference memory, reliable versioned editing, and self- coherence. MINT (Wang et al.,2024) restructures reasoning, code, and decision-making tasks into a user-tool-LLM interaction format. Long-term memory.LoCoMo (Maharana et al., 2024) evaluates long-term conversational mem- ory through extended multi-session dialogues grounded in personas and temporal event graphs. LongMemEval (Wu et al.,2025b) tests informa- tion extraction, multi-session reasoning, knowl- edge updates, temporal reasoning, and absten- tion. DialSim ( Kim et al.,2024b) evaluates long- term multi-party dialogue using TV-series tran- scripts and response-time constraints. Person- aMem ( Jiang et al.,2025) evaluates dynamic user profiling and preference tracking across multi- session histories. Mem-Gallery ( Bei et al.,2026) extends long-memory evaluation to multimodal settings. Training corpora.UltraChat (Ding et al., 2023) provides synthetic multi-turn instruction- following conversations. WildChat ( Zhao et al., 2024b) contains opt-in real ChatGPT conver- sations with broad topic and language diver- sity. LMSYS-Chat-1M (Zheng et al.,2024) col- lects real user conversations with multiple LLMs through public chat interfaces. DocTalk (Lee et al. ,2025a) synthesizes multi-turn information- seeking dialogues from Wikipedia and is mainly designed as a pre-training resource. Persona, role-play, and social interaction. PIPPA (Gosling et al.,2023) contains long role- play conversations and supports evaluation of in-character consistency. PRODIGy ( Occhipinti et al.,2024) grounds personas in fictional char- acter profiles and studies the effect of backstory and speaking style. SOTOPIA (Zhou et al.,2024a) 26 evaluates social intelligence through role-play sce- narios where agents pursue social goals while maintaining appropriate behavior. Tool use, retrieval, and proactive assistance. Ď-Bench (Yao et al.,2025) evaluates agents in stateful user-tool interaction settings with task and policy constraints. ToolSandbox (Lu et al., 2025) evaluates stateful tool use and interaction consistency across turns. mtRAG (Katsis et al., 2025) evaluates multi-turn retrieval-augmented generation, where the agent must decide when and what to retrieve as the conversation evolves. ProMISe (Butala et al.,2024) evaluates proactive information seeking by asking agents to generate useful follow-up suggestions during a dialogue. A.1.2 Spoken Dialogue Resources Foundational resources.TurnGPT (Ekstedt and Skantze,2020) applies a Transformer lan- guage model to turn-taking prediction in spoken dialogue. It frames turn completion as a language- modeling problem over transcribed conversational text, providing an early trainable baseline for predicting when a speakerâs turn is likely to end. SpokenWOZ ( Si et al.,2023) provides 249 hours of human-human task-oriented speech across eight domains and 203K turns. It preserves audio, transcripts, and dialogue-state annotations, allowing evaluation of how recognition errors affect downstream state tracking and task success. Multi-turn speech-to-speech evaluation.Au- dio MultiChallenge (Gosai et al.,2025) extends multi-turn evaluation to spoken dialogue systems using 452 conversations and 1,712 instance-level rubrics. It evaluates the four axes of MultiChal- lenge together with a Voice Editing axis for mid- utterance speech repairs, and compares turn-level and conversation-level scores. URO-Bench ( Yan et al.,2025b) evaluates end-to-end speech-to- speech dialogue models across multilingual, multi- round, and paralinguistic settings, with basic and pro tracks covering 20 S2S task types. MTalk- Bench ( Du et al.,2025b) constructs multi-turn speech-to-speech test dialogues through an LLM- human pipeline, then adds human recording, voice conversion, and ambient sound mixing. It evalu- ates semantic content, vocal cues, and background sound using both arena-style pairwise compari- son and rubric-based absolute scoring. MULTI- Bench ( Deng et al.,2025) focuses on emotional intelligence in spoken dialogue and tests whether models can track and respond to emotional-state changes across turns. FD-Bench (Peng et al., 2025) evaluates simultaneous bidirectional speech through simulated five-round conversations, mea- suring interruption handling and response timing with metrics such as SIR, SRIR, EIR, NIR, IRD, FSED, ERT, and EIT, as referred in Table9. Related audio-agent benchmarks.ADU- Bench (Gao et al.,2025a) provides open-ended audio dialogue tasks covering multiple scenarios, skills, languages, and ambiguity types such as intonation and homophones. Because it is not explicitly multi-turn in the current categorization, it is best treated as a related audio-dialogue benchmark rather than a core multi-turn resource. Full-Duplex-Bench-v3 ( Lin et al.,2026) extends full-duplex benchmarking to tool-use scenarios, measuring whether voice agents invoke tools correctly while handling disfluencies such as false starts and repairs. Since it is marked as work in progress, it can be mentioned as an emerging direction rather than a central benchmark. Training corpora and specialized spoken re- sources.DeepDialogue (Koudounas et al.,2025) provides synthetic multi-turn dialogues paired with emotionally consistent speech, making it useful for training spoken dialogue models with emotion-control signals. ASK-QA ( Chen et al., 2025d) contains spoken conversational QA ses- sions with multiple speakers and deliberately am- biguous user requests, supporting research on clar- ification and spoken question answering. The MENASpeechBank ( Ali et al.,2026) provides persona-conditioned conversations in Arabic and regional dialects across the MENA region, com- bining voice-bank design with multi-turn spoken Arabic dialogue. A.1.3 Multimodal Dialogue Resources Foundational and open-domain resources. VisDial (Das et al.,2017) pairs COCO images with ten rounds of question answering and established the standard format for multi-turn visual dialogue. MMDialog ( Feng et al.,2023a) provides large-scale real-world multimodal con- versations with images across thousands of topics and introduces M-Relevance for modality- aware response evaluation. DialogCC ( Lee et al., 2024) and IMAD (Moskvoretskii et al.,2024) construct image-augmented dialogues by match- ing images to existing text conversations using 27 vision-language similarity, making them scal- able resources for multimodal dialogue training. InfoVisDial (Wen et al.,2023) adds Wikipedia- grounded knowledge to visual dialogue, requiring models to combine visual evidence with external factual information across turns. Instruction-tuning corpora.MMDU-45K (Liu et al.,2024d) contains multi-turn, multi-image di- alogues generated with GPT-4o and human refine- ment, with sessions that can include many images and long image-text contexts. MMCR (Yan et al., 2025a) provides both an instruction-tuning corpus and a diagnostic benchmark covering multiple do- mains and subtopics, supporting evaluation under single-image and multi-image conversational set- tings. Capability evaluation.ConvBench (Liu et al., 2024b) evaluates multi-turn LVLM dialogues across perception, reasoning, and creation tasks, showing that early visual perception errors can propagate into later turns. MultiVerse (Lee et al. ,2025b) evaluates multi-turn multimodal dialogue across many task types and explic- itly tests in-context learning during conversation. MMRC (Xue et al.,2025) evaluates information extraction, multi-turn reasoning, information up- date, image management, memory recall, and re- fusal behavior, identifying recurring failures such as memory degradation, update failure, and er- ror propagation. MMMT-IF ( Epstein et al.,2024) focuses on multi-turn instruction following with image inputs, making it a multimodal counter- part to text-only instruction-retention benchmarks. Chatterbox/CB-300K ( Tian et al.,2025b) targets multi-round referring and grounding, with empha- sis on referential ambiguity, spatial relations, and grounding consistency across turns. Long-term memory and context management. Mem-Gallery (Bei et al.,2026) evaluates multi- session multimodal memory across visual and textual histories, including extraction, reasoning, and knowledge management. DiagNote/M- Diag ( Liu et al.,2025a) provides multi-turn dia- logues with explicit note-taking annotations and studies whether learned notes can help VLMs maintain grounded context over long interactions. Emerging tri-modal, generative, cultural, and video resources.InteractiveOmni (Tong et al., 2025) extends multi-turn dialogue to tri-modal audio-visual-language interaction, using cross- modal context and historical memory. Di- alogGen/DialogBen (Huang et al.,2025b) evalu- ates dialogue-conditioned image generation and editing, where the model must preserve both conversational context and output-modality coher- ence. AlignMMBench (Wu et al.,2025d) evalu- ates cultural alignment and factual consistency in single- and multi-turn Chinese and bilingual multi- modal dialogues. SVBench (Yang et al.,2025) in- troduces temporal multi-turn dialogue over stream- ing video, where later questions depend on both earlier answers and specific video segments. A.1.4 Cultural and Linguistic Resources Arabic and dialectal multi-turn resources. Shawarma Chats ( Zeinalipour et al.,2025) pro- vides 30K six-turn Arabic conversations grounded in Wikipedia content, covering MSA, Egyptian Arabic, and Maghrebi Arabic. It is one of the few large-scale multi-turn resources with explicit coverage of Maghrebi Arabic, making it useful for dialect-aware dialogue modeling and evalua- tion. Naous et al.(2020) introduced an empathy- driven Arabic conversational chatbot and an ac- companying message-response corpus spanning Levantine, Egyptian, and Gulf Arabic, providing an early resource for Arabic dialectal conversa- tion modeling. Alexandria ( El Mekki et al.,2026) is a multi-turn conversational machine translation dataset for English-Dialectal Arabic, covering 13 Arab countries and 11 domains with city-level di- alect metadata and speaker-addressee gender infor- mation. It supports evaluation of dialect natural- ness, style and formality, semantic accuracy, and gender-sensitive translation. Cross-lingual and culturally grounded multi- turn benchmarks.C3 ( Ma et al.,2025) is a bilingual English-Chinese spoken dialogue bench- mark targeting phonological ambiguity, corefer- ence, omission, and multi-turn dependency.Ka- math et al.(2025) introduce a Hindi evaluation suite consisting of MT-Bench-Hi for multi-turn instruction following, ChatRAG-Hi for conversa- tional retrieval-augmented generation, and BFCL- Hi for function calling. cuDialog ( Cao et al., 2024) conditions response prediction on Hofst- ede cultural-value dimensions, making cultural be- havior an explicit variable in multi-turn dialogue evaluation. IndoToD ( Kautsar et al.,2023) pro- vides an Indonesian multi-domain task-oriented 28 ResourceYear LanguagesRegion / CulturesMT M Dialect Venue Arabic / MENA OASIS / EVERYDAYMMQA (Alam et al.,2025) 2025 EN, MSA, Egyptian, Levantine 18 Arab countries7 â 33arXiv M 2 CQA (Mousi et al.,2026)2025 MSA + dialects17 MENA countries733arXiv DALLAH(Alwajih et al.,2024)2024 6 Arabic dialectsMENA733ArabicNLP ARABICDIALECTDIALOGUE(Naous et al.,2020) 2020 Levantine, Egyptian, GulfMENA373WANLP ARABICAQA (Abdallah et al.,2024)2024 Arabic (MSA)MENA777SIGIR MENA SPEECHBANK(Ali et al.,2026)2026 Arabic + dialectsMENA333arXiv SHAWARMACHATS(Zeinalipour et al.,2025)2025 MSA, Egyptian, MaghrebiMENA373ArabicNLP ALEXANDRIA(El Mekki et al.,2026)2026 en + 13 Arabic dialects13 Arab countries373ACL Chinese ALIGNMMBENCH(Wu et al.,2025d)2025 zhChina337ACL CMT-EVAL(Tian et al.,2025a)2025 zhChina373arXiv FB-BENCH( Li et al.,2025d)2025 zh + enChina377EMNLP Multilingual CULTURALGROUND(Nyandwi et al.,2025)2025 39 langs42 countries73â EMNLP CVQA (Mogrovejo et al.,2024)2024 26 langs30 countries73â NeurIPS MMA-ASIA (Weihua et al.,2025)2025 10 langs (8 Asian countries) Asia73â arXiv MSTEB (Beyene et al.,2025)2025 40+ langsâ73â arXiv M2LINGUAL(Maheshwary et al.,2025)2025 MLâ37â arXiv WILDCHAT(Zhao et al.,2024b)2024 ML (incl. code-switch)200+ countries37â ICLR MLQA (Lewis et al.,2020)2020 en, ar, de, es, hi, vi, zhCross-lingual77â ACL ADU-BENCH( Gao et al.,2025a)2025 ar, zh, en, fr, de, ja, ko, ru, es Multilingual audio dialogue737ACL MT-BENCH-HI(Kamath et al.,2025)2025 HindiIndia377ACL-W CUDIALOG(Cao et al.,2024)2024 enMultiple cultures377EACL-F INDOTOD ( Kautsar et al.,2023)2023 IndonesianSoutheast Asia377SEALP TYPHOON-AUDIO(Manakul et al.,2025)2025 Thai, EnglishThailand / Thai speech737Interspeech SEA-VQA (Urailertprasert et al.,2024)2024 enSoutheast Asia / 8 countries, 53 cultures737ALVR Table 4:Cultural and linguistic resources relevant to multi-turn / multimodal dialogue. MT:3= multi-turn structure;7= single-turn only.M:3= more than one modality (vision and/or speech).Dialect:3= dialect- aware (beyond MSA or a single variety). dialogue benchmark with native-speaker transla- tions of CamRest and SMD, covering NLU, dia- logue state tracking, and response generation. Single-turn cultural and linguistic anchors. OASIS, developed through the EverydayMMQA framework (Alam et al.,2025), is a culturally grounded spoken visual QA resource covering En- glish, MSA, and Arabic dialectal varieties across 18 Arab countries. It combines real images, text questions, spoken questions, and image-grounded answers, but its current task format is single- turn rather than conversational. M2CQA ( Mousi et al.,2026) provides a single-turn multimodal QA benchmark spanning 17 MENA countries in MSA and multiple Arabic dialects, and intro- duces the Cultural Hallucination and Factual Re- call metric for measuring culturally grounded cor- rectness. ArabicaQA (Abdallah et al.,2024) sup- ports Arabic question answering, Dallah (Alwajih et al. ,2024) focuses on dialect-aware Arabic multi- modal modeling, MLQA (Lewis et al.,2020) pro- vides a cross-lingual extractive QA baseline, and mSTEB (Beyene et al.,2025) extends multilingual evaluation to speech and text tasks. Broader cultural multimodal resources.SEA- VQA ( Urailertprasert et al.,2024) evaluates cul- tural visual question answering across South- east Asian settings. CVQA (Mogrovejo et al., 2024), CulturalGround (Nyandwi et al.,2025), and MMA-ASIA (Weihua et al.,2025) extend culturally focused multimodal evaluation across broader language and regional settings. These re- sources are mostly single-turn, but they provide useful baselines and design signals for future cul- turally grounded multi-turn dialogue benchmarks. A.2 Modeling Paradigms and Training Strategies A.2.1 Classical Dialogue Systems Classical dialogue systems treated dialogue as a pipeline of separate modules: natural language un- derstanding, dialogue state tracking, policy learn- ing, and natural language generation. Early rule-based systems used hand-written templates and finite-state managers. Statistical approaches later modeled dialogue as a partially observable Markov decision process ( GaĹĄi Ě c et al.,2013). Neural sequence-to-sequence models (Sutskever et al.,2014;Serban et al.,2016) replaced parts of the response-generation pipeline with end- to-end models, but the broader system design often remained modular. These systems re- main useful baselines for understanding how later Transformer- and LLM-based dialogue models ab- sorbed functions that were previously handled by separate components. A.2.2 Transformer Dialogue Architectures Pre-LLM Transformer dialogue architectures es- tablished many capabilities that were later sub- 29 ModelYear Mod. Ctx Stream Key noveltyVenue Pre-LLM transformer dialogue managers DIALOGPT (Zhang et al.,2020b)2020 TCC7147M Reddit pre-trainACL PLATO (Bao et al.,2020)2020 TCC7Discrete latent for diversityACL TOD-BERT (Wu et al.,2020)2020 TCC7Unified TOD pre-train (9 corpora)EMNLP BLENDERBOT(Roller et al.,2021)2021 TCC7Blended persona / knowledge / empathyEACL TURNGPT (Ekstedt and Skantze,2020)2020 T+SCC7Turn-taking transformerSIGDIAL PPTOD (Su et al.,2022)2022 TCC7Plug-and-play multi-task TODACL KA-GPT2 (Lin et al.,2021a)2021 TCC7KG-aware GPT-2 for dialogue state trackingEMNLP COMMONSENSE+NER (Varshney et al.,2022) 2022 TCC7KG + NER grounding for open-domain dialogueNAACL DLGNET(Oluwatobi and Mueller,2020)2020 TCC7Contextual Transformer encoding for response generationACL-W MT-BERT-DST (Kapelonis et al.,2022)2022 TCC7Shared BERT encoder; multi-task intent + slot predictionInterspeech LLM-based, text-only INSTRUCTGPT (Ouyang et al.,2022)2022 TCC7RLHF (PPO) alignmentNeurIPS COMEDY (Chen et al.,2025f)2025 TMA7Compressive session memoryCOLING FNCTOD (Li et al.,2024c)2024 TCC7Zero-shot DST via function calling; no task-specific trainingACL UNICONV(Mo et al.,2025)2025 TCC7Joint retrieval-generation model for multi-turn conversational searchACL CHATQA (Liu et al.,2024c)2024 TCC7LLM + dedicated dense retriever jointly tuned for multi-turn conversational QANeurIPS Long-context / recurrent architectures TRANSFORMER-XL (Dai et al.,2019)2019 TRNN7Segment-level recurrenceACL RMT (Bulatov et al.,2022)2022 TMA7Memory tokensNeurIPS RWKV (Peng et al.,2023)2023 TRNN7Linear-attn RNN-likearXiv MEMBART (Wu and Yu,2024)2024 TMA7Dual-attn read/write per turnarXiv LLM-DST (Feng et al.,2023b)2023 TCC7LLM-driven dialogue state tracking;EMNLP CCM (Kim et al.,2024a)2024 TMA7KV-cache compression into compact context memory for online LM interactionICLR AudioLLMs (speech + text) SPEECHGPT (Zhang et al.,2023)2023 T+SCC7Discrete speech tokensEMNLP-F SALMONN (Tang et al.,2024)2024 T+SCC7Dual audio encoderICLR QWEN2-AUDIO(Chu et al.,2024)2024 T+SCC7Multi-task audio chatarXiv MOSHI(DĂŠfossez et al.,2024)2024 T+SCC3Full-duplex; 160ms; inner monologuearXiv AUDIOPALM (Rubenstein et al.,2023)2023 T+SCC7PaLM extended with audio tokens for speech + textarXiv QWEN-AUDIO(Chu et al.,2023)2023 T+SCC7Universal audio understanding via unified audio-language modelarXiv STYLE-TALKER(Li et al.,2024b)2024 T+SCC7Audio-LLM + style-based TTS for fast spoken dialoguearXiv MINI-OMNI(Xiezhifei,2024)2024 T+SCC3End-to-end real-time speech interaction with parallel text-audio streaming outputarXiv FREEZE-OMNI(Wang et al.,2025b)2025 T+SCC3Low-latency speech dialogue with frozen LLMICML MINMO(Chen et al.,2025g)2025 T+SCC3Full-duplex multimodal LLM for seamless spoken interactionarXiv DGSLM (Nguyen et al.,2023)2023 SCA7First textless spoken dialogue model; dual-tower cross-attention on raw two-channel audio TACL LLAMA-OMNI2 ( Fang et al.,2025)2025 T+SCC3Autoregressive streaming speech synthesis for real-time spoken chatbotACL SPIRIT-LM ( Nguyen et al.,2025)2025 T+SCC7Interleaved speech-text tokens; word-level alignment; prosody preservedTACL GLM-4-VOICE(Zeng et al.,2024)2024 T+SCC3End-to-end spoken chatbot; controllable emotion, rate, and dialectarXiv SLAM-OMNI(Chen et al.,2025h)2025 T+SKV3Historical text prompting; timbre control; single-stage SDM trainingACL-F STEP-AUDIO2 (Wu et al.,2025a)2025 T+SCC3Large-scale audio-LM; streaming speech, RAG, and tool usearXiv Omni-modal (text + speech + vision) GPT-4O(Hurst et al.,2024)2024 T+S+V C3End-to-end omni baselineOpenAI QWEN2.5-OMNI(Xu et al.,2025b)2025 T+S+V C3ThinkerâTalker + TMRoPEarXiv VITA-1.5 (Fu et al.,2026)2026 T+S+V C31.5s end-to-end latencyNeurIPS IXC2.5-OMNILIVE(Zhang et al.,2024c)2024 T+S+V MA3Streaming + long-term memoryarXiv MING-OMNI(AI et al.,2025)2025 T+S+V C3Unified perception + generationarXiv INTERACTIVEOMNI(Tong et al.,2025)2025 T+S+V MA3Tri-modal audio-visual multi-turn dialogue; cross-modal memoryarXiv MINI-OMNI2 (Xie and Wu,2024)2024 T+S+V C3Adds vision to Mini-Omni; full duplexarXiv EMOVA (Chen et al.,2025c)2025 T+S+V C3Disentangled speech tokeniser + vision encoder; vivid emotionsCVPR STREAM-OMNI(Zhang et al.,2025c)2025 T+S+V C3Simultaneous multimodal interactions; streamingarXiv M2-OMNI( Guo et al.,2025)2025 T+S+V C3Comprehensive modality support with competitive performancearXiv BAICHUAN-OMNI( Li et al.,2024a)2024 T+S+V C3Omni-modal technical reportarXiv MIO (Wang et al.,2025d)2025 T+S+V C3Foundation model on multimodal tokensEMNLP VISION-SPEECH(Royer et al.,2025)2025 T+S+V C7Spoken dialogue grounded in vision; no text intermediaryarXiv MGM-OMNI(Wang et al.,2025a)2025 T+S+V C3Dual-track omni LLM; long-audio understanding and personalized long speecharXiv OLA(Liu et al.,2025b)2025 T+S+V C7Omni-modal input understanding with progressive modality alignmentarXiv Multimodal multi-turn-aware (vision + text, context-managed) CONTEXTQFORMER(Lei et al.,2025)2025 T+V QC7Memory queue + cross-attentionarXiv MADAKV ( Li et al.,2025c)2025 T+V KV7Modality-aware KV-cache evictionACL DIAGNOTE(Liu et al.,2025a)2025 T+V MA7Note-taking module for VLMsEMNLP DALLAH(Alwajih et al.,2024)2024 T+V C7Dialect-aware Arabic VLMACL VIDEOLLAMA 2 (Cheng et al.,2024)2024 T+S+V C7Audio branch extension of LLaMA for video+audioarXiv CAT ( Ye et al.,2024b)2024 T+S+V C7Dynamic audio-visual scene QA alignment under temporal changeECCV LOOPSERVE(Li et al.,2025a)2025 T+V KV7Adaptive sparsification for multi-turn KV pressure at inferencearXiv AIRCACHE(Huang et al.,2025a)2025 T+V KV7Inter-modal relevancy KV compression; elite observation windowICCV CLARIFYWN (Zhang and Choi,2025)2025 TCA7Clarification-as-policy; when/what to ask across turnsNAACL DIALOGGEN( Huang et al.,2025b)2025 T+V C7Bilingual text-image generation and editing conditioned on multi-turn conversational history NAACL-F Agentic / tool-augmented (adjacent scope) REACT(Yao et al.,2022)2023 TCAâReasoning + acting interleavedICLR REFLEXION(Shinn et al.,2023)2023 TMAâVerbal RL via self-reflectionNeurIPS METAGPT (Hong et al.,2024)2024 TCAâRole-specialised multi-agentICLR SAPIENT (Du et al.,2025a)2025 TCAâMCTS planner for conversational rec.NAACL TOOLPLANNER(Wu et al.,2024)2024 TCAâPath planning + feedback RLEMNLP CHATCOT ( Chen et al.,2023)2023 TCAâMulti-turn tool-augmented CoT reasoningEMNLP WEBLINX (Lu et al.,2024)2024 TCCâReal-website navigation; 2,300 demos, 150+ sitesICML VIDEOMIND(Liu et al.,2026)2026 T+Vid CAâPlannergrounderverifieranswerer with Chain-of-LoRAICLR VIDEO-MTR ( Xie et al.,2025)2025 T+Vid CAâIterative segment selection; gated bi-level rewardarXiv Table 5:Representative multi-turn dialogue models by paradigm(full taxonomy).Ctx: C = full-context con- cat; SW = sliding window; CA = cross-attention; MA = memory-augmented; QC = Q-Former / token compression; KV = KV-cache mgmt.; RNN = recurrent / SSM.Streaming:3= real-time / full-duplex capable. sumed by LLMs. Open-domain systems such as DialoGPT (Zhang et al.,2020b), Blender- Bot (Roller et al.,2021), and DLGNet (Oluwa- tobi and Mueller ,2020) improved response gen- eration by using pretrained or contextual dia- logue encoders. In task-oriented dialogue, TOD- BERT (Wu et al.,2020), PPTOD (Su et al.,2022), and the shared BERT encoder proposed byKapelo- nis et al. (2022) advanced unified and multi-task transformer modeling for intent recognition, slot 30 MethodYear TypeMT-aware Multi-turn noveltyVenue (i) Supervised fine-tuning WILDCHAT(Zhao et al.,2024b)2024 SFT31M real opt-in convs; multilingualICLR PARROT(Sun et al.,2024)2024 SFT+PO3Parrot-Ask elicits anaphora/ellipsis in follow-ups; SFT + context-aware preference opt. ACL AQUILA-MED(Zhao et al.,2024a)2024 SFT+RLHF+DPO3Full-process medical dialogue recipearXiv QILIN-MED(Ye et al.,2024a)2024 SFT+DPO3Multi-stage knowledge injection for medical multi-turn dialoguearXiv ZHONGJING(Yang et al.,2024)2024 SFT+RLHF3Expert-feedback RLHF + real-world multi-turn medical dialogueAAAI (i) Reinforcement learning & preference optimisation INSTRUCTGPT (Ouyang et al.,2022)2022 PPO7RLHF preference recipeNeurIPS DMPO (Shi et al.,2024)2024 DPO (MT)3Direct multi-turn preference optimisationEMNLP M-DPO / M-KTO (Xiong et al.,2025)2025 DPO/KTO3Trajectory-level reward assignment (multi-turn math agents)ICLR ITPO (Wang et al.,2026)2026 RL3Implicit turn-wise policy optimisationarXiv SDPO (Kong et al.,2025)2025 DPO (segment)3Segment-level preference optimisation for social dialogueACL DIATOOL-DPO (Jung et al.,2025)2025 DPO3Multi-turn dialogue-control optimisation for tool useSIGDIAL ARCHER(Zhou et al.,2024b)2024 HRL3Two-level (on and off-policy RL)ICML REFUEL (Gao et al.,2025b)2025 RLHF3Efficient multi-turn RLHF; regresses relative-future rewardsICLR SWEET-RL (Zhou et al.,2025)2025 RL3Turn-wise advantage critic for collaborative reasoningarXiv ACT (Chen et al.,2025e)2025 DPO / self-training3Action-level contrastive self-training (clarify vs. answer)ICLR SCORE(Kumar et al.,2025)2025 RL3Online RL for self-correction over multi-turn tracesICLR MT-RLHF (Shani et al.,2024)2024 RL3Mirror-descent policy opt.; Nash from conversation-level prefsNeurIPS JOSH (Lattimer et al.,2025)2025 Self-alignment/RL3Sparse-reward self-training in MultiWOZ-derived tool envACL-F DOCTORAGENT-RL (Feng et al.,2026)2026 RL3Multi-agent collaborative RL for multi-turn clinical dialogueICASSP OFFLINERL PERSONA(Shea and Yu,2023) 2023 Offline RL3Persona-break penalty, Persona consistency criticEMNLP STUDENT-SIM(Scarlatos et al.,2025)2025 SFT+DPO3Tutor training via distillation + DPO; simulated-student model scores preference pairs AIED (i) Multi-task learning DAMSEL (Chen et al.,2025d)2025 Multi-task learning3Auxiliary task design for spoken QA with limited speech dataACL-F (iv) Synthetic data generation U LTRA C HAT (Ding et al.,2023) 2023 Synthetic data 3 1.5M synthetic multi-turn convs via paired-LLM self-chat EMNLP MMDU-45K(Liu et al.,2024d)2024 Synthetic (+bench)3GPT-4o-generated multi-turn multi-image; doubles as benchmarkNeurIPS TMDIALOG(Lei et al.,2025)2025 Synthetic data3GPT-4-generated multi-turn multimodal dialogue corpusarXiv MMDIAG(Liu et al.,2025a)2025 Synthetic data3GPT-assisted multi-turn multimodal dialogue with note annotationsEMNLP ARABICSFT CORPUS( Misbah et al.,2026) 2026 Synthetic data343K Arabic synthetic multi-turn dialogues; blueprint for Arabic SFTPLoS ONE DEEPDIALOGUE(Koudounas et al.,2025) 2025 Synthetic data3Synthetic multi-turn spoken dialogue with emotion-consistency controlarXiv TOOLDIAL(Shim et al.,2025)2025 Synthetic data3Multi-turn dialogue generation pipeline for tool-augmented LMsICLR CONSISTENTCHAT(Chen et al.,2025a)2025 Synthetic data3Skeleton-guided generation with nine intent trajectoriesEMNLP DOCTALK(Lee et al.,2025a)2025 Synthetic (PT)3Graph-based synthesis of 730k multi-turn dialogues from Wikipedia (pre-training)SIGDIAL REVIEW-INSTRUCT( Wu et al.,2025c)2025 Synthetic data3Ask-Respond-Review three-role pipeline for high-consistency dataACL-F SELF-TALK( Ulmer et al.,2024)2024 Synthetic data3LLM-simulated user-agent task dialogue generationACL-F APIGEN-MT (Prabhakar et al.,2026)2026 Synthetic data3Verifiable multi-turn agent trajectory generationNeurIPS (v) Conversational retrieval-augmented training CHATQA (Liu et al.,2024c)2024 SFT (retrieval)3Two-stage tuning of retriever + LLM for multi-turn conversational QANeurIPS CHATQA-2 ( Xu et al.,2025c)2025 SFT (retrieval)3Three-stage context-expansion tuning (8K128K) for long conversational RAGICLR HACONVDR (Mo et al.,2024)2024 Dense retrieval3Context-denoised reformulation + turn-impact supervisionACL-F CONVAUG (Chen et al.,2024a)2024 Dense retrieval3LLM-cognition augmentation + difficulty-adaptive contrastive learningACL Table 6:Multi-turn training strategies, grouped by the five families of §5.Type: SFT = supervised fine- tuning; DPO/KTO = direct-preference / KahnemanâTversky optimisation (MT = multi-turn, segment = segment- level); RLHF/PPO, RL, HRL = hierarchical RL;retrieval= conversational-retrieval training;Synthetic= data- generation pipeline; Combined types (e.g. SFT+DPO) mark hybrids that span families; each work is listed under itsprimarycontribution.MT-aware:3if the method explicitly optimises across turns rather than treating each turn independently. filling, dialogue state tracking, policy prediction, and response generation. A related line incor- porated structured knowledge and dialogue state tracking. Lin et al.(2021b) inject schema and knowledge-graph structure into GPT-2 for dia- logue state tracking, while Feng et al.(2023b) show that LLMs can directly track slot-value pairs across turns without a dedicated DST module. An- other line addresses long context and memory. Transformer-XL ( Dai et al.,2019) reuses hidden states across segments, Recurrent Memory Trans- former (Bulatov et al.,2022) adds recurrent mem- ory tokens, RWKV ( Peng et al.,2023) reformu- lates attention as a linear recurrence, and Mem- BART (Wu and Yu,2024) adds a gated memory module to preserve dialogue history without sim- ply enlarging the input window. A.2.3 LLM-based Dialogue Models The LLM-based phase shifted dialogue modeling away from task-specific managers toward general- purpose instruction-following models. Instruct- GPT ( Ouyang et al.,2022) helped define this phase through supervised fine-tuning, reward mod- eling from human preferences, and reinforcement learning from human feedback. FnCTOD (Li et al. ,2024c) shows that task-oriented structure can be recovered inside an LLM by treating each dialogue slot as a callable function and letting the model fill slot values zero-shot. LLM-centric sys- tems for conversational QA and RAG, including ChatQA, ChatQA-2, and UniConv, are discussed in Section 5, since their primary contribution is training rather than architecture. A.2.4 AudioLLMs AudioLLMs extend conversational models to speech by jointly modeling language and au- dio. Early systems such as SpeechGPT and Au- dioPaLM represent speech through discrete units or audio tokens within an LLM, while Qwen- Audio, Qwen2-Audio, SALMONN, Kimi-Audio, 31 StrategyTraining signalRole across turnsBest suited toKey trade-offModality Objective-level strategies Supervised fine- tuning Reference dialogue turns; token-level likelihood Implicit:learns re- sponses conditioned on dialogue history General dialogue ability;domain, format, and safety adaptation Stable and data- efficient, but pro- vides no explicit delayed or cross-turn credit T / S / M Multi-task learning Dialogue loss with auxiliary objectives such as persona, knowledge,re- trieval, or dialogue state Indirect:auxiliary tasks encourage consis- tency and grounding Persona and knowl- edgegrounding; dialogue-state track- ing; transfer across skills Improves transfer, but may cause neg- ative transfer and requires loss balanc- ing T / M Preference opti- mization Preference labels or rewards over responses or trajec- tories Explicit:can optimize outcomes over multiple turns Preferencealign- ment; long-horizon task success; safety Supports delayed credit, but can be costly and sensitive to reward misspecifi- cation T Data and grounding strategies Synthetic dataGenerated, self- play, or distilled multi-turndia- logues Coverage:expands the diversity and length of dialogue trajectories Low-resource set- tings; rare behaviors; safety and robustness cases Scalable and control- lable, but may am- plify errors and bi- ases and requires fil- tering T / (S, M) Conversational RAG History-aware retrievaland grounded-response supervision Grounding:teaches evidence use across di- alogue turns Knowledge-intensive dialogue; long his- tories;changing knowledge Improves grounding, but depends on re- trieval quality and adds latency T / (M) Table 7: Comparison ofstrategies for training multi-turn conversational systems. The role across turns in- dicates how each strategy supports learning beyond isolated turns. T = text; S = speech; M = multimodal; parentheses indicate emerging use. Strategies are complementary rather than mutually exclusive. and SpiRit-LM broaden this direction toward au- dio understanding, spoken dialogue, and cross- modal learning (Zhang et al.,2023;Rubenstein et al. ,2023;Chu et al.,2023,2024;Tang et al., 2024;Team,2024;Nguyen et al.,2025). A second line focuses on real-time and full- duplex spoken interaction. Moshi supports low- latency full-duplex dialogue and uses time-aligned text tokens as an inner monologue before speech generation (DĂŠfossez et al.,2024). Subsequent systems explore streaming speech generation, con- trollable voice, style and timbre modeling, and si- multaneous listening and speaking (Chen et al., 2025g;Fang et al.,2025;Li et al.,2024b;Zeng et al.,2024;Xiezhifei,2024;Chen et al.,2025h). Freeze-Omni keeps the LLM backbone frozen while training speech input and output modules, whereas SALMONN-omni jointly models user speech and system output in a codec-free architec- ture that supports turn-taking, barge-ins, and echo cancellation ( Wang et al.,2025b;Yu et al.,2026). Other work separates duplex control from the dialogue backbone. FlexDuo uses a Speak-Listen- Idle state machine to reduce false interruptions, while FireRedChat combines personalized VAD with semantic end-of-turn detection for control- lable barge-in ( Liao et al.,2025;Chen et al., 2025b). In a different direction, dGSLM models two-channel raw audio without text supervision, generating speech, laughter, and other paralinguis- tic signals ( Nguyen et al.,2023).Lu et al.(2026) provide a broader review of full-duplex architec- tures, interaction ontologies, and decision states. Overall, AudioLLMs have substantially im- proved speech-native and real-time interaction. However, most systems still rely on the underlying language model to maintain dialogue history, with limited explicit modeling of long-horizon spoken memory and cross-turn acoustic grounding. A.2.5 Omni-modal Models Omni-modal models extend AudioLLMs by in- tegrating vision with text and speech, and in some cases video. GPT-4o provides an end-to- end omni-modal system with real-time stream- ing, while Mini-Omni2 extends Mini-Omni with 32 visual input (Hurst et al.,2024;Xie and Wu, 2024;Xiezhifei,2024). Other systems explore low-latency vision-speech interaction, simultane- ous multimodal processing, unified tokenization, and multimodal pretraining (Fu et al.,2026;Chen et al.,2025c;Zhang et al.,2025c;Guo et al.,2025; Wang et al.,2025d).Royer et al.(2025) directly connect speech and vision encoders without using text as an intermediate representation. Recent models further broaden this design space through instruction tuning, multimodal gen- eration, video-audio alignment, and long-horizon speech modeling (Li et al.,2024a;AI et al.,2025; Xu et al.,2025b;Wang et al.,2025a). Audio-video models such as VideoLLaMA 2 and CAT addition- ally target reasoning over dynamic audio-visual content ( Cheng et al.,2024;Ye et al.,2024b). A separate line focuses explicitly on multi-turn interaction. Most omni-modal models treat dia- logue history as a flat input sequence, whereas newer methods introduce mechanisms for mem- ory and context management. DiagNote uses ex- plicit note-taking, CoLVLM follows a memory- perception-planning-execution loop, and Contex- tQFormer maintains a dedicated contextual mem- ory block ( Liu et al.,2025a;Han et al.,2025;Lei et al.,2025). LoopServe and MadaKV reduce in- ference cost in long multimodal sessions, while DialogGen conditions image generation and edit- ing on conversational history ( Li et al.,2025a,c; Huang et al.,2025b). Collectively, omni-modal models have broad- ened multimodal interaction; however, explicit cross-turn memory and long-horizon grounding re- main less developed. A.2.6 Agentic and Tool-augmented Dialogue Agentic dialogue systems use conversation as an interface for planning, tool use, observation, and revision. ReAct interleaves reasoning with exter- nal actions, Reflexion stores verbal feedback from failed attempts for later use, and MetaGPT coor- dinates multiple role-based LLM agents through structured messages ( Yao et al.,2022;Shinn et al., 2023;Hong et al.,2024). Tool-augmented systems connect this process to external environments. ToolPlanner combines path planning with process- and outcome-level feedback, while ChatCoT integrates external tools into reasoning ( Wu et al.,2024;Chen et al.,2023). Ď-Bench evaluates agents in retail and airline set- tings using APIs and simulated users, whereas ToolSandbox focuses on stateful tool use over variable-length interactions (Yao et al.,2025;Lu et al.,2025). xLAM-2 represents open models for multi-turn function calling (Prabhakar et al., 2026). Other systems extend agentic interaction to web navigation, recommendation, video, and stream- ing environments. WebLINX evaluates conver- sational web navigation, while SAPIENT applies Monte Carlo Tree Search to conversational rec- ommendation ( Lu et al.,2024;Du et al.,2025a). Zhang et al.(2024d) incorporate usersâ emo- tional states into recommendation. VideoMind and Video-MTR support iterative reasoning over video, while IXC2.5-OmniLive combines stream- ing perception, multimodal long-term memory, and reasoning for live interaction (Liu et al.,2026; Xie et al.,2025;Zhang et al.,2024c). Overall, agentic systems highlight three multi- turn requirements: tracking external state across actions, replanning as goals or environments change, and maintaining memory over extended interactions. A.2.7 Supervised Fine-tuning Supervised fine-tuning adapts models to multi- turn interaction using conversations that preserve cross-turn dependencies. UltraChat provides large-scale synthetic instructional dialogues, while WildChat contains real opt-in conversations with broad topic and language coverage. Parrot ex- plicitly targets anaphora and ellipsis in follow-up turns ( Ding et al.,2023;Zhao et al.,2024b;Sun et al.,2024). Domain-specific work extends SFT to clinical dialogue and adversarial safety ( Zhao et al.,2024a;Ye et al.,2024a;Yang et al.,2024; Rahman et al.,2025). SFT provides a practi- cal foundation for multi-turn training, with perfor- mance largely shaped by the coverage and quality of training data. A.2.8 Reinforcement Learning and Preference Optimization Multi-turn reinforcement learning and preference optimization address delayed rewards and credit assignment across turns. InstructGPT estab- lished the standard RLHF pipeline, while ArCHer and MT-RLHF extend optimization toward turn- and conversation-level feedback ( Ouyang et al., 2022;Zhou et al.,2024b;Shani et al.,2024). Preference-based methods such as DMPO, Multi- turn DPO/KTO, Parrot, and SDPO further opti- 33 mize dialogue histories, interaction sequences, or segments rather than individual responses (Shi et al.,2024;Xiong et al.,2025;Sun et al.,2024; Kong et al.,2025). A second line focuses on actions and out- comes across turns. Action-Based Contrastive Self-Training distinguishes actions such as answer- ing and clarifying, while DiaTool-DPO and JOSH optimize tool-use behavior from interaction se- quences (Chen et al.,2025e;Jung et al.,2025; Lattimer et al.,2025). Other approaches target self-correction, collaborative reasoning, proactive interaction, and sparse long-horizon rewards (Ku- mar et al.,2025;Zhou et al.,2025;Gao et al., 2025b;Wang et al.,2026;Abdulhai et al.,2025). Domain-specific work further applies these meth- ods to clinical and tutoring dialogue (Scarlatos et al.,2025;Feng et al.,2026). More broadly, these methods extend optimiza- tion from individual responses to cross-turn behav- ior and long-horizon outcomes. A.2.9 Multi-task Learning Multi-task learning improves dialogue modeling by sharing representations across related objec- tives. PPTOD jointly trains response genera- tion, dialogue state tracking, and policy prediction, while TOD-BERT learns dialogue-aware represen- tations across multiple task-oriented dialogue cor- pora ( Su et al.,2022;Wu et al.,2020). DAMSEL extends this approach to spoken conversational QA, using auxiliary objectives to improve learning when labeled speech data is limited ( Chen et al., 2025d). Taken together, multi-task learning improves transfer across related dialogue skills, although most approaches do not explicitly optimize session-level dependencies. A.2.10 Synthetic Data Generation Synthetic data generation scales multi-turn train- ing data while allowing control over dialogue structure, topic, and modality. UltraChat simulates user-assistant conversations, while MMDU-45K, MMDiag, and TMDialog extend synthetic genera- tion to multimodal dialogue, note-taking, and con- text modeling ( Ding et al.,2023;Liu et al.,2024d, 2025a;Lei et al.,2025). Other pipelines target Arabic and emotional spoken dialogue, tool use, coherent intent sequences, and multi-topic infor- mation seeking ( Misbah et al.,2026;Koudounas et al.,2025;Shim et al.,2025;Chen et al.,2025a; Lee et al.,2025a). Recent approaches also use reviewer-guided regeneration, task-oriented dia- logue bootstrapping, and simulated API interac- tions to improve data quality and task cover- age (Wu et al.,2025c;Ulmer et al.,2024;Prab- hakar et al.,2026). In summary, synthetic generation offers a scal- able way to target multi-turn behaviors, shifting from general dialogue scaling toward targeted sim- ulation across modalities and tasks. A.2.11 Training for Conversational RAG Conversational RAG trains models to retrieve and use evidence as dialogue context evolves. PK-ICR jointly retrieves persona and external knowledge, while commonsense retrieval helps fill knowledge gaps in open-domain dialogue ( Oh et al.,2023; Varshney et al.,2022). ChatQA and ChatQA-2 jointly train retrieval and generation for conversa- tional and long-context RAG (Liu et al.,2024c; Xu et al.,2025c). Other methods improve history selection, retrieval supervision, and query refor- mulation: HAConvDR filters noisy dialogue his- tory, UniConv jointly optimizes retrieval and gen- eration, ConvAUG strengthens retrieval through LLM-cognition data augmentation, and IterCQR iteratively reformulates queries using retrieval feedback ( Mo et al.,2024,2025;Chen et al., 2024a;Jang et al.,2024). KEDiT compresses re- trieved evidence into adapter parameters, while CORAL incorporates citation labeling into conver- sational RAG training and evaluation ( Zhang et al., 2025a;Cheng et al.,2025). Overall, conversational RAG has progressed to- ward history-aware retrieval and evidence integra- tion across turns. In Table8, we map common multi-turn training goals to the primary strategies discussed above, providing a practical guide for strategy selection. A.3 Details on Evaluation In Table9, we summarize representative metrics and frameworks across the five evaluation families. In Table10, we compare these families by their evaluation focus, suitable settings, and main limi- tations. We discuss each family in detail below. A.3.1 Evaluation Families Surface-form and task-oriented metrics. Surface-form metrics such as BLEU, ROUGE, and BERTScore remain useful baselines; however, they mainly score individual responses and 34 Training goalPrimary strategy General conversational and instruction-following ability SFT on curated multi-turn dialogues Domain-specific or safety- focused behavior SFT with targeted synthetic augmentation Human preference and long- horizon outcomes Preference optimization or multi-turn RL Knowledge-intensive multi- turn QA Conversational RAG train- ing Low-resource settings or rare behaviors Synthetic data generation Persona/knowledge ground- ing and skill transfer Multi-task learning Table 8: Mapping common multi-turn training goals to their primary training strategies. These strategies can also be combined depending on the target setting. provide limited insight into cross-turn behavior. They do not capture whether a model preserves instructions, persona, grounding, or coherence across a session. Task-oriented metrics such as joint goal accuracy, Slot-F1, Inform, and Success better reflect multi-turn task completion by tracking dialogue states and outcomes. However, they remain task-specific, can be sensitive to paraphrasing, and often reduce success to binary outcomes. M-Relevance extends response relevance to imageâtext dialogue, although it is limited to two-modal settings. Session-level dialogue metrics.Session-level metrics directly assess behaviors that depend on earlier turns, including memory recall, instruction retention, constraint satisfaction, feedback integra- tion, and dialogue-level hallucination. Representa- tive approaches measure rubric compliance (APR and ARS), positional consistency (PWC), mul- timodal memory and reasoning (MMRC), inter- action patterns (MT-Eval), structural constraints (WCSR), instruction hierarchy (IHEval), program- matic instruction following (PIF), feedback-based reasoning (TurnBench-MS), and dialogue-level hallucination (DiaHalu). Together, these metrics provide a broader view of cross-turn competence, although their definitions and scoring remain frag- mented across benchmarks. Agentic and retrieval-grounded evaluation. Agentic and retrieval-grounded systems require evaluation of both responses and actions across an interaction. A fluent answer can still fail if the model uses an incorrect tool state, calls an API with missing information, or cites the wrong source. CORAL jointly measures retrieval, gen- eration, and citation attribution, whileĎ-Bench measures repeated-trial task success throughĎ- pass k . ToolSandbox evaluates stateful interactions using milestone and minefield scoring, and Agent- Board tracks progress through subgoal comple- tion. These settings make external state, side ef- fects, and repeated-trial reliability central evalua- tion concerns. Speech-native and full-duplex evaluation. Spoken dialogue evaluation must consider both semantic content and interaction dynamics. ADU- Bench covers open-ended audio dialogue across skills, languages, and ambiguity types, while MTalk-Bench combines pairwise and rubric- based scoring for semantic content, vocal cues, and ambient sound. ASK-QA focuses on spoken clarification, and URO-Bench separates correct- ness, speech quality, speechâtext matching, and first-packet latency. FD-Bench adds interruption and timing measures for full-duplex interaction, while Full-Duplex-Bench-v3 extends evaluation to tool use and disfluency handling. These benchmarks show that transcript correctness alone is insufficient for evaluating spoken interaction. LLM-as-a-judge and human evaluation. LLM-as-a-judge and human evaluation support di- mensions that are difficult to score automatically, including open-ended coherence, persona consis- tency, multimodal grounding, and human-likeness. MT-Bench, GPTScore, JudgeLM, BotChat, Con- textualJudgeBench, LLM-Eval Analysis, and Multi-Judge Evaluator represent judge-based approaches, while ABC-Eval, MMDU, MTalk- Bench, and WildBench incorporate human or hybrid assessment. These methods offer flexible evaluation, although their reliability depends on rubric design and can be affected by judge bias, rater inconsistency, cost, and small ranking differences. Overall, these evaluation methods capture com- plementary aspects of multi-turn performance, but no single approach fully measures session-level competence across modalities and interaction set- tings. A.4 Search Keywords We organized the literature search into six query groups covering core multi-turn dialogue, spo- ken and audio interaction, multimodal and omni- modal dialogue, training and alignment, evalua- tion, and cultural and multilingual settings. Within 35 Metric / FrameworkTypeWhat it measuresMT-spec. Core limitationApplied in Surface-form and task-oriented BLEU (Papineni et al.,2002)NLGn-gram precision7No semantics; no coherence signalMT-Bench, CoQA ROUGE (Lin,2004)NLGn-gram recall7Surface-form only; ignores contextSpokenWOZ BERTScore (Zhang et al.,2020a)NLGSemantic embedding similarity7Single utterance; history-blindMMDU M-Relevance (Feng et al.,2023a)NLGModality-aware response relevance3Defined for two-modal onlyMMDialog JGADSTJoint goal accuracy across slots3Brittle to surface paraphraseMultiWOZ Slot-F1DSTPer-slot precision and recall7Slot-type bias; misses long-range contextSpokenWOZ Inform / SuccessDSTEntity provision and task success3Binary; misses partial successMultiWOZ Session-level metrics APR & ARS (Deshpande et al.,2025;Gosai et al.,2025) CONSAverage pass rate and rubric score3Turn-level and session-level scores can divergeMultiChallenge PWC (Li et al.,2025e)CONSPosition-weighted consistency3Requires adversarial probe setMT-Eval style MMRC 6-axis (Xue et al.,2025)MEMExtract, reason, update, manage, recall, refuse3Image management is rare in text-only settingsMMRC TURNWISE gap (Graf et al.,2026)CONSMatched single-turn and multi-turn gap3Requires paired single-turn referencesGeneral LLMs WCSR (Li et al.,2025b)CONSStructural and intra-turn constraint satisfaction3Depends on constraint extraction and LLM judging StructFlowBench MT-Eval patterns (Kwan et al.,2024)CONSRecollection, expansion, refinement, follow-up3Single-session; no cross-session coverageMT-Eval IHEval (Zhang et al.,2025e)CONSInstruction hierarchy across system, user, history, and tool outputs3Conflict resolution remains difficultIHEval Feedback-loop rule inference (Zhang et al.,2025d)CONSIntegration of structured feedback across turns3Game setting may not cover open-domain dialogue TurnBench-MS DiaHalu taxonomy ( Chen et al.,2024b)CONSDialogue-level hallucination subtypes3Subtype labels require human annotationDiaHalu PIF (Epstein et al.,2024)CONSProgrammatic instruction following across accumulated constraints3Checks format constraints, not answer correctness MMMT-IF Agentic and retrieval-grounded CORAL citation labeling (Cheng et al.,2025)AGENTRetrieval, generation, and citation attribution3Citation annotation requires human labelingCORAL Ď-pass k (Yao et al.,2025)AGENTTask pass rate across repeated trials3Requires live user simulator; high costĎ-Bench ToolSandbox scoring (Lu et al.,2025)AGENTStateful trajectory with milestone and minefield scoring3Simulator-dependent; limited domainsToolSandbox AgentBoard progress rate (Ma et al.,2024)AGENTSubgoal completion across agent trajectories3Requires task-specific progress decompositionAgentBoard Speech-native and full-duplex ADU-Bench (Gao et al.,2025a)JUDGEOpen-ended audio dialogue skills and ambiguities3Judge-dependent; rubric design centralADU-Bench MTalk-Bench protocol (Du et al.,2025b)JUDGE+HUMAN Pairwise and rubric-based S2S scoring3Small ranking gaps unstableMTalk-Bench ASK-QA outcome similarity (Chen et al.,2025d)CONSSemantic similarity after spoken clarification turns3Uses simulated users and TTS follow-upsASK-QA URO-Bench (Yan et al.,2025b)CONSCorrectness, speech quality, speech-text match, latency3Dimensions are not always separableURO-Bench FD-Bench interruption set (Peng et al.,2025)FDSIR, SRIR, EIR, NIR, and SRR3Synthetic user speech; limited systemsFD-Bench FD-Bench timing set (Peng et al.,2025)FDIRD, FSED, ERT, and EIT3Requires reliable VAD, ASR, and audio logsFD-Bench Full-Duplex-Bench-v3 (Lin et al.,2026)FD+AGENTTool-use correctness and disfluency tolerance3Emerging full-duplex tool-use settingFull-Duplex-Bench-v3 LLM-as-judge and human evaluation MT-Bench (Zheng et al.,2023)JUDGEPairwise preference and 1--10 rating3Self-enhancement and verbosity biasMT-Bench GPTScore (Fu et al.,2024)JUDGEScoring through instruction-prompt likelihood3Prompt-sensitive; weak dialogue groundingMulti-turn dialogue quality LLM-Eval Analysis (Zhang et al.,2024a)JUDGECoherence, engagement, informativeness3No multi-turn-specific axesOpen-domain dialogue BotChat (Duan et al.,2024)JUDGEHuman-likeness discrimination3Confounded by GPT-4 dominanceOpen LLMs JudgeLM (Zhu et al.,2025)JUDGEFine-tuned open judge3Knowledge and position biasJudgeLM bench ContextualJudgeBench (Xu et al.,2025a)JUDGEJudge consistency under context and RAG3Contextual sensitivity; reference dependenceRAG + multi-turn eval Multi-Judge Evaluator (Tang et al.,2025)JUDGEDistilled agreement from multiple judges3Dependent on judge pool qualityMulti-turn dialogue ABC-Eval (Zhao et al.,2026a)HUMANFine-grained behavioural criteria3Labour-intensive; rater driftOpen-domain chat MMDU rubric judging ( Liu et al.,2024d)HUMANLong-form multimodal rubric judgment3Expensive; hard to scaleMMDU WildBench (Lin et al.,2025)JUDGEPairwise evaluation on real user conversations3Scores individual responses pairwiseWildBench Table 9:Representative evaluation metrics and frameworksfor multi-turn dialogue.Type: NLG = surface- form overlap; DST = task-state tracking; CONS = consistency; MEM = memory; AGENT = agentic/stateful; JUDGE = LLM-as-judge; HUMAN = human rubric; FD = full-duplex speech.MT-spec.:3if the metric is un- defined or degenerate in single-turn settings. FD-Bench metric abbreviations:SRR = successful reply rate; SIR = successful interrupt rate; SRIR = successful reply-to- interrupt rate; EIR = early interrupt rate; NIR = noise interrupt rate; IRD = interrupt response delay; FSED = first speech emit delay; ERT = early reply time; EIT = early interrupt time. each group,ORconnected alternative keywords, whileANDcombined complementary concepts. This strategy aimed to provide broad coverage across modalities, methods, benchmarks, and lan- guage settings. Core multi-turn dialogue.We used âmulti-turnâ, âmulti turnâ, âmultiturnâ, âmulti-roundâ, âconver- sational AIâ, âdialogue systemâ, âdialog systemâ, âtask-oriented dialogueâ, âtask oriented dialogâ, âopen-domain dialogueâ, âchit-chatâ, âconversa- tional agentâ, âdialogue state trackingâ, âconversa- tion historyâ, âcontext-aware dialogueâ, and âdia- logue managementâ. These terms were combined with âlanguage modelâ, âLLMâ, âlarge language modelâ, âtransformerâ, âBERTâ, âGPTâ, âT5â, âinstruction tuningâ, âfine-tuningâ, âpre-trained modelâ, and âneural dialogueâ. Spoken and audio dialogue.Speech-related keywords included âspoken dialogueâ, âspeech dialogueâ, âvoice assistantâ, âspoken language understandingâ, âSLUâ, âspoken QAâ, âend-to- end spoken dialogueâ, âspeech-to-speechâ, âspo- ken conversationalâ, âaudio dialogueâ, âASR di- alogueâ, âautomatic speech recognitionâ, âtext- to-speechâ, âTTSâ, âspeech synthesisâ, âspoken language modelâ, and âaudio language modelâ. We combined these with âmulti-turnâ, âconversa- tionalâ, âdialogueâ, or âmulti-roundâ. Multimodal and omni-modal dialogue.We used âmultimodal dialogueâ, âvisual dialogueâ, âvisual QA dialogueâ, âvideo dialogueâ, âimage- grounded conversationâ, âvision-language dia- logueâ, âVQA multi-turnâ, âomni-modalâ, âany- to-anyâ, âspeech-visionâ, âmultimodal conver- sational agentâ, âembodied dialogueâ, âaudio- visual dialogueâ, âmulti-modal chatâ, and âvi- sual chatbotâ. We also included representative model families such as âLLaVAâ, âInstructBLIPâ, âFlamingoâ, âGPT-4Vâ, âGeminiâ, âQwen-VLâ, and âCogVLMâ. These terms were combined with âmulti-turnâ, âconversationalâ, âcontextâ, or âhis- toryâ. Training and alignment.Training-related keywords included âinstruction tuningâ, âin- struction followingâ, âreinforcement learning from human feedbackâ, âRLHFâ, âdirect prefer- ence optimisationâ, âDPOâ, âLoRAâ, âlow-rank adaptationâ, âQLoRAâ, âparameter-efficient fine- tuningâ, âPEFTâ, âalignmentâ, âconstitutional AIâ, âRLAIFâ, âcontext windowâ, âlong contextâ, 36 Evaluation family What it measuresBest suited forMain limitation Surface-form& task-oriented Response overlap, semantic similarity, dialogue-state ac- curacy, and task success Task-oriented dialogue and fast automatic scor- ing Limited signal on session-level be- havior and sensitivity to valid para- phrases Session-levelInstruction retention, mem- ory, constraint satisfaction, feedback integration, and dialogue-level hallucination Memory, consistency, and long-horizon interac- tion Metric definitions vary across benchmarks, limiting direct com- parison Agentic & retrieval- grounded Tool use, stateful execution, evidence retrieval, and source attribution Tool-augmented agents and conversational RAG Limited support for hidden state, side effects, and repeated-trial reli- ability Speech-native & full-duplex Semantic quality, speech qual- ity, paralinguistic cues, inter- ruption handling, latency, and timing Spoken and full-duplex dialogue systems Semantic correctness, speech qual- ity, and interaction timing are often evaluated separately LLM-as-judge & hu- man Rubric-based and pairwise judgments of open-ended response and dialogue quality Open-ended or subjective evaluation Sensitive to judge bias, rater incon- sistency, cost, and reproducibility Table 10: Comparison of five evaluation families for multi-turn dialogue. SurveyConceptual framingModalityMT Cult. EvalâKey distinction Yi et al.(2025)LLM-based multi-turn dialogue; methods, tasks, resources, and evaluation T373Broad synthesis centered on text; no common framework for comparing session-level require- ments across modalities. Wang et al.(2023) Language-model dialogue; evolution of task-oriented and open-domain systems T377Historical view of dialogue systems across language-model generations, primarily in text. Zhang et al.(2025b) Multi-turn LLM capabilities; instruction following, memory, planning, reasoning, and evaluation T373Capability-centered analysis with limited compari- son across speech, vision/video, and omni-modal interaction. Li et al.(2025f)Multi-turn LLM interaction; task families and model-, memory-, and agent-based methods T3â˘3Broad multi-turn taxonomy; multimodal session be- havior is not a central comparison dimension. Guan et al.(2026) Conversational-agent evaluation; tools, memory, planning, and interaction quality T + Tool373Detailed agent-evaluation framework without joint analysis of data, models, training, and evaluation across modalities. Zhang et al.(2024b) Multimodal LLMs; architectures, train- ing, model taxonomy, and benchmarks T + V (+S)777Broad multimodal coverage, with dialogue treated mainly as an application rather than the primary unit of analysis. This surveyComplete multi-turn session; modal- ity breadthĂsession-level competence across data, models, training, and eval- uation T + S + V + Vid + O + Tool 333Cross-modal analysis using common session- level capabilities, including memory, revision, grounding, state tracking, timing, safety, and culturalâlinguistic competence. Table 11: Comparison with closely related surveys by conceptual framing and scope.Modality: T = text, S = speech/audio, V = vision, Vid = video, O = omni-modal, Tool = tool-augmented.MT,Cult., andEvalâindicate multi-turn, culturalâlinguistic, and multi-turn evaluation coverage, respectively.â˘denotes partial coverage. âretrieval augmented generationâ, âRAGâ, and âmemory-augmentedâ. We combined these with âdialogueâ, âconversationalâ, âmulti-turnâ, âchat- botâ, or âassistantâ. Evaluation frameworks and metrics. Evaluation-related searches used âdialogue eval- uationâ, âconversation evaluationâ, âturn-level evaluationâ, âmulti-turn evaluationâ, âcoherenceâ, âengagementâ, âconsistencyâ, âfactual ground- ingâ, âhallucinationâ, âfaithfulnessâ, âdialogue benchmarkâ, âconversational benchmarkâ, âhu- man evaluation dialogueâ, âautomatic evaluation dialogueâ, âBLEU dialogueâ, âBERTScore dia- logueâ, âperplexity dialogueâ, âFEDâ, âUSRâ, âGRADEâ, âDialEvalâ, and âCGPRâ. Cultural and multilingual grounding. We used âmultilingual dialogueâ, âcross-lingual dialogueâ, âArabic dialogueâ, âArabic conversationalâ, âlow- resource dialogueâ, âculturally groundedâ, âcul- tural dialogueâ, âdialectâ, âcode-switchingâ, âmul- tilingual QAâ, âcross-cultural conversationâ, âAra- bic NLPâ, âArabic language modelâ, âArabiziâ, âcultural benchmarkâ, âcultural biasâ, âcultural adaptationâ, âGulf dialectâ, âMSA dialogueâ, and âArabic ASRâ. These terms were combined with âmulti-turnâ, âdialogueâ, âconversationalâ, âchat- botâ, or âassistantâ. B Comparison with Related Surveys We extend the comparison in Table1in Table11, where we examine the conceptual framing, modal- ity coverage, and key distinctions of closely re- lated surveys. We show how our survey centers the complete multi-turn session and analyzes session- level competence across modalities, training, eval- uation, and culturalâlinguistic settings. 37