Paper deep dive
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data
Subhankar Ghosh, Jason Li, Paarth Neekhara, Shehzeen Hussain, Ryan Langman, Xuesong Yang, Roy Fejgin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 2:30:42 AM
Summary
MagpieTTS-LF is an inference-time approach designed to enable long-form speech generation without model retraining. It addresses common issues in long-form TTS like prosodic drift, speaker inconsistency, and boundary artifacts. The method introduces three key innovations: soft attention priors to maintain monotonic alignment while preserving distant context, a stateful inference algorithm to maintain encoder and attention states across sentence chunks, and history-aware text encoding for better prosodic planning. Experimental results on the Long-Form HifiTTS benchmark show that MagpieTTS-LF outperforms state-of-the-art models like VibeVoice, XTTS, and Qwen3-TTS in terms of intelligibility (WER/CER), prosodic boundary continuity, and speaker consistency.
Entities (9)
Relation Signals (5)
MagpieTTS-LF → evaluatedon → Long-Form HifiTTS
confidence 100% · We introduce Long-Form HifiTTS dataset, a benchmark for evaluating long-form speech synthesis... We evaluate MagpieTTS-LF long-form generation against the state-of-the-art baselines...
MagpieTTS-LF → isanextensionof → MagpieTTS
confidence 100% · enables MagpieTTS to produce coherent long-form speech without model retraining.
MagpieTTS-LF → uses → Soft Attention Priors
confidence 100% · Our method introduces three key innovations: (1) soft attention priors...
MagpieTTS-LF → uses → Stateful Chunk Generation
confidence 100% · a stateful inference algorithm that maintains context across sentence chunks
NVIDIA Corporation → developed → MagpieTTS-LF
confidence 90% · Authors are listed with NVIDIA Corporation affiliation.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence boundary artifacts. Existing approaches either compress sequences, increase context length or naively concatenate independently synthesized chunks. We present an inference-time approach called MagpieTTS-LF that enables MagpieTTS to produce coherent long-form speech without model retraining. Our method introduces three key innovations: (1) soft attention priors to guide monotonic alignment while preserving past and future context; (2) a stateful inference algorithm that maintains context across sentence chunks, ensuring prosodic continuity; (3) history-aware text encoding that uses past text for discourse-level prosodic planning. Experiments on long texts show significant improvements in long-range intelligibility, prosodic coherence, speaker consistency, and boundary naturalness compared to other baselines.
Tags
Links
- Source: https://arxiv.org/abs/2606.18485v1
- Canonical: https://arxiv.org/abs/2606.18485v1
Trouble viewing inline? Open PDF directly →
Full Text
25,095 characters extracted from source content.
Expand or collapse full text
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data Subhankar Ghosh 1 , Jason Li 1 , Paarth Neekhara 1 , Shehzeen Hussain 1 , Ryan Langman 1 , Xuesong Yang 1 , Roy Fejgin 1 1 NVIDIA Corporation, USA subhankarg,jasoli,pneekhara,shehzeenh,rlangman,xueyang,rfejgin@nvidia.com Abstract Neural Text-to-Speech (TTS) systems achieve remarkable qual- ity on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence boundary artifacts. Existing approaches either compress sequences, in- crease context length or naively concatenate independently syn- thesized chunks. We present an inference-time approach called MagpieTTS-LF that enables MagpieTTS to produce coher- ent long-form speech without model retraining. Our method introduces three key innovations: (1) soft attention priors to guide monotonic alignment while preserving past and future context; (2) a stateful inference algorithm that maintains con- text across sentence chunks, ensuring prosodic continuity; (3) history-aware text encoding that uses past text for discourse- level prosodic planning. Experiments on long texts show signif- icant improvements in long-range intelligibility, prosodic coher- ence, speaker consistency, and boundary naturalness compared to other baselines. Index Terms: Text-to-Speech, Speech Synthesis, Speech LLM, Long-form Generation 1. Introduction While advancements in large-scale generative modeling in TTS has enabled unprecedented naturalness and speaker similarity, yet most of the methods suffer from hallucinations, prosodic drift, and boundary artifacts as generation length grows. State- of-the-art models like Tortoise TTS [1], VALL-E 2 [2], VALL- E R [3], NaturalSpeech 2/3 [4, 5], VoiceBox [6], XTTS [7], Qwen3-TTS [8], MagpieTTS [9, 10] and CosyVoice [11] pro- duce extremely natural speech on 2-20 second utterances but offer no native mechanism for paragraph length speech gener- ation. When generating longer length speech, these systems default to sentence-level chunking followed by concatenation. This strategy that produces characteristic artifacts including energy discontinuities and warbles at boundaries, inconsistent speaking rate across segments, and loss of prosodic patterns such as intonation shift. The literature shows three paradigms for extending genera- tion length, each with distinct limitations. Sequence compres- sion approaches reduce the number of tokens per second of au- dio to fit more speech within fixed context windows. VibeVoice [12] achieves a compression at 7.5 Hz, enabling up to 90 min- utes of speech generation in a single pass, however it sacrifices temporal resolution, representing each∼ 133ms of audio with a single token. SpeechSSM [13] uses state-space models for theoretically infinite extrapolation. Streaming and block-wise methods such as CosyVoice 2 [14] employ block-wise attention masks to enable chunk-wise generation with bounded memory. However, these binary masks create hard information cutoffs rather than graceful gradients. Moreover, all streaming ap- proaches require training time architectural modifications; the chunking strategy is baked into the model. Cross-sentence prosody models like HiGNN-TTS [15], and Context-Aware Memory [16] demonstrate that text context from neighboring sentences improves prosodic naturalness. Yet these approaches require dedicated graph networks, or memory modules trained jointly with the TTS system, making these improvements diffi- cult to apply on existing deployments. We present long-form speech generation algorithm MagpieTTS-LF 1 , addressing four converging gaps in the literature. First, our approach operates entirely at inference time, requiring no architectural changes or retraining existing MagpieTTS model. Second, we introduce soft attention priors that maintain non-zero weights on past and future tokens, preserving a gradient of attention, which does not completely suppress distant context. Our soft priors allow the model to retain long-range information while still focusing computation on locally relevant positions.Third, our stateful inference algorithm maintains attention prior states and encoder context across independently generated chunks, creating continuity between segments. Fourth, we leverage past text history by passing them through text encoder for prosodic planning, utilizing the model’s native text representations rather than requiring specialized context modules. Our key contributions can be summarized as follows: • We propose an inference-time soft attention prior mechanism that guides the auto-regressive generation toward monotonic alignment while preserving useful context from distant to- kens. Unlike binary masking approaches, our priors main- tain non-zero weights on past and future positions, enabling graceful information decay rather than hard cutoffs. • We introduce a stateful chunk generation algorithm that car- ries attention prior states, encoder hidden states, and text history across sentence boundaries. This enables coherent prosody in long-form synthesis without the memory over- head of single-pass generation or the discontinuities of naive chunk concatenation. • We demonstrate that inference-time soft priors combined with cross-chunk state propagation achieve significant im- provements in long-range prosodic coherence, speaker con- sistency over multi-minute durations, and boundary natural- ness. • We present a structured comparison against state-of-the-art models, spanning neural codec language models, large-scale multimodal architectures, and streaming diffusion frame- works. • We introduce Long-Form HifiTTS dataset, a benchmark 1 Open-source code https://github.com/NVIDIA-NeMo/NeMo arXiv:2606.18485v1 [cs.SD] 16 Jun 2026 Figure 1: Stateful chunk generation for sentence s i . History tokens H i are prepended to s i to form encoder input ̃ X i and the encoder output is concatenated with cached states H enc to produce ̃ H i . A soft attention prior encourages monotonic alignment during decoding while preserving long-range context across chunk boundaries. dataset for evaluating long-form speech synthesis, designed to measure prosodic continuity, speaker persistence, and boundary robustness under extended generation settings. 2. Methodology In this section, we present an inference-time approach for long- form speech generation that enables any chunk-based encoder- decoder TTS system to produce coherent speech without model retraining. In section 2.1, we briefly go over the MagpieTTS model architecture. In section 2.2, we describe the soft atten- tion priors that guide generation toward monotonic alignment while preserving context from distant tokens and section 2.3 talks about a stateful chunk generation algorithm that maintains attention states and encoder context across sentence boundaries that is at the core of this approach. 2.1. MagpieTTS Architecture The MagpieTTS model follows Koel-TTS [10], which is an encoder-decoder transformer architecture that operates on dis- crete audio tokens produced by a neural audio codec [17]. The encoder processes text through a stack of self-attention layers, producing contextual representations that guide audio genera- tion. The decoder auto-regressively generates discrete audio to- kens by attending to both the encoded text and any provided audio context for voice cloning. During training, MagpieTTS employs Connectionist Temporal Classification (CTC) loss and learned attention priors to enforce monotonic cross-attention between text and audio, preventing hallucinations such as word repetition, skipping, or misalignment. The standard model operates on single utterances; for longer inputs exceeding the generation time of 20 seconds, the text must be split into chunks and processed independently. This causes the cross-sentence coherence to be lost and bound- ary artifacts. 2.2. Inference-Time Soft Attention Priors We introduce soft attention priors that maintain non-zero weights across the full context while focusing on the locally relevant positions, this does not suppress distant positions com- pletely unlike binary attention masks. At each decoding step t, we compute (1) the text position T t that has the highest cross- attention score at decoding timestept−1, (2) a prior distribution P t ∈R N over encoder positions based on the expected mono- tonic alignment: P t [i] = w j ,if i∈T t − 1,T t ,T t + 1,T t + 2,T t + 3 P t [i] = eps,if i /∈T t − 1,T t ,T t + 1,T t + 2,T t + 3 Here i is the index of the text tokens, w= (w −1 ,w 0 ,w 1 ,w 2 ,w 3 ) are the experimentally determined fixed weight vector that defines the attention of the current segment of the text, and eps is the epsilon attention value for distant po- sitions. This ensures that distant encoder positions still receive non-zero weight, preserving gradual context rather than hard cutoffs. The modified attention ̃ A t becomes: ̃ A t = softmax Q t K ⊤ √ d + λ logP t where λ controls the prior strength and is experimentally determined. This formulation allows the model to leverage its learned attention patterns while being gently guided toward monotonic alignment. Unlike the binary masks, our soft pri- ors preserve information from distant tokens—useful for long- range dependencies 2.3. Stateful Chunk Generation Algorithm To generate long-form speech, we process long text in sentence- level chunks while maintaining information across chunk boundaries. The state comprises three components: • History Text Tokens H text ∈R B×K : The final K text tokens from the previous chunk, prepended to the current chunk’s input to provide linguistic context for prosodic plan- ning • History Encoder Context H enc ∈R B×K×d : The corre- sponding encoder hidden states for the history tokens, con- catenated with the current chunk’s encoder output to provide continuous representations across boundaries. The history text tokens are first appended, then the encoder outputs at the corresponding positions are discarded, and finally history encoder context is appended • Attention Tracking: A record of the last attended text posi- tions from the previous chunk, used to initialize the soft at- tention prior in the current chunk, ensuring smooth prosodic transitions. τ is the initial soft attention prior. The generation algorithm, also shown in Figure 1, works as follows. Given a long input text, we use punctuation-aware splitting to split into sentences S = s 0 ,s 1 ,...,s M . We ini- tialize an empty state and process each sentence iteratively: • Prepare Context: Concatenate history text tokens with current sentence tokens: ̃ X i = [H text ;s i ].Encode and concatenate with history encoder context: ̃ H i = [H enc ; Encoder(s i )]. • Prior-Guided Generation: Generate audio tokens auto- regressively using soft attention priors initialized from τ . The prior guides attention to start where the previous chunk ended, ensuring prosodic continuity. At every auto-regressive step the prior is updated to guide the model to maintain monotonicity. • Update and Maintain State: Save the final K tokens and their encoder states as new history. Record final attention prior weights as new τ . • Save Generated Code: Once end of speech token has been detected for the given sentence, save the generated codes in the current iteration. Once all the sentences have been gen- erated, these codes would be concatenated to get the entire audio code sequence for the total input text. 2.4. Long-Form HifiTTS dataset We construct a long-form evaluation benchmark concatenating paragraphs from Multilingual LibriSpeech (MLS) [18] to form 20 passages of approximately 3-4 minutes each. We estimate the duration by assuming 135 words per minute rate of speech in English. We normalize and clean the text resolving acronyms, numerical values. 3. Experiments and Results We evaluate MagpieTTS-LF 2 long-form generation against the state-of-the-art baselines across three dimensions: alignment robustness over long sequences, prosodic continuity at chunk boundaries, and speaker identity consistency, naturalness over long sequences. Together, these metrics capture the major chal- lenges of long-form synthesis identified in our literature review. 3.1. Evaluation Setup We evaluate on a curated dataset comprising of 20 long texts in English. We compare against Qwen3-TTS [8], VibeVoice- TTS [12], X-TTS [7] inference. All inference are run on single A6000 GPU. We use Whisper-Large [19] for ASR. For Long- form speech generation algorithm with MagpieTTS-LF we use the 0.1 as the value of eps, w = (0.2, 0.8, 1.0, 0.8, 0.2), tem- perature of 0.7, λ = 1.0 and cfg scale of 2.5. For VibeVoice and Qwen3-TTS, we synthesize speech either by segmenting text into sentences using punctuation-aware segmentation and concatenating the resulting audio, or by processing the text un- segmented up to the model’s maximum input length, and report whichever configuration performs better. 3.2. Intelligibility and Speaker Similarity To measure alignment robustness over time, we compute Word Error Rate (WER) and Character Error Rate (CER). We gen- erate transcriptions of the generated audio using Whisper- Large. We also compute speaker similarity (SSIM) by extract- ing speaker embeddings from reference speaker audio and gen- erated speech using Titanet [20] and WavLM [21], and then cal- culating cosine similarity. Results: As shown in Table 1, our proposed method in MagpieTTS-LF achieves significantly lower WER and CER compared to all other systems.While XTTS and Qwen3- TTS demonstrate moderate degradation, most likely due to er- ror accumulated at chunk boundaries and lack of past context, VibeVoice shows significant degradations. This might suggest that extreme token compression might lead to low intelligibil- ity over extended sequences. Through these results, we can see that our inference-time stateful approach maintains robustness 2 Website: https://magpietts-lf.github.io/ Table 1: Intelligibility evaluation on the long-form HiFiTTS 1- hour subset. Word Error Rate (WER) and Character Error Rate (CER) are computed using Whisper-Large. Lower values indi- cate higher intelligibility. ModelWER↓CER↓SSIM (TitaNet)↑SSIM (WavLM)↑ MagpieTTS-LF0.0250.0120.79± 0.020.979± 0.002 XTTS [7]0.0510.0350.69± 0.060.929± 0.042 Qwen3-TTS [8]0.0450.0280.80± 0.090.958± 0.025 VibeVoice[12]0.1150.1050.53± 0.150.848± 0.162 Table 2: Prosodic Boundary Deviation (PBD) metrics across models. Lower values indicate better cross-sentence prosodic coherence. Composite is calculated by taking an average of ∆ F0 and ∆ Energy. Model∆ F0 (Hz)↓∆ Energy (dB)↓Composite↓ MagpieTTS-LF69.1914.040.4646 XTTS [7]67.1330.620.734 Qwen3-TTS [8]65.5417.910.5169 VibeVoice [12]69.0828.900.712 across long sequences while models relying simply on naive chunking and aggressive compression compound errors as se- quence length increases. We also find that SSIM (WavLM) is the highest for MagpieTTS-LF, Qwen3-TTS is a close second, and XTTS suf- fers from the naive chunking and concatenation leads to in- consistent speaker characteristics across boundaries. VibeVoice suffer from low SSIM showing that token compression might lead to loss of speaker characteristics. Although Qwen3-TTS achieves the highest SSIM (TitaNet) score, its margin over MagpieTTS-LF is not statistically significant, so we do not dis- cuss this difference further. 3.3. Prosodic Boundary Discontinuity (PBD) To quantify prosodic continuity at chunk boundaries, we ex- tract F0 and energy in±1000ms regions around each sentence boundary and compute: (1) ∆ F0 (Hz jump) - absolute differ- ence in mean F0 (Hz) before and after sentence boundaries, and (2) ∆ Energy (dB) - energy discontinuity before and after sen- tence boundaries. We aggregate across all boundaries per pas- sage and min-max normalize each of the aggregated metrics. Results: From Table 2 we find that our method achieves the best prosodic continuity, with boundary energy discontinuity of just 14.04 dB—roughly half that of competing models. While F0 jumps remain comparable across systems (67–69 Hz), en- ergy consistency proves the dominant differentiator. XTTS, de- spite smooth pitch transitions, suffers from severe loudness in- consistency due to independent per-chunk gain normalization. Qwen3-TTS occupy the best middle ground with smooth tran- sition in both pitch and energy. Our method’s advantage stems from stateful inference maintaining coherence across chunks, which is the most perceptually salient aspect of boundary arti- facts. 3.4. Naturalness and Speaker Consistency To measure speaker consistency, we first divide the generated audio into non-overlapping chunks of 10 seconds each. We ex- tract Titanet [20] and WavLM embeddings [21] from each of the chunks and compute cosine similarity with reference speaker Figure 2: Speaker similarity (TitaNet, top; WavLM, bottom) across relative position in long-form utterances. Shaded regions denote standard deviation. MagpieTTS-LF maintains the most stable similarity throughout generation, while other models exhibit higher variance and drift over sequence length. Figure 3: Plot of UTMOSv2 scores vs relative position in long-form utterances. Shaded regions denote standard deviation. MagpieTTS- LF achieves the highest quality with consistent scores throughout generation out-performing other baselines in UTMOSv2 score. audio. We plot the speaker similarity of these chunks against their relative positions in the long text. For example, a 10 sec- ond chunk close to the starting of the speech has a relative po- sition closer to 0 on the x-axis and one closer to the end of the speech is closer to 1. To measure naturalness consistency we measure UTMOSv2 [22] from 10 second windows and visual- ize the change in UTMOSv2 score relative to position in longer sequences. Results: As shown in Figure 2, MagpieTTS-LF maintains the highest and most stable speaker similarity throughout the sequence for both TitaNet and WavLM similarity. It also shows minimal variance and no observable drift from start to end, in- dicating that speaker characteristics are preserved across chunk boundaries. All competing models exhibit either higher vari- ance, or drift showing speaker characteristics were not pre- served across boundaries and over long distances. Interest- ingly, VibeVoice shows a downward trend, indicating incon- sistent speaker representation despite single-pass generation. Magpietts-LF has the highest UTMOSv2 and lowest variance, which shows that MagpieTTS-LF produces the most natural and stable audio. VibeVoice degrades the most over the course of long-form synthesis, XTTS produces the least natural sounding speech. 4. Conclusion We present MagpieTTS-LF, an inference-time approach to syn- thesize robust, coherent and natural sounding long-form speech without retraining on long-form data. Our method uses soft attention prior to guide monotonicity, a history-aware stateful chunk generation that helps maintain prosodic continuity and speaker consistency over the entirety of the generated speech. We also demonstrate that MagpieTTS-LF achieves much lower word error rates than competing systems, the lowest prosodic boundary discontinuity, and stable speaker similarity through- out generation compared to the state-of-the-art long-form TTS systems. MagpieTTS-LF also maintains the highest and most stable naturalness. We show these by experimenting with mul- tiple dimensions of generated speech. By operating entirely at inference time, our approach can be extended to any chunk- based encoder-decoder TTS system for long-form synthesis. 5. Generative AI Use Disclosure Generative AI was used for checking grammar and spelling of the entire paper. It was very minimally used to refine the lan- guage at some parts of the paper and with L A T E Xsyntax. 6. References [1] J.Betker,“Betterspeechsynthesisthroughscaling,” https://github.com/neonbjb/tortoise-tts, 2023. [2] S. Chen, S. Liu, L. Zhou, Y. Liu, X. Tan, J. Li, S. Zhao, Y. Qian, and F. Wei, “VALL-E 2: Neural codec language models are hu- man parity zero-shot text to speech synthesizers,” arXiv preprint arXiv:2406.05370, 2024. [3] C. Zhang, S. Wang et al., “Vall-e r: Robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,” arXiv preprint arXiv:2406.07855, 2024. [Online]. Available: https://arxiv.org/abs/2406.07855 [4] K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian, “Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,” in ICLR 2024, April 2023. [5] Z. Ju, Y. Wang, K. Shen, X. Tan, D. Xin, D. Yang, Y. Liu, Y. Leng, K. Song, S. Tang, Z. Wu, T. Qin, X.-Y. Li, W. Ye, S. Zhang, J. Bian, L. He, J. Li, and S. Zhao, “Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,” in ICML, June 2024. [6] M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, and W.-N. Hsu, “Voicebox:Text-guided multilingual universal speech generation at scale,” in Advances in Neural Informa- tion Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36.Curran Associates, Inc., 2023, p. 14 005–14 034. [Online]. Avail- able:https://proceedings.neurips.c/paper files/paper/2023/file/ 2d8911db9ecedf866015091b28946e15-Paper-Conference.pdf [7] E. Casanova, C. Shulby, A. Aljafari, J. Meyer, R. Morais, S. Olayemi, and J. Weber, “XTTS: a massively multilingual zero- shot text-to-speech model,” in Proc. Interspeech, 2024. [8] H. Hu et al., “Qwen3-TTS technical report,” arXiv preprint arXiv:2601.15621, 2026. [9] P. Neekhara, S. Hussain, S. Ghosh, J. Li, R. Valle, R. Badlani, and B. Ginsburg, “Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment,” in Proc. INTER- SPEECH 2024, 2024, p. –. [10] S. S. Hussain, P. Neekhara, X. Yang, E. Casanova, S. Ghosh, R. Fejgin, M. T. Desta, R. Valle, and J. Li, “Koel-TTS: Enhancing LLM based speech generation with preference alignment and classifier free guidance,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds.Suzhou, China: Association for Computational Linguistics, Nov. 2025. [Online]. Available: https://aclanthology. org/2025.emnlp-main.1076/ [11] Z. Du, Q. Chen, S. Zhang, K. Hu, H. Lu, Y. Yang, H. Hu, S. Zheng, Y. Gu, Z. Ma et al., “CosyVoice: A scalable multi- lingual zero-shot text-to-speech synthesizer based on supervised semantic tokens,” arXiv preprint arXiv:2407.05407, 2024. [12] Z. Peng, J. Yu, W. Wang, Y. Chang, Y. Sun, L. Dong, Y. Zhu, W. Xu, H. Bao, Z. Wang et al., “VibeVoice technical report,” arXiv preprint arXiv:2508.19205, 2025. [13] K. Miyazaki, M. Murata, and T. Kosaka, “Structured state space decoder for speech recognition and synthesis,” in APSIPA Annual Summit and Conference, 2022. [14] Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang et al., “CosyVoice 2: Scalable stream- ing speech synthesis with large language models,” arXiv preprint arXiv:2412.10117, 2024. [15] D. Guo, X. Zhu, L. Xue, T. Li, Y. Lv, Y. Jiang, and L. Xie, “HiGNN-TTS: Hierarchical prosody modeling with graph neural networks for expressive long-form TTS,” in Proc. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023, p. 1–7. [16] Z. Li, X. Xing, J. Xing, H. Hu, H. Lu, and X. Xu, “Long-Context Speech Synthesis with Context-Aware Memory,” in Interspeech 2025, 2025, p. 2455–2459. [17] E. Casanova, P. Neekhara, R. Langman, S. Hussain, S. Ghosh, X. Yang, A. Jukic, J. Li, and B. Ginsburg, “Nanocodec: Towards high-quality ultra fast speech llm inference,” in Proc. Interspeech 2025, 2025. [18] V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Proc. Interspeech, 2020, p. 2757–2761. [19] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” in Proc. International Conference on Machine Learn- ing (ICML), 2023. [20] N. R. Koluguri, T. Park, and B. Ginsburg, “TitaNet: Neural model for speaker representation with 1D depth-wise separable convolu- tions and global context,” in ICASSP 2022 – IEEE International Conference on Acoustics, Speech and Signal Processing, 2022, p. 8102–8106. [21] S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, p. 1505–1518, Oct. 2022. [Online]. Available: http://dx.doi.org/10.1109/JSTSP.2022.3188113 [22] T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab system for Voice- MOS challenge 2022,” in Proc. Interspeech, 2022, p. 4521– 4525.