Paper deep dive
Diffusion Language Models for Speech Recognition
Davyd Naveriani, Albert Zeyer, Ralf Schlüter, Hermann Ney
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/18/2026, 1:31:43 AM
Summary
This paper explores the application of discrete diffusion language models, specifically Masked Diffusion Language Models (MDLM) and Uniform-State Diffusion Models (USDM), for Automatic Speech Recognition (ASR). The authors introduce novel rescoring strategies for MDLM and a joint-decoding framework that integrates USDM with Connectionist Temporal Classification (CTC) to improve ASR accuracy by leveraging bidirectional context and parallel generation.
Entities (5)
Relation Signals (3)
USDM → integratedwith → CTC
confidence 100% · we design a new joint-decoding method that combines CTC and USDM
MDLM → usedfor → ASR Rescoring
confidence 95% · we systematically investigate diffusion language models for ASR rescoring, comparing masked diffusion language models (MDLM)
MDLM → improves → ASR
confidence 90% · Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporating masked diffusion language models (MDLM) and uniform-state diffusion models (USDMs) for rescoring ASR hypotheses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise probability distributions derived from CTC with the labelwise probability distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong language knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We publish all our code and recipes.
Tags
Links
- Source: https://arxiv.org/abs/2604.14001v1
- Canonical: https://arxiv.org/abs/2604.14001v1
Trouble viewing inline? Open PDF directly →
Full Text
27,969 characters extracted from source content.
Expand or collapse full text
Diffusion Language Models for Speech Recognition Davyd Naveriani 1,∗ , Albert Zeyer ID 1,2,∗ , Ralf Schlüter ID 1,2 , Hermann Ney 1,2 1 Machine Learning and Human Language Technology Group, RWTH Aachen University, Germany 2 AppTek, Germany davyd.naveriani@rwth-aachen.de, zeyer,schlueter,ney@cs.rwth-aachen.de Abstract Diffusion language models have recently emerged as a lead- ing alternative to standard language models, due to their ability for bidirectional attention and parallel text generation. In this work, we explore variants for their use in speech recognition. Specifically, we introduce a comprehensive guide to incorporat- ing masked diffusion language models (MDLM) and uniform- state diffusion models (USDMs) for rescoring ASR hypothe- ses. Additionally, we design a new joint-decoding method that combines CTC and USDM by integrating the framewise proba- bility distributions derived from CTC with the labelwise proba- bility distributions computed by USDM at each decoding step, thereby generating new candidates that combine strong lan- guage knowledge from USDM and acoustic information from CTC. Our findings reveal that USDM, as well as MDLM, can significantly improve the accuracy of recognized text. We pub- lish all our code and recipes. Index Terms: Speech Recognition, Diffusion Language Model 1. Introduction Autoregressive language models (LMs) are commonly used to improve automatic speech recognition (ASR) systems due to their strong linguistic capabilities and the ability to incorpo- rate external textual knowledge [1–3]. However, applying tra- ditional autoregressive LMs in joint decoding inherently limits the speed due to their strictly left-to-right decoding structure. Non-autoregressive language models [4] can operate in a par- allel manner, potentially resolving this bottleneck and enabling faster decoding. Recently, discrete diffusion models, such as large language diffusion with masking (LLaDA) [5] and masked diffusion lan- guage models (MDLM) [6] have emerged as powerful non- autoregressive alternatives [5–10]. Prior work has explored diffusion-based models in ASR as audio-conditioned decoders [11–13]. Other non-autoregressive speech recognition models have been proposed [14–16]. Here, we focus on keeping a separate language model to better leverage large amounts of text data and to allow more flexible integration of the language model into the ASR sys- tem [17, 18]. An investigation of diffusion language models as standalone models for joint ASR decoding has not yet been conducted and their integration into token-level joint decoding remains unexplored. In this work, we systematically investigate diffusion lan- guage models for ASR rescoring, comparing masked diffusion language models (MDLM) and uniform-state diffusion models * These authors contributed equally. (USDM). Furthermore, we develop a novel token-level decod- ing method that allows for the integration of USDM with a non-autoregressive connectionist temporal classification (CTC) speech recognition model [19]. Because USDM corrupts se- quences using uniform transitions without artificial mask to- kens, it provides a full vocabulary probability distribution for every token at each denoising step, which enables a direct com- bination of framewise CTC probabilities and labelwise diffu- sion distributions during hypothesis construction, as detailed in Section 2 and Section 3. To the best of our knowledge, this is the first work that (i) systematically compares masked and uniform-state diffusion language models for ASR rescor- ing, and (i) integrates a uniform-state diffusion language model into CTC-based token-level joint decoding. 2. Diffusion Language Models Masked diffusion language model. MDLM corrupts text by randomly masking tokens and learns to reconstruct the sequence during the reverse generative pass. During the forward process, tokens are independently masked based on a monotonically decreasing noise schedule α t ∈ [0, 1]. Essentially, α t represents the probability of a token retaining its original value, while (1− α t ) is the probability of it being masked. As the process reaches the final step T , α T ap- proaches 0, meaning the sequence becomes completely masked with probability 1. The marginal distribution of this forward process is defined as: q(z t | w) = Cat(z t ; α t 1 w + (1− α t )1 m )(1) where Cat is the categorical distribution, w is the original clean token, z t ∈ V is the token at diffusion step t, 1 w ∈ 0, 1 |V | denotes the one-hot vector with a 1 at position w ∈ V and indi- cates a clean token, and m is the index of the [MASK] token. During the reverse process, starting from a fully masked sequence, an MDLM iteratively denoises the text to recover the original tokens. The theoretical objective is to align the reverse transition with the true posterior distribution, which is defined as: q(z s | z t ,w) = Cat(z s ; 1 z t )z t ̸= m, Cat z s ; (1− α s )1 m + (α s − α t )1 w 1− α t z t = m. (2) However, because the target token w is unknown during gen- eration, a parameterized model w θ (z t ,t) is trained to directly predict the best estimate of the unmasked tokens from the noisy arXiv:2604.14001v1 [cs.CL] 15 Apr 2026 state z t . This implies that the training objective is computed exclusively over the masked tokens, taking the form of a cross- entropy loss weighted by the noise schedule [5, 6]. The bidirec- tional context modeling of MDLM makes it particularly well- suited for ASR hypothesis rescoring [11, 12]. Uniform-state diffusion model. USDM works similarly to MDLM but uses a different corruption strategy. During the forward process, tokens are replaced with ran- dom samples from the vocabulary rather than a mask to- ken. The marginal distribution is defined as q(z t | w) = Cat(z t ; α t 1 w + (1− α t )π), whereπ = 1 |V | 1 is the uniform distribution over the vocabulary V [20]. During the reverse process, these forward dynamics allow for continual token updates. Since corrupted tokens are indis- tinguishable from clean ones, the model must re-evaluate ev- ery position at each denoising step, enabling a self-correcting property where previously predicted tokens can be updated or fixed [7, 21–23]. Consequently, at each denoising step, the model w θ (z t ,t) produces a full probability distribution over the entire vocabulary for every token in the sequence, regardless of its current state. For ASR, this dense output is particularly ad- vantageous as it provides a continuous stream of vocabulary- wide probabilities that can be directly aligned and combined with frame-wise CTC scores during joint decoding. 3. Methodology 3.1. Rescoring We rescoren-best CTC hypotheses ̃a ̃ S 1 = ( ̃a 1 ,..., ̃a ̃ S ) by com- bining the CTC log-probability with a diffusion language model (DiffLM) score and a prior correction term: S( ̃a ̃ S 1 ) = λ CTC logP CTC ( ̃a ̃ S 1 | x T 1 ) + λ DiffLM logP DiffLM ( ̃a ̃ S 1 ) − λ prior logP prior ( ̃a ̃ S 1 )(3) where x T 1 denotes the sequence of T acoustic feature frames, λ CTC , λ DiffLM , and λ prior are tunable interpolation weights, and the prior term compensates for the implicit language model bias learned by the CTC model [24, 25]. Since logP DiffLM ( ̃a ̃ S 1 ) is intractable, we approximate it via K Monte Carlo samples by applying forward noise and computing the reconstruction log- likelihood. MDLM. A naive sequence-length normalization estimate av- erages over the full sequence length ̃ S: logP DiffLM ( ̃a ̃ S 1 )≈ 1 K K X k=1 1 ̃ S X j∈M k logP θ ( ̃a j | z t,k )(4) where t is a fixed noise step that determines the masking prob- ability, M k is the set of randomly masked positions in the k- th sample, ̃a j is the j-th token of the hypothesis, and z t,k de- notes the noisy sequence obtained by masking positions M k with masking probability (1 − α t ). This overweights heav- ily corrupted samples—steps with many masks contribute many log-prob terms, dominating the score regardless of model con- fidence. We propose three alternative scoring strategies to ad- dress this issue. Sample-level mask normalization. We normalize each sam- ple by its own mask count before averaging over K: logP DiffLM ( ̃a ̃ S 1 )≈ 1 K K X k=1 1 |M k | X j∈M k logP θ ( ̃a j | z t,k ) (5) Global mask normalization. We pool all masked predictions across all K samples and divide by the total mask count: logP DiffLM ( ̃a ̃ S 1 )≈ P K k=1 P j∈M k logP θ ( ̃a j | z t,k ) P K k=1 |M k | (6) Sample-level normalization weights each Monte Carlo sample equally, whereas global normalization weights each predicted token equally. As shown in Figure 3, they significantly improve WER over sequence-length normalization. Coupled scoring. Additionally, inspired by the coupled- sampling scheme from [26], we construct K pairs of comple- mentary masksM (1) k andM (2) k = M (1) k , such that every token is masked in exactly one of the two forward passes per pair: logP DiffLM ( ̃a ̃ S 1 ) ≈ 1 K K X k=1 1 ̃ S X j∈M (1) k logP θ ( ̃a j | z (1) t,k ) + X j∈M (2) k logP θ ( ̃a j | z (2) t,k ) ! (7) This guarantees that every token contributes to the score. USDM. We adopt the ELBO from [7] as the scoring objec- tive. Unlike MDLM, where only masked tokens contribute to the score, USDM corrupts positions with uniform noise and ev- ery token participates in the score regardless of the noise level. 3.2. Joint-Decoding USDM exhibits several properties that make it particularly suit- able for integration with CTC in a joint decoding framework (see Figure 1): its self-correcting nature allows the model to continuously refine all positions rather than committing to pre- dictions early; and since tokens are corrupted with uniform noise rather than a mask token, the model maintains a well- defined probability distribution over the full vocabulary at every position throughout the entire denoising process. These proper- ties motivate us to extend USDM to a novel joint CTC-USDM decoding framework. We initialize the denoising process from the CTC greedy sequence at noise level t start [13], which we treat as a tunable hyperparameter. Since CTC operates on frames while USDM operates on tokens, we align the two by extracting the log- probability distribution of the first frame corresponding to each collapsed token and renormalizing it over the non-blank vo- cabulary. At each denoising step l, and each token position i, USDM produces a token-level distribution P θ (· | z t l ), which we combine with the CTC distribution: logP comb,i (v j ) = λ CTC logP CTC,τ i (v j | x T 1 ) + λ DiffLM logP θ,i (v j | z t l )(8) Here, v j ∈ V denotes the j-th token from the vocabulary V . τ i denotes the first CTC frame aligned with token position i in 퐶표푛푓표푟푚푒푟 퐸푛푐표푑푒푟 Audio 풕풐풌풆풏 풎 풕풐풌풆풏 풎 퐷푖푓푢푠푖표푛 푀표푑푒푙 푈푛푖푓표푟푚−푆푡푎푡푒 풕풐풌풆풏 풎 풕풐풌풆풏 풌 풕풐풌풆풏 풌 풃풍풂풏풌 푧 # =푦 $%$&'())*+ 푡=푡 ,#-(# log푃 ./01,3 푣 4 =휆 CTC log푃 $%$,5 ! 푣 4 푥 6 % +휆 73889: log푃 ;,3 푣 4 푧 # " 퐼푛푝푢푡 log푃 $%$,5 # 푣 3 푥 6 % log푃 $%$,5 $ 푣 3 푥 6 % 푆푡푒푝 0 푦 !"#$% 푙=푙−1 푧 # "%# ,3 ∼푃 ./01,3 Figure 1: Overview of the proposed joint CTC-USDM decoding. At each denoising step, the USDM token-level distribution is combined with the CTC frame-level distribution to sample input to the next denoising step. the collapsed greedy sequence. The term P CTC,τ i (v j | x T 1 ) corresponds to the frame-level CTC probability of token v j at frame τ i , obtained from the encoder output and renormalized over the non-blank vocabulary. The term P θ,i (v j | z t l ) de- notes the token-level probability predicted by the USDM at po- sition i given the current noisy sequence z t l . The combined score P comb,i (v j ) therefore integrates acoustic evidence from the CTC model with contextual information from the diffusion language model during each denoising step. The next input z t l−1 is then sampled from this combined distribution. We use ancestral sampling from [20, 21, 27], which draws the next state directly from the USDM reverse posterior. 4. Experiments 4.1. Experimental Setup We trained MDLM and USDM on a combined corpus of nor- malized LibriSpeech LM data and train-other transcriptions [28]. For our experiments, we leveraged the training frame- works from [6, 7].Models were trained for 5, 10 and 25 epochs using AdamW (0.1 weight decay) [29], a piecewise lin- ear LR scheduler, and a 20,000 token batch size. Our primary "medium" architecture for both models is a 24-layer Diffusion Transformer (DiT) [30] with 16 attention heads and a 1024- dimensional hidden state. We also evaluated a "small" variant (12 layers, 12 heads, 768-dim, 0.1 dropout [31]). Both configu- rations used 128-dimensional diffusion time embeddings. Text was tokenized via SentencePiece into 10,240 subwords [32]. 4.2. Results Language model training. Table 1 shows the perplexity up- per bounds for USDM and MDLM trained with the same con- figuration. MDLM achieves lower PPL at 5 and 10 epochs (see Figure 2, e.g. 36.6 vs. 40.2 on dev at 5 epochs), but USDM surpasses it at 25 epochs (34.0 vs. 32.3 on dev). This can be explained by the fact that USDM corrupts tokens with uniform noise rather than explicit mask tokens, making the task inher- ently harder since the model must evaluate every position. Both models show improvement with longer training. Rescoring. Figure 3 compares MDLM and USDM rescoring strategies across varying numbers of Monte Carlo samples K. As illustrated in Figure 3, MDLM rescoring consistently outper- forms the CTC baseline (5.08% WER) and USDM. While stan- 020406080100 Subepoch 40 60 80 100 120 140 PPL USDM 5ep MDLM 5ep Figure 2: Dev PPL learning curves for MDLM and USDM trained for 5 full epochs. USDM learning curve is shown af- ter the curriculum learning phase. Table 1: Perplexity (PPL) upper bounds on train, dev, and de- vtrain splits for MDLM and USDM trained for 5, 10 and 25 epochs on LibriSpeech LM data. Diffusion LMNumber of Epochs PPL traindevdevtrain MDLM 5≤ 36.2≤ 36.6≤ 47.9 10≤ 32.7≤ 37.0≤ 44.6 25≤ 30.0≤ 34.0≤ 42.2 USDM 5≤ 40.5≤ 40.2≤ 59.5 10≤ 37.0≤ 39.4≤ 48.0 25≤ 33.5≤ 32.3≤ 44.1 Table 2: MDLM rescoring WER [%] on dev-other across 5, 10 and 25 training epochs, varying the number of Monte Carlo samples (K). Standard deviations computed over 5 random seeds. K WER [%] MDLM 5 epMDLM 10 epMDLM 25 ep 14.97± .034.94± .014.95± .01 24.94± .014.94± .024.91± .02 164.79± .034.78± .024.75± .03 324.72± .034.67± .024.65± .03 644.65± .044.60± .024.56± .02 1284.60± .024.58± .024.55± .02 2564.59± .024.56± .034.52± .01 dard sequence-length normalization achieves a WER of 4.73% (K = 256), our proposed global mask normalization yields a further reduction to 4.59%. Additionally, coupled scoring improves results with fewer samples, reaching 4.71% WER at K = 46. Furthermore, as shown in Table 2, extending MDLM training to 10 and 25 epochs provides additional gains, reaching a rescoring performance of 4.56% and 4.52% WER at K = 256. While USDM rescoring does not match the per- formance of MDLM, it still yields improvements over the CTC baseline. As shown in Table 3, USDM reduces the WER from 5.08% to 4.82% at K = 256. With an increased number of epochs, the WER decreased to 4.80% for 25 epochs. CTC-USDM Joint Decoding. As shown in Table 4, joint de- coding yields better results than USDM rescoring, suggesting that the active participation of USDM in hypothesis construc- 050100150200250 K 4.6 4.7 4.8 4.9 5.0 WER [%] MDLM Seq Level MDLM Global Mask USDM MDLM Coupled MDLM Sample Mask Figure 3: WER [%] on dev-other comparing MDLM rescor- ing (sequence-level, global-mask and sample-mask score nor- malization), MDLM (5 ep) coupled scoring, and USDM (5 ep) rescoring, across different numbers of Monte Carlo samples (K). Stars mark the best WER for each method. Table 3: USDM rescoring WER [%] on dev-other across 5, 10 and 25 training epochs, varying the number of Monte Carlo samples (K). Standard deviations computed over 5 random seeds. K WER [%] USDM 5 epUSDM 10 epUSDM 25 ep 14.99± .034.97± .024.97± .01 24.96± .034.99± .034.99± .02 164.94± .034.92± .024.92± .02 324.91± .024.89± .044.89± .03 644.87± .044.85± .044.86± .02 1284.85± .034.83± .024.80± .01 2564.82± .024.82± .024.80± .02 Table 4: WER [%] on dev-other for CTC+USDM joint decoding with different initial noise level t start , and denoising steps K (for all experiments, λ DiffLM = 0.3 shows the best performance). WER [%] K t start 181216324864 0.34.794.784.794.774.784.774.77 0.54.814.814.804.784.774.774.77 0.84.824.794.784.784.774.774.77 tion produces better hypotheses. A lower initial noise level (t start = 0.3) allows the model to reach optimal WER with fewer denoising steps, while all configurations converge to the same final WER. As shown in Table 5, extending USDM training to 10 and 25 epochs further improves joint decoding, reaching a peak WER of 4.73% and 4.71%, respectively. Comparison with Autoregressive LMs. Table 6 provides an overview of the results across the different modeling ap- proaches. As expected, the autoregressive language models achieve the lowest overall WER, with 4.19% for rescoring and 3.86% for joint decoding (Table 6). However, scaling behav- ior differs across architectures: for MDLM, increasing from Table 5: WER [%] on dev-other for CTC+USDM joint decoding comparing models trained on different numbers of epochs with λ DiffLM = 0.3 and t start = 0.3. WER [%] K Epochs181216324864 54.794.784.794.774.784.774.77 104.744.734.744.764.744.744.73 254.744.724.744.724.734.724.71 Table 6: WER [%] on dev-other comparing CTC without LM, with an autoregressive LM, with MDLM (10 ep, global mask norm., K = 256) and USDM (10 ep, K = 256 rescoring and K = 64 joint-decoding). LM Model Dim Num Layers WER [%] dev-other None–greedy5.08 Auto- regressive LM 76812 rescoring4.10 joint-decoding3.98 102424 rescoring4.19 joint-decoding3.86 MDLM 76812 rescoring 4.64 1024244.56 USDM102424 rescoring4.82 joint-decoding4.73 12 to 24 layers reduces rescoring WER from 4.64% to 4.56%, whereas for the autoregressive model, the same scaling de- grades rescoring performance from 4.10% to 4.19%, while still improving joint decoding from 3.98% to 3.86%. 5. Conclusions In this work, we systematically explored the integration of dis- crete diffusion language models into ASR systems. While tra- ditional autoregressive models are constrained by a strictly se- quential, left-to-right decoding structure, diffusion LMs lever- age bidirectional context and parallel generation, offering a more flexible and theoretically faster alternative for ASR. We introduced new methods to rescore ASR hypotheses using MDLM, namely Global Mask Normalization and Sample Mask Normalization. By utilizing the mask length for normalization, these methods significantly improved performance in compari- son to standard sequence-level normalization. Most notably, we noticed unique properties of Uniform-State Diffusion Models, specifically their lack of artificial mask tokens and their full- vocabulary probability distribution for each position, and de- veloped a CTC-USDM joint decoding framework, which suc- cessfully outperformed static rescoring with USDM. Evalua- tion shows MDLM achieves better rescoring accuracy on lim- ited data than USDM. This is likely because MDLM’s explicit mask tokens provide a clearer reconstruction signal, whereas USDM’s uniform noise forces the model to implicitly distin- guish clean from noisy tokens at every position. Consequently, USDM’s more complex objective may require additional data or scaling to reach peak performance. Additionally, we com- pared our methods with standard autoregressive language mod- els. While the autoregressive baselines achieved better results, increasing the model capacity significantly improved MDLM’s rescoring performance, whereas it degraded the rescoring re- sults of the autoregressive models. In future work, we plan to evaluate these models on larger datasets and scale the model capacities further to close the performance gap with autoregres- sive baselines. We also plan to investigate the joint decoding framework in greater detail, including extending it to MDLM. 6. Acknowledgements This work was partially supported by NeuroSys, which as part of the initiative “Clusters4Future” is funded by the Fed- eral Ministry of Education and Research BMBF (funding IDs 03ZU2106DA and 03ZU2106D), and by the project RESCALE within the program AI Lighthouse Projects for the Environment, Climate, Nature and Resources funded by the Federal Ministry for the Environment, Nature Conservation, Nuclear Safety and Consumer Protection (BMUV), funding ID: 67KI32006A. The authors gratefully acknowledge the com- puting time provided to them at the NHR Center NHR4CES at RWTH Aachen University (project number p0023565 and p0023999). This is funded by the Federal Ministry of Edu- cation and Research, and the state governments participating on the basis of the resolutions of the GWK for national high performance computing at universities (w.nhr-verein. de/unsere-partner). 7. Generative AI Use Disclosure We use LLMs to improve the formulations and grammar of the paper. 8. References [1] F. Jelinek, Statistical Methods for Speech Recognition.MIT press, 1998. [2] K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Language Modeling with Deep Transformers,” in Interspeech, Graz, Austria, Sep. 2019, p. 3905–3909, iSCA Best Student Paper Award. [slides]. [Online]. Available: http://arxiv.org/pdf/1905.04226.pdf [3] R. Prabhavalkar, T. Hori, T. N. Sainath, R. Schlüter, and S. Watanabe, “End-to-End Speech Recognition: A Survey,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, vol. 32, p. 325–351, 2023. [4] Y. Su, D. Cai, Y. Wang, D. Vandyke, S. Baker, P. Li, and N. Col- lier, “Non-Autoregressive Text Generation with Pre-trained Lan- guage Models,” in Proceedings of the 16th Conference of the Eu- ropean Chapter of the Association for Computational Linguistics: Main Volume, 2021, p. 234–243. [5] S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J.-R. Wen, and C. Li, “Large Language Diffusion Models,” in The Thirty-ninth Annual Conference on Neural Information Process- ing Systems, 2025. [6] S. S. Sahoo, M. Arriola, A. Gokaslan, E. M. Marroquin, A. M. Rush, Y. Schiff, J. T. Chiu, and V. Kuleshov, “Simple and Effec- tive Masked Diffusion Language Models,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [7] S. S. Sahoo, J. Deschenaux, A. Gokaslan, G. Wang, J. T. Chiu, and V. Kuleshov, “The Diffusion Duality,” in Forty-second Inter- national Conference on Machine Learning, 2025. [8] T. Li, M. Chen, B. Guo, and Z. Shen, “A Survey on Diffusion Language Models,” arXiv preprint arXiv:2508.10875, 2025. [9] J. Ni, Q. Liu, L. Dou, C. Du, Z. Wang, H. Yan, T. Pang, and M. Q. Shieh, “Diffusion Language Models are Super Data Learn- ers,” arXiv preprint arXiv:2511.03276, 2025. [10] S. Khanna, S. Kharbanda, S. Li, H. Varma, E. Wang, S. Birn- baum, Z. Luo, Y. Miraoui, A. Palrecha, S. Ermon, A. Grover, and V. Kuleshov, “Mercury: Ultra-Fast Language Models Based on Diffusion,” arXiv preprint arXiv:2506.17298, 2025. [11] M. Wang, Z. Liu, Z. Jin, G. Sun, C. Zhang, and P. C. Woodland, “Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing,” arXiv preprint arXiv:2509.16622, 2025. [12] T. Kwon, J. Ahn, T. Yun, H. Jwa, Y. Choi, S. Park, N.-J. Kim, J. Kim, H. G. Ryu, and H.-J. Lee, “Whisfusion: Paral- lel ASR Decoding via a Diffusion Transformer,” arXiv preprint arXiv:2508.07048, 2025. [13] W. Tian, B. Mu, G. Ma, X. Geng, Z. Zhao, and L. Xie, “dLLM- ASR: A Faster Diffusion LLM-based Framework for Speech Recognition,” arXiv preprint arXiv:2601.17902, 2026. [14] Y. Higuchi, N. Chen, Y. Fujita, H. Inaguma, T. Komatsu, J. Lee, J. Nozaki, T. Wang, and S. Watanabe, “A Comparative Study on Non-autoregressive Modelings for Speech-to-Text Generation,” in IEEE Automatic Speech Recognition and Understanding Work- shop (ASRU), 2021, p. 47–54. [15] K. Deng, Z. Yang, S. Watanabe, Y. Higuchi, G. Cheng, and P. Zhang, “Improving Non-Autoregressive End-to-End Speech Recognition with Pre-trained Acoustic and Language Models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, p. 8522–8526. [16] A. Navon, A. Shamsian, N. Glazer, Y. Segal-Feldman, G. Hetz, J. Keshet, and E. Fetaya, “Drax: Speech Recognition with Dis- crete Flow Matching,” arXiv preprint arXiv:2510.04162, 2025. [17] C. Gulcehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, F. Bougares, H. Schwenk, and Y. Bengio, “On Using Monolin- gual Corpora in Neural Machine Translation,” Computer Speech & Language, vol. 45, p. 137–148, 2015. [18] S. Toshniwal, A. Kannan, C.-C. Chiu, Y. Wu, T. N. Sainath, and K. Livescu, “A Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognition,” in IEEE Spoken Language Technology Workshop (SLT), 2018, p. 369– 375. [19] A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Con- nectionist Temporal Classification: Labelling Unsegmented Se- quence Data with Recurrent Neural Networks,” in Twenty-third International Conference on Machine Learning, 2006, p. 369– 376. [20] J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. van den Berg, “Structured Denoising Diffusion Models in Discrete State- Spaces,” in The Thirty-fifth Annual Conference on Neural Infor- mation Processing Systems, 2021. [21] J. Deschenaux, C. Gulcehre, and S. S. Sahoo, “The Diffusion Duality, Chapter I: Ψ-Samplers and Efficient Curriculum,” in The Fourteenth International Conference on Learning Represen- tations, 2026. [22] D. von Rütte, A. Orvieto, J. Fluri, O. Pooladzandi, B. Schölkopf, and T. Hofmann, “Scaling Behavior of Discrete Diffusion Lan- guage Models,” in The Fourteenth International Conference on Learning Representations, 2026. [23] S. S. Sahoo, J.-M. Lemercier, Z. Yang, J. Deschenaux, J. Liu, J. Thickstun, and A. Jukic, “Scaling Beyond Masked Diffusion Language Models,” arXiv preprint arXiv:2602.15014, 2026. [24] A. Zeyer, A. Merboldt, W. Michel, R. Schlüter, and H. Ney, “Lib- rispeech Transducer Model with Internal Language Model Prior Correction,” in Interspeech, 2021, p. 2052–2056. [25] Z. Meng, S. Parthasarathy, E. Sun, Y. Gaur, N. Kanda, L. Lu, X. Chen, R. Zhao, J. Li, and Y. Gong, “Internal Language Model Estimation for Domain-Adaptive End-to-End Speech Recogni- tion,” in IEEE Spoken Language Technology Workshop (SLT), 2021, p. 243–250. [26] S. Gong, R. Zhang, H. Zheng, J. Gu, N. Jaitly, L. Kong, and Y. Zhang, “DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation,” in The Fourteenth Inter- national Conference on Learning Representations, 2026. [27] A. Campbell, J. Benton, V. D. Bortoli, T. Rainforth, G. Deligian- nidis, and A. Doucet, “A Continuous Time Framework for Dis- crete Denoising Models,” in The Thirty-sixth Annual Conference on Neural Information Processing Systems, 2022. [28] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: An ASR Corpus Based on Public Domain Audio Books,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, p. 5206–5210. [29] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regular- ization,” in The Seventh International Conference on Learning Representations, 2019. [30] W. Peebles and S. Xie, “Scalable Diffusion Models with Trans- formers,” in IEEE/CVF International Conference on Computer Vision (ICCV), 2023, p. 4172–4182. [31] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Re- search, vol. 15, no. 56, p. 1929–1958, 2014. [32] T. Kudo and J. Richardson, “SentencePiece: A Simple and Lan- guage Independent Subword Tokenizer and Detokenizer for Neu- ral Text Processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2018, p. 66–71.