Paper deep dive
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding
Jonathan Lys, Vincent Gripon, Bastien Pasdeloup, Axel Marmoret, Lukas Mauch, Fabien Cardinaux, Ghouthi Boukli Hacene
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 6:13:11 AM
Summary
D5P4 is a novel decoding algorithm for discrete diffusion models that addresses the challenge of in-batch diversity by formulating candidate selection as MAP inference over a Partition Determinantal Point Process (DPP). By leveraging a scalable greedy solver, D5P4 enables an explicit trade-off between model probability and target diversity with minimal computational overhead, outperforming standard beam search and existing diversity-promoting baselines in text generation and question answering tasks.
Entities (5)
Relation Signals (3)
D5P4 â uses â Determinantal Point Process
confidence 100% ¡ D5P4, which formulates the selection step as MAP inference over a Determinantal Point Process.
D5P4 â improves â in-batch diversity
confidence 95% ¡ Experimental results show consistent improvements in diversity over strong decoding baselines
MDLM â isa â Discrete Diffusion Model
confidence 95% ¡ We review the discrete diffusion framework underlying MDLM
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Discrete diffusion models are promising alternatives to autoregressive approaches for text generation, yet their decoding methods remain under-studied. Standard decoding methods for autoregressive models, such as beam search, do not directly apply to iterative denoising, and existing diffusion decoding techniques provide limited control over in-batch diversity. To bridge this gap, we introduce a generalized beam-search framework for discrete diffusion that generates candidates in parallel and supports modular beam-selection objectives. As a diversity-focused instantiation, we propose D5P4, which formulates the selection step as MAP inference over a Determinantal Point Process. Leveraging a scalable greedy solver, D5P4 maintains multi-GPU compatibility and enables an explicit trade-off between model probability and target diversity with near-zero compute overhead. Experiments on free-form generation and question answering demonstrate that D5P4 improves diversity over strong baselines while maintaining competitive generation quality.
Tags
Links
- Source: https://arxiv.org/abs/2603.19146v1
- Canonical: https://arxiv.org/abs/2603.19146v1
Trouble viewing inline? Open PDF directly â
Full Text
57,049 characters extracted from source content.
Expand or collapse full text
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Jonathan Lys 1 Vincent Gripon 1 Bastien Pasdeloup 1 Axel Marmoret 1 Lukas Mauch 2 Fabien Cardinaux, 2 Ghouthi Boukli Hacene 2 Abstract Discrete diffusion models are promising alterna- tives to autoregressive approaches for text gener- ation, yet their decoding methods remain under- studied. Standard decoding methods for autore- gressive models, such as beam search, do not directly apply to iterative denoising, and exist- ing diffusion decoding techniques provide lim- ited control over in-batch diversity. To bridge this gap, we introduce a generalized beam-search framework for discrete diffusion that generates candidates in parallel and supports modular beam- selection objectives. As a diversity-focused in- stantiation, we propose D5P4, which formulates the selection step as MAP inference over a De- terminantal Point Process. Leveraging a scal- able greedy solver, D5P4 maintains multi-GPU compatibility and enables an explicit trade-off between model probability and target diversity with near-zero compute overhead. Experiments on free-form generation and question answering demonstrate that D5P4 improves diversity over strong baselines while maintaining competitive generation quality. 1 1. Introduction Discrete diffusion models have recently emerged as a competitive alternative to autoregressive language models (ARMs), achieving strong performance on a range of gen- erative tasks (Sahoo et al., 2024; Nie et al., 2025; Ye et al., 2025). By refining sequences in parallel through iterative denoising, they depart from left-to-right decoding but intro- duce new structural inference challenges. In practice, discrete diffusion models rely on simple sam- 1 IMT Atlantique, Lab-STICC, UMR CNRS 6285, F-29238 Brest, France 2 Sony Europe Ltd. Stuttgart Technology Cen- ter, EUREC, Germany.Correspondence to: Jonathan Lys <jonathan.lys@imt-atlantique.fr>. Preprint. March 20, 2026. 1 https://github.com/jonathanlys01/d5p4 pling procedures, whereas ARMs benefit from mature de- coding algorithms such as beam search to obtain high- quality outputs (Lowerre, 1976). An analogous decoding paradigm for diffusion models remains largely unexplored, as the non-monotonic and parallel nature of diffusion trajec- tories prevents the application of ARM-specific algorithms. Beyond this algorithmic gap, recent evidence highlights growing limitations in output diversity. In ARMs, Yue et al. (2025) shows that reinforcement learning and supervised fine-tuning substantially improve top-1 performance while saturating output coverage (pass@k), indicating reduced diversity. Similarly, in conditioned diffusion models, strong guidance sharpens fidelity at the expense of diversity (Sadat et al., 2024). Together, these observations motivate decod- ing algorithms that explicitly reason about in-batch diversity rather than relying on independent sequence scores. To address these challenges, we introduce Partition Deter- minantal Point Processes for Diversity in Parallel Discrete Diffusion Decoding (D5P4), a novel decoding algorithm tailored for the iterative and parallel nature of discrete diffu- sion. Rather than scoring candidates independently, D5P4 performs set-level selection, modeling interactions among hypotheses to balance generation quality and diversity. At each diffusion step, D5P4 formulates candidate selec- tion as sampling from a Determinantal Point Process (DPP), a probabilistic model that favors high-quality candidates while repelling similar ones (Kulesza et al., 2012). To com- plement the diversity induced by the DPP objective, we introduce a structural partition constraint on candidate selec- tion that prevents lineage collapse, a phenomenon in which hypotheses degenerate toward a single ancestry, as identi- fied in diverse beam search (Vijayakumar et al., 2016). The resulting selection procedure leverages an efficient greedy MAP solver for the partition DPP. It operates on represen- tations already computed by the diffusion model, therefore incurring negligible overhead and remaining scalable. We evaluate D5P4 on challenging discrete generation tasks, including open-ended text generation and question answer- ing. Experimental results show consistent improvements in diversity over strong decoding baselines, including beam search and prior diversity-promoting methods, while main- 1 arXiv:2603.19146v1 [cs.AI] 19 Mar 2026 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding taining competitive or superior quality. 2. Related Work 2.1. Discrete diffusion language models Discrete diffusion models originate from early formula- tions of diffusion processes over discrete variables (Sohl- Dickstein et al., 2015), later extended to categorical state spaces via multinomial diffusion and argmax-based flows (Hoogeboom et al., 2021). D3PM (Austin et al., 2021) formalized discrete diffusion as a structured Markov pro- cess over vocabulary elements, yielding a unified ELBO- based training objective and accommodating diverse cor- ruption mechanisms, including uniform, Gaussian, and absorbing-state transitions. Absorbing-state diffusion under- lies masked diffusion language models (MDLMs), which leverage advances in continuous diffusion (Ho et al., 2020; Rombach et al., 2022) and transformer-based architec- tures (Peebles & Xie, 2023) to narrow the performance gap with autoregressive models. Sahoo et al. (2024) intro- duce a simplified and stable training recipe, while subse- quent work shows that large-scale MDLMs can approach autoregressive performance under suitable scaling and fine- tuning regimes (von R Ě utte et al., 2025; Nie et al., 2024; Ye et al., 2025). More recently, models such as LLaDA (Nie et al., 2025) demonstrate competitive instruction-following and in-context learning behavior. A defining characteristic of MDLMs is parallel decoding, where tokens are refined jointly rather than generated left-to-right. While this decou- ples decoding depth from sequence length, naive parallel refinement often degrades generation quality, as token up- dates fail to explicitly model inter-token dependencies. As a result, prior work typically frames parallel decoding as a speedâquality trade-off. Several approaches accelerate in- ference through speculative execution or diffusion-specific keyâvalue caching (Wu et al., 2025; Israel et al., 2025; Ma et al., 2025; Wang et al., 2025), but largely focus on through- put, not joint quality or diversity across parallel hypotheses. 2.2. Sampling and selection for text generation In autoregressive models, decoding can be interpreted as approximate search over full sequences. Beam search (Low- erre, 1976) maintains multiple partial hypotheses, but shared prefixes often lead to rapid collapse into a single ancestral path. Nucleus sampling (Holtzman et al., 2020) promotes diversity via stochastic truncation, yet lacks explicit coor- dination or comparison across hypotheses. Other meth- ods introduce diversity at selection time, including Diverse Beam Search (Vijayakumar et al., 2016), stochastic beam search with Gumbel-Top-ksampling (Kool et al., 2019), and reranking approaches such as Maximal Marginal Rele- vance (Carbonell & Goldstein, 1998), which explicitly pe- nalize similarity among outputs. Recent work on test-time compute scaling reinforces a âgenerate many, then selectâ paradigm, showing that performance can improve substan- tially by selecting from large candidate pools (Beeching et al., 2024). For example, Kang et al. (2025) propose self- certainty as a lightweight selection criterion that does not rely on external verifiers. In the context of discrete diffusion, Dang et al. (2025a) introduce Particle Gibbs sampling for diffusion language models, enabling reward-guided infer- ence via MCMC/SMC-style resampling over full denoising trajectories. However, this approach does not model inter- actions between candidate solutions and is therefore not directly comparable to our method. While an adaptation to incorporate such interactions may be possible, the result- ing procedure would incur substantial additional test-time cost, scaling with the number of particles, diffusion steps, sampling iterations, and repeated reward evaluations. As a consequence, any direct quantitative comparison would conflate fundamentally different compute regimes and opti- mization objectives. 2.3. Diversity-aware decoding Recent analyses of reinforcement learning and supervised fine-tuning (SFT) in ARMs (Dang et al., 2025b; Yue et al., 2025) observe a consistent reduction in output diversity, particularly in reasoning-focused settings. Li et al. (2025) study this phenomenon in the context of SFT and propose a mitigation strategy at training time that improves cover- age and benefits test-time scaling. These findings suggest that gains in alignment and fidelity are not necessarily ac- companied by broader coverage, further motivating explicit diversity-aware decoding mechanisms. In diffusion mod- els, strong classifier-free guidance (CFG) is known to in- duce mode collapse, an effect that also appears in discrete diffusion for text (Schiff et al., 2024). Existing strategies to recover diversity in continuous diffusion typically rely on noise injection (Sadat et al., 2024) or partial guidance schedules (Kynk Ě a Ě anniemi et al., 2024), but do not address diversity at the set level. Most closely related to our work, Determinantal Beam Search (Meister et al., 2021) formu- lates beam selection using DPPs, encouraging diverse yet high-scoring hypotheses via a log-determinant objective and a string-based measure of similarity. However, this approach is inherently left-to-right and does not naturally extend to parallel diffusion trajectories. 3. Methods 3.1. Preliminary on Discrete Diffusion Algorithms We review the discrete diffusion framework underlying MDLM (Sahoo et al., 2024) and LLaDA (Nie et al., 2025), introducing a unified notation for training and inference. Notation. LetVdenote a vocabulary of tokens represented 2 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding as âone-hotâ column vector, including a special mask token, denotedmâV. A sequence of lengthLis denotedxâV L . The diffusion process operates on latent variablesz t âV L , indexed by a continuous timetâ [0, 1]. The noise schedule Îą t â [0, 1]defines the probability that a token remains unmasked at timet, such thatz 1 is fully masked (m L ) and z 0 corresponds to the clean data x 0 . Training Objective. Following Sahoo et al. (2024), the Rao-Blackwellized variational bound yields a simplified ob- jective dependent only on the transition rates. Letp θ (x|z t ,t) denote the modelâs predicted categorical distribution over clean tokens given the noisy statez t . Since the variational bound is invariant to the specific noise schedule, we assume a linear scheduleÎą t = 1â t. Under this choice, the loss function reduces to: L(θ) =âE t,x,z t   1 t X i:z i t =m logp θ (x i | z t ,t)   .(1) This objective satisfies the variational bound on the data log-likelihood: âE p data (x 0 ) [logp θ (x 0 )]â¤L(θ). Inference Dynamics. Both MDLM and LLaDA share a high-level generative process starting from a fully masked sequencez 1 = m L . In MDLM, the lengthLis fixed, while in LLaDA (instruction mode) it is a tunable hyperparame- ter. Notably, LLaDAâs denoising process excludes prompt tokens from the remasking budget, though they remain part of the conditioning context. For an intermediate transition from timestepttos, where 0 ⤠s < t ⤠1, the model generates logitsp θ (¡ | z t )over V L without requiring explicit timestep embeddings. The subsequent state z s is obtained via a projection operator: z s = Î t,s p θ (¡| z t ) .(2) The operatorÎ encapsulates the specific sampling and re- masking strategies. LLADA samples a fully denoised se- quence from the logits, then re-masks exactlyâL ¡ s/tâ tokens, selecting positions either uniformly or based on low confidence. MDLM enforces the masking ratio in logit space such that the expected number of masked tokens matches Îą s L. A defining advantage of this inference scheme is its inherent parallelism across all sequence positions. By decoupling the generation process from the strict left-to-right constraints of traditional autoregressive models, this approach facili- tates highly efficient batched decoding. Consequently, it enables the simultaneous processing of multiple candidate sequences, significantly enhancing throughput and maximiz- ing hardware utilization during large-scale inference. 3.2. Beam-Style Decoding for Discrete Diffusion While parallelism enables scalable sampling, it does not by itself address the intractability of sequence-level search. Effective approximations therefore require maintaining and selectively refining a limited set of high-quality hypotheses. We therefore formulate a beam-style decoding approach for discrete diffusion, in which parallel sampling is structured through intermediate selection steps. Branching and Scoring. Letkdenote the number of re- tained beams andwthe branching factor, yielding a can- didate pool of sizen = k ¡ w. At each diffusion stept, we retainkbeams and generatewdescendants from each by applying the stochastic projection operatorÎ t,s to the denoising logits p θ (¡| z t ). Candidates are evaluated using a scoring function q : V L â R . Since diffusion models do not admit monotonic prefix likelihoods, we rely on sequence-level entropy and self-certainty (Kang et al., 2025), defined as the KL divergence between token-level logits and the uniform distribution, as proxies for generation quality. This yields a score vector Qâ R n over the candidate pool. Selecting the top-kcandidates based solely onQcan lead to ancestral collapse, in which near-duplicate sequences are selected, thereby underutilizing the available parallel budget (Vijayakumar et al., 2016). To mitigate this issue, we impose a transversal partition constraint. Specifically, thencandidates are partitioned intokgroups, each with the w descendants of a common parent beam. A simple selection strategy chooses the highest-scoring can- didate within each group, ensuring that retained beams orig- inate from distinct parents. However, this approach remains agnostic to interactions between candidates across groups and therefore provides only a limited form of diversity. 3.3. DPP Sampling for Diversity Structured selection alone provides only a coarse, binary form of diversity control. To achieve a finer-grained trade- off between diversity and quality, we introduce a selection mechanism that explicitly accounts for both candidate qual- ity and pairwise interactions within the selected set, rather than treating candidates independently. To this end, we employ Determinantal Point Processes (DPPs) as a prin- cipled mathematical framework for diversity-aware beam selection (Kulesza et al., 2012). A DPP defines a probability distribution over subsets S â 1,...,n according to: P(S)â det(L S ) ,(3) whereL â R nĂn is a positive semidefinite kernel andL S 3 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Lineage independence Seq z 1,1 Seq z 1,2 Seq z 1,3 Seq z 2,1 Seq z 2,2 Seq z 2,3 Step t Group 1 Group 2 p θ p θ p θ p θ p θ p θ Joint Scoring & Groupwise Selection Seq z 1,1 Seq z 2,3 Î Î Seq z Ⲡ1,2 Seq z Ⲡ1,1 Seq z Ⲡ1,3 Seq z Ⲡ2,2 Seq z Ⲡ2,1 Seq z Ⲡ2,3 Step s Selected 1 Selected 2 Figure 1. Overview of the discrete diffusion beam search algorithm. Partially denoised sequences at timetare fed to the modelp θ . The model produces logits and embeddings that are then used in the Joint Scoring & Groupwise Selection block. The logits of the selected sequences for each group are expanded through independent applications of the projection operatorÎ , resulting in the sequence at times. denotes the principal submatrix indexed byS. This formula- tion corresponds to anL-ensemble (Tremblay et al., 2023), in which the DPP is parameterized by an unnormalized positive semidefinite kernelL. Intuitively, the determinant penalizes redundancy: selecting highly similar candidates induces near-linear dependence inL S , thereby reducing its value. By appropriately designing the kernel matrix, DPPs provide a natural and theoretically grounded mechanism for balancing candidate quality and diversity at the set level. The discrete diffusion beam-search, including joint scoring with interaction modeling, is illustrated in Figure 1. Kernel Construction.Building on the denoising process described earlier, we construct theL-ensemble kernel from the model outputs. Sequence-level quality scores are de- rived from the token logits, while pairwise interactions are computed using the hidden representations immediately pre- ceding the unembedding layer, which are available at no additional computational cost. The kernel is defined as: L = ( diag(Q) + βK,(additive) diag(e Q/β )¡K¡ diag(e Q/β )(multiplicative) , (4) whereβis a diversity coefficient that allows for a direct and interpretable quality-diversity trade-off at the set level. Here, Qrepresents the vector per-sequence quality scores and the matrixKencodes the pairwise interactions between candi- dates. While the multiplicative formulation is more com- mon, the additive is adapted from (Meister et al., 2021). In practice,Kis a kernel defined over the normalized sequence embeddings e i , using either a cosine or RBF formulation: K ij =â¨e i , e j ⊠orK ij = exp(âÎłâĽe i â e j ⼠2 ) . (5) Fixed-Size Selection and Sampling.In the general case, sampling from a DPP yields a subset whose cardinality is controlled only in expectation and does not guarantee selecting exactlykelements. Ak-DPP addresses this by restricting the distribution to subsets of fixed cardinality k. Exact sampling from either a DPP or ak-DPP can be performed via eigendecomposition (Kulesza et al., 2012), with a computational complexity ofO(n 3 ). 3.4. Greedy Estimation of the MAP Interestingly, the standard beam search baseline (top-k) is equivalent to computing the MAP solution of a DPP with a quality-only kernel, i.e., it computes the argmax of the diagonal-only kernel. In this regime, directly maximizing the subdeterminant can be preferable to sampling from the induced distribution. This reduces to a discrete optimization problem that is NP-hard due to the combinatorial nature of selecting k elements from n. To tackle this issue, we leverage a fast greedy algo- rithm (Chen et al., 2018). We extend it to handle the parti- tion constraint and incorporate a multi-initialization strategy, initializing the algorithm from the argmax of each group to reduce sensitivity to initialization and exploring all ini- tializations in parallel. To our knowledge, no closed-form solution exists for DPPs under such partition constraints. The complexity of the original algorithm isO(k 2 n). Ours has complexityO(k 3 n); however, because the additional computation is fully parallelized on the GPU, the effective wall-clock cost remains almost identical to that of the origi- nal method in practice. In this regime, it is also more scal- able than eigendecomposition-based approaches for directly 4 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding sampling from a DPP. The combination of the discrete diffusion framework, our kernel, and the fast MAP solver constitutes D5P4. 4. Experimental setup The exact configuration used in each experiment can be found in Section B.1 of the Appendix. 4.1. Metrics We employ a combination of reference-free and reference- based metrics to measure generation quality and diversity. Quality Metrics. For open-ended generation, we report Perplexity (PPL), which measures sequence likelihood un- der an external autoregressive evaluator, using GPT-2 (Rad- ford et al., 2019) and Llama-3 (Dubey et al., 2024). We also report MAUVE (Pillutla et al., 2021), a distribution- level metric that quantifies similarity between generated and reference text, computed over large-scale sample sets, as well as MAUVE*, a robust variant. For question answering, we evaluate correctness using BLEU (Papineni et al., 2002) and token-level F1-score. When multiple valid references are available, we additionally compute the Wasserstein dis- tance between generated and reference answer clusters via Optimal Transport (Flamary et al., 2021), which penalizes semantic deviation across reference sets. Diversity Metrics. To assess semantic diversity indepen- dently of surface form, we compute the average in-batch cosine similarity (COS) of Jina embeddings (Sturua et al., 2024), where lower similarity indicates higher semantic diversity. To measure lexical diversity, we report Self- BLEU (Zhu et al., 2018), quantifying inter-sample simi- larity, and Distinct-n(Li et al., 2016), measuringn-gram uniqueness. We additionally report EAD (Expectation- Adjusted Distinct), a normalized variant of Distinct-nthat is robust to variations in sequence length (Liu et al., 2022). 4.2. Baselines 4.2.1. QUALITY-CENTRIC BASELINES Best-of-n(Oversample) is a brute-force baseline where a large candidate pool (n > k) is generated via independent sampling, and the top-ksequences are selected according to the final evaluation metric. This configuration allows for a controlled FLOP matching experiment. Standard Beam Search selects the group-wise argmax of the score Q, ignoring redundancy across beams. 4.2.2. DIVERSITY-PROMOTING BASELINES Diverse Beam Search (Transversal MMR). Since discrete diffusion lacks the partial hypotheses used in autoregressive diverse beam search (Vijayakumar et al., 2016), we adapt this baseline as a partition-constrained Maximal Marginal Relevance (MMR) (Carbonell & Goldstein, 1998) objective. We employ a greedy subset selection strategy where candi- date scores are penalized by their similarity to the currently selected set. To ensure robustness, we run the greedy proce- dure from multiple initializations and select the subset with the highest cumulative score. This strategy also matches the initialization strategy in our method. Standard DPP Sampling (Non-Transversal Only). To isolate the impact of our proposed partition constraints, we use theDPPylibrary (Gautier et al., 2019) to sample exactly kitems from the unconstrained kernelL. This serves as a baseline for pure DPP sampling partition constraints: exact transversal sampling is not possible in this setting. 5. Results and Analysis 5.1. Preliminary Alignment Analysis Table 1. Alignment metrics between the models used to generate and evaluate text. Model Entropy ĎEntropy ĎCKA (GPT-2)(LLaMA-3)(Jina) MDLM0.906-0.821 LLADA-0.8920.667 We first evaluate alignment of internal representations be- tween MDLM and LLADA with ARMs used as external evaluators in our subsequent experiments, respectively GPT- 2 and LLaMA-3. This preliminary experiment aims to show that internal signals of the model can be used without the need for external validation. We rely on two proxies: entropy-based scores to estimate generation quality, and intermediate diffusion embeddings to capture diversity. In parallel, we assess representation- level alignment by measuring the average Centered Kernel Alignment (CKA) (Kornblith et al., 2019) between diffusion- model latent states and the Jina embedding space used for semantic diversity, across several masking ratios. Results in Table 1 show that diffusion entropy estimates ex- hibit strong agreement with autoregressive log-likelihoods, achieving a Spearman correlation ofĎ > 0.89. This trend is further illustrated in Figure 5, where MDLMâs Monte-Carlo likelihood estimates reach a91.1% correlation with GPT-2 log-likelihood. The CKA scores, reaching up to0.821, in- dicate that diffusion representations are semantically well structured. This analysis is conducted on FineWeb (Penedo et al., 2024) text for MDLM and on the TruthfulQA train- 5 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding ing set for LLADA, reflecting the respective domains and training regimes of the two diffusion models. Together, these results justify our architectural choices: the diffusion modelâs internal estimates of uncertainty and se- mantic structure are sufficiently aligned with external bench- marks to reliably drive the DPP kernel, eliminating the need for external evaluators at inference time. 5.2. Open-Ended Generation In Figure 2, we compare several diversity-oriented methods with our decoding algorithm D5P4. We generate text using MDLM, starting from fully masked sequences, and evaluate both the additive and multiplicative variants of D5P4. As baselines, we consider independent generation ofkse- quences (Baseline), standard beam search with temperature scaling applied to the categorical distributionÎ (CAT), as well as the Diverse Beam Search algorithm (DivBS). We compare with our approach integrating the multiplicative version of MAP (D5P4Ă) and the additive version (D5P4+). To rigorously assess the diversity-quality trade-off, we per- form systematic parameter sweeps across all methods: we vary the temperature in CAT, the diversity penaltyÎą div in DivBS, and the interaction parameter β in D5P4. All search-based methods dominate the Baseline, achieving lower perplexity for the same cosine similarity, or higher cosine similarity at comparable perplexity. This confirms that even lightweight search or reweighting mechanisms are effective at navigating the diversity-quality trade-off be- yond naive sampling. Across all methods, increasing the diversity-controlling parameter eventually leads to a sharp degradation in perplexity. This behavior is consistent with over-penalizing high-probability modes, forcing the model into low-likelihood regions of the distribution. The main distinction between methods lies in where this breakdown occurs along the cosine similarity axis. More specifically, temperature scaling exhibits an abrupt transition: perplexity remains competitive up to a relatively high cosine similarity, after which it increases sharply. This suggests that it pro- vides limited granularity and tends to fail catastrophically once pushed beyond a narrow operating range. DivBS shows a smoother but earlier degradation compared to CAT. While it encourages diversity effectively at mod- erate cosine similarity, its reliance on explicit MMR-style penalties leads to higher perplexity sooner, indicating a stronger bias toward diversity at the expense of fluency. Both D5P4Ăand D5P4+achieve the most favorable Pareto fronts overall. D5P4Ămaintains low perplexity across a wider cosine range, suggesting that multiplicative reweight- ing preserves relative likelihood structure better under in- creasing diversity pressure. D5P4+consistently dominates in the high-diversity regime, delaying perplexity blow-up the furthest, which aligns with additive kernels offering more stable control when aggressively reshaping the dis- tribution. Overall, kernel-based MAP approximations in D5P4 provide a more controlled and extensible trade-off, with failure modes that are delayed and more predictable compared to temperature-based or heuristic search methods. Table 2 summarizes how each diversity parameter influ- ences the qualityâdiversity trade-off and demonstrates the effectiveness of our method in controlling this balance. Fig- ure 6 in the appendix provides a more detailed view of these relationships. Table 2. Correlation between diversity control parameters and quality (PPL) and diversity (COS) metrics across methods. MethodParameter Correlation PPLCOS CATlog temperature0.940-0.438 DivBSlogÎą div 0.950-0.846 D5P4+ logβ inter 0.902-0.875 D5P4Ă0.959-0.628 In Figure 3, we report MAUVE with the FineWeb dataset (Penedo et al., 2024), for fixed values of interac- tion parameter,β. This result suggests that the parameter admits an optimal intermediate regime, balancing genera- tion quality (β = 0) and diversity (β ââ). 10 0 10 1 10 2 10 3 0.85 0.9 0.95 β parameter MAUVE | MAUVE* Baseline MAUVEBaseline MAUVE* MAUVE MAUVE* Figure 3. Effect of diversity control parameterβon distribution fidelity (MAUVE and MAUVE*), using D5P4+. 5.3. Question Answering We examine the effect of classifier-free guidance (CFG) strength on output diversity by prompting LLADA to answer questions from the TruthfulQA (Lin et al., 2021) dataset under increasing CFG values. As shown in Figure 4 and complementary figures in Appendix C, higher CFG levels are consistently associated with reduced diversity at both the lexical and semantic levels, as measured by Distinct-2, 6 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding 0.720.730.740.750.760.770.780.790.80 Cosine Similarity 10 15 20 25 30 35 40 Perplexity Better Baseline CAT DivBS D5P4 Ă D5P4 + Figure 2. Pareto front comparison of diversity-encouraging methods in open-ended generation. Lower perplexity indicates higher quality, and lower cosine similarity indicates higher diversity. The single point for baseline corresponds to an independent sampling baseline. CAT is the categorical temperature modulation. DivBS is the transversal MMR search. Our methods approximate the MAP of an additive and multiplicative kernel. D5P4 consistently achieves better diversity-quality trade-offs. Self-BLEU, and in-batch cosine similarity. These results are consistent with the findings of Schiff et al. (2024). In con- trast, even at high CFG values, D5P4 preserves substantially higher lexical and semantic diversity across all diversity met- rics while maintaining comparable F1, Wasserstein distance and perplexity. This indicates that structured selection at decoding time can effectively counteract guidance-induced mode collapse without sacrificing generation quality. 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 Expectation Adjusted Distinct D5P4 Independent Figure 4. Mitigation of CFG diversity collapse. Increasing CFG strength typically reduces diversity (higher EAD). D5P4 counter- acts this collapse, maintaining higher diversity across CFG values. We next compare the best-of-kbaseline against two diversity-promoting approaches: D5P4+and an augmented variant that applies classifier-free guidance only during the first half of the denoising steps, inspired by continuous dif- fusion methods (Kynk Ě a Ě anniemi et al., 2024) and referred to as P-CFG (partial CFG). We match the computational budget (FLOPs) of the best-of-kbaseline and report mul- tiple quality, alignment, and diversity metrics in Table 3 to illustrate the resulting trade-offs across methods. Our methods substantially increase output diversity while main- taining competitive alignment and generation quality. We evaluate the methods on TruthfulQA (Lin et al., 2021) and CommonSenseQA (Talmor et al., 2019). 5.4. Ablations Diversity and Quality Estimation.We conduct ablation studies to evaluate alternative strategies for estimating diver- sity and generation quality during decoding. These experi- ments assess the design choices underlying the DPP kernel and isolate the contribution of its individual components. For quality estimation, we compare our entropy-based scor- ing function with the self-certainty metric proposed by Kang et al. (2025). The results are reported in Figure 9. For diversity estimation, we evaluate several embedding pooling strategies used to construct the DPP kernel: average pooling over all tokens, masked tokens pooling, non-masked tokens pooling, and flattened sequence-level embeddings. This analysis characterizes the sensitivity of diversity esti- mation to both representation choice and token subset. We report CKA and pairwise cosine similarity of diffusion em- beddings to monitor representation collapse, which may not be captured by CKA alone. Results are presented in Table 4, and support the rationale behind our modeling choices. 7 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Table 3. Qualityâdiversity trade-off in question answering with LLaDA. We compare decoding strategies using answer quality and diversity metrics. D5P4 achieves comparable answer quality to the independent baseline while consistently increasing output diversity. SettingBest-of-kD5P4+D5P4+ with P-CFG DescriptionDatasetTruthfulCommonSenseTruthfulCommonSenseTruthfulCommonSense QualityPerplexity17.44627.46215.72525.96915.01526.084 AccuracyF1-score0.2120.0190.1840.0120.1950.013 AccuracyMax F1-score0.2340.0240.2210.020.2410.02 Accuracy BLEU5.0010.0154.5820.0144.6750.016 AlignmentMax COS0.8740.7440.8730.7450.8750.743 AlignementWasserstein0.5780.7240.5790.7250.5770.728 DiversityAverage COS0.9630.9690.9460.920.9180.859 Diversity Distinct-20.5940.5690.6320.6260.6160.622 DiversityDistinct-30.6870.6570.7290.7180.710.701 Diversity EAD0.3630.3890.3850.4310.3890.455 DiversitySelf-BLEU47.10252.5540.40443.01842.7845.814 Table 4. Representation alignment (CKA) with the reference model for different pooling methods MethodMDLMLLADA Mean0.7770.482 Non-masked0.6600.536 Masked0.7100.435 Flatten0.8210.667 Table 5. Correlation of sequence-level scoring methods with GPT- 2 perplexity (reference quality signal) MethodCorrelation Self-certainty-0.290 Entropy-0.776 Sub-Determinant Maximization Algorithms Table 6 reports the speed and objective value of the different se- lection algorithms. It corresponds to one operating point of the more detailed scaling study presented in Figure 10 of the appendix. For this experiment, we generate syn- thetic DPP kernels and evaluate each method on a highly complex task of selecting elements across 32 groups of 32 items. The results demonstrate that our Greedy MAP al- gorithm achieves the highest normalized sub-determinant value (1.0214) while maintaining an extremely low average runtime of 0.0023,s. Compared to Diverse Beam Search, our method not only yields a substantially better objective value but is also over 10Ăfaster. Furthermore, while the standard DPP baseline (included solely as a reference, as it does not support transversal sampling) yields poor perfor- mance and takes over half a second, the minimal selection cost of Greedy MAP introduces negligible slowdown rel- ative to independent random sampling, with the overhead largely dominated by GPU synchronization. Table 6. Speedâaccuracy trade-off for subset selection with 32 groups of 32 elements. (*) Does not support transversal sampling. MethodValueTime (s) Random-0.90740.0001 DPP (*)-0.87520.5478 Diverse Beam Search0.66450.0295 Greedy MAP (ours)1.02140.0023 6. Conclusion In this work, we introduce D5P4, a decoding framework for discrete diffusion language models that couples parallel denoising with principled, diversity-aware selection. D5P4 casts beam selection as MAP inference in a Partition Deter- minantal Point Process, yielding a simple and interpretable mechanism for trading off model-implied quality and in- batch diversity. Crucially, both signals are computed from the diffusion model itself (sequence-level entropy for quality and hidden-state representations for semantic similarity) to avoid reliance on costly external scorers at inference time. Across open-ended generation with MDLM and question answering with LLADA, D5P4 improves diversity without sacrificing competitive quality, and consistently provides a stronger qualityâdiversity Pareto trade-off than widely used alternatives such as diverse beam search and temperature- based sampling. These gains come with minimal overhead: our greedy MAP solver is efficient, multi-GPU friendly, and preserves the throughput advantages of parallel decoding. More broadly, D5P4 suggests a practical route to test-time scaling for diffusion LMs that increases coverage through structured candidate selection rather than additional model calls. Future work could extend this principle to other diffu- sion formulations and modalities, and explore richer kernel designs or task-aware objectives while retaining the same efficient selection backbone. 8 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Impact Statement This work introduces an algorithmic framework to improve the diversity of outputs in discrete diffusion language mod- els. As language models are increasingly deployed in real- world applications, the tendency for these models to suffer from mode collapse or reduced output coverage (particularly after supervised fine-tuning) presents a significant challenge for creative generation and unbiased reasoning. By provid- ing a mechanism to explicitly encourage diversity at the set level, our method can help mitigate the risks of repetitive or overly homogenized content, potentially reducing the reinforcement of narrow biases found in training data. While D5P4 improves the technical control over model outputs, it does not inherently prevent the generation of harmful content. The ethical implications of this work are largely aligned with broader research into generative AI, where increased diversity could, in some contexts, lead to the exploration of lower-probability but harmful regions of the distribution if not properly gated. We encourage practitioners to combine our decoding strategy with robust safety filters and alignment techniques to ensure that the increased diversity serves to enrich rather than compromise the safety and reliability of the generated text. References Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and Van Den Berg, R. Structured denoising diffusion models in discrete state-spaces. Advances in neural information processing systems, 34:17981â17993, 2021. Beeching, E., Tunstall, L., and Rush, S. Scaling test-time compute with open models, 2024. URLhttps:// huggingface.co/spaces/HuggingFaceH4/ blogpost-scaling-test-time-compute. Carbonell, J. and Goldstein, J. The use of mmr, diversity- based reranking for reordering documents and produc- ing summaries. In Proceedings of the 21st Annual In- ternational ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR â98, p. 335â336, New York, NY, USA, 1998. Association for Computing Machinery. ISBN 1581130155. doi: 10. 1145/290941.291025. URLhttps://doi.org/10. 1145/290941.291025. Chen, L., Zhang, G., and Zhou, E. Fast Greedy MAP Infer- ence for Determinantal Point Process to Improve Recom- mendation Diversity. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018.URLhttps://proceedings.neurips. c/paper_files/paper/2018/hash/ dbbf603f0e99629dda5d75b6f75f966-Abstract. html. Dang, M., Han, J., Xu, M., Xu, K., Srivastava, A., and Er- mon, S. Inference-time scaling of diffusion language models with particle gibbs sampling. arXiv preprint arXiv:2507.08390, 2025a. Dang, X., Baek, C., Kolter, J. Z., and Raghunathan, A. Assessing Diversity Collapse in Reasoning. In Scaling Self-Improving Foundation Models without Human Super- vision, 2025b. URLhttps://openreview.net/ forum?id=AMiKsHLjQh. Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. Flamary, R., Courty, N., Gramfort, A., Alaya, M. Z., Bois- bunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., Gautheron, L., Gayraud, N. T. H., Janati, H., Rakotomamonjy, A., Redko, I., Rolet, A., Schutz, A., Seguy, V., Sutherland, D. J., Tavenard, R., Tong, A., and Vayer, T. POT: Python Optimal Transport. Journal of Ma- chine Learning Research, 22(78):1â8, 2021. URLhttp: //jmlr.org/papers/v22/20-451.html. Gautier, G., Polito, G., Bardenet, R., and Valko, M. DPPy: DPP Sampling with Python. Journal of Ma- chine Learning Research - Machine Learning Open Source Software (JMLR-MLOSS), 2019. URLhttp: //jmlr.org/papers/v20/19-179.html. Code at http://github.com/guilgautier/DPPy/ Documentation at http://dppy.readthedocs.io/. Ho, J., Jain, A., and Abbeel, P. Denoising diffusion proba- bilistic models. Advances in neural information process- ing systems, 33:6840â6851, 2020. Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In International Conference on Learning Representations, 2020. URLhttps://openreview.net/forum? id=rygGQyrFvH. Hoogeboom, E., Nielsen, D., Jaini, P., Forr Ě e, P., and Welling, M. Argmax flows and multinomial diffusion: Learning categorical distributions. Advances in neural information processing systems, 34:12454â12465, 2021. Israel, D., Broeck, G. V. d., and Grover, A. Accelerating dif- fusion llms via adaptive parallel decoding. arXiv preprint arXiv:2506.00413, 2025. Kang, Z., Zhao, X., and Song, D.Scalable Best- of-N Selection for Large Language Models via Self- Certainty. In 2nd AI for Math Workshop @ ICML 2025, 2025. URLhttps://openreview.net/forum? id=nddwJseiiy. 9 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Kool, W., Van Hoof, H., and Welling, M. Stochastic beams and where to find them: The gumbel-top-k trick for sam- pling sequences without replacement. In International conference on machine learning, p. 3499â3508. PMLR, 2019. Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. Sim- ilarity of Neural Network Representations Revisited. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Pro- ceedings of the 36th International Conference on Ma- chine Learning, volume 97 of Proceedings of Machine Learning Research, p. 3519â3529. PMLR, June 2019. URLhttps://proceedings.mlr.press/v97/ kornblith19a.html. Kulesza, A., Taskar, B., et al. Determinantal point pro- cesses for machine learning. Foundations and TrendsÂŽ in Machine Learning, 5(2â3):123â286, 2012. Kynk Ě a Ě anniemi, T., Aittala, M., Karras, T., Laine, S., Aila, T., and Lehtinen, J. Applying guidance in a limited interval improves sample and distribution quality in diffusion models. In Proc. NeurIPS, 2024. Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B. A Diversity-Promoting Objective Function for Neural Conversation Models.In Knight, K., Nenkova, A., and Rambow, O. (eds.), Proceedings of the 2016 Con- ference of the North American Chapter of the Asso- ciation for Computational Linguistics: Human Lan- guage Technologies, p. 110â119, San Diego, Cali- fornia, June 2016. Association for Computational Lin- guistics. doi: 10.18653/v1/N16-1014. URLhttps: //aclanthology.org/N16-1014/. Li, Z., Chen, C., Xu, T., Qin, Z., Xiao, J., Luo, Z.-Q., and Sun, R. Preserving diversity in supervised fine- tuning of large language models. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/forum? id=NQEe7B7bSw. Lin, S., Hilton, J., and Evans, O. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021. Liu, S., Sabour, S., Zheng, Y., Ke, P., Zhu, X., and Huang, M. Rethinking and refining the distinct metric. In Mure- san, S., Nakov, P., and Villavicencio, A. (eds.), Proceed- ings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), p. 762â770, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022. acl-short.86. URLhttps://aclanthology.org/ 2022.acl-short.86/. Lowerre, B. T. The Harpy Speech Recognition System. PhD thesis, Carnegie Mellon University, Pittsburgh, PA, 1976. Ma, X., Yu, R., Fang, G., and Wang, X. dkv-cache: The cache for diffusion language models. arXiv preprint arXiv:2505.15781, 2025. Meister, C., Forster, M., and Cotterell, R. Determinantal beam search. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 6551â6562, 2021. Nie, S., Zhu, F., Du, C., Pang, T., Liu, Q., Zeng, G., Lin, M., and Li, C. Scaling up masked diffusion models on text. arXiv preprint arXiv:2410.18514, 2024. Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R., and Li, C. Large language diffusion models. arXiv preprint arXiv:2502.09992, 2025. Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine transla- tion. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, p. 311â318, 2002. Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, p. 4195â4205, 2023. Penedo, G., Kydl Ě Äą Ë cek, H., allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L. V., and Wolf, T. The fineweb datasets: Decanting the web for the finest text data at scale, 2024. URLhttps://arxiv.org/abs/ 2406.17557. Pillutla, K., Swayamdipta, S., Zellers, R., Thickstun, J., Welleck, S., Choi, Y., and Harchaoui, Z. MAUVE: Mea- suring the gap between neural text and human text using divergence frontiers. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021. URLhttps: //openreview.net/forum?id=Tqx7nJp7PR. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models, April 2022. URLhttp://arxiv. org/abs/2112.10752. arXiv:2112.10752 [cs]. Sadat, S., Buhmann, J., Bradley, D., Hilliges, O., and Weber, R. M. CADS: Unleashing the diversity of diffusion mod- els through condition-annealed sampling. In The Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum? id=zMoNrajk2X. 10 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Sahoo, S. S., Arriola, M., Gokaslan, A., Marroquin, E. M., Rush, A. M., Schiff, Y., Chiu, J. T., and Kuleshov, V. Simple and effective masked diffusion language mod- els. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URLhttps: //openreview.net/forum?id=L4uaAR4ArM. Schiff, Y., Sahoo, S. S., Phung, H., Wang, G., Boshar, S., Dalla-torre, H., de Almeida, B. P., Rush, A., Pierrot, T., and Kuleshov, V. Simple guidance mechanisms for dis- crete diffusion models. arXiv preprint arXiv:2412.10193, 2024. Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequi- librium thermodynamics. In International conference on machine learning, p. 2256â2265. PMLR, 2015. Sturua, S., Mohr, I., Akram, M. K., G Ě unther, M., Wang, B., Krimmel, M., Wang, F., Mastrapas, G., Koukounas, A., Koukounas, A., Wang, N., and Xiao, H. jina-embeddings- v3: Multilingual embeddings with task lora, 2024. URL https://arxiv.org/abs/2409.10173. Talmor, A., Herzig, J., Lourie, N., and Berant, J. Common- senseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In Burstein, J., Doran, C., and Solorio, T. (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 4149â4158, Min- neapolis, Minnesota, June 2019. Association for Compu- tational Linguistics. doi: 10.18653/v1/N19-1421. URL https://aclanthology.org/N19-1421/. Tremblay, N., Barthelm Ě e, S., Usevich, K., and Amblard, P.-O. Extended l-ensembles: a new representation for determinantal point processes. The Annals of Applied Probability, 33(1):613â640, 2023. Vijayakumar, A. K., Cogswell, M., Selvaraju, R. R., Sun, Q., Lee, S., Crandall, D., and Batra, D. Diverse beam search: Decoding diverse solutions from neural sequence models. arXiv preprint arXiv:1610.02424, 2016. von R Ě utte, D., Fluri, J., Pooladzandi, O., Sch Ě olkopf, B., Hofmann, T., and Orvieto, A.Scaling behavior of discrete diffusion language models. arXiv preprint arXiv:2512.10858, 2025. Wang, X., Xu, C., Jin, Y., Jin, J., Zhang, H., and Deng, Z. Diffusion llms can do faster-than-ar inference via dis- crete diffusion forcing. arXiv preprint arXiv:2508.09192, 2025. Wu, C., Zhang, H., Xue, S., Liu, Z., Diao, S., Zhu, L., Luo, P., Han, S., and Xie, E. Fast-dllm: Training-free acceleration of diffusion llm by enabling kv cache and parallel decoding, 2025. URLhttps://arxiv.org/ abs/2505.22618. Ye, J., Xie, Z., Zheng, L., Gao, J., Wu, Z., Jiang, X., Li, Z., and Kong, L. Dream 7b: Diffusion large language models. arXiv preprint arXiv:2508.15487, 2025. Yue, Y., Chen, Z., Lu, R., Zhao, A., Wang, Z., Yue, Y., Song, S., and Huang, G. Does Reinforcement Learn- ing Really Incentivize Reasoning Capacity in LLMs Be- yond the Base Model?In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URLhttps://openreview.net/forum? id=4OsgYD7em5. Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y. Texygen: A Benchmarking Platform for Text Generation Models, February 2018. URLhttp:// arxiv.org/abs/1802.01886. arXiv:1802.01886 [cs]. 11 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding A. Algorithmic Details of D5P4 Selection Algorithm 1 Fast Greedy MAP Inference for Partition DPPs (adapted from Chen et al. (2018)) 1: Input: Kernel L, Target size K, Group constraints 2: Initialize: c i =â , d 2 i = L i for all iâ Z 3: Select first item: j = arg max iâZ log(d 2 i ) 4: Y =j 5: for t = 2 to K do 6: Z valid âiâ Z | group(i) /â groups(Y ) 7:for iâ Z valid do 8:Compute projection: e i = (L ji ââ¨c j , c i âŠ)/ q d 2 j 9:Update history: c i â [c i e i ] 10:Update marginal: d 2 i â d 2 i â e 2 i 11:end for 12:Select next item: j = arg max iâZ valid log(d 2 i ) 13: Y â Y âŞj 14: end for 15: Return: Y Algorithm 2 Multi-init Parallel Transversal Fast Greedy MAP 1: Input: Kernel L, Target size K, Group constraints 2: Initialize: Create|Z| parallel trajectories indexed by bâ Z 3: Parallel Initialization for all bâ Z: 4:c (b) i =â , (d (b) i ) 2 = L i 5:Force start item: j (b) = b 6: Y (b) =j (b) ,L (b) = log((d (b) j (b) ) 2 ) 7: for t = 2 to K do 8:Parallel Update for all bâ Z: 9: Z (b) valid âiâ Z | group(i) /â groups(Y (b) ) 10:for iâ Z (b) valid do 11:e (b) i = (L j (b) i ââ¨c (b) j (b) , c (b) i âŠ)/ q (d (b) j (b) ) 2 12:c (b) i â [c (b) i e (b) i ] 13:(d (b) i ) 2 â (d (b) i ) 2 â (e (b) i ) 2 14:end for 15:Select next item: j (b) = arg max iâZ (b) valid log((d (b) i ) 2 ) 16: Y (b) â Y (b) âŞj (b) 17: L (b) âL (b) + log((d (b) j (b) ) 2 ) 18: end for 19: Return: Y (b â ) where b â = arg max b L (b) B. Extended Experimental Setup and Results B.1. Detailed Experimental Setup In the open-ended generation experiment (Figure 2), we use4elements per group and8groups, yielding a total batch size of32. In practice, groups can be distributed across GPUs or nodes to enable parallel execution. Each reported point corresponds to8independent runs, while the baseline aggregates results over100independent runs. The MAUVE study (Figure 3) uses the same configuration. In the question answering experiments (Table 3), we evaluate on the full datasets (817runs for TruthfulQA and1.14k for CommonsenseQA) with a global batch size of 9, split into 3 groups of 3 elements each. 12 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding B.2. Extended Quality and Diversity Analysis We report the correlation between MDLM Monte Carlo estimates of the log-likelihood versus that of GPT-2. The strong linear relationship (Pearson r = 0.911) indicates close alignment between the two quality measures. â6â5â4â3â2 â6 â5 â4 â3 â2 MDLM log-likelihood GPT2 log-likelihood Samples Regression Figure 5. Correlation between MDLM-estimated log-likelihood and GPT-2 log-likelihood across samples. 0.760.770.780.790.80 Cosine Similarity 10 15 20 25 30 35 40 45 50 Perplexity CAT (log scale) 0.01 0.02 0.03 0.04 0.05 0.06 0.07 log 10 (temp + 10 â 8 ) (a) Categorical temperature scaling 0.740.750.760.770.78 Cosine Similarity 10 15 20 25 30 Perplexity DivBS (log scale) 0 1 2 3 log 10 ( Îą div + 10 â 8 ) (b) DivBS with varying Îą div 0.720.730.740.750.760.770.780.79 Cosine Similarity 10 15 20 25 30 35 40 Perplexity Additive MAP (log scale) 0 1 2 3 log 10 ( β +10 â 8 ) (c) D5P4+ with varying β 0.720.730.740.750.760.770.780.79 Cosine Similarity 10 15 20 25 30 35 40 45 Perplexity Multiplicative MAP (log scale) 5 4 3 2 1 0 1 log 10 ( β + 10 â 8 ) (d) D5P4Ă with varying β Figure 6. Dynamics of the parameters controlling the diversity 13 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding C. Comprehensive Metrics for CFG Collapse Mitigation using D5P4 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 BLEU@K D5P4 Independent (a) BLEU@k 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 Cosine Similarity D5P4 Independent (b) Cosine 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 Distinct-2 D5P4 Independent (c) Distinct-2 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 Expectation Adjusted Distinct D5P4 Independent (d) EAD 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 F1@K D5P4 Independent (e) F1@k 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 Perplexity D5P4 Independent (f) PPL 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 Self Bleu D5P4 Independent (g) Self-BLEU 1.002.00 CFG 0.0 0.2 0.4 0.6 0.8 1.0 Wasserstein Distance D5P4 Independent (h) Wasserstein Figure 7. Impact of CFG across evaluation metrics, baseline vs D5P4. 14 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding D. Detailed Ablation Studies 00.20.40.60.81 0.2 0.4 0.6 0.8 Mask Ratio CKA Similarity Mean Pool Non-Masked Pool MaskedFlatten (a) MDLM CKA 00.20.40.60.81 0.9 0.95 Mask Ratio Average Cosine Similarity Mean Pool Non-Masked Pool MaskedFlatten (b) MDLM ACS 0 0.20.40.60.81 0.2 0.4 0.6 0.8 Mask Ratio CKA Similarity Mean Pool Non-Masked Pool MaskedFlatten (c) LLADA CKA 00.20.40.60.81 0 0.2 0.4 0.6 Mask Ratio Average Cosine Similarity Mean Pool Non-Masked Pool MaskedFlatten (d) LLADA ACS Figure 8. CKA and Average Cosine Similarity (ACS) of pooling methods for different masking ratios across MDLM and LLADA. 0.60.650.70.75 0.8 2 3 4 5 â0.4â0.35â0.3 2 3 4 5 Entropy Score Reference PPL Self-Certainty Score Reference PPL Figure 9. Correlation between sequence scoring methods and the reference perplexity target. 15 D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding E. Scaling Behavior and Computational Efficiency In this section, we study the scaling behavior of the subset selection algorithms underlying D5P4. We construct synthetic DPP instances following the kernel construction used in the main paper. Concretely, we estimate first- and second-order statistics from MDLM representations, sample embeddings and quality scores from these statistics, and build additive DPP kernels combining quality and similarity terms. We evaluate subdeterminant maximization across methods. For each instance, we report: (i) the normalized objective value, defined as the achieved normalized subdeterminant across all methods, (i) the average rank of each method across trials, and (i) the execution time of the optimization step. We vary both the group size and the number of groups from 4 to 64, corresponding to increasing log combinatorial complexity, given bylog(|S|) = log(group size number of groups ). We also sweep the interaction parameter controlling the strength of pairwise repulsion. All results are averaged over 500 independent trials. With the exception of the DPPy baseline, which relies on NumPy and executes on the CPU, all methods are implemented on the GPU. Notably, DPPy does not support transversal partition constraints and is included strictly as a reference. We also include a Triton-optimized D5P4 implementation to assess the impact of kernel-level speedups. Regarding objective maximization, DPPy and Greedy Beam Search perform comparably to the random baseline. While Diverse Beam Search matches our method at low complexity, our approach achieves superior performance at higher combinatorial complexities. For computational efficiency, Random and Greedy Beam exhibit near-constant runtimes due to minimal overhead, whereas the CPU-bound DPPy scales poorly. Our standard GPU method scales more favorably than Diverse Beam Search, and our Triton-optimized variant further reduces execution time, underscoring the benefit of specialized kernels for large-scale subset selection. 10 1 10 2 â1 0 1 10 1 10 2 2 3 4 5 10 1 10 2 10 â4 10 â3 10 â2 10 â1 10 0 Log Combinatorial Complexity Normalized value Log Combinatorial Complexity Average Rank Log Combinatorial Complexity Time (s) Greedy MAP (ours)Greedy MAP triton (ours)Greedy BeamDiverse BeamRandom DPPy Figure 10. Comparison of different subsampling methods on Time and Normalized value against Log Combinatorial Complexity. 16