arXiv · 2606.10829
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
Abstract
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-$k$, Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty. Across LLaDA-8B-Base and Dream-7B-Base on the reasoning benchmarks GSM8K and MATH500 and the code benchmarks HumanEval and MBPP, plugging ADAS into all three samplers improves low-NFE performance at matched denoiser evaluations by $9.11$ and $10.46$ percentage points on average, respectively, with $3.1\%$ per-forward runtime overhead.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro. 2026-06-09. Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models. https://arxiv.org/abs/2606.10829
Cite the original work for its findings. Save a collection to share your selection of sources.