Continuous denoising at 8B scale
DSL-LLaDA
Scaling continuous denoising to 8B scale.
Adapt LLaDA-8B to continuous embedding-space denoising in 1,000 steps. All response positions remain editable until final decoding, avoiding severe low-step repetition and early termination failures.
University of California, Riverside · Georgia Institute of Technology · Microsoft
Interactive demo
Watch both decoding paths evolve.
Frozen matched-protocol traces compare continuous SDE refinement with discrete iterative unmasking under the same prompt, response length, and NFE budget.
Loading trace data...
LLaDA
Method
Bridge continuous diffusion to a pretrained 8B backbone.
Standard masked diffusion forces a binary input: a clean token or
[MASK]. DSL-LLaDA adds continuous per-token Gaussian
noise at varying SNR, then uses a lightweight softmax converter to
feed soft embedding mixtures into the 8B backbone without
architectural changes.
Continuous Noise State
Inject per-token Gaussian noise at varying SNR into unit-sphere embeddings instead of using binary mask tokens.
Softmax Converter
Map continuous noisy states into soft mixtures of backbone embeddings without modifying the 8B architecture.
Joint SDE Denoising
Update all response positions jointly in continuous embedding space using second-order Heun integration.
Soft Updates, Final Readout
Predict token probabilities at every denoising step and use their weighted embedding to update the continuous state. Select hard token IDs only at the end.
Results
Few steps, fewer failure modes.
We evaluate fluency, repetition, and output length together so neither early termination nor fluent repetition can masquerade as generation quality.
Few-step summarization
Best ROUGE-1 across XSum, CNN/DM, PubMed, and arXiv at NFE 8 and 16 without length heuristics.
Adjacent-word repetition
Exact adjacent duplicates stay below 9% of word transitions from NFE 8 to 128, while the length-controlled EOS+block baseline exceeds 80% at NFE 8 and 16.
Clean-token preservation
Selectively corrects random-token corruptions while preserving more than 98% of uncorrupted positions.
Selective OCR repair
Repair the errors. Preserve the document.
Training on continuously noised token states teaches DSL-LLaDA selective token-to-token editing: repair uncertain tokens while preserving clean ones. Compared with uniform random-token corruption, this behavior retains substantially more clean text and transfers to real OCR repair.
Original Impresso English test record 40 of 100
GLCO-1850-08-13-a-p0002_par58
Raw OCR input
The rates obtained for
Euglish wheat were equal to
last Mon¬
day's prices, but millers
evice little desire to
operate, but the
supply is limited.
There was heavy rain, with thander; but there are symptoms of change.
One-pass DSL output
The rates obtained for
English wheat were equal to
last Mon¬
day's prices, but millers
evince little desire to
operate, but the
supply is limited.
There was heavy rain, with thunder; but there are symptoms of change.
All official English test domains
DSL reduces macro word matching error rate on every domain.
| Domain | Docs | Raw wMER | DSL wMER | Word PCIS, 95% CI |
|---|---|---|---|---|
| Impresso | 100 | 7.23 | 6.57 | 0.0075 [0.0045, 0.0105] |
| ICDAR2017 | 101 | 12.83 | 11.62 | 0.0150 [0.0112, 0.0190] |
| Overproof | 32 | 23.30 | 21.31 | 0.0306 [0.0203, 0.0415] |
Lower wMER is better. Every word-level PCIS confidence interval is strictly above zero.
Impresso selectivity
DSL improves OCR while retaining nearly all clean tokens.
| Method | wMER | Clean kept |
|---|---|---|
| Raw OCR | 7.23 | 100.00% |
| DSL-LLaDA | 6.57 | 99.85% |
| XDLM | 20.60 | 88.01% |
Clean kept is the share of originally correct words preserved.
Scope
Where continuous 8B diffusion excels.
The experiments identify three reliable gains from adapting an 8B masked diffusion model to continuous denoising.
Few-step grounded generation
At NFE 8 and 16, DSL-LLaDA-SDE achieves the best ROUGE-1 on XSum, CNN/DM, PubMed, and arXiv without length-control heuristics.
Stable full-length decoding
Across 200 open-ended prompts, SDE generates more than 150 words with less than 9% adjacent-word repetition from NFE 8 to 128, avoiding both early termination and low-step loops.
Selective token correction
A single forward pass repairs corrupted tokens while preserving more than 98% of clean positions. The same selective behavior reduces word-level OCR error across all three English test domains.
Resources
Paper, model, code, and traces.
Citation
@inproceedings{yu2026dsllada,
title = {DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs},
author = {Yu, Longxuan and Wu, Yunshu and Fu, Yu and Xiong, Siheng and Brekelmans, Rob and Liu, Hui and Dong, Yue and Ver Steeg, Greg},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
year = {2026},
url = {https://arxiv.org/abs/2606.01024}
}