DSL-LLaDA
Accepted to Findings of EMNLP 2026

Continuous denoising at 8B scale

DSL-LLaDA

Scaling continuous denoising to 8B scale.

Adapt LLaDA-8B to continuous embedding-space denoising in 1,000 steps. All response positions remain editable until final decoding, avoiding severe low-step repetition and early termination failures.

Longxuan Yu* · Yunshu Wu* · Yu Fu · Siheng Xiong · Rob Brekelmans · Hui Liu · Yue Dong · Greg Ver Steeg†

University of California, Riverside · Georgia Institute of Technology · Microsoft

* Equal contribution · † Corresponding author

8B LLaDA-Instruct backbone
1,000 DSL adaptation steps
≤16 Low-NFE evaluation focus
4 / 4 Best ROUGE-1 at NFE 8 and 16

Watch both decoding paths evolve.

Frozen matched-protocol traces compare continuous SDE refinement with discrete iterative unmasking under the same prompt, response length, and NFE budget.

NFE -- -- tokens seed --
Prompt

Loading trace data...

0%

DSL-LLaDA

Continuous state / Heun SDE
-- metric
Noise state SNR 0.01
Current decode

LLaDA

Discrete low-confidence remasking
-- metric
All positions masked 0 visible
Current decode

Source document

Full prompt


          

Bridge continuous diffusion to a pretrained 8B backbone.

Standard masked diffusion forces a binary input: a clean token or [MASK]. DSL-LLaDA adds continuous per-token Gaussian noise at varying SNR, then uses a lightweight softmax converter to feed soft embedding mixtures into the 8B backbone without architectural changes.

Original LLaDA uses discrete masked diffusion while DSL-LLaDA uses continuous per-token noise and a softmax converter
The converter exposes the existing backbone to three uncertainty phases: mask-like at low SNR, unreliable at intermediate SNR, and clear token evidence at high SNR.
01

Continuous Noise State

Inject per-token Gaussian noise at varying SNR into unit-sphere embeddings instead of using binary mask tokens.

02

Softmax Converter

Map continuous noisy states into soft mixtures of backbone embeddings without modifying the 8B architecture.

03

Joint SDE Denoising

Update all response positions jointly in continuous embedding space using second-order Heun integration.

04

Soft Updates, Final Readout

Predict token probabilities at every denoising step and use their weighted embedding to update the continuous state. Select hard token IDs only at the end.

Few steps, fewer failure modes.

We evaluate fluency, repetition, and output length together so neither early termination nor fluent repetition can masquerade as generation quality.

Generation perplexity, repetition rate, and output length across NFE budgets for DSL-LLaDA-SDE and discrete LLaDA baselines
Joint quality diagnostics (200 prompts, 256 tokens). At low NFE, iterative unmasking trades early termination for fluent looping. In this evaluation, DSL-LLaDA-SDE avoids that tradeoff, maintaining sub-9% repetition, more than 150 output words, and lower perplexity than vanilla LLaDA across all tested budgets.
4 / 4

Few-step summarization

Best ROUGE-1 across XSum, CNN/DM, PubMed, and arXiv at NFE 8 and 16 without length heuristics.

<9%

Adjacent-word repetition

Exact adjacent duplicates stay below 9% of word transitions from NFE 8 to 128, while the length-controlled EOS+block baseline exceeds 80% at NFE 8 and 16.

>98%

Clean-token preservation

Selectively corrects random-token corruptions while preserving more than 98% of uncorrupted positions.

Repair the errors. Preserve the document.

Training on continuously noised token states teaches DSL-LLaDA selective token-to-token editing: repair uncertain tokens while preserving clean ones. Compared with uniform random-token corruption, this behavior retains substantially more clean text and transfers to real OCR repair.

Original Impresso English test record 40 of 100

GLCO-1850-08-13-a-p0002_par58

3positions changed 3 / 3match the transcript 9 to 6probe word edits

Raw OCR input

The rates obtained for Euglish wheat were equal to last Mon¬
day's prices, but millers evice little desire to operate, but the
supply is limited.

There was heavy rain, with thander; but there are symptoms of change.

One-pass DSL output

The rates obtained for English wheat were equal to last Mon¬
day's prices, but millers evince little desire to operate, but the
supply is limited.

There was heavy rain, with thunder; but there are symptoms of change.

Single forward pass. All three repaired words exactly match the human transcript; every other model token is unchanged.

All official English test domains

DSL reduces macro word matching error rate on every domain.

Domain Docs Raw wMER DSL wMER Word PCIS, 95% CI
Impresso 100 7.23 6.57 0.0075 [0.0045, 0.0105]
ICDAR2017 101 12.83 11.62 0.0150 [0.0112, 0.0190]
Overproof 32 23.30 21.31 0.0306 [0.0203, 0.0415]

Lower wMER is better. Every word-level PCIS confidence interval is strictly above zero.

Impresso selectivity

DSL improves OCR while retaining nearly all clean tokens.

Method wMER Clean kept
Raw OCR 7.23 100.00%
DSL-LLaDA 6.57 99.85%
XDLM 20.60 88.01%

Clean kept is the share of originally correct words preserved.

Where continuous 8B diffusion excels.

The experiments identify three reliable gains from adapting an 8B masked diffusion model to continuous denoising.

Few-step grounded generation

At NFE 8 and 16, DSL-LLaDA-SDE achieves the best ROUGE-1 on XSum, CNN/DM, PubMed, and arXiv without length-control heuristics.

Stable full-length decoding

Across 200 open-ended prompts, SDE generates more than 150 words with less than 9% adjacent-word repetition from NFE 8 to 128, avoiding both early termination and low-step loops.

Selective token correction

A single forward pass repairs corrupted tokens while preserving more than 98% of clean positions. The same selective behavior reduces word-level OCR error across all three English test domains.

Paper, model, code, and traces.

Citation

@inproceedings{yu2026dsllada,
  title     = {DSL-LLaDA: Scaling Continuous Denoising to 8B Masked Diffusion LMs},
  author    = {Yu, Longxuan and Wu, Yunshu and Fu, Yu and Xiong, Siheng and Brekelmans, Rob and Liu, Hui and Dong, Yue and Ver Steeg, Greg},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026},
  url       = {https://arxiv.org/abs/2606.01024}
}