Language Modeling on PTB zero-shot
82.05PerplexityAR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ARZero-shot=true, Training Tokens=524B, Training Dataset=OWT, Status=Retrained2024.06 | 82.05 | — | |
| Autoregressive TransformerZero-shot=true2026.02 | 82.05 | — | |
| AR†2026.06 | 82.05 | — | |
| RADDModel Size=Medium, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | 82.08 | — | |
| Uni-EDLM-NCE2026.06 | 86.18 | — | |
| SEDDModel Size=Medium, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | 87.12 | — | |
| Uni-EDLM-AR2026.06 | 87.36 | — | |
| DuoZero-shot=true, Number of Parameters=138M, Model Type=Diffusion2026.02 | 89.35 | — | |
| Duo‡2026.06 | 89.35 | — | |
| EDLM-CoAR†2026.06 | 89.73 | — | |
| Duo++Zero-shot=true, Number of Parameters=138M, Model Type=Diffusion, k=32026.02 | 91.94 | — | |
| EDLM-NCE†2026.06 | 93.21 | — | |
| Duo++Zero-shot=true, Number of Parameters=138M, Model Type=Diffusion, k=52026.02 | 94.46 | — | |
| Duo++Zero-shot=true, Number of Parameters=138M, Model Type=Diffusion, k=22026.02 | 94.96 | — | |
| MDLMZero-shot=true, Training Tokens=524B, Training Dataset=OWT2024.06 | 95.26 | — | |
| MDLM†2026.06 | 95.26 | — | |
| DCD*2026.06 | 95.69 | — | |
| Block Diffusion*2026.06 | 97.23 | — | |
| ARMDModel Size=Medium, Training Steps=300K, Evaluation Protocol=Zero-shot2026.01 | 97.75 | — | |
| GaussianTraining steps=1M, Training dataset=OpenWebText, Evaluation protocol=zero-shot2026.05 | 98.16 | — | |
| SEDDZero-shot=true, Training Tokens=524B, Training Dataset=OWT, Status=Retrained2024.06 | 100.09 | — | |
| SEDD†2026.06 | 100.09 | — | |
| ARMDModel Size=Medium, Training Steps=120K, Evaluation Protocol=Zero-shot2026.01 | 105.09 | — | |
| SEDD UniformZero-shot=true, Number of Parameters=138M, Model Type=Diffusion2026.02 | 105.51 | — | |
| RADDModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | 107.85 | — | |
| UDLMZero-shot=true, Number of Parameters=138M, Model Type=Diffusion2026.02 | 112.82 | — | |
| SEDDModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | 114.24 | — | |
| BaseTraining steps=1M, Training dataset=OpenWebText, Evaluation protocol=zero-shot2026.05 | 115.99 | — | |
| GPT-2Model Size=Medium, Evaluation Protocol=Zero-shot2026.01 | 123.14 | — | |
| ARMDModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | 123.43 | — | |
| ARMDModel Size=Small, Training Steps=180K, Evaluation Protocol=Zero-shot2026.01 | 130.31 | — | |
| GPT-2Model Size=Small, Evaluation Protocol=Zero-shot2026.01 | 138.43 | — | |
| SEDD-UniformModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | 140.12 | — | |
| PLAIDModel Size=Small, Training Steps=600K, Evaluation Protocol=Zero-shot2026.01 | 142.6 | — | |
| D3PMModel Size=Small, Evaluation Protocol=Zero-shot2026.01 | 200.82 | — | |
| BD3-LMEvaluation methodology=DUEL, Training data=OWT, Evaluation protocol=Zero-shot, Block size (L')=42026.03 | — | 81.8 | |
| MDLMEvaluation methodology=DUEL, Training data=OWT, Evaluation protocol=Zero-shot2026.03 | — | 34.3 | |
| SEDDEvaluation methodology=DUEL, Training data=OWT, Evaluation protocol=Zero-shot2026.03 | — | 31.3 |