Language Modeling on WikiText (val)
12.51PerplexityMDM-Prime-v2
Evaluation Results
| Method | Links | |
|---|---|---|
| MDM-Prime-v2zero-shot=true, ℓ=16, variant=*2026.03 | 12.51 | |
| Block Diffusion + Gumbel DistillationTraining Dataset=OWT, Evaluation Protocol=Zero-shot2026.03 | 13.86 | |
| MDLM + Gumbel DistillationTraining Dataset=OWT, Evaluation Protocol=Zero-shot2026.03 | 15.57 | |
| MDM-Primezero-shot=true, ℓ=6, variant=*2026.03 | 19.46 | |
| AsyncMeshMesh Configuration=4 x 2, Iterations=30k, AsyncPP=true2026.01 | 21.14 | |
| FullSyncMesh Configuration=4 x 2, Iterations=30k, AsyncPP=false2026.01 | 21.23 | |
| DPAvgMesh Configuration=4 x 2, Iterations=30k, AsyncPP=true2026.01 | 21.26 | |
| SPARTAMesh Configuration=4 x 2, Iterations=30k, AsyncPP=true2026.01 | 21.3 | |
| ARMzero-shot=true, variant=*2026.03 | 22.28 | |
| MDM-Primezero-shot=true, ℓ=42026.03 | 23.86 | |
| AsyncSPARTAMesh Configuration=4 x 2, Iterations=30k, AsyncPP=true2026.01 | 24.8 | |
| ARParam=0.1B, Evaluation Protocol=Zero-shot2026.05 | 25.32 | |
| ARMzero-shot=true2026.03 | 25.75 | |
| MDM-Primezero-shot=true, ℓ=82026.03 | 25.77 | |
| MDM-Prime-v2zero-shot=true, ℓ=162026.03 | 26.05 | |
| AR TransformerTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 26.54 | |
| MDMzero-shot=true, variant=*2026.03 | 26.65 | |
| MDM-Primezero-shot=true, ℓ=62026.03 | 26.87 | |
| MDM-Primezero-shot=true, ℓ=22026.03 | 27.93 | |
| EDLM-coARzero-shot=true2026.03 | 28.31 | |
| EDLM-NCEzero-shot=true2026.03 | 30.77 | |
| Block DiffusionTraining Dataset=OWT, Evaluation Protocol=Zero-shot2026.03 | 30.82 | |
| MASA-QKVModel size=729M, Attn CR=50.0%, Zero-shot=true2025.08 | 30.83 | |
| Transformer-LModel size=729M, Attn CR=0%, Zero-shot=true2025.08 | 30.88 | |
| BD3-LMzero-shot=true2026.03 | 31.31 | |
| BDLMParam=0.1B, Training Tokens=256B, Evaluation Protocol=Zero-shot2026.05 | 31.31 | |
| MASA-QKVOModel size=729M, Attn CR=66.7%, Zero-shot=true2025.08 | 31.34 | |
| GQAModel size=729M, Attn CR=41.7%, Zero-shot=true2025.08 | 31.74 | |
| DCDMParam=0.1B, Training Tokens=128B, Evaluation Protocol=Zero-shot2026.05 | 31.94 | |
| Repeat-all-overModel size=729M, Attn CR=66.7%, Zero-shot=true2025.08 | 32.27 | |
| LoopMDMS (inference-time loop count)=12, Zero-shot protocol=true, Model parameters=170M2026.05 | 32.4 | |
| Seq-SharingModel size=729M, Attn CR=66.7%, Zero-shot=true2025.08 | 32.43 | |
| MDMzero-shot=true2026.03 | 32.83 | |
| BD3-LMTraining steps=250K, Training Dataset=OWT, Zero-shot=true, L'=162025.06 | 32.88 | |
| MDLMParam=0.1B, Training Tokens=256B, Evaluation Protocol=Zero-shot2026.05 | 33.22 | |
| Low-RankModel size=729M, Attn CR=66.7%, Zero-shot=true2025.08 | 33.28 | |
| Duozero-shot=true2026.03 | 33.57 | |
| LoopMDMS (inference-time loop count)=6, Zero-shot protocol=true, Model parameters=170M2026.05 | 34.2 | |
| SEDDzero-shot=true2026.03 | 34.28 | |
| MDLMTraining Dataset=OWT, Evaluation Protocol=Zero-shot2026.03 | 34.52 | |
| MDMZero-shot protocol=true, Model parameters=170M2026.05 | 35.1 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.1252025.06 | 35.65 | |
| MDLMTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 37.08 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.252025.06 | 37.32 | |
| SEDD AbsorbTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 38.55 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.52025.06 | 39.57 | |
| LoopMDMS (inference-time loop count)=1, Zero-shot protocol=true, Model parameters=170M2026.05 | 41.8 | |
| MASA-QKVModel size=335M, Attn CR=50.0%, Zero-shot=true2025.08 | 42.31 | |
| Transformer-MModel size=335M, Attn CR=0%, Zero-shot=true2025.08 | 44.49 | |
| MASA-QKVOModel size=335M, Attn CR=66.7%, Zero-shot=true2025.08 | 45 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=12025.06 | 45.08 | |
| GQAModel size=335M, Attn CR=43.8%, Zero-shot=true2025.08 | 46.21 | |
| Seq-SharingModel size=335M, Attn CR=66.7%, Zero-shot=true2025.08 | 47.36 | |
| Low-RankModel size=335M, Attn CR=66.7%, Zero-shot=true2025.08 | 47.48 | |
| Repeat-all-overModel size=335M, Attn CR=66.7%, Zero-shot=true2025.08 | 47.63 | |
| MASA-QKVModel size=110M, Attn CR=50.0%, Zero-shot=true2025.08 | 72.08 | |
| MASA-QKVOModel size=110M, Attn CR=66.7%, Zero-shot=true2025.08 | 72.82 | |
| Transformer-SModel size=110M, Attn CR=0%, Zero-shot=true2025.08 | 76.11 | |
| GQAModel size=110M, Attn CR=41.7%, Zero-shot=true2025.08 | 78.41 | |
| Repeat-all-overModel size=110M, Attn CR=66.7%, Zero-shot=true2025.08 | 78.97 | |
| Seq-SharingModel size=110M, Attn CR=66.7%, Zero-shot=true2025.08 | 80.35 | |
| Low-RankModel size=110M, Attn CR=66.7%, Zero-shot=true2025.08 | 83.25 |