Language Modeling on Lambada (val)
10.14PerplexityGlauber-UL2-L
Evaluation Results
| Method | Links | |
|---|---|---|
| Glauber-UL2-LZero-shot=true, Model size=Large, N value=3, Note=upper bound2026.05 | 10.14 | |
| GPT-2-LZero-shot=true, Model size=Large2026.05 | 10.87 | |
| Glauber-UL2-LZero-shot=true, Model size=Large, N value=1, Note=upper bound2026.05 | 11.25 | |
| MDM-Prime-v2zero-shot=true, ℓ=16, variant=*2026.03 | 12.37 | |
| GPT-2-MZero-shot=true, Model size=Medium2026.05 | 15.6 | |
| Glauber-UL2-MZero-shot=true, Model size=Medium, N value=3, Note=upper bound2026.05 | 17.14 | |
| Glauber-UL2-MZero-shot=true, Model size=Medium, N value=1, Note=upper bound2026.05 | 17.89 | |
| MDM-Primezero-shot=true, ℓ=42026.03 | 24.44 | |
| MDM-Primezero-shot=true, ℓ=82026.03 | 25.23 | |
| MDM-Primezero-shot=true, ℓ=62026.03 | 25.8 | |
| MDM-Prime-v2zero-shot=true, ℓ=162026.03 | 28.4 | |
| MDM-Primezero-shot=true, ℓ=22026.03 | 30.91 | |
| MDM-Primezero-shot=true, ℓ=6, variant=*2026.03 | 36.15 | |
| ARMzero-shot=true, variant=*2026.03 | 37.52 | |
| MDMzero-shot=true, variant=*2026.03 | 40.43 | |
| SEDD-MZero-shot=true, Model size=Medium, Note=upper bound2026.05 | 42.77 | |
| LoopMDMS (inference-time loop count)=6, Zero-shot protocol=true, Model parameters=170M2026.05 | 44.6 | |
| LoopMDMS (inference-time loop count)=12, Zero-shot protocol=true, Model parameters=170M2026.05 | 45.8 | |
| DCDMParam=0.1B, Training Tokens=128B, Evaluation Protocol=Zero-shot2026.05 | 46.43 | |
| EDLM-NCEzero-shot=true2026.03 | 46.92 | |
| MDMzero-shot=true2026.03 | 47.52 | |
| MDLMParam=0.1B, Training Tokens=256B, Evaluation Protocol=Zero-shot2026.05 | 48.29 | |
| MDMZero-shot protocol=true, Model parameters=170M2026.05 | 49.6 | |
| Duozero-shot=true2026.03 | 49.78 | |
| SEDDzero-shot=true2026.03 | 49.86 | |
| BD3-LMzero-shot=true2026.03 | 50.03 | |
| BDLMParam=0.1B, Training Tokens=256B, Evaluation Protocol=Zero-shot2026.05 | 50.03 | |
| EDLM-coARzero-shot=true2026.03 | 50.04 | |
| BD3-LMTraining steps=250K, Training Dataset=OWT, Zero-shot=true, L'=162025.06 | 50.05 | |
| ARMzero-shot=true2026.03 | 51.28 | |
| AR TransformerTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 51.69 | |
| MDLMTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 52.06 | |
| ARParam=0.1B, Evaluation Protocol=Zero-shot2026.05 | 52.13 | |
| SEDD AbsorbTraining steps=250K, Training Dataset=OWT, Zero-shot=true2025.06 | 52.16 | |
| LoopMDMS (inference-time loop count)=1, Zero-shot protocol=true, Model parameters=170M2026.05 | 53.3 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.52025.06 | 57.33 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.252025.06 | 60.15 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=12025.06 | 61.37 | |
| Eso-LMsTraining steps=250K, Training Dataset=OWT, Zero-shot=true, alpha_0=0.1252025.06 | 69.13 |