Language Modeling on LAMBADA zero-shot (test)
69.12Accuracy (zero-shot)GPTQ
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPTQBackbone=OPT-30B, Bit-width=4 bits, Hardware=A100, Protocol=Zero-shot2026.01 | 69.12 | — | |
| Original GPT-2Model Size=1.5B, Data=different data, Zero-shot=true2021.08 | 63.24 | — | |
| GPTQBackbone=OPT-2.7B, Bit-width=4 bits, Hardware=A100, Protocol=Zero-shot2026.01 | 62.97 | — | |
| Megatron-LM GPT-2Model Size=2.5B, Data=same data, Zero-shot=true2021.08 | 61.73 | — | |
| SLWModel Size=1.5B, Pre-training parameters=bsz4K-seqlen1K, Training steps=58.8K, Training tokens=157B, Training time=155Hr, Zero-shot=true, SLW Version=45K2021.08 | 58.47 | — | |
| SLWModel Size=1.5B, Pre-training parameters=bsz4K-seqlen1K, Training steps=50K, Training tokens=121B, Training time=121Hr, Zero-shot=true, SLW Version=45K2021.08 | 58.2 | — | |
| SLWModel Size=1.5B, Pre-training parameters=bsz512-seqlen1K, Training steps=428K, Training tokens=157B, Training time=364Hr, Zero-shot=true, SLW Version=270K2021.08 | 57.89 | — | |
| SLWModel Size=1.5B, Pre-training parameters=bsz512-seqlen1K, Training steps=360K, Training tokens=122B, Training time=286Hr, Zero-shot=true, SLW Version=270K2021.08 | 57.38 | — | |
| BaselineModel Size=1.5B, Pre-training parameters=bsz512-seqlen1K, Training steps=300K, Training tokens=157B, Training time=341Hr, Zero-shot=true2021.08 | 57.29 | — | |
| ShortformerModel Size=1.5B, Pre-training parameters=bsz4K-seqlen1K, Training steps=55K, Training tokens=157B, Training time=162Hr, Zero-shot=true2021.08 | 57.23 | — | |
| GPTQBackbone=OPT-1.3B, Bit-width=4 bits, Hardware=A100, Protocol=Zero-shot2026.01 | 56.45 | — | |
| Bsz WarmupModel Size=1.5B, Pre-training parameters=bsz4K-seqlen1K, Training steps=58.8K, Training tokens=157B, Training time=165Hr, Zero-shot=true2021.08 | 56.36 | — | |
| BaselineModel Size=1.5B, Pre-training parameters=bsz4K-seqlen1K, Training steps=37.5K, Training tokens=157B, Training time=151Hr, Zero-shot=true2021.08 | 55.06 | — | |
| Original GPT-2Model Size=117M, Data=different data, Zero-shot=true2021.08 | 45.99 | — | |
| Megatron-LM GPT-2Model Size=355M, Data=same data, Zero-shot=true2021.08 | 45.18 | — | |
| Uniform-RTNBackbone=QWEN-7B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 38.75 | — | |
| FP16Backbone=QWEN-7B, Bit-width=16 bits, Protocol=Zero-shot2026.01 | 38.36 | — | |
| FP16Backbone=QWEN3-14B, Bit-width=16 bits, Protocol=Zero-shot2026.01 | 38.36 | — | |
| Uniform-RTNBackbone=QWEN3-14B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 38.01 | — | |
| Benford-QuantBackbone=OPT-30B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 36.77 | — | |
| Uniform-RTNBackbone=OPT-30B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 36.67 | — | |
| Benford-QuantBackbone=QWEN3-14B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 35.52 | — | |
| SLWModel Size=117M, Pre-training parameters=bsz512-seqlen1K, Training steps=200K, Training tokens=89B, Training time=20Hr, Zero-shot=true, SLW Version=60K2021.08 | 34.78 | — | |
| SLWModel Size=117M, Pre-training parameters=bsz512-seqlen2K, Training steps=205K, Training tokens=157B, Training time=31Hr, Zero-shot=true, SLW Version=110K2021.08 | 34.58 | — | |
| SLWModel Size=117M, Pre-training parameters=bsz512-seqlen1K, Training steps=330K, Training tokens=157B, Training time=33Hr, Zero-shot=true, SLW Version=60K2021.08 | 34.41 | — | |
| SLWModel Size=117M, Pre-training parameters=bsz4K-seqlen1K, Training steps=52.5K, Training tokens=157B, Training time=16Hr, Zero-shot=true, SLW Version=30K2021.08 | 34.16 | — | |
| SLWModel Size=117M, Pre-training parameters=bsz4K-seqlen1K, Training steps=37K, Training tokens=92B, Training time=10Hr, Zero-shot=true, SLW Version=30K2021.08 | 33.4 | — | |
| SLWModel Size=117M, Pre-training parameters=bsz512-seqlen2K, Training steps=122.5K, Training tokens=71B, Training time=15Hr, Zero-shot=true, SLW Version=110K2021.08 | 33.24 | — | |
| BaselineModel Size=117M, Pre-training parameters=bsz512-seqlen1K, Training steps=300K, Training tokens=157B, Training time=37Hr, Zero-shot=true2021.08 | 33.19 | — | |
| BaselineModel Size=117M, Pre-training parameters=bsz512-seqlen2K, Training steps=150K, Training tokens=157B, Training time=32Hr, Zero-shot=true2021.08 | 32.99 | — | |
| Benford-QuantBackbone=QWEN-7B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 32.56 | — | |
| BaselineModel Size=117M, Pre-training parameters=bsz4K-seqlen1K, Training steps=37.5K, Training tokens=157B, Training time=16Hr, Zero-shot=true2021.08 | 32.54 | — | |
| Benford-QuantBackbone=OPT-2.7B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 32 | — | |
| Uniform-RTNBackbone=OPT-2.7B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 31.59 | — | |
| SwitchHead MAC-matchedtotal params=376M2023.12 | 30.2 | — | |
| SwitchHeadtotal params=262M2023.12 | 29.4 | — | |
| Uniform-RTNBackbone=OPT-1.3B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 29.2 | — | |
| Benford-QuantBackbone=OPT-1.3B, Bit-width=4 bits, Hardware=H200, Protocol=Zero-shot2026.01 | 29.05 | — | |
| SwitchHead Shared selectiontotal params=262M2023.12 | 28.6 | — | |
| Transformertotal params=262M2023.12 | 28.2 | — | |
| SwitchHead MAC-matchedtotal params=63M2023.12 | 23.5 | — | |
| SwitchHeadtotal params=47M2023.12 | 20.4 | — | |
| Transformertotal params=47M2023.12 | 20.4 | — | |
| SwitchHead Shared selectiontotal params=47M2023.12 | 20 | — | |
| Absorbing Uniform DiffusionLoss=ELBO, Param.=Denoiser, Zero-shot transfer=true2026.05 | — | 45.87 | |
| ARZero-shot=true, Training Tokens=524B, Training Dataset=OWT, Status=Retrained2024.06 | — | 51.28 | |
| AR†2026.06 | — | 51.28 | |
| ARMDModel Size=Small, Training Steps=180K, Evaluation Protocol=Zero-shot2026.01 | — | 44.66 | |
| ARMDModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | — | 45.35 | |
| ARMDModel Size=Medium, Training Steps=120K, Evaluation Protocol=Zero-shot2026.01 | — | 40.62 | |
| ARMDModel Size=Medium, Training Steps=300K, Evaluation Protocol=Zero-shot2026.01 | — | 39.08 | |
| BaseTraining steps=1M, Training dataset=OpenWebText, Evaluation protocol=zero-shot2026.05 | — | 48.18 | |
| Block Diffusion*2026.06 | — | 49.5 | |
| D3PMModel Size=Small, Evaluation Protocol=Zero-shot2026.01 | — | 93.47 | |
| DCD*2026.06 | — | 46.71 | |
| Duo‡2026.06 | — | 49.78 | |
| EDLM-CoAR†2026.06 | — | 50.04 | |
| EDLM-NCE†2026.06 | — | 46.92 | |
| GaussianTraining steps=1M, Training dataset=OpenWebText, Evaluation protocol=zero-shot2026.05 | — | 46.39 | |
| GPT-2Model Size=Small, Evaluation Protocol=Zero-shot2026.01 | — | 45.04 | |
| GPT-2Model Size=Medium, Evaluation Protocol=Zero-shot2026.01 | — | 35.66 | |
| Masked DiffusionLoss=ELBO, Param.=Denoiser, Zero-shot transfer=true2026.05 | — | 48.79 | |
| Max Coupling Uniform DiffusionLoss=ELBO, Param.=Denoiser, Zero-shot transfer=true2026.05 | — | 64.71 | |
| Max Coupling Uniform DiffusionLoss=ELBO, Param.=LOO-Denoiser, Zero-shot transfer=true2026.05 | — | 62.44 | |
| MDLMZero-shot=true, Training Tokens=524B, Training Dataset=OWT2024.06 | — | 47.52 | |
| MDLM†2026.06 | — | 47.52 | |
| PLAIDModel Size=Small, Training Steps=600K, Evaluation Protocol=Zero-shot2026.01 | — | 57.28 | |
| RADDModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | — | 51.7 | |
| RADDModel Size=Medium, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | — | 44.1 | |
| SEDDZero-shot=true, Training Tokens=524B, Training Dataset=OWT, Status=Retrained2024.06 | — | 49.86 | |
| SEDDModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | — | 50.92 | |
| SEDDModel Size=Medium, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | — | 42.77 | |
| SEDD-UniformModel Size=Small, Training Steps=400K, Evaluation Protocol=Zero-shot2026.01 | — | 65.4 | |
| SEDD†2026.06 | — | 49.86 | |
| Uni-EDLM-AR2026.06 | — | 46.62 | |
| Uni-EDLM-NCE2026.06 | — | 48.26 | |
| Uniform DiffusionLoss=Cross-Entropy, Param.=Denoiser, Zero-shot transfer=true2026.05 | — | 49.04 | |
| Uniform DiffusionLoss=Cross-Entropy, Param.=LOO-Denoiser, Zero-shot transfer=true2026.05 | — | 46.91 | |
| Uniform DiffusionLoss=ELBO, Param.=Denoiser, Zero-shot transfer=true2026.05 | — | 46.8 | |
| Uniform DiffusionLoss=ELBO, Param.=LOO-Denoiser, Zero-shot transfer=true2026.05 | — | 46.22 |