Language Modeling on 1BW
21.8PerplexityADEPT
Evaluation Results
| Method | Links | |
|---|---|---|
| ADEPTArchitecture=Llama 7B, Reduction=10%2026.01 | 21.8 | |
| ADEPTArchitecture=Llama 7B, Reduction=20%2026.01 | 22.7 | |
| PABEEArchitecture=Llama 7B, Reduction=20%2026.01 | 26.3 | |
| PABEEArchitecture=Llama 7B, Reduction=10%2026.01 | 27 | |
| DeeBERTArchitecture=Llama 7B, Reduction=10%2026.01 | 28.5 | |
| DeeBERTArchitecture=Llama 7B, Reduction=20%2026.01 | 30.4 | |
| Base ModelArchitecture=Llama 7B, Reduction=None2026.01 | 31.6 | |
| ADEPTArchitecture=GPT2 XL, Reduction=10%2026.01 | 31.8 | |
| ADEPTArchitecture=GPT2 XL, Reduction=20%2026.01 | 33.4 | |
| PABEEArchitecture=GPT2 XL, Reduction=10%2026.01 | 34.5 | |
| PABEEArchitecture=GPT2 XL, Reduction=20%2026.01 | 34.8 | |
| DeeBERTArchitecture=GPT2 XL, Reduction=10%2026.01 | 37.8 | |
| DeeBERTArchitecture=GPT2 XL, Reduction=20%2026.01 | 38.9 | |
| Base ModelArchitecture=GPT2 XL, Reduction=None2026.01 | 42.2 | |
| GPT-2Model Size=medium, Evaluation Protocol=Zero-shot2024.06 | 55.72 | |
| RADD-AOModel Size=medium, Evaluation Protocol=Zero-shot, Training Iterations=400k2024.06 | 57.07 | |
| RADD-DSEModel Size=medium, Evaluation Protocol=Zero-shot, Training Iterations=400k2024.06 | 57.45 | |
| RADD-t-DCEModel Size=medium, Evaluation Protocol=Zero-shot, Training Iterations=400k2024.06 | 57.95 | |
| RADD-λ-DCEModel Size=medium, Evaluation Protocol=Zero-shot, Training Iterations=400k2024.06 | 60.32 | |
| SEDD-ScaleModel Size=medium, Evaluation Protocol=Zero-shot, Training Iterations=400k2024.06 | 61.19 | |
| Manta-LMZero-shot=true, Protocol=ELBO-based proxy, Model parameters=101M2026.05 | 62.55 | |
| SEDD-UnscaleModel Size=medium, Evaluation Protocol=Zero-shot, Training Iterations=400k2024.06 | 67.91 | |
| MD4Zero-shot=true, Protocol=ELBO-based proxy2026.05 | 68.1 | |
| RADD-DSEModel Size=small, Zero-shot=true, Iterations=400k2024.06 | 72.35 | |
| RADD-t-DCEModel Size=small, Zero-shot=true, Iterations=400k2024.06 | 72.6 | |
| RADD-λ-DCEModel Size=small, Zero-shot=true, Iterations=400k2024.06 | 72.99 | |
| RADDZero-shot=true, Protocol=ELBO-based proxy2026.05 | 72.99 | |
| RADD-AOModel Size=small, Zero-shot=true, Iterations=400k2024.06 | 74.28 | |
| GPT-2Model Size=small, Zero-shot=true2024.06 | 75.2 | |
| GPT-2Zero-shot=true2026.05 | 75.2 | |
| AO-GPTZero-shot=true2026.05 | 76.73 | |
| SEDD-ScaleModel Size=small, Zero-shot=true, Iterations=400k2024.06 | 79.29 | |
| SEDDZero-shot=true, Protocol=ELBO-based proxy2026.05 | 79.29 | |
| SEDD-UnscaleModel Size=small, Zero-shot=true, Iterations=400k2024.06 | 80.7 | |
| PLAIDModel Size=small, Zero-shot=true2024.06 | 91.12 | |
| PLAIDZero-shot=true, Protocol=ELBO-based proxy2026.05 | 91.12 | |
| SEDD-UniformModel Size=small, Zero-shot=true, Iterations=400k2024.06 | 101.37 | |
| D3PMModel Size=small, Zero-shot=true2024.06 | 138.92 | |
| D3PMZero-shot=true, Protocol=ELBO-based proxy2026.05 | 138.92 |