Language Modeling on arXiv
2.46PerplexitySHAREDLLM
Evaluation Results
| Method | Links | |
|---|---|---|
| SHAREDLLMContext Length=32K2026.03 | 2.46 | |
| CEPEContext Length=32K2026.03 | 2.51 | |
| YaRNContext Length=32K2026.03 | 2.58 | |
| LLaMA-3.1Context Length=32K2026.03 | 2.63 | |
| PIContext Length=32K2026.03 | 2.77 | |
| SHAREDLLMContext Length=128K2026.03 | 2.91 | |
| LLaMA-2-32KContext Length=32K2026.03 | 2.96 | |
| CEPEContext Length=128K2026.03 | 2.97 | |
| SHAREDLLMContext Length=8K2026.03 | 2.97 | |
| SHAREDLLMContext Length=4K2026.03 | 2.99 | |
| CEPEContext Length=8K2026.03 | 3.02 | |
| CEPEContext Length=4K2026.03 | 3.03 | |
| YaRNContext Length=8K2026.03 | 3.09 | |
| LLaMA-3.1Context Length=128K2026.03 | 3.12 | |
| LLaMA-3.1Context Length=4K2026.03 | 3.17 | |
| PIContext Length=8K2026.03 | 3.21 | |
| LLaMA-3.1Context Length=8K2026.03 | 3.26 | |
| LLaMA-2-32KContext Length=8K2026.03 | 3.34 | |
| YaRNContext Length=4K2026.03 | 3.35 | |
| PIContext Length=4K2026.03 | 3.49 | |
| LLaMA-2-32KContext Length=4K2026.03 | 3.58 | |
| (w+a)kNN-LMa (rescore)LMa=domain-adapted GPT-2, datastore=combined, rescore=true, context_representation=LN22022.11 | 17.47 | |
| (a)kNN-LMa (rescore)LMa=domain-adapted GPT-2, datastore=adaptation, rescore=true, context_representation=LN22022.11 | 17.49 | |
| (a)kNN-LMaLMa=domain-adapted GPT-2, datastore=adaptation, context_representation=LN22022.11 | 17.81 | |
| (w+a)kNN-LMaLMa=domain-adapted GPT-2, datastore=combined, context_representation=LN22022.11 | 17.85 | |
| ARMzero-shot=true, number of parameters=1B, training tokens=300B2026.01 | 18.15 | |
| CARDzero-shot=true, number of parameters=1B, training tokens=300B2026.01 | 20.34 | |
| MDLMzero-shot=true, number of parameters=1B, training tokens=300B2026.01 | 23.58 | |
| (w)kNN-LMa (rescore)LMa=domain-adapted GPT-2, datastore=pretraining, rescore=true, context_representation=LN22022.11 | 24.19 | |
| (a)kNN-LMLM=GPT-2, datastore=adaptation, context_representation=original2022.11 | 24.38 | |
| (w)kNN-LMaLMa=domain-adapted GPT-2, datastore=pretraining, context_representation=LN22022.11 | 24.42 | |
| LMa (only)LMa=domain-adapted GPT-22022.11 | 24.97 | |
| ABDZero-shot=true, Block Size (BS)=42026.06 | 37.11 | |
| XDLMSetting=Zero-shot2026.02 | 37.232 | |
| MDLMModel Category=Diffusion (Absorbing-state), Zero-shot=true, Training Data=OWT, Training Steps=1M2026.04 | 37.37 | |
| MDLMSetting=Zero-shot2026.02 | 37.457 | |
| SCMDMAddZero-shot=true, Training tokens=262B + 3.25B, Pre-training dataset=OWT2026.04 | 37.67 | |
| SCMDMConcatZero-shot=true, Training tokens=262B + 3.25B, Pre-training dataset=OWT2026.04 | 37.75 | |
| MDLMZero-shot=true2026.06 | 37.89 | |
| LangFlowModel Category=Diffusion (Uniform-state / Continuous), Zero-shot=true, Training Data=OWT, Training Steps=1M2026.04 | 38.47 | |
| SEDD AbsorbModel Category=Diffusion (Absorbing-state), Zero-shot=true, Training Data=OWT, Training Steps=1M2026.04 | 38.48 | |
| GIDDSetting=Zero-shot2026.02 | 39.019 | |
| ABDZero-shot=true, Block Size (BS)=12026.06 | 39.11 | |
| MDLMZero-shot=true, Training tokens=262B + 3.25B, Pre-training dataset=OWT2026.04 | 39.14 | |
| BD3LMZero-shot=true, Block Size (BS)=42026.06 | 39.2 | |
| SEDDZero-shot=true2026.06 | 40.03 | |
| DuoModel Category=Diffusion (Uniform-state / Continuous), Zero-shot=true, Training Data=OWT, Training Steps=1M2026.04 | 40.39 | |
| ARZero-shot=true2026.06 | 41.22 | |
| Autoregressive TransformerModel Category=Autoregressive, Zero-shot=true, Training Data=OWT, Training Steps=1M2026.04 | 41.73 | |
| UDLMSetting=Zero-shot2026.02 | 42.671 | |
| UDLMModel Category=Diffusion (Uniform-state / Continuous), Zero-shot=true, Training Data=OWT, Training Steps=1M2026.04 | 44.08 | |
| BD3LMzero-shot=true, number of parameters=1B, training tokens=300B2026.01 | 44.6 | |
| SEDD UniformModel Category=Diffusion (Uniform-state / Continuous), Zero-shot=true, Training Data=OWT, Training Steps=1M2026.04 | 50.86 | |
| (w)kNN-LMLM=GPT-2, datastore=pretraining, context_representation=LN22022.11 | 51.89 | |
| (w)kNN-LMLM=GPT-2, datastore=pretraining, context_representation=original2022.11 | 53.03 | |
| LM (only)LM=GPT-22022.11 | 56.83 | |
| DID-FSize=Medium, Zero-shot=true, Alignment=FLOPs-aligned2026.03 | 61.77 | |
| DID-SSize=Medium, Zero-shot=true, Alignment=Steps-aligned2026.03 | 63.95 | |
| RADDSize=Medium, Zero-shot=true2026.03 | 66.28 | |
| DID-FSize=Small, Zero-shot=true, Alignment=FLOPs-aligned2026.03 | 78.38 | |
| DID-SSize=Small, Zero-shot=true, Alignment=Steps-aligned2026.03 | 82.41 | |
| RADDSize=Small, Zero-shot=true2026.03 | 85.95 | |
| (a)kNN (only)datastore=adaptation2022.11 | 87.58 | |
| (w)kNN (only)datastore=pretraining2022.11 | 513.45 |