Language Modeling on LAMBADA (Accuracy)
79.4AccuracyLLaMA-2 70B
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaMA-2 70BModel=LLaMA-2 70B2026.03 | 79.4 | |
| Focus (LLaMA-2 70B)Backbone=LLaMA-2 70B, K=4, dg=162026.03 | 79.4 | |
| LLaMA-2 13BModel=LLaMA-2 13B2026.03 | 76.6 | |
| Focus (LLaMA-2 13B)Backbone=LLaMA-2 13B, K=2, dg=162026.03 | 76.6 | |
| Gemma 2 9BModel=Gemma 2 9B2026.03 | 75.5 | |
| Focus (Gemma 2 9B)Backbone=Gemma 2 9B, K=4, dg=162026.03 | 75.5 | |
| Mistral 7BModel=Mistral 7B2026.03 | 75.3 | |
| Focus (Mistral 7B)Backbone=Mistral 7B, K=4, dg=162026.03 | 75.3 | |
| OLMo-2 7BModel=OLMo-2 7B2026.03 | 73.2 | |
| Focus (OLMo-2 7B)Backbone=OLMo-2 7B, K=4, dg=162026.03 | 73.2 | |
| Qwen 2.5 7BModel=Qwen 2.5 7B2026.03 | 70.7 | |
| Focus (Qwen 2.5 7B)Backbone=Qwen 2.5 7B, K=4, dg=162026.03 | 70.7 | |
| Fullw_eff=∞2026.04 | 64.6 | |
| Stochasticw_eff=1282026.04 | 64.6 | |
| SWAw_eff=2562026.04 | 64.6 | |
| Stochasticw_eff=2562026.04 | 64.6 | |
| MoBA (k=2)w_eff=2562026.04 | 64.6 | |
| SWAw_eff=1282026.04 | 64.5 | |
| MoBA (k=2)w_eff=1282026.04 | 64.5 | |
| NITPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 64.49 | |
| IntAttentionModel=Llama-3.2-1B2025.11 | 63.61 | |
| FP16Model=Llama-3.2-1B2025.11 | 62.95 | |
| Stochasticw_eff=642026.04 | 62.9 | |
| Quant-OnlyModel=Llama-3.2-1B2025.11 | 62.62 | |
| MoHGE-3BGPU Utilization=balanced, Parameter Scale=3B2026.04 | 62.37 | |
| NTPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 62.32 | |
| HMoE-3BGPU Utilization=unbalanced, Parameter Scale=3B2026.04 | 62.25 | |
| Llama-3.2-1BModel Size Category=XL, Parameters=908M-1.2B, Zero-shot=true, Closed-book=true2025.07 | 62.1 | |
| MoDSE-3BGPU Utilization=balanced, Parameter Scale=3B2026.04 | 61.98 | |
| AdamWBackbone=Llama3-1B, Protocol=Finetuning2026.06 | 61.85 | |
| MuonpBackbone=Llama3-1B, Protocol=Finetuning2026.06 | 61.81 | |
| MuonBackbone=Llama3-1B, Protocol=Finetuning2026.06 | 61.46 | |
| OLMo-1B-0724Model Size Category=XL, Parameters=908M-1.2B, Zero-shot=true, Closed-book=true2025.07 | 61 | |
| IntAttentionModel=OPT-1.3B2025.11 | 58.78 | |
| Ettin-Dec-1BModel Size Category=XL, Parameters=908M-1.2B, Zero-shot=true, Closed-book=true2025.07 | 58.4 | |
| Quant-OnlyModel=OPT-1.3B2025.11 | 58 | |
| FP16Model=OPT-1.3B2025.11 | 57.85 | |
| U-NorMuonModel Size=1.1B, Evaluation Protocol=0-shot2026.06 | 56.9 | |
| MuonModel Size=1.1B, Evaluation Protocol=0-shot2026.06 | 55.7 | |
| NTPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 55.45 | |
| NITPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 55.45 | |
| NorMuonModel Size=1.1B, Evaluation Protocol=0-shot2026.06 | 55.1 | |
| AuroraModel Size=1.1B, Evaluation Protocol=0-shot2026.06 | 54.5 | |
| MoHGE-1BTotal Parameters=0.891B, Activated Parameters of Experts=0.122B2026.04 | 53.75 | |
| SmolLM2-360mModel Size Category=Large, Parameters=360-410M, Zero-shot=true, Closed-book=true2025.07 | 53.5 | |
| MoE-1BTotal Parameters=1.098B, Activated Parameters of Experts=0.163B2026.04 | 53.2 | |
| Ettin-Dec-400mModel Size Category=Large, Parameters=360-410M, Zero-shot=true, Closed-book=true2025.07 | 52.3 | |
| DenseTotal Parameters=0.570B2026.04 | 51.87 | |
| Pythia-410mModel Size Category=Large, Parameters=360-410M, Zero-shot=true, Closed-book=true2025.07 | 51.5 | |
| GPT-2 1.5BModel=GPT-2 1.5B2026.03 | 51.2 | |
| Focus (GPT-2 1.5B)Backbone=GPT-2 1.5B, K=8, Centroids=full-rank2026.03 | 51.2 | |
| FP16Model=Qwen3-1.7B2025.11 | 50.73 | |
| NITPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 50.52 | |
| NTPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 50.36 | |
| SWAw_eff=642026.04 | 50.2 | |
| BF16Model=Llama3.1-8B, Evaluation=Zero-shot2026.04 | 49.64 | |
| LoRA (r=640)Model=Pythia 410M, Params=94.4M, Evaluation protocol=Frozen probe2026.04 | 48.8 | |
| AdaHOP-Lv2Model=Llama3.1-8B, Evaluation=Zero-shot2026.04 | 48.57 | |
| Pre-proj + skipModel=Pythia 410M, Params=88.1M, Evaluation protocol=Frozen probe2026.04 | 48.4 | |
| HALOModel=Llama3.1-8B, Evaluation=Zero-shot2026.04 | 48.04 | |
| HALOModel=Llama3.2-3B, Evaluation=Zero-shot2026.04 | 47.95 | |
| GPT-2 774MModel=GPT-2 774M2026.03 | 47.7 | |
| Focus (GPT-2 774M)Backbone=GPT-2 774M, K=8, Centroids=full-rank2026.03 | 47.7 | |
| BF16Model=Llama3.2-3B, Evaluation=Zero-shot2026.04 | 47.53 | |
| AdaHOP-Lv1Model=Llama3.2-3B, Evaluation=Zero-shot2026.04 | 47.41 | |
| AdaHOP-Lv2Model=Llama3.2-3B, Evaluation=Zero-shot2026.04 | 47.28 | |
| Pre-projModel=Pythia 410M, Params=62.9M, Evaluation protocol=Frozen probe2026.04 | 47 | |
| Pre-proj + LoRAModel=Pythia 410M, Params=72.4M, Evaluation protocol=Frozen probe2026.04 | 47 | |
| IntAttentionModel=Qwen3-1.7B2025.11 | 46.94 | |
| BaselineModel=Pythia 410M, Params=—, Evaluation protocol=Frozen probe2026.04 | 46.6 | |
| BF16Model=Instella-3B, Evaluation=Zero-shot2026.04 | 46.26 | |
| Quant-OnlyModel=Qwen3-1.7B2025.11 | 46.24 | |
| HALOModel=Instella-3B, Evaluation=Zero-shot2026.04 | 45.9 | |
| AdaHOP-Lv2Model=Instella-3B, Evaluation=Zero-shot2026.04 | 45.42 | |
| Tseng et al.Model=Instella-3B, Evaluation=Zero-shot2026.04 | 45.12 | |
| AdaHOP-Lv1Model=Instella-3B, Evaluation=Zero-shot2026.04 | 45.02 | |
| MoBA (k=2)w_eff=642026.04 | 44.7 | |
| AuroraModel Size=340M, Evaluation Protocol=0-shot2026.06 | 43.5 | |
| MXFP4+HadamardModel=Instella-3B, Evaluation=Zero-shot2026.04 | 43.39 | |
| Ettin-Dec-150mModel Size Category=Base, Parameters=135-160M, Zero-shot=true, Closed-book=true2025.07 | 43.2 | |
| SmolLM2-135mModel Size Category=Base, Parameters=135-160M, Zero-shot=true, Closed-book=true2025.07 | 42.9 | |
| U-NorMuonModel Size=340M, Evaluation Protocol=0-shot2026.06 | 41.5 | |
| MuonModel Size=340M, Evaluation Protocol=0-shot2026.06 | 40.2 | |
| NorMuonModel Size=340M, Evaluation Protocol=0-shot2026.06 | 39.1 | |
| HALOModel=Llama3.2-1B, Evaluation=Zero-shot2026.04 | 39.03 | |
| AdaHOP-Lv2Model=Llama3.2-1B, Evaluation=Zero-shot2026.04 | 38.55 | |
| AdaHOP-Lv1Model=Llama3.2-1B, Evaluation=Zero-shot2026.04 | 38.21 | |
| BF16Model=Llama3.2-1B, Evaluation=Zero-shot2026.04 | 38.17 | |
| MXFP4+HadamardModel=Llama3.2-1B, Evaluation=Zero-shot2026.04 | 37.68 | |
| Tseng et al.Model=Llama3.2-1B, Evaluation=Zero-shot2026.04 | 36.81 | |
| Ettin-Dec-68mModel Size Category=Small, Parameters=68-82M, Zero-shot=true, Closed-book=true2025.07 | 35.2 | |
| Stochasticw_eff=322026.04 | 33.2 | |
| Pythia-160mModel Size Category=Base, Parameters=135-160M, Zero-shot=true, Closed-book=true2025.07 | 32.9 | |
| GPT-2 124MModel=GPT-2 124M2026.03 | 32.6 | |
| Focus (GPT-2 124M)Backbone=GPT-2 124M, K=2, dg=162026.03 | 32.6 | |
| Ettin-Dec-32mModel Size Category=XS, Parameters=32M, Zero-shot=true, Closed-book=true2025.07 | 28.5 | |
| MXFP4Model=Instella-3B, Evaluation=Zero-shot2026.04 | 27.71 | |
| DistilGPTModel Size Category=Small, Parameters=68-82M, Zero-shot=true, Closed-book=true2025.07 | 25 | |
| Ettin-Dec-17mModel Size Category=XXS, Parameters=14-17M, Zero-shot=true, Closed-book=true2025.07 | 23 | |
| Pre-proj + skipModel=Pythia 160M, Params=24.8M, Evaluation protocol=Frozen probe2026.04 | 18 |