Language Modeling on LAMBADA (Acc, Ppl)
86.9AccuracyPaLM-2 L
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| PaLM-2 Lprotocol=1-shot2023.11 | 86.9 | — | — | — | — | — | — | |
| Falcon-180Bprotocol=1-shot2023.11 | 84.4 | — | — | — | — | — | — | |
| PaLM-2 Mprotocol=1-shot2023.11 | 83.7 | — | — | — | — | — | — | |
| PaLMprotocol=1-shot2023.11 | 81.8 | — | — | — | — | — | — | |
| PaLM-2 Sprotocol=1-shot2023.11 | 80.7 | — | — | — | — | — | — | |
| FalconModel Size=180B2023.11 | 79.8 | — | — | — | — | — | — | |
| Inflection-12023.11 | 78.5 | — | — | — | — | — | — | |
| PaLM2023.11 | 77.9 | — | — | — | — | — | — | |
| Chinchilla2023.11 | 77.4 | — | — | — | — | — | — | |
| LLaMA-2Model Size=7B2023.11 | 77.4 | — | — | — | — | — | — | |
| FalconModel Size=40B2023.11 | 77.3 | — | — | — | — | — | — | |
| MT-NLG2023.11 | 76.6 | — | — | — | — | — | — | |
| GPT-32023.11 | 76.2 | — | — | — | — | — | — | |
| Llama3 8BCompression Ratio=Baseline2025.09 | 75.6 | 3.1 | — | — | — | — | — | |
| Llama3 8B2026.02 | 75.6 | 3.1 | — | — | — | — | — | |
| Llama 3.1 8BAttn CR=N/A2025.08 | 75.4 | 3.13 | — | — | — | — | — | |
| FalconModel Size=7B2023.11 | 74.9 | — | — | — | — | — | — | |
| LoRA-Drop (50%)Backbone=Qwen2.5-14B, Evaluation Protocol=Zero-shot, Drop Ratio=50%, Speedup (x)=1.682026.01 | 74.6 | — | — | — | — | — | — | |
| Gopher2023.11 | 74.5 | — | — | — | — | — | — | |
| Matrix PCABase Model=Llama 3.1 8B, Attn CR=20%2025.08 | 73.9 | 3.35 | — | — | — | — | — | |
| CoSpaDiCompression Ratio=0.22025.09 | 73.8 | 4.3 | — | — | — | — | — | |
| FlexiDepthBackbone=LLaMA3-8B, Evaluation Protocol=Zero-shot, Speedup (x)=1.502026.01 | 72.1 | — | — | — | — | — | — | |
| COMPOTCR=0.22026.02 | 70.5 | 4.6 | — | — | — | — | — | |
| Llama 3.2 3BAttn CR=N/A2025.08 | 70.5 | 3.94 | — | — | — | — | — | |
| SVD-LLMBase Model=Llama 3.1 8B, Attn CR=20%2025.08 | 70.5 | 4.63 | — | — | — | — | — | |
| Matrix PCABase Model=Llama 3.2 3B, Attn CR=20%2025.08 | 69 | 4.39 | — | — | — | — | — | |
| LoRA-Drop (25%)Backbone=LLaMA2-7B, Evaluation Protocol=Zero-shot, Drop Ratio=25%, Speedup (x)=1.372026.01 | 68.4 | — | — | — | — | — | — | |
| DLMTokens=100B, Model Scale=1.5B, Initialization=Autoregressive LLM (Qwen2.5)2026.01 | 66.58 | — | — | — | — | — | — | |
| DLM + extra tokenTokens=100B, Model Scale=1.5B, Initialization=Autoregressive LLM (Qwen2.5)2026.01 | 66.41 | — | — | — | — | — | — | |
| Densezero-shot=true2026.02 | 65.55 | — | — | — | — | — | — | |
| SVD-LLMBase Model=Llama 3.2 3B, Attn CR=20%2025.08 | 65.1 | 5.57 | — | — | — | — | — | |
| COMPOTCR=0.32026.02 | 64.3 | 6.6 | — | — | — | — | — | |
| SpanNormParam=5B, Tokens=200B, Architecture=Dense2026.01 | 64.2 | 5.8 | — | — | — | — | — | |
| SpanNormParam=A2.4B-16B, Tokens=200B, Architecture=MoE2026.01 | 63 | 6.5 | — | — | — | — | — | |
| Llama 3.2 1BAttn CR=N/A2025.08 | 62.9 | 5.73 | — | — | — | — | — | |
| SparseGPTzero-shot=true2026.02 | 62.76 | — | — | — | — | — | — | |
| PreNormParam=A2.4B-16B, Tokens=200B, Architecture=MoE2026.01 | 61.7 | 6.7 | — | — | — | — | — | |
| PreNormParam=5B, Tokens=200B, Architecture=Dense2026.01 | 61.4 | 6.9 | — | — | — | — | — | |
| CoSpaDiCompression Ratio=0.32025.09 | 61.3 | 9.2 | — | — | — | — | — | |
| Matrix PCABase Model=Llama 3.2 1B, Attn CR=20%2025.08 | 59.9 | 6.65 | — | — | — | — | — | |
| Taylorzero-shot=true2026.02 | 59.23 | — | — | — | — | — | — | |
| SVD-LLMBase Model=Llama 3.2 1B, Attn CR=20%2025.08 | 55.4 | 9.55 | — | — | — | — | — | |
| HGRN2State size=128 × Ld, Model size=2.7B, Training tokens=100B, Number of layers (L)=32, Model dimension (d)=2,560, Evaluation protocol=Zero-shot2024.09 | 55.4 | 8.8 | — | — | — | — | — | |
| GPT2-Medium + ACDModel Size=355M parameters, Decoding=ACD, Contrast Layers=24 and 122023.05 | 55 | 15.4 | — | — | — | — | — | |
| Matrix PCABase Model=Llama 3.2 1B, Attn CR=30%2025.08 | 54.5 | 8.79 | — | — | — | — | — | |
| NoPEModel Architecture=GLA, Parameters=1.3B, Training Tokens=26B2025.11 | 53.8 | 8.59 | — | — | — | — | — | |
| Selective RoPEModel Architecture=GLA, Parameters=1.3B, Training Tokens=26B2025.11 | 53.8 | 8.5 | — | — | — | — | — | |
| Learnable ResFormerParameters=468M, Training Tokens=20B, Evaluation Protocol=Zero-shot2024.10 | 53.5 | 21.2 | — | — | — | — | — | |
| Learnable ResFormer plusParameters=468M, Training Tokens=20B, Evaluation Protocol=Zero-shot2024.10 | 53.5 | 21.2 | — | — | — | — | — | |
| RoPEModel Architecture=GLA, Parameters=1.3B, Training Tokens=26B2025.11 | 53 | 8.88 | — | — | — | — | — | |
| GSAState size=128 × Ld, Model size=2.7B, Training tokens=100B, Number of layers (L)=32, Model dimension (d)=2,560, Evaluation protocol=Zero-shot2024.09 | 52.7 | 9.8 | — | — | — | — | — | |
| LDDM-M2025.10 | 52.4 | — | — | — | — | — | — | |
| Xfmr++State size=N/A, Model size=2.7B, Training tokens=100B, Number of layers (L)=32, Model dimension (d)=2,560, Evaluation protocol=Zero-shot2024.09 | 52.3 | 10.7 | — | — | — | — | — | |
| SpanNormParam=1.3B, Tokens=100B, Architecture=Dense2026.01 | 52.1 | 12 | — | — | — | — | — | |
| GHOSTzero-shot=true2026.02 | 51.76 | — | — | — | — | — | — | |
| Mamba-3-MIMO-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B, MIMO Rank=42026.03 | 51.7 | 10.2 | — | — | — | — | — | |
| LIDASModel Size=1.7B, Cooldown Protocol=Standard2025.12 | 51.41 | — | — | — | — | — | — | |
| PreNormParam=1.3B, Tokens=100B, Architecture=Dense2026.01 | 51.1 | 12.6 | — | — | — | — | — | |
| DLM + GATokens=100B, Model Scale=1.5B, Initialization=Autoregressive LLM (Qwen2.5)2026.01 | 51.1 | — | — | — | — | — | — | |
| GPT2-XLModel Size=1.5B parameters2023.05 | 51 | 10.6 | — | — | — | — | — | |
| LLM-PrunerCompression Ratio=0.22025.09 | 51 | 11 | — | — | — | — | — | |
| LLM-PrunerCR=0.22026.02 | 51 | 11 | — | — | — | — | — | |
| Gated DeltaNet-2Architecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 50.9 | 10.43 | — | — | — | — | — | |
| MIDASModel Size=1.7B, Cooldown Protocol=Standard2025.12 | 50.81 | — | — | — | — | — | — | |
| GPT-2 smallARACH=true, Logit offset (b)=-0.52026.03 | 50.42 | — | — | — | — | — | — | |
| GLAState size=256 × Ld, Model size=2.7B, Training tokens=100B, Number of layers (L)=32, Model dimension (d)=2,560, Evaluation protocol=Zero-shot2024.09 | 50.4 | 12.4 | — | — | — | — | — | |
| Oryx-TM (Mamba-2)Parameter Scale=1.4B, Family=Oryx-TM2026.05 | 50.4 | 10.5 | — | — | — | — | — | |
| Nirvana (Ours)Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 50.37 | 11.56 | — | — | — | — | — | |
| Transformer-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 50.3 | 11.1 | — | — | — | — | — | |
| ParallaxSize=1.7B, Optimizer=Muon, RoPE on ρ=✓2026.05 | 50.26 | 10.8 | — | — | — | — | — | |
| GPT-2 smallARACH=true, Logit offset (b)=-0.42026.03 | 50.24 | — | — | — | — | — | — | |
| Oryx-TM (Transformer)Parameter Scale=1.4B, Family=Oryx-TM2026.05 | 50.2 | 11.1 | — | — | — | — | — | |
| BaselineModel Size=1.7B, Cooldown Protocol=Standard2025.12 | 50.05 | — | — | — | — | — | — | |
| Oryx-TG (Gated DeltaNet)Parameter Scale=1.4B, Family=Oryx-TG2026.05 | 50 | 10.6 | — | — | — | — | — | |
| GPT-2 smallARACH=true, Logit offset (b)=-0.32026.03 | 49.93 | — | — | — | — | — | — | |
| TransformerParameter Scale=1.4B, Family=Baseline2026.05 | 49.9 | 11.4 | — | — | — | — | — | |
| Gated DeltaNetParameter Scale=1.4B, Family=Baseline2026.05 | 49.9 | 10.6 | — | — | — | — | — | |
| Oryx-TG (Transformer)Parameter Scale=1.4B, Family=Oryx-TG2026.05 | 49.9 | 10.9 | — | — | — | — | — | |
| Mamba-3 (MIMO)Architecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 49.82 | 10.92 | — | — | — | — | — | |
| Gated DeltaNetArchitecture Group=Recurrent models, Evaluation Protocol=zero-shot2026.05 | 49.62 | 11.89 | — | — | — | — | — | |
| RetNetState size=512 × Ld, Model size=2.7B, Training tokens=100B, Number of layers (L)=32, Model dimension (d)=2,560, Evaluation protocol=Zero-shot2024.09 | 49.6 | 11.9 | — | — | — | — | — | |
| SoftmaxQuantization=2-bit2025.04 | 49.56 | 11.38 | — | — | — | — | — | |
| SoftmaxQuantization=3-bit2025.04 | 49.56 | 11.38 | — | — | — | — | — | |
| ParallaxSize=1.7B, Optimizer=Muon, RoPE on ρ=✗2026.05 | 49.54 | 10.85 | — | — | — | — | — | |
| Mamba-3-MIMO-880MTraining Tokens=100B, Context Length=2K, Scale=880M, MIMO Rank=42026.03 | 49.5 | 11.8 | — | — | — | — | — | |
| SoftmaxModel Scale=1.8B2025.04 | 49.43 | 11.38 | — | — | — | — | — | |
| HGRN2State size=128 × Ld, Model size=1.3B, Training tokens=100B, Number of layers (L)=24, Model dimension (d)=2,048, Evaluation protocol=Zero-shot2024.09 | 49.4 | 11.8 | — | — | — | — | — | |
| Mamba-3-SISO-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 49.4 | 10.9 | — | — | — | — | — | |
| Nirvana-noTriggerSetting=Zero-shot, Model Size=1.3B parameters2025.10 | 49.4 | 12.25 | — | — | — | — | — | |
| KDAArchitecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 49.21 | 10.66 | — | — | — | — | — | |
| NeuTRENOParameters=468M, Training Tokens=20B, Evaluation Protocol=Zero-shot2024.10 | 49.2 | 20.7 | — | — | — | — | — | |
| GDN-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 49.2 | 10.9 | — | — | — | — | — | |
| Mamba-3 (SISO)Architecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 49.19 | 10.65 | — | — | — | — | — | |
| LIME+1Model Size=500M, Shot Count=5-shot, Evaluation Protocol=generative greedy sampling2025.12 | 49 | — | — | — | — | — | — | |
| LN-ScalingModel Size=1.7B, Cooldown Protocol=Standard2025.12 | 48.94 | — | — | — | — | — | — | |
| Gated DeltaNet-H2Setting=Zero-shot, Model Size=1.3B parameters2025.10 | 48.76 | 12.55 | — | — | — | — | — | |
| Gated DeltaNetArchitecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 48.71 | 10.82 | — | — | — | — | — | |
| SoftmaxQuantization=8-bit2025.04 | 48.65 | 11.74 | — | — | — | — | — | |
| Mamba-2Parameter Scale=1.4B, Family=Baseline2026.05 | 48.6 | 11.2 | — | — | — | — | — | |
| TransformerArchitecture Group=Attention or hybrid models, Evaluation Protocol=zero-shot2026.05 | 48.32 | 13.72 | — | — | — | — | — |