Commonsense Reasoning on HellaSwag (accuracy, acc_norm)
90.92Normalized AccuracyKimi-K2 Base
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Kimi-K2 Baseshots=10-shot2026.06 | 90.92 | — | |
| Nemotron-3-Ultra 550B-A55B-Baseshots=10-shot2026.06 | 90.51 | — | |
| GLM-4.5 Baseshots=10-shot2026.06 | 90.17 | — | |
| DeepSeek-V3.2 Exp-Baseshots=10-shot2026.06 | 89.44 | — | |
| Mistral-Large-3 675B-Base-2512shots=10-shot2026.06 | 88.88 | — | |
| RecurrentGemma-9BArchitecture=Subquadratic (Pretrained), Tokens (B)=2000, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 80.1 | — | |
| Falcon3-Mamba-7BArchitecture=Subquadratic (Pretrained), Tokens (B)=7300, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 79.8 | — | |
| LoLA-8B (ours)Architecture=Subquadratic (Distilled), Tokens (B)=0.04, Evaluation Protocol=zero-shot, Normalization=normalized logits, η=64, λ=642025.05 | 79.8 | — | |
| Llama-3.1-8BArchitecture=Transformer, Tokens (B)=15000, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 79.3 | — | |
| LoLCATs-8BArchitecture=Subquadratic (Distilled), Tokens (B)=0.04, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 79.1 | — | |
| Griffin 7BArchitecture=Subquadratic (Pretrained), Tokens (B)=300, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 78.6 | — | |
| Mamba2-8BArchitecture=Subquadratic (Pretrained), Tokens (B)=3500, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 77.7 | — | |
| Hawk 7BArchitecture=Subquadratic (Pretrained), Tokens (B)=300, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 77.6 | — | |
| Llamba-8BArchitecture=Subquadratic (Distilled), Tokens (B)=12, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 77.6 | — | |
| Mamba-8BArchitecture=Subquadratic (Pretrained), Tokens (B)=1100, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 75.6 | — | |
| RWKV-6 (W2.1) 7BArchitecture=Subquadratic (Pretrained), Tokens (B)=1420, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 75.1 | — | |
| Mamba2-Llama3-8B (L3.1-Instr.)Architecture=Subquadratic (Distilled), Tokens (B)=20, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 70.8 | — | |
| Hedgehog-8B (Llama-3)Architecture=Subquadratic (Distilled), Tokens (B)=0.04, Evaluation Protocol=zero-shot, Normalization=normalized logits2025.05 | 66.5 | — | |
| LoLA-1BTokens (B)=0.04, Zero-shot=true, Training Strategy=Subquadratic, Distillation Source=Llama-3.2-1B2025.05 | 64.1 | — | |
| Llama-3.2-1BTokens (B)=9000, Zero-shot=true, Training Strategy=Transformers2025.05 | 63.7 | — | |
| LoLCATs-Llama-1BTokens (B)=0.04, Zero-shot=true, Training Strategy=Subquadratic, Distillation Source=Llama-3.2-1B2025.05 | 63.7 | — | |
| Phi-1.5-1.3BTokens (B)=150, Zero-shot=true, Training Strategy=Transformers2025.05 | 62.6 | — | |
| LoLCATs-Phi-1.3BTokens (B)=0.04, Zero-shot=true, Training Strategy=Subquadratic, Distillation Source=Phi-1.5-1.3B2025.05 | 62.3 | — | |
| Llamba-1BTokens (B)=8, Zero-shot=true, Training Strategy=Subquadratic, Distillation Source=Llama-3.2-1B2025.05 | 61.2 | — | |
| xLSTM-1.4BTokens (B)=300, Zero-shot=true, Training Strategy=Subquadratic, Training Source=various sources2025.05 | 60.9 | — | |
| RecurrentGemma-2BTokens (B)=2000, Zero-shot=true, Training Strategy=Subquadratic, Training Source=various sources2025.05 | 60.3 | — | |
| Phi-Mamba-1.5BTokens (B)=3, Zero-shot=true, Training Strategy=Subquadratic, Distillation Source=Phi-1.5-1.3B2025.05 | 60.2 | — | |
| Mamba2-1.3BTokens (B)=315, Zero-shot=true, Training Strategy=Subquadratic, Training Source=various sources2025.05 | 59.9 | — | |
| CARVE + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 59.83 | — | |
| Mamba1-1.4BTokens (B)=315, Zero-shot=true, Training Strategy=Subquadratic, Training Source=various sources2025.05 | 59.1 | — | |
| Mamba-3 SISO + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 58.75 | — | |
| GDN-2 + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 58.46 | — | |
| Mamba-3 MIMO + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 58.19 | — | |
| Mamba-2 + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 57.52 | — | |
| Gated DeltaNet + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 57.5 | — | |
| CARVEModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 57.31 | — | |
| Finch-1.6BTokens (B)=1100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=various sources2025.05 | 57.3 | — | |
| KDA + SWAModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.89 | — | |
| GDN-2Model Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.84 | — | |
| Gated DeltaNetModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.5 | — | |
| Mamba-3 MIMOModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.49 | — | |
| TransformerModel Category=Hybrid, Parameter Scale=1.3B, Token Scale=100B2026.06 | 56.12 | — | |
| Gated-DeltaNet-1.3BTokens (B)=100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=FineWeb-Edu2025.05 | 55.8 | — | |
| KDAModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 55.75 | — | |
| Mamba2-1.3BTokens (B)=100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=FineWeb-Edu2025.05 | 55.7 | — | |
| Mamba-3 SISOModel Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 55.58 | — | |
| Mamba-2Model Category=Recurrent, Parameter Scale=1.3B, Token Scale=100B2026.06 | 55.51 | — | |
| Samba-1.3BTokens (B)=100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=various sources2025.05 | 54.7 | — | |
| Mamba1-1.3BTokens (B)=100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=FineWeb-Edu2025.05 | 52.9 | — | |
| DeltaNet-1.3BTokens (B)=100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=FineWeb-Edu2025.05 | 50.9 | — | |
| HGRN2-1.3BTokens (B)=100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=FineWeb-Edu2025.05 | 49.5 | — | |
| RetNet-1.3BTokens (B)=100, Zero-shot=true, Training Strategy=Subquadratic, Training Source=FineWeb-Edu2025.05 | 49.2 | — | |
| MHF (H4B16K)Model Scale=3B, Embedding Parameters (Emb)=138M, Decoder Parameters (Dec)=2.8B, Evaluation Protocol=Zero-shot2026.06 | 42.89 | — | |
| MHF (H4B16K)Model Scale=1B, Embedding Parameters (Emb)=138M, Decoder Parameters (Dec)=0.9B, Evaluation Protocol=Zero-shot2026.06 | 41.32 | — | |
| MHF (H3B10K)Model Scale=1B, Embedding Parameters (Emb)=65M, Decoder Parameters (Dec)=0.9B, Evaluation Protocol=Zero-shot2026.06 | 40.45 | — | |
| Standard+2LModel Scale=1B, Embedding Parameters (Emb)=65M, Decoder Parameters (Dec)=1.0B, Evaluation Protocol=Zero-shot2026.06 | 39.85 | — | |
| Dense 1.3BArchitecture=Dense, Total Parameters=1.3B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 39.71 | — | |
| StandardModel Scale=3B, Embedding Parameters (Emb)=65M, Decoder Parameters (Dec)=2.8B, Evaluation Protocol=Zero-shot2026.06 | 39.65 | — | |
| StandardModel Scale=1B, Embedding Parameters (Emb)=65M, Decoder Parameters (Dec)=0.9B, Evaluation Protocol=Zero-shot2026.06 | 38.99 | — | |
| MoL 0.61B/2.08BArchitecture=MoL, Total Parameters=2.08B, Active Parameters=0.61B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 35.1 | — | |
| Dense 0.7BArchitecture=Dense, Total Parameters=0.7B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 34.1 | — | |
| Standard+4LModel Scale=100M, Embedding Parameters (Emb)=25M, Decoder Parameters (Dec)=116M, Evaluation Protocol=Zero-shot2026.06 | 29.99 | — | |
| StandardModel Scale=100M, Embedding Parameters (Emb)=25M, Decoder Parameters (Dec)=87M, Evaluation Protocol=Zero-shot2026.06 | 28.59 | — | |
| MHF (H4B16K)Model Scale=100M, Embedding Parameters (Emb)=52M, Decoder Parameters (Dec)=87M, Evaluation Protocol=Zero-shot2026.06 | 27.99 | — | |
| MHF (H3B10K)Model Scale=100M, Embedding Parameters (Emb)=25M, Decoder Parameters (Dec)=87M, Evaluation Protocol=Zero-shot2026.06 | 27.81 | — | |
| Rnd. GuessModel Scale=N/A, Embedding Parameters (Emb)=-, Decoder Parameters (Dec)=-, Evaluation Protocol=Random Guessing2026.06 | 25 | — |