Common Sense Reasoning on PIQA
91.89AccuracyLlama-3.3-70B-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama-3.3-70B-InstructShots=52026.04 | 91.89 | — | |
| Qwen3-14BShots=52026.04 | 89.88 | — | |
| Qwen3.5-9BShots=52026.04 | 88.41 | — | |
| SecGPT-14BShots=52026.04 | 87.21 | — | |
| Qwen3-8BShots=52026.04 | 87.11 | — | |
| XekRung-8BShots=52026.04 | 86.4 | — | |
| Qwen2.5-7B-InstructModel=Qwen2.5-7B-Instruct2026.03 | 83 | — | |
| LLaMA-3-8B-LizardTraining Tokens (B)=0.04, Attention configuration=Linearized (Keep 50% Full Attn.)2025.07 | 82.2 | — | |
| Mamba2-LLaMA-3Training Tokens (B)=20, Attention configuration=Linearized (Keep 50% Full Attn.)2025.07 | 81.5 | — | |
| Zamba-7BTraining Tokens (B)=1000, Attention configuration=Hybrid Softmax2025.07 | 81.4 | — | |
| DiVA-Llama3.1-8BInput Modality=Text, Parameters=8B2025.10 | 80.8 | 0.6 | |
| Mistral-7BBase Model=Mistral-7B2026.04 | 80.6 | — | |
| SUPRABase Model=Mistral-7B, Healing Tokens (B)=1002026.04 | 80.4 | — | |
| Liger-GLABase Model=Llama-3-8B, Healing Tokens (B)=0.022026.04 | 80.3 | — | |
| Llama-3.1-8B-InstructShots=52026.04 | 80.3 | — | |
| SALAD-7BInput Modality=Text, Parameters=7B2025.10 | 80.2 | -0.3 | |
| Liger-GLABase Model=Mistral-7B, Healing Tokens (B)=0.022026.04 | 80.1 | — | |
| LoLCATsBase Model=Llama-3-8B, Healing Tokens (B)=0.042026.04 | 80.1 | — | |
| LLaMA-3-8BTraining Tokens (B)=15000, Attention configuration=Softmax2025.07 | 79.9 | — | |
| LoLCATsBase Model=Mistral-7B, Healing Tokens (B)=0.042026.04 | 79.7 | — | |
| Llama-3-8BBase Model=Llama-3-8B2026.04 | 79.4 | — | |
| SUPRABase Model=Llama-3-8B, Healing Tokens (B)=202026.04 | 78.9 | — | |
| Qwen2.5-Omni-7BInput Modality=Text, Parameters=7B2025.10 | 78.8 | 1.1 | |
| StripedHyena-Nous-7BTraining Tokens (B)=–, Attention configuration=Hybrid Softmax2025.07 | 78.8 | — | |
| SALAD-3BInput Modality=Text, Parameters=3B2025.10 | 78.6 | 0 | |
| Dense BaselineBackbone=Llama-2-7B2026.03 | 78.5 | — | |
| Standard Top-K (Oracle)Backbone=Llama-2-7B2026.03 | 78.2 | — | |
| AFBS-BOBackbone=Llama-2-7B2026.03 | 78.1 | — | |
| GLM-4-Voice-9BInput Modality=Text, Parameters=9B2025.10 | 77.9 | -3.4 | |
| H2OBackbone=Llama-2-7B2026.03 | 77.9 | — | |
| PlaintextBase Model=Qwen2.5-14B-Instruct2026.03 | 77.45 | — | |
| Routing TransformerBackbone=Llama-2-7B2026.03 | 77.1 | — | |
| Ministral-3-8B-Instruct-2512Model=Ministral-3-8B-Instruct-25122026.03 | 77 | — | |
| Mamba2-LlamaBase Model=Llama-3-8B, Healing Tokens (B)=202026.04 | 76.8 | — | |
| Qwen2-Audio-7BInput Modality=Text, Parameters=7B2025.10 | 76 | 2.8 | |
| Llama-3.1-8B-InstructModel=Llama-3.1-8B-Instruct2026.03 | 76 | — | |
| Llama-EstLLM-8B-Instruct-CVModel=Llama-EstLLM-8B-Instruct-CV2026.03 | 76 | — | |
| LayerBoostBase Model=Qwen3-4B, Healing Tokens (B)=0.042026.04 | 75.7 | — | |
| MLAEvaluation Mode=Zero-shot2026.03 | 75.68 | — | |
| MLRA-2Evaluation Mode=Zero-shot2026.03 | 75.52 | — | |
| GLA-2Evaluation Mode=Zero-shot2026.03 | 75.41 | — | |
| AloePriBase Model=Qwen2.5-14B-Instruct2026.03 | 75.23 | — | |
| MFAEvaluation Mode=Zero-shot2026.03 | 75.19 | — | |
| GTAEvaluation Mode=Zero-shot2026.03 | 75.14 | — | |
| GQAEvaluation Mode=Zero-shot2026.03 | 75.08 | — | |
| Qwen3-4BBase Model=Qwen3-4B2026.04 | 74.9 | — | |
| MHAEvaluation Mode=Zero-shot2026.03 | 74.86 | — | |
| GLA-4Evaluation Mode=Zero-shot2026.03 | 74.65 | — | |
| TPAEvaluation Mode=Zero-shot2026.03 | 74.54 | — | |
| MQAEvaluation Mode=Zero-shot2026.03 | 74.48 | — | |
| MLRA-4Evaluation Mode=Zero-shot2026.03 | 74.48 | — | |
| Window AttentionBackbone=Llama-2-7B2026.03 | 74.2 | — | |
| StateXArchitecture Category=State Space Model, Base Model=Mamba2, Evaluation Protocol=zero-shot2025.09 | 73.6 | — | |
| In-Place TTTArchitecture=Full Attention, Parameter Count=4B2026.04 | 73.29 | — | |
| Apertus-8B-Instruct-2509Model=Apertus-8B-Instruct-25092026.03 | 73 | — | |
| VanillaArchitecture Category=State Space Model, Base Model=Mamba2, Evaluation Protocol=zero-shot2025.09 | 73 | — | |
| RoPEModel scale=Llama-3B2026.02 | 72.74 | — | |
| Full AttentionModel Group=Baselines, Parameter Count=4B2026.04 | 72.63 | — | |
| Sliding-Window Attention (SWA)Model Group=Baselines, Parameter Count=4B2026.04 | 72.58 | — | |
| RoPE-IDModel scale=Llama-3B2026.02 | 72.25 | — | |
| YaRNModel scale=Llama-3B2026.02 | 72.2 | — | |
| HalfRoPEModel scale=Llama-3B2026.02 | 72.03 | — | |
| In-Place TTTArchitecture=Sliding-Window Attention (SWA), Parameter Count=4B2026.04 | 72.03 | — | |
| High frequencyModel scale=Llama-3B2026.02 | 71.87 | — | |
| Llama-EstLLM-8B-InstructModel=Llama-EstLLM-8B-Instruct2026.03 | 71 | — | |
| StateXArchitecture Category=Linear Attention, Base Model=GLA, Evaluation Protocol=zero-shot2025.09 | 69.7 | — | |
| VanillaArchitecture Category=Linear Attention, Base Model=GLA, Evaluation Protocol=zero-shot2025.09 | 69.6 | — | |
| RoPEModel scale=Llama-1B2026.02 | 69.26 | — | |
| High frequencyModel scale=Llama-1B2026.02 | 69.26 | — | |
| Glauber-MN=32026.05 | 68.9 | — | |
| RoPE-IDModel scale=Llama-1B2026.02 | 68.88 | — | |
| HalfRoPEModel scale=Llama-1B2026.02 | 68.77 | — | |
| YaRNModel scale=Llama-1B2026.02 | 68.61 | — | |
| GPT-2-M2026.05 | 67.4 | — | |
| Glauber-MN=12026.05 | 66.8 | — | |
| Mamba-2Model Scale=440M, Evaluation Protocol=Zero-shot2026.04 | 65.3 | — | |
| Mamba-2 + PoSTModel Scale=440M, Evaluation Protocol=Zero-shot2026.04 | 65.3 | — | |
| HADESZero-shot=true, Parameters=218M2026.03 | 63.93 | — | |
| DeltaNetZero-shot=true, Parameters=370M2026.03 | 63.6 | — | |
| Mamba2Zero-shot=true, Parameters=370M2026.03 | 63.44 | — | |
| MoMArchitecture Category=Sparse Model, Evaluation Protocol=zero-shot2025.09 | 63.3 | — | |
| Mamba-2 + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 63.2 | — | |
| RWKV-7Model Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 63.1 | — | |
| salamandra-7b-instructModel=salamandra-7b-instruct2026.03 | 63 | — | |
| Mamba-2Model Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 62.9 | — | |
| RWKV-7 + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 62.9 | — | |
| Gated DeltaNet + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 62.9 | — | |
| Gated DeltaNetModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 62.7 | — | |
| RetNetZero-shot=true, Parameters=370M2026.03 | 62.4 | — | |
| Mamba1Zero-shot=true, Parameters=370M2026.03 | 60.72 | — | |
| Linear TransformerZero-shot=true, Parameters=370M2026.03 | 60.55 | — | |
| (Gong et al., 2025)-M2026.05 | 59.6 | — | |
| EuroLLM-9B-InstructModel=EuroLLM-9B-Instruct2026.03 | 58 | — | |
| SEDD-M2026.05 | 56.1 | — | |
| Apertus-EstLLM-InstructModel=Apertus-EstLLM-Instruct2026.03 | 56 | — | |
| SANTEXTBase Model=Qwen2.5-14B-Instruct2026.03 | 49.72 | — | |
| Llama-Primus-Reasoning-8BShots=52026.04 | 46.57 | — | |
| SGTBase Model=Qwen2.5-14B-Instruct2026.03 | 26.06 | — | |
| RANTEXTBase Model=Qwen2.5-14B-Instruct2026.03 | 18.3 | — | |
| LlammasModel=Llammas2026.03 | 0 | — |