Question Answering on ARC Easy (Accuracy, Normalized Accuracy)
84.1Normalized AccuracySwimba-14B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Swimba-14BParameters=14B, Experts=4, FLOPs / token=1.51282 × 10^102026.03 | 84.1 | 84.5 | — | |
| Nemotron-H-8BParameters=8B, FLOPs / token=1.51263 × 10^102026.03 | 83.7 | 84 | — | |
| GLA-HedgehogScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 67.76 | — | — | |
| SoftmaxModel Scale=1.8B2025.04 | 67.21 | 72.73 | — | |
| SoftmaxQuantization=2-bit2025.04 | 67.21 | 72.64 | — | |
| SoftmaxQuantization=3-bit2025.04 | 67.21 | 72.64 | — | |
| SoftmaxQuantization=8-bit2025.04 | 66.96 | 72.77 | — | |
| GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 66.54 | — | — | |
| CCQ-Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 66.46 | — | — | |
| Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 65.91 | — | — | |
| SoftmaxQuantization=4-bit2025.04 | 65.74 | 70.2 | — | |
| Mamba2Scale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 64.6 | — | — | |
| CCQ-GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 64.23 | — | — | |
| SoftpickModel Scale=1.8B2025.04 | 62.04 | 68.6 | — | |
| SoftpickQuantization=2-bit2025.04 | 61.99 | 68.6 | — | |
| SoftpickQuantization=3-bit2025.04 | 61.99 | 68.6 | — | |
| SoftpickQuantization=8-bit2025.04 | 61.95 | 68.6 | — | |
| SoftpickQuantization=4-bit2025.04 | 61.87 | 68.73 | — | |
| TransformerScale=1.3B, Training tokens=40B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 61.49 | — | — | |
| SoftmaxQuantization=3-bit2025.04 | 61.28 | 64.9 | — | |
| SoftpickQuantization=3-bit2025.04 | 57.49 | 64.27 | — | |
| CCQ-GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 56.82 | — | — | |
| SoftpickModel Scale=340M2025.04 | 56.73 | 61.11 | — | |
| SoftmaxModel Scale=340M2025.04 | 56.61 | 60.35 | — | |
| SoftpickQuantization=8-bit2025.04 | 56.36 | 60.82 | — | |
| CCQ-Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 56.14 | — | — | |
| Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 55.56 | — | — | |
| SoftmaxQuantization=8-bit2025.04 | 55.51 | 59.81 | — | |
| Mamba2Scale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 55.13 | — | — | |
| TransformerScale=500M, Training tokens=15B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 54.8 | — | — | |
| SoftpickQuantization=4-bit2025.04 | 53.87 | 59.6 | — | |
| SoftmaxQuantization=4-bit2025.04 | 53.66 | 58.88 | — | |
| GLA-HedgehogScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 53.37 | — | — | |
| MoBiQuantCalibration Dataset=PTB, Avg. Bits=3.02, Base Model=Llama3.2-1B2026.02 | 53.2 | 58.8 | — | |
| Dense 1.3BArchitecture=Dense, Total Parameters=1.3B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 53.07 | — | — | |
| GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 52.23 | — | — | |
| MoBiQuantCalibration Dataset=Mix, Avg. Bits=3.03, Base Model=Llama3.2-1B2026.02 | 50.2 | 53.4 | — | |
| Dense 0.7BArchitecture=Dense, Total Parameters=0.7B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 48.86 | — | — | |
| MoL 0.61B/2.08BArchitecture=MoL, Total Parameters=2.08B, Active Parameters=0.61B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 48.15 | — | — | |
| MoBiQuantCalibration Dataset=Wiki, Avg. Bits=3.01, Base Model=Llama3.2-1B2026.02 | 48.1 | 52.1 | — | |
| OmniQuantCalibration Dataset=Mix, Base Model=Llama3.2-1B2026.02 | 47.4 | 50.2 | — | |
| OmniQuantCalibration Dataset=Wiki, Base Model=Llama3.2-1B2026.02 | 47.2 | 51 | — | |
| OmniQuantCalibration Dataset=PTB, Base Model=Llama3.2-1B2026.02 | 47.1 | 50.3 | — | |
| OmniQuantCalibration Dataset=C4, Base Model=Llama3.2-1B2026.02 | 47 | 50.3 | — | |
| MoBiQuantCalibration Dataset=C4, Avg. Bits=3.03, Base Model=Llama3.2-1B2026.02 | 45.2 | 52.2 | — | |
| RFMoEScale=L, Size=870.6M, FLOPs=613.2M, Zero-shot=true2026.04 | 37.46 | — | — | |
| MoEScale=L, Size=808.4M, FLOPs=608.4M, Zero-shot=true2026.04 | 36.41 | — | — | |
| MoEScale=M, Size=289.9M, FLOPs=248.0M, Zero-shot=true2026.04 | 36.2 | — | — | |
| RFMoEScale=S, Size=95.32M, FLOPs=91.08M, Zero-shot=true2026.04 | 35.77 | — | — | |
| RFMoEScale=M, Size=307.3M, FLOPs=249.2M, Zero-shot=true2026.04 | 35.27 | — | — | |
| SoftpickQuantization=2-bit2025.04 | 34.6 | 33.88 | — | |
| AoEScale=S, Size=93.85M, FLOPs=88.57M, Zero-shot=true2026.04 | 33.84 | — | — | |
| MoEScale=S, Size=92.44M, FLOPs=90.93M, Zero-shot=true2026.04 | 33.42 | — | — | |
| ReMoEScale=S, Size=92.44M, FLOPs=90.93M, Zero-shot=true2026.04 | 33.38 | — | — | |
| SoftmaxQuantization=2-bit2025.04 | 28.62 | 29.67 | — | |
| FLAPModel=LLaMA-7B, Parameter Retention=40%2026.04 | — | 27.99 | 38.25 | |
| FLAPModel=LLaMA-2-7B, Parameter Retention=40%2026.04 | — | 36.57 | 39.64 | |
| FLAPModel=Vicuna-7B, Parameter Retention=40%2026.04 | — | 35.82 | 38.99 | |
| FLAPModel=LLaMA-2-13B, Parameter Retention=40%2026.04 | — | 40.57 | 43.44 | |
| GRASPruneModel=LLaMA-7B, Parameter Retention=40%2026.04 | — | 35.31 | 40.74 | |
| GRASPruneModel=LLaMA-2-7B, Parameter Retention=40%2026.04 | — | 37.04 | 40.6 | |
| GRASPruneModel=Vicuna-7B, Parameter Retention=40%2026.04 | — | 38.05 | 41.61 | |
| GRASPruneModel=LLaMA-2-13B, Parameter Retention=40%2026.04 | — | 39.94 | 44.12 | |
| LLM-PrunerModel=LLaMA-7B, Parameter Retention=40%2026.04 | — | 28.91 | 37.58 | |
| LLM-PrunerModel=LLaMA-2-7B, Parameter Retention=40%2026.04 | — | 27.78 | 36.54 | |
| LLM-PrunerModel=Vicuna-7B, Parameter Retention=40%2026.04 | — | 30.77 | 38.29 | |
| LLM-PrunerModel=LLaMA-2-13B, Parameter Retention=40%2026.04 | — | 30.43 | 38.81 | |
| SliceGPTModel=LLaMA-7B, Parameter Retention=40%2026.04 | — | 26.64 | 35.28 | |
| SliceGPTModel=LLaMA-2-7B, Parameter Retention=40%2026.04 | — | 28.24 | 36.52 | |
| SliceGPTModel=Vicuna-7B, Parameter Retention=40%2026.04 | — | 29.92 | 36.88 | |
| SliceGPTModel=LLaMA-2-13B, Parameter Retention=40%2026.04 | — | 30.43 | 37.25 |