Question Answering on ARC Challenge (acc_norm)
59Normalized AccuracyDCR
Evaluation Results
| Method | Links | |
|---|---|---|
| DCRBase Model=Qwen2.5-7B2026.02 | 59 | |
| STLBase Model=LLaMA-3-8B2026.02 | 56 | |
| SurgicalBase Model=LLaMA-3-8B2026.02 | 56 | |
| SCANSBase Model=LLaMA-3-8B2026.02 | 56 | |
| STL-augBase Model=LLaMA-3-8B2026.02 | 55 | |
| RAISEModel=Qwen-2.5-3B2025.04 | 51.28 | |
| STLBase Model=Qwen2.5-7B2026.02 | 51 | |
| SurgicalBase Model=Qwen2.5-7B2026.02 | 51 | |
| DCRBase Model=LLaMA-3-8B2026.02 | 51 | |
| SSPLModel=Qwen-2.5-3B2025.04 | 50.68 | |
| IFDModel=Qwen-2.5-3B2025.04 | 50.43 | |
| RANDModel=Qwen-2.5-3B2025.04 | 50.18 | |
| STL-augBase Model=Qwen2.5-7B2026.02 | 50 | |
| SCANSBase Model=Qwen2.5-7B2026.02 | 50 | |
| AlpaGasusModel=Qwen-2.5-3B2025.04 | 49.91 | |
| DEITAModel=Qwen-2.5-3B2025.04 | 49.66 | |
| Full Alpaca (100%)Model=Qwen-2.5-3B2025.04 | 49.12 | |
| STLBase Model=Qwen2.5-1.5B2026.02 | 48 | |
| STL-augBase Model=Qwen2.5-1.5B2026.02 | 48 | |
| SurgicalBase Model=Qwen2.5-1.5B2026.02 | 48 | |
| Base Model (0%)Model=Qwen-2.5-3B2025.04 | 47.13 | |
| SCANSBase Model=Qwen2.5-1.5B2026.02 | 47 | |
| DCRBase Model=Qwen2.5-1.5B2026.02 | 47 | |
| RAISEModel=Llama-3.2-3B2025.04 | 46.59 | |
| IFDModel=Llama-3.2-3B2025.04 | 46.42 | |
| TOPSize=7B, Zero-shot=true2025.08 | 46.42 | |
| NTPSize=7B, Zero-shot=true2025.08 | 45.65 | |
| MTPSize=7B, Zero-shot=true2025.08 | 45.56 | |
| DEITAModel=Llama-3.2-3B2025.04 | 44.88 | |
| DS-MTPSize=7B, Zero-shot=true2025.08 | 44.37 | |
| AlpaGasusModel=Llama-3.2-3B2025.04 | 44.11 | |
| Full Alpaca (100%)Model=Llama-3.2-3B2025.04 | 43.77 | |
| RANDModel=Llama-3.2-3B2025.04 | 42.32 | |
| TOPSize=1.8B, Zero-shot=true2025.08 | 42.32 | |
| Base Model (0%)Model=Llama-3.2-3B2025.04 | 42.15 | |
| SSPLModel=Llama-3.2-3B2025.04 | 41.64 | |
| MTPSize=1.8B, Zero-shot=true2025.08 | 40.61 | |
| DS-MTPSize=1.8B, Zero-shot=true2025.08 | 40.44 | |
| GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 38.91 | |
| CCQ-Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 38.82 | |
| NTPSize=1.8B, Zero-shot=true2025.08 | 38.65 | |
| GLA-HedgehogScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 38.65 | |
| Baichuan-7bRatio=0%, Zero-shot=true2024.03 | 38.14 | |
| CCQ-GLAScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 36.95 | |
| Full Alpaca (100%)Model=Llama-3.2-1B2025.04 | 36.86 | |
| SSPLModel=Llama-3.2-1B2025.04 | 36.6 | |
| Mamba2Scale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 36.52 | |
| Gated DeltaNetScale=1.3B, Training tokens=40B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 36.35 | |
| TransformerScale=1.3B, Training tokens=40B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 36.01 | |
| RAISEModel=Llama-3.2-1B2025.04 | 35.58 | |
| RANDModel=Llama-3.2-1B2025.04 | 34.81 | |
| HyWIARatio=25%, Zero-shot=true2024.03 | 34.76 | |
| IFDModel=Llama-3.2-1B2025.04 | 34.47 | |
| AlpaGasusModel=Llama-3.2-1B2025.04 | 33.87 | |
| LLM-Pruner Vector⋆Ratio=25%, Zero-shot=true2024.03 | 33.87 | |
| Base Model (0%)Model=Llama-3.2-1B2025.04 | 33.76 | |
| DEITAModel=Llama-3.2-1B2025.04 | 33.45 | |
| LLM-Pruner Element2⋆Ratio=25%, Zero-shot=true2024.03 | 33.45 | |
| Composite: StretchedModel Architecture=Composer (Stretched), Model Parameter Scale=1B, Pre-training Dataset=DCLM, Pre-training Tokens=37.5B2025.10 | 32.25 | |
| Sand. TransformerModel Architecture=Sandwich Transformer, Model Parameter Scale=1B, Pre-training Dataset=DCLM, Pre-training Tokens=37.5B2025.10 | 30.8 | |
| 1:8 Striped Attn.Model Architecture=Striped Attention (1:8 ratio), Model Parameter Scale=1B, Pre-training Dataset=DCLM, Pre-training Tokens=37.5B2025.10 | 30.7 | |
| CCQ-Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 30.55 | |
| TransformerScale=500M, Training tokens=15B, Model architecture=Attention, Evaluation protocol=Zero-shot2026.05 | 30.12 | |
| MTPSize=340M, Zero-shot=true2025.08 | 29.86 | |
| Llama 3.2Model Architecture=Llama 3.2, Model Parameter Scale=1B, Pre-training Dataset=DCLM, Pre-training Tokens=37.5B2025.10 | 29.8 | |
| 1:4 Striped Attn.Model Architecture=Striped Attention (1:4 ratio), Model Parameter Scale=1B, Pre-training Dataset=DCLM, Pre-training Tokens=37.5B2025.10 | 29.8 | |
| C-AdamWModel Scale=1.2B, Pre-training Budget=1x Chinchilla, Evaluation Framework=LM Eval-Harness2024.11 | 29.78 | |
| CCQ-GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 29.78 | |
| TOPSize=340M, Zero-shot=true2025.08 | 29.35 | |
| Mamba2Scale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 29.27 | |
| 1:2 Striped Attn.Model Architecture=Striped Attention (1:2 ratio), Model Parameter Scale=1B, Pre-training Dataset=DCLM, Pre-training Tokens=37.5B2025.10 | 29 | |
| NTPSize=340M, Zero-shot=true2025.08 | 28.84 | |
| Composite: StackedModel Architecture=Composer (Stacked), Model Parameter Scale=1B, Pre-training Dataset=DCLM, Pre-training Tokens=37.5B2025.10 | 28.84 | |
| AdamWModel Scale=1.2B, Pre-training Budget=1x Chinchilla, Evaluation Framework=LM Eval-Harness2024.11 | 28.75 | |
| Gated DeltaNetScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 28.67 | |
| GLAScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 28.58 | |
| GLA-HedgehogScale=500M, Training tokens=15B, Model architecture=Recurrent, Evaluation protocol=Zero-shot2026.05 | 28.33 | |
| ABL-SOC-110M-3Zero-shot=true2026.03 | 28.16 | |
| STAR*Model Architecture=STAR, Model Parameter Scale=1B, Pre-training Dataset=Original Paper Dataset2025.10 | 27.9 | |
| DS-MTPSize=340M, Zero-shot=true2025.08 | 27.56 | |
| Dense 1.3BArchitecture=Dense, Total Parameters=1.3B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 26.88 | |
| SUB-SOC-110M-1Zero-shot=true2026.03 | 26.79 | |
| MoL 0.61B/2.08BArchitecture=MoL, Total Parameters=2.08B, Active Parameters=0.61B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 26.02 | |
| GatedFWA-NSAClass=Attention, Implementation=Triton, Linear=true2025.12 | 25.52 | |
| HGRN2Class=RNN-Like, Implementation=Triton, Linear=true2025.12 | 25.51 | |
| Dense 0.7BArchitecture=Dense, Total Parameters=0.7B, Precision=bf16, Shot count=0-shot, Evaluation Framework=lm-eval-harness 0.4.72026.05 | 25.43 | |
| SUB-SOC-110M-2Zero-shot=true2026.03 | 25.26 | |
| GatedFWAClass=Attention, Implementation=Triton, Linear=true2025.12 | 25.14 | |
| Transformer (LLaMA) + SWA + NSAClass=Attention, Implementation=Triton, Linear=true2025.12 | 24.93 | |
| ABL-SOC-110M-2Zero-shot=true2026.03 | 24.91 | |
| Transformer (LLaMA)Class=Attention, Implementation=CUDA, Linear=false2025.12 | 24.49 | |
| Transformer (LLaMA) + SWAClass=Attention, Implementation=CUDA, Linear=true2025.12 | 24.4 | |
| DeltaNetClass=RNN-Like, Implementation=Triton, Linear=true2025.12 | 24.32 | |
| Softmaxzero-shot=true2025.01 | 23.72 | |
| MambaClass=RNN-Like, Implementation=CUDA, Linear=true2025.12 | 23.6 | |
| RetNetClass=RNN-Like, Implementation=CUDA, Linear=true2025.12 | 23.4 | |
| GPT-Neo-125MZero-shot=true2026.03 | 23.12 | |
| PLDRv51-SOC-110M-1Zero-shot=true2026.03 | 22.95 | |
| PLDRv51-SOC-110M-5Zero-shot=true2026.03 | 22.95 | |
| GLAClass=RNN-Like, Implementation=Triton, Linear=true2025.12 | 22.7 |