Reasoning on ARC Easy
96.63AccuracyGPT-4
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GPT-4Zero-shot=true2023.11 | 96.63 | — | — | — | — | — | |
| EMoEBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=12 + 22025.09 | 95.94 | — | — | — | — | — | |
| Top-kBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=6 + 22025.09 | 95.24 | — | — | — | — | — | |
| Top-kBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=12 + 22025.09 | 94.89 | — | — | — | — | — | |
| EMoEBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=6 + 22025.09 | 94.71 | — | — | — | — | — | |
| Top-kBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=3 + 22025.09 | 94.18 | — | — | — | — | — | |
| EMoEBase Model=ERNIE-4.5-21B-A3B, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=3 + 22025.09 | 94.18 | — | — | — | — | — | |
| ChatGPTZero-shot=true2023.11 | 93.73 | — | — | — | — | — | |
| KALEBackbone=Qwen3-32B-thinking2026.01 | 93.43 | — | — | — | — | — | |
| Orca-2-13BZero-shot=true, System Message=empty2023.11 | 92.85 | — | — | — | — | — | |
| Fine-tuned SOTASetting=Fine-tuned2020.05 | 92 | — | — | — | — | — | |
| OLMo 2Parameters=7B, Shots=252026.01 | 92 | — | — | — | — | — | |
| OLMo 3Parameters=7B, Shots=252026.01 | 91 | — | — | — | — | — | |
| Engram-40BShots=25-shot2026.01 | 90.1 | — | — | — | — | — | |
| Engram-27BShots=25-shot2026.01 | 89 | — | — | — | — | — | |
| Orca-2-7BZero-shot=true, System Message=empty2023.11 | 87.79 | — | — | — | — | — | |
| SFTBackbone=Qwen3-32B-thinking2026.01 | 87.54 | — | — | — | — | — | |
| BayesLoRAr=8, k=0, Params (M)=4.482025.06 | 86.96 | — | — | — | 12.36 | 0.94 | |
| BayesLoRAk=0, Params (M)=2.40–3.872025.06 | 86.73 | — | — | — | 12.79 | 0.91 | |
| MoE-27BShots=25-shot2026.01 | 86.5 | — | — | — | — | — | |
| BayesLoRAk=10, Params (M)=2.40–3.872025.06 | 86.32 | — | — | — | 7.99 | 0.55 | |
| Orca-1-13BZero-shot=true2023.11 | 86.24 | — | — | — | — | — | |
| MCDParams (M)=4.482025.06 | 86.21 | — | — | — | 12.2 | 1 | |
| BayesLoRAr=8, k=10, Params (M)=4.482025.06 | 86.14 | — | — | — | 7.7 | 0.55 | |
| BBBParams (M)=9.982025.06 | 85.86 | — | — | — | 12.28 | 0.91 | |
| LoRAParams (M)=4.482025.06 | 85.65 | — | — | — | 13.12 | 1.17 | |
| Orca-2-13BZero-shot=true, System Message=cautious2023.11 | 85.31 | — | — | — | — | — | |
| Orca-2-7BZero-shot=true, System Message=cautious2023.11 | 85.1 | — | — | — | — | — | |
| UnfilteredShots=252026.01 | 85 | — | — | — | — | — | |
| Misalignment UpsampledStrategy=CPT, Shots=252026.01 | 85 | — | — | — | — | — | |
| ASCTotal AI agents=30, Cooperative agents=10, GPU=NVIDIA Tesla P1002026.02 | 85 | — | — | — | — | — | |
| VanillaBackbone=Qwen3-32B-thinking2026.01 | 84.6 | — | — | — | — | — | |
| ENSParams (M)=44.802025.06 | 84.4 | — | — | — | 12.57 | 0.82 | |
| Filtered + Alignment UpsampledStrategy=E2E, Shots=252026.01 | 84 | — | — | — | — | — | |
| Misalignment UpsampledStrategy=Mid, Shots=252026.01 | 84 | — | — | — | — | — | |
| EMoEBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=12 + 22025.09 | 83.95 | — | — | — | — | — | |
| LizardTok. (B)=0.04, Cache Size=1322026.07 | 83.5 | — | — | — | — | — | |
| FilteredShots=252026.01 | 83 | — | — | — | — | — | |
| Alignment UpsampledStrategy=Mid, Shots=252026.01 | 83 | — | — | — | — | — | |
| Alignment UpsampledStrategy=CPT, Shots=252026.01 | 83 | — | — | — | — | — | |
| STILL (concurrent)Tok. (B)=0.04, Cache Size=NA2026.07 | 83 | — | — | — | — | — | |
| EMoEBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=6 + 22025.09 | 82.54 | — | — | — | — | — | |
| LlambaTok. (B)=12, Cache Size=0, Param.=8B2026.07 | 82.5 | — | — | — | — | — | |
| LoLaTok. (B)=0.04, Cache Size=1282026.07 | 82.5 | — | — | — | — | — | |
| LolCaTTok. (B)=0.04, Cache Size=642026.07 | 82.37 | — | — | — | — | — | |
| GenKnowSubSetting=Avg, Zero-shot=true, Base Model=Phi-32025.05 | 82.28 | — | — | — | — | — | |
| LLAMA-2-Chat-70BZero-shot=true2023.11 | 82.2 | — | — | — | — | — | |
| LLaMA 3.1 8BCache Size=∞2026.07 | 82.15 | — | — | — | — | — | |
| OURS (GDN)Tok. (B)=0.01, Cache Size=128, Param.=10.3M2026.07 | 82.15 | — | — | — | — | — | |
| GenKnowSubSetting=En, Zero-shot=true, Base Model=Phi-32025.05 | 82.11 | — | — | — | — | — | |
| OURS (GDN)Tok. (B)=0.01, Cache Size=64, Param.=10.3M2026.07 | 81.97 | — | — | — | — | — | |
| Liger-GLATok. (B)=0.02, Cache Size=642026.07 | 81.8 | — | — | — | — | — | |
| GenKnowSubSetting=De, Zero-shot=true, Base Model=Phi-32025.05 | 81.75 | — | — | — | — | — | |
| GenKnowSubSetting=Fr, Zero-shot=true, Base Model=Phi-32025.05 | 81.75 | — | — | — | — | — | |
| ARMADALoss Function=L_cosine, Teacher Model=Stable Diffusion, Student Model=LLaMA-7B, Zero-shot=true2026.03 | 81.1 | — | — | — | — | — | |
| ClusCompBackbone=Llama-3-8B, #Bit=4.13, Zero-shot=true2025.03 | 80.9 | — | 79.6 | — | — | — | |
| LLaMA-7BDistillation type=undistilled, Zero-shot=true2026.03 | 80.9 | — | — | — | — | — | |
| ARMADALoss Function=L_euclid, Teacher Model=Stable Diffusion, Student Model=LLaMA-7B, Zero-shot=true2026.03 | 80.8 | — | — | — | — | — | |
| ARMADALoss Function=L_elementwise, Teacher Model=Stable Diffusion, Student Model=LLaMA-7B, Zero-shot=true2026.03 | 80.8 | — | — | — | — | — | |
| WizardLM-70BZero-shot=true2023.11 | 80.68 | — | — | — | — | — | |
| ArrowZero-shot=true, Base Model=Phi-32025.05 | 80.53 | — | — | — | — | — | |
| Llama-3-8BBackbone=Llama-3-8B, #Bit=16.00, Zero-shot=true2025.03 | 80.1 | — | — | — | — | — | |
| LLaMAParameters=33B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 80 | — | — | — | — | — | |
| Misalignment UpsampledStrategy=E2E, Shots=252026.01 | 80 | — | — | — | — | — | |
| SliM-LLMBackbone=Llama-3-8B, #Bit=4.13, Zero-shot=true2025.03 | 79.9 | — | — | — | — | — | |
| EMoEBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=EMoE, Inference Budget (k')=42025.09 | 79.72 | — | — | — | — | — | |
| AWQBackbone=Llama-3-8B, #Bit=4.13, Zero-shot=true2025.03 | 79.7 | — | — | — | — | — | |
| EMoEBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=EMoE, Inference Budget (k')=62025.09 | 79.37 | — | — | — | — | — | |
| Filtered + Alignment UpsampledStrategy=Mid, Shots=252026.01 | 79 | — | — | — | — | — | |
| Filtered + Alignment UpsampledStrategy=CPT, Shots=252026.01 | 79 | — | — | — | — | — | |
| LLaMAParameters=65B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 78.9 | — | — | — | — | — | |
| GPTQBackbone=Llama-3-8B, #Bit=4.13, Zero-shot=true2025.03 | 78.8 | — | — | — | — | — | |
| SharedZero-shot=true, Base Model=Phi-32025.05 | 78.77 | — | — | — | — | — | |
| EMoEBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=EMoE, Inference Budget (k')=22025.09 | 78.66 | — | — | — | — | — | |
| EMoEBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=EMoE, Inference Budget (k')=3 + 22025.09 | 78.66 | — | — | — | — | — | |
| LLAMA-3-8BModel=LLAMA-3-8B, Bits=FP16, Zero-shot=true2026.04 | 77.48 | — | — | — | — | — | |
| Llama 2Parameters=7B, Shots=252026.01 | 77 | — | — | — | — | — | |
| Dense-4BShots=25-shot2026.01 | 76.8 | — | — | — | — | — | |
| Top-kBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=Top-k, Inference Budget (k')=42025.09 | 76.65 | — | — | — | — | — | |
| PaLMParameters=540B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 76.6 | — | — | — | — | — | |
| EMoEBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=EMoE, Inference Budget (k')=162025.09 | 76.37 | — | — | — | — | — | |
| Llama2-7BZero-shot=true, Number of Parameters=7B2023.09 | 76.3 | — | — | — | — | — | |
| LLAMA-2-Chat-13BZero-shot=true2023.11 | 76.26 | — | — | — | — | — | |
| FP16Model=Hymba-Instruct-1.5B, Compression Ratio=1x2026.01 | 76.14 | — | — | — | — | — | |
| phi-1.5-web (1.3B)Zero-shot=true, Number of Parameters=1.3B2023.09 | 76.1 | — | — | — | — | — | |
| ClusCompBackbone=Llama-3-8B, #Bit=2.87, Zero-shot=true2025.03 | 76 | — | 74.5 | — | — | — | |
| LAWTotal AI agents=30, Cooperative agents=10, GPU=NVIDIA Tesla P1002026.02 | 76 | — | — | — | — | — | |
| Top-kBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=Top-k, Inference Budget (k')=22025.09 | 75.98 | — | — | — | — | — | |
| QMCModel=Hymba-Instruct-1.5B, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 75.96 | — | — | — | — | — | |
| FP16Model=Qwen2.5-1.5B-Instruct, Compression Ratio=1x2026.01 | 75.8 | — | — | — | — | — | |
| phi-1.5 (1.3B)Zero-shot=true, Number of Parameters=1.3B2023.09 | 75.6 | — | — | — | — | — | |
| Top-kBase Model=LoRAMoE, Trained Activated Experts=2, MoE Routing Method=Top-k, Inference Budget (k')=62025.09 | 75.6 | — | — | — | — | — | |
| Top-kBase Model=DeepSeek-V2-Lite, Trained Activated Experts=6+2, MoE Routing Method=Top-k, Inference Budget (k')=6 + 22025.09 | 75.49 | — | — | — | — | — | |
| QMCModel=Hymba-Instruct-1.5B, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 75.42 | — | — | — | — | — | |
| Vicuna-13B (v1.1)Zero-shot=true, Number of Parameters=13B, Version=v1.12023.09 | 75.4 | — | — | — | — | — | |
| PaLMParameters=62B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 75.2 | — | — | — | — | — | |
| EMoEBase Model=OLMoE-1B-7B-0924, Trained Activated Experts=8, MoE Routing Method=EMoE, Inference Budget (k')=82025.09 | 75.13 | — | — | — | — | — | |
| OriginalCompression Ratio=0%, Memory (GB)=26.95 GB2026.02 | 75 | — | — | — | — | — | |
| MPT-7BZero-shot=true, Number of Parameters=7B2023.09 | 74.9 | — | — | — | — | — | |
| LLaMAParameters=13B, Zero-shot=true, Likelihood Normalization=Per-character2023.02 | 74.8 | — | — | — | — | — |