Science Question Answering on ARC
98.8ARC AccuracyLlama-4-Maverick
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-4-MaverickCategory=Candidate Models2026.03 | 98.8 | |
| Claude-Haiku-4.5Category=Candidate Models2026.03 | 98.8 | |
| DeepSeek-R1Category=Candidate Models2026.03 | 98.8 | |
| GraphRouterCategory=Routers2026.03 | 98.8 | |
| FineRouterCategory=Routers2026.03 | 98.8 | |
| IPRCategory=Routers2026.03 | 97.7 | |
| Claude-Sonnet-4.5Category=Candidate Models2026.03 | 97.6 | |
| GPT-OSS-120BCategory=Candidate Models2026.03 | 97.6 | |
| kNNCategory=Routers2026.03 | 97.6 | |
| RouteLLMCategory=Routers2026.03 | 97.6 | |
| Qwen3-32BCategory=Candidate Models2026.03 | 96.4 | |
| Mistral-LargeCategory=Candidate Models2026.03 | 92.9 | |
| DeepSeek-v3Category=Candidate Models2026.03 | 91.7 | |
| Qwen3-235B-A22BCategory=Candidate Models2026.03 | 91.7 | |
| RouterDCCategory=Routers2026.03 | 91.7 | |
| Llama-3.3-70BCategory=Candidate Models2026.03 | 90.5 | |
| FPModel=LLADA-1.5, Precision=Full Precision2026.06 | 88.5 | |
| STaR-QuantModel=LLADA-1.5, Quantization=W8A82026.06 | 86.5 | |
| AutoAdaptSetting=Template Aware (TA), Technique=SFT2026.03 | 83.79 | |
| HF DefaultsSetting=Template Aware (TA), Technique=SFT2026.03 | 78.24 | |
| MLCopilotSetting=Template Aware (TA), Technique=SFT2026.03 | 77.56 | |
| DS-AgentSetting=Template Aware (TA), Technique=SFT2026.03 | 77.47 | |
| AutoMLAgentSetting=Template Aware (TA), Technique=SFT2026.03 | 77.39 | |
| MLPCategory=Routers2026.03 | 70.2 | |
| Mistral-SmallCategory=Candidate Models2026.03 | 69.1 | |
| FullBackbone=Mixtral-8x7B-Instruct, Expert Sparsity=0%2025.12 | 62.54 | |
| FullModel=Mixtral-8x7B-Instruct, Expert Sparsity=0%2025.12 | 62.54 | |
| STaR-QuantModel=DREAM, Quantization=W8A82026.06 | 60.28 | |
| FPModel=DREAM, Precision=Full Precision2026.06 | 59.8 | |
| FullModel=Mixtral-8x7B, Expert Sparsity=0%2025.12 | 56.48 | |
| L3-8B (Base)Model=L3-8B, dqk=128, dvo=128, Method=Base2025.07 | 55.5 | |
| Mistral-7B-INSTBase Model=Mistral-7B-INST, Alignment Strategy=Base, Alignment Data Prompts=HelpSteer22024.08 | 55.29 | |
| RLHFBase Model=Mistral-7B-INST, Alignment Strategy=RLHF, Alignment Data Prompts=HelpSteer22024.08 | 55.2 | |
| UNABase Model=Mistral-7B-INST, Alignment Strategy=UNA, Alignment Data Prompts=HelpSteer22024.08 | 55.2 | |
| KV-Latent Train (L3-8B)Model=L3-8B, dqk=64, dvo=64, Method=Train2025.07 | 53.8 | |
| MoE PathfinderBackbone=Mixtral-8x7B-Instruct, Expert Sparsity=50%2025.12 | 53.18 | |
| MoE PathfinderModel=Mixtral-8x7B-Instruct, Expert Sparsity=50%2025.12 | 53.18 | |
| NAMEx-FullRouting Strategy=Stable-MoE, Evaluation Protocol=Fine-tuned (SmolTalk), Disagreement Point=02025.10 | 50.64 | |
| NAMEx-FullRouting Strategy=Stable-MoE, Evaluation Protocol=Fine-tuned (SmolTalk), Disagreement Point=mean2025.10 | 50.63 | |
| EP-CAMExRouting Strategy=Stable-MoE, Evaluation Protocol=Fine-tuned (SmolTalk)2025.10 | 50.45 | |
| Deepseek-MoERouting Strategy=Stable-MoE, Evaluation Protocol=Fine-tuned (SmolTalk)2025.10 | 50.28 | |
| NAMEx-FullRouting Strategy=Stable-MoE, Evaluation Protocol=Zero-Shot, Disagreement Point=mean2025.10 | 50.19 | |
| NAMEx-FullRouting Strategy=Stable-MoE, Evaluation Protocol=Zero-Shot, Disagreement Point=02025.10 | 50.15 | |
| EP-CAMExRouting Strategy=Stable-MoE, Evaluation Protocol=Zero-Shot2025.10 | 50 | |
| NAMEx-FullRouting Strategy=Cosine, Evaluation Protocol=Fine-tuned (SmolTalk), Disagreement Point=02025.10 | 49.92 | |
| NAMEx-FullRouting Strategy=Cosine, Evaluation Protocol=Fine-tuned (SmolTalk), Disagreement Point=mean2025.10 | 49.92 | |
| Deepseek-MoERouting Strategy=Stable-MoE, Evaluation Protocol=Zero-Shot2025.10 | 49.9 | |
| NAMEx-FullRouting Strategy=Linear, Evaluation Protocol=Fine-tuned (SmolTalk), Disagreement Point=02025.10 | 49.85 | |
| NAMEx-FullRouting Strategy=Linear, Evaluation Protocol=Fine-tuned (SmolTalk), Disagreement Point=mean2025.10 | 49.84 | |
| EP-CAMExRouting Strategy=Cosine, Evaluation Protocol=Fine-tuned (SmolTalk)2025.10 | 49.73 | |
| MoE PathfinderModel=Mixtral-8x7B, Expert Sparsity=50%2025.12 | 49.66 | |
| EP-CAMExRouting Strategy=Linear, Evaluation Protocol=Fine-tuned (SmolTalk)2025.10 | 49.62 | |
| NAMEx-FullRouting Strategy=Cosine, Evaluation Protocol=Zero-Shot, Disagreement Point=02025.10 | 49.6 | |
| Deepseek-MoERouting Strategy=Cosine, Evaluation Protocol=Fine-tuned (SmolTalk)2025.10 | 49.6 | |
| NAMEx-FullRouting Strategy=Cosine, Evaluation Protocol=Zero-Shot, Disagreement Point=mean2025.10 | 49.58 | |
| NAMEx-FullRouting Strategy=Linear, Evaluation Protocol=Zero-Shot, Disagreement Point=mean2025.10 | 49.52 | |
| NAMEx-FullRouting Strategy=Linear, Evaluation Protocol=Zero-Shot, Disagreement Point=02025.10 | 49.51 | |
| Deepseek-MoERouting Strategy=Linear, Evaluation Protocol=Fine-tuned (SmolTalk)2025.10 | 49.5 | |
| EP-CAMExRouting Strategy=Cosine, Evaluation Protocol=Zero-Shot2025.10 | 49.4 | |
| Deepseek-MoERouting Strategy=Cosine, Evaluation Protocol=Zero-Shot2025.10 | 49.3 | |
| EP-CAMExRouting Strategy=Linear, Evaluation Protocol=Zero-Shot2025.10 | 49.26 | |
| Deepseek-MoERouting Strategy=Linear, Evaluation Protocol=Zero-Shot2025.10 | 49.15 | |
| UNABase Model=Qwen2-1.5B-INST, Alignment Strategy=UNA, Alignment Data Prompts=HelpSteer22024.08 | 44.28 | |
| STaR-QuantModel=LLADA, Quantization=W8A82026.06 | 44.25 | |
| FPModel=LLADA, Precision=Full Precision2026.06 | 44.03 | |
| Qwen2-1.5B-INSTBase Model=Qwen2-1.5B-INST, Alignment Strategy=Base, Alignment Data Prompts=HelpSteer22024.08 | 43.94 | |
| LLaMA-3.2-3B-InstructNumber of shots=0-shot2025.12 | 43.6 | |
| RLHFBase Model=Qwen2-1.5B-INST, Alignment Strategy=RLHF, Alignment Data Prompts=HelpSteer22024.08 | 42.83 | |
| LLaMA-3.2-3B-Instruct + SFT-v3.4 + DPONumber of shots=0-shot2025.12 | 41.81 | |
| LLaMA-3.2-3B-Instruct + SFT-v3.4 + DPO + DTFT-v3Number of shots=0-shot2025.12 | 41.13 | |
| LLaMA-3.2-3B-Instruct + SFT-v3.4 + DPO + DTFT-v5Number of shots=0-shot2025.12 | 40.78 | |
| LLaMA-3.2-3B-Instruct + SFT-v3.4 + DPO + DTFT-v4Number of shots=0-shot2025.12 | 40.7 | |
| StatLLaMA (DTFT-v2)Number of shots=0-shot2025.12 | 40.61 | |
| LLaMA-3.2-3B-Instruct + SFT-v3.4 + DPO + DTFT-v1Number of shots=0-shot2025.12 | 40.27 | |
| KV-Latent Distill (L3-8B)Model=L3-8B, dqk=64, dvo=64, Method=Distill2025.07 | 39.1 | |
| KV-Latent Train (L3-8B, d=16)Model=L3-8B, dqk=16, dvo=16, Method=Train2025.07 | 38.5 | |
| L2-7B (Base)Model=L2-7B, dqk=128, dvo=128, Method=Base2025.07 | 30.7 | |
| RandomBackbone=Mixtral-8x7B-Instruct, Expert Sparsity=50%2025.12 | 29.61 | |
| RandomModel=Mixtral-8x7B-Instruct, Expert Sparsity=50%2025.12 | 29.61 | |
| KV-Latent Train (L2-7B)Model=L2-7B, dqk=64, dvo=64, Method=Train2025.07 | 27.5 | |
| KV-Latent Distill (L2-7B)Model=L2-7B, dqk=64, dvo=64, Method=Distill2025.07 | 27 | |
| RandomModel=Mixtral-8x7B, Expert Sparsity=50%2025.12 | 23.03 |