Knowledge on ARC Easy
99.54ARC-E ScoreDeepSeek-R1-Distill-Qwen-32B (Reasoning)
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-R1-Distill-Qwen-32B (Reasoning)Model Family=Qwen2.5-32B2026.01 | 99.54 | |
| DeepSeek-R1-Distill-Qwen-14B (Reasoning)Model Family=Qwen2.5-14B2026.01 | 99.5 | |
| Qwen2.5-32B-Instruct (Safety)Model Family=Qwen2.5-32B2026.01 | 98.94 | |
| ReasonAnyModel Family=Qwen2.5-32B2026.01 | 98.75 | |
| TIESModel Family=Qwen2.5-32B2026.01 | 98.59 | |
| DAREModel Family=Qwen2.5-32B2026.01 | 98.59 | |
| FuseLLMModel Family=Qwen2.5-32B2026.01 | 98.59 | |
| FuseLLMModel Family=Qwen2.5-14B2026.01 | 97.71 | |
| ReasonAnyModel Family=Qwen2.5-14B2026.01 | 97.71 | |
| Task ArithmeticModel Family=Qwen2.5-14B2026.01 | 97.53 | |
| LinearModel Family=Qwen2.5-32B2026.01 | 97.45 | |
| LinearModel Family=Qwen2.5-14B2026.01 | 97.35 | |
| Task ArithmeticModel Family=Qwen2.5-32B2026.01 | 97.22 | |
| Qwen2.5-14B-Instruct (Safety)Model Family=Qwen2.5-14B2026.01 | 97.18 | |
| LEDModel Family=Qwen2.5-14B2026.01 | 97.18 | |
| LEDModel Family=Qwen2.5-32B2026.01 | 96.88 | |
| TIESModel Family=Qwen2.5-14B2026.01 | 96.83 | |
| DAREModel Family=Qwen2.5-14B2026.01 | 91.53 | |
| Task ArithmeticMerging Protocol=Task Arithmetic2026.01 | 90.3 | |
| LEDMerging Protocol=LED2026.01 | 89.77 | |
| ReasonAnyMerging Protocol=ReasonAny2026.01 | 89.77 | |
| Qwen3-1.7B-ALLMEMModel Size=1.7B, Architecture=ALLMEM, Sliding window length=2048, Attention sinks=128, Few-shot setting=0-shot, Max tokens=327682026.02 | 84.7 | |
| Qwen3-1.7BModel Size=1.7B, Architecture=Standard, Few-shot setting=0-shot, Max tokens=327682026.02 | 84.5 | |
| ReasoningBase Model=Llama-3.1-8B2026.01 | 83.95 | |
| DAREMerging Protocol=DARE2026.01 | 77.78 | |
| FuseLLMMerging Protocol=FuseLLM2026.01 | 74.78 | |
| TIESMerging Protocol=TIES2026.01 | 70.37 | |
| Qwen3-0.6B-ALLMEMModel Size=0.6B, Architecture=ALLMEM, Sliding window length=2048, Attention sinks=128, Few-shot setting=0-shot, Max tokens=327682026.02 | 70.2 | |
| Qwen3-0.6BModel Size=0.6B, Architecture=Standard, Few-shot setting=0-shot, Max tokens=327682026.02 | 70 | |
| SafetyBase Model=Llama-3.1-8B2026.01 | 40.74 | |
| LinearMerging Protocol=Linear2026.01 | 31.39 |