Reasoning on ARC Challenge (P@1, P@32, #Tok)
95.5P@1DeepSeek-V4-pro
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DeepSeek-V4-proModel Scale=Reference, Model Category=Open-source frontier2026.06 | 95.5 | — | — | |
| Qwen3-maxModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 93.9 | — | — | |
| Qwen3-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 92 | — | — | |
| GPT-5(high)Model Scale=Reference, Model Category=Closed-source frontier2026.06 | 91.8 | — | — | |
| GPT-o1-miniModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 90.4 | — | — | |
| DeepSeek-R1Model Scale=Reference, Model Category=Open-source frontier2026.06 | 89.6 | — | — | |
| DiScOModel Scale=32B, Model Category=Ours2026.06 | 88 | — | — | |
| DeepSeek-R1-Distill-Qwen-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 87.3 | — | — | |
| QwQ-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 87.1 | — | — | |
| GPT-o1-previewModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 86.8 | — | — | |
| DiScOw/o ITDEModel Scale=32B, Model Category=Ours2026.06 | 86.4 | — | — | |
| LLaMA3.1-405B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 85.6 | — | — | |
| Qwen2.5-32B-InstructModel Scale=32B, Model Category=Instruction-tuned models2026.06 | 85.5 | — | — | |
| LLaMA3.1-70B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 85.2 | — | — | |
| DiScOModel Scale=7B, Model Category=Ours2026.06 | 84.8 | — | — | |
| DeepSeek-V3Model Scale=Reference, Model Category=Open-source frontier2026.06 | 84.7 | — | — | |
| Qwen2.5-Math-72B-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 84.4 | — | — | |
| ReasonFlux-32BModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 84.3 | — | — | |
| Sky-T1-32B-previewModel Scale=32B, Model Category=Reasoning-specialized models2026.06 | 83.7 | — | — | |
| Qwen2.5-Math-7B-InstructModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 83.6 | — | — | |
| DIVERModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 83 | — | — | |
| DeepSeek-Coder-V2-InstructModel Scale=Reference, Model Category=Open-source frontier2026.06 | 82.5 | — | — | |
| DiScOw/o ITDEModel Scale=7B, Model Category=Ours2026.06 | 82.3 | — | — | |
| Qwen3-8BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 82 | — | — | |
| GRPO w/ Clip-higherModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 81.6 | — | — | |
| rStar-Math-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 79.3 | — | — | |
| Pass@k TrainingModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 78.5 | — | — | |
| Claude3.5-SonnetModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 76.7 | — | — | |
| HRPOBackbone=Qwen2.5-3B-Instruct2026.06 | 75.06 | 97.44 | 277.8 | |
| GRPOBackbone=Qwen2.5-3B-Instruct2026.06 | 74.78 | 98.12 | 272.4 | |
| CoTBackbone=Qwen2.5-3B-Instruct2026.06 | 74.05 | 98.21 | 106.2 | |
| TARPOBackbone=Qwen2.5-3B-Instruct2026.06 | 74.01 | 98.89 | 189.4 | |
| ReasonFlux-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 74 | — | — | |
| NuminaMath-72B-CoTModel Scale=Reference, Model Category=Open-source frontier2026.06 | 73.9 | — | — | |
| GPT-4oModel Scale=Reference, Model Category=Closed-source frontier2026.06 | 72.8 | — | — | |
| Entropy-RoutedBackbone=Qwen2.5-3B-Instruct2026.06 | 71.09 | 96.25 | 95.4 | |
| Pure LatentBackbone=Qwen2.5-3B-Instruct2026.06 | 70.57 | 78.16 | 97.1 | |
| SuperCorrect-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 67.3 | — | — | |
| Qwen2.5-Math-7BModel Scale=7B, Model Category=Math instruction-tuned models2026.06 | 65.3 | — | — | |
| DeepSeek-R1-Distill-Qwen-7BModel Scale=7B, Model Category=Reasoning-specialized models2026.06 | 58.8 | — | — | |
| Entropy-RLModel Scale=7B, Model Category=Diversity-oriented RL models2026.06 | 52 | — | — | |
| LLaMA3.1-8B-InstructModel Scale=7B, Model Category=General instruction-tuned models2026.06 | 22.7 | — | — | |
| Mathstral-7B-v0.1Model Scale=7B, Model Category=General instruction-tuned models2026.06 | 22.4 | — | — |