Mathematical Reasoning on AIME 2024 (Accuracy and Reasoning Analysis)
0.8148AccuracyQwen3-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-8BPruning strategy=80% lowest-entropy steps2025.08 | 0.8148 | 11,533.57 | |
| Qwen3-8BPruning strategy=None2025.08 | 0.7931 | 20,936.57 | |
| DeepSeek-R1-14BPruning strategy=None2025.08 | 0.6552 | 15,414.83 | |
| DeepSeek-R1-7BPruning strategy=None2025.08 | 0.6333 | 15,843.43 | |
| DeepSeek-R1-14BPruning strategy=80% lowest-entropy steps2025.08 | 0.5862 | 8,705.57 | |
| DeepSeek-R1-7BPruning strategy=80% lowest-entropy steps2025.08 | 0.5667 | 10,092.8 | |
| VanillaBase Model=DeepSeek-R1-Distill-Llama-8B2026.02 | 0.4854 | 10,774 | |
| ARLCPBase Model=DeepSeek-R1-Distill-Llama-8B2026.02 | 0.4458 | 7,436 | |
| ARLCPBase Model=Qwen3-1.7B2026.02 | 0.4292 | 9,466 | |
| VanillaBase Model=Qwen3-1.7B2026.02 | 0.3875 | 13,138 | |
| Qwen2.5-Math-72B-InstructEvaluating Model=Qwen2.5-Math-72B-Instruct2025.03 | 0.3 | — | |
| DeepSeek-R1-Distill-Qwen-7BEvaluating Model=DeepSeek-R1-Distill-Qwen-7B2025.03 | — | 4,159 |