Mathematical Reasoning on MATH Easy
97.8AccuracyRPDI-EE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RPDI-EEBackbone=Qwen3-235B-Thinking, Reasoning Strategy=RPDI-EE2026.03 | 97.8 | 5,869 | |
| Vanilla CoTBackbone=Qwen3-235B-Thinking, Reasoning Strategy=Vanilla CoT2026.03 | 97.7 | 5,805 | |
| DEERBackbone=Qwen3-235B-Thinking, Reasoning Strategy=DEER2026.03 | 97.7 | 4,591 | |
| RPDI-EEBackbone=Qwen3-30B-A3B-Thinking-2507, Reasoning Strategy=RPDI-EE2026.03 | 97.4 | 5,039 | |
| DEERBackbone=Qwen3-30B-A3B-Thinking-2507, Reasoning Strategy=DEER2026.03 | 97.2 | 4,477 | |
| ThinkLessBackbone=Qwen3-30B-A3B-Thinking-2507, Reasoning Strategy=ThinkLess2026.03 | 96.9 | 4,409 | |
| Vanilla CoTBackbone=Qwen3-30B-A3B-Thinking-2507, Reasoning Strategy=Vanilla CoT2026.03 | 96.7 | 5,178 | |
| RPDI-EEBackbone=DeepSeek-R1-Distill-Qwen-14B, Reasoning Strategy=RPDI-EE2026.03 | 93.5 | 3,426 | |
| RPDI-EEBackbone=DeepSeek-R1-Distill-Qwen-32B, Reasoning Strategy=RPDI-EE2026.03 | 93.5 | 3,210 | |
| NoThinkingBackbone=Qwen3-30B-A3B-Thinking-2507, Reasoning Strategy=NoThinking2026.03 | 92.2 | 2,820 | |
| DEERBackbone=DeepSeek-R1-Distill-Llama-70B, Reasoning Strategy=DEER2026.03 | 91.7 | 2,200 | |
| DEERBackbone=DeepSeek-R1-Distill-Qwen-14B, Reasoning Strategy=DEER2026.03 | 90.9 | 2,665 | |
| DEERBackbone=DeepSeek-R1-Distill-Qwen-32B, Reasoning Strategy=DEER2026.03 | 90.9 | 2,757 | |
| Vanilla CoTBackbone=DeepSeek-R1-Distill-Qwen-14B, Reasoning Strategy=Vanilla CoT2026.03 | 90 | 3,549 | |
| RPDI-EEBackbone=DeepSeek-R1-Distill-Llama-70B, Reasoning Strategy=RPDI-EE2026.03 | 89.4 | 3,876 | |
| Vanilla CoTBackbone=DeepSeek-R1-Distill-Llama-70B, Reasoning Strategy=Vanilla CoT2026.03 | 89 | 3,639 | |
| Vanilla CoTBackbone=DeepSeek-R1-Distill-Qwen-32B, Reasoning Strategy=Vanilla CoT2026.03 | 88.7 | 3,753 | |
| RPDI-EEBackbone=DeepSeek-R1-Distill-Qwen-7B, Reasoning Strategy=RPDI-EE2026.03 | 88 | 3,825 | |
| Vanilla CoTBackbone=DeepSeek-R1-Distill-Qwen-7B, Reasoning Strategy=Vanilla CoT2026.03 | 87.9 | 3,801 | |
| DEERBackbone=DeepSeek-R1-Distill-Qwen-7B, Reasoning Strategy=DEER2026.03 | 85.4 | 2,733 | |
| RPDI-EEBackbone=DeepSeek-R1-Distill-Llama-8B, Reasoning Strategy=RPDI-EE2026.03 | 82.8 | 4,271 | |
| NoThinkingBackbone=DeepSeek-R1-Distill-Qwen-32B, Reasoning Strategy=NoThinking2026.03 | 82.1 | 797 | |
| NoThinkingBackbone=DeepSeek-R1-Distill-Qwen-14B, Reasoning Strategy=NoThinking2026.03 | 80.5 | 1,187 | |
| NoThinkingBackbone=DeepSeek-R1-Distill-Qwen-7B, Reasoning Strategy=NoThinking2026.03 | 77.6 | 912 | |
| ThinkLessBackbone=DeepSeek-R1-Distill-Qwen-32B, Reasoning Strategy=ThinkLess2026.03 | 77.1 | 485 | |
| ThinkLessBackbone=DeepSeek-R1-Distill-Qwen-7B, Reasoning Strategy=ThinkLess2026.03 | 76.8 | 758 | |
| Dynasor-CoTBackbone=DeepSeek-R1-Distill-Qwen-14B, Reasoning Strategy=Dynasor-CoT2026.03 | 76.4 | 1,535 | |
| Dynasor-CoTBackbone=DeepSeek-R1-Distill-Qwen-32B, Reasoning Strategy=Dynasor-CoT2026.03 | 75.1 | 1,390 | |
| Dynasor-CoTBackbone=Qwen3-30B-A3B-Thinking-2507, Reasoning Strategy=Dynasor-CoT2026.03 | 74.6 | 1,504 | |
| Dynasor-CoTBackbone=DeepSeek-R1-Distill-Qwen-7B, Reasoning Strategy=Dynasor-CoT2026.03 | 73.8 | 1,673 | |
| DEERBackbone=DeepSeek-R1-Distill-Llama-8B, Reasoning Strategy=DEER2026.03 | 73.6 | 3,383 | |
| Vanilla CoTBackbone=DeepSeek-R1-Distill-Llama-8B, Reasoning Strategy=Vanilla CoT2026.03 | 70.8 | 5,036 | |
| RPDI-EEBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Reasoning Strategy=RPDI-EE2026.03 | 69.4 | 5,177 | |
| Vanilla CoTBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Reasoning Strategy=Vanilla CoT2026.03 | 65.7 | 5,615 | |
| DEERBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Reasoning Strategy=DEER2026.03 | 64.3 | 3,202 | |
| ThinkLessBackbone=DeepSeek-R1-Distill-Qwen-14B, Reasoning Strategy=ThinkLess2026.03 | 63.8 | 1,394 |