Mathematical Reasoning on AIME 25 (Acc, TR)
83.3AccuracyFull CoT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Full CoTModel=Qwen3-30B-A3B-Thinking2026.05 | 83.3 | 0 | |
| DEERModel=Qwen3-30B-A3B-Thinking2026.05 | 80 | 7.3 | |
| PUMAModel=Qwen3-30B-A3B-Thinking2026.05 | 80 | 14.7 | |
| DynasorModel=Qwen3-30B-A3B-Thinking2026.05 | 76.7 | 2.8 | |
| DEERModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 50 | 13.2 | |
| Full CoTModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 50 | 0 | |
| DEERModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 50 | -35.9 | |
| PUMAModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 50 | 18.1 | |
| No-ThinkModel=Qwen3-30B-A3B-Thinking2026.05 | 50 | 50 | |
| PUMAModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 46.7 | 29.8 | |
| Full CoTModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 43.3 | 0 | |
| DynasorModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 43.3 | -14.4 | |
| Plan&BudgetModel=Qwen3-30B-A3B-Thinking2026.05 | 36.7 | 60.5 | |
| CCoTModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 33.3 | 60.2 | |
| CoDModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 33.3 | 58.7 | |
| DynasorModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 33.3 | 9.1 | |
| CoDModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 33.3 | 39.4 | |
| Plan&BudgetModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 33.3 | 31.6 | |
| No-ThinkModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 23.3 | 57.9 | |
| Plan&BudgetModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 23.3 | 53.9 | |
| Ans. Conv.Model=DeepSeek-R1-Distill-Qwen-7B2026.05 | 20 | 82.8 | |
| CCoTModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 20 | 36.1 | |
| CCoTModel=Qwen3-30B-A3B-Thinking2026.05 | 20 | 60.3 | |
| No-ThinkModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 16.7 | -43.4 | |
| CoDModel=Qwen3-30B-A3B-Thinking2026.05 | 10 | 57.8 | |
| Ans. Conv.Model=Llama-3.1-Nemotron-Nano-8B2026.05 | 3.3 | 90.4 | |
| Ans. Conv.Model=Qwen3-30B-A3B-Thinking2026.05 | 0 | 90.6 |