Mathematical Reasoning on AIME 24 (Acc, TR)
90AccuracyPUMA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PUMAModel=Qwen3-30B-A3B-Thinking2026.05 | 90 | 21 | |
| DynasorModel=Qwen3-30B-A3B-Thinking2026.05 | 86.7 | 12.1 | |
| Full CoTModel=Qwen3-30B-A3B-Thinking2026.05 | 83.3 | 0 | |
| DEERModel=Qwen3-30B-A3B-Thinking2026.05 | 83.3 | 10.5 | |
| PUMAModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 70 | 12.3 | |
| Full CoTModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 66.7 | 0 | |
| PUMAModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 60 | 30.3 | |
| No-ThinkModel=Qwen3-30B-A3B-Thinking2026.05 | 60 | 56.5 | |
| DEERModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 53.3 | -22.5 | |
| Full CoTModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 50 | 0 | |
| CoDModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 50 | 40.9 | |
| DynasorModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 50 | -1 | |
| Plan&BudgetModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 43.3 | 52.9 | |
| Plan&BudgetModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 43.3 | 32.9 | |
| Plan&BudgetModel=Qwen3-30B-A3B-Thinking2026.05 | 43.3 | 56.6 | |
| CoDModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 40 | 56.9 | |
| DEERModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 40 | 14.1 | |
| CCoTModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 36.7 | 59.6 | |
| CCoTModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 36.7 | 36 | |
| CCoTModel=Qwen3-30B-A3B-Thinking2026.05 | 36.7 | 55.9 | |
| Ans. Conv.Model=DeepSeek-R1-Distill-Qwen-7B2026.05 | 26.7 | 83.3 | |
| No-ThinkModel=Llama-3.1-Nemotron-Nano-8B2026.05 | 26.7 | -32.3 | |
| DynasorModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 23.3 | 21.7 | |
| CoDModel=Qwen3-30B-A3B-Thinking2026.05 | 23.3 | 52.5 | |
| No-ThinkModel=DeepSeek-R1-Distill-Qwen-7B2026.05 | 20 | 67.7 | |
| Ans. Conv.Model=Llama-3.1-Nemotron-Nano-8B2026.05 | 6.7 | 91.9 | |
| Ans. Conv.Model=Qwen3-30B-A3B-Thinking2026.05 | 0 | 94.4 |