Mathematical Reasoning on AIME 24 (Accuracy, Inference Cost, Cost Reduction)
93.3AccuracyThink@n
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Think@nBackbone=OSS-120B-medium, Aggregation Method=Majority voting over highest DTR scores, Early Stopping Prefix Length=502026.02 | 93.3 | 121.3 | -48 | |
| Cons@nBackbone=Qwen3-4B-Thinking, Aggregation Method=Standard self-consistency2026.02 | 93.3 | 950.1 | — | |
| Think@nBackbone=Qwen3-4B-Thinking, Aggregation Method=Majority voting over highest DTR scores, Early Stopping Prefix Length=502026.02 | 93.3 | 482.2 | -49 | |
| Cons@nBackbone=OSS-120B-medium, Aggregation Method=Standard self-consistency2026.02 | 92.7 | 235.1 | — | |
| Self-Certainty@nBackbone=OSS-120B-medium, Aggregation Method=Majority voting over highest Self-Certainty score, Early Stopping Prefix Length=502026.02 | 91.3 | 119.3 | -49 | |
| Short@nBackbone=Qwen3-4B-Thinking, Aggregation Method=Majority voting over shortest responses2026.02 | 90 | 871 | -8 | |
| Self-Certainty@nBackbone=Qwen3-4B-Thinking, Aggregation Method=Majority voting over highest Self-Certainty score, Early Stopping Prefix Length=502026.02 | 90 | 480.9 | -49 | |
| Short@nBackbone=OSS-120B-medium, Aggregation Method=Majority voting over shortest responses2026.02 | 88 | 200.9 | -15 | |
| Long@nBackbone=OSS-120B-medium, Aggregation Method=Majority voting over longest responses2026.02 | 86.7 | 235.1 | — | |
| Long@nBackbone=Qwen3-4B-Thinking, Aggregation Method=Majority voting over longest responses2026.02 | 86.7 | 950.1 | — | |
| Mean@nBackbone=Qwen3-4B-Thinking, Aggregation Method=Average accuracy2026.02 | 86.3 | 950.1 | — | |
| Mean@nBackbone=OSS-120B-medium, Aggregation Method=Average accuracy2026.02 | 81.6 | 235.1 | — | |
| LaPha / sg@128Base model=Qwen2.5-Math-7B, Tool=✓2026.02 | 60 | — | — | |
| LaPha / sg@128Base model=Qwen2.5-Math-1.5B, Tool=✓2026.02 | 56.7 | — | — | |
| GPT-o1-miniTool=✗2026.02 | 56.7 | — | — | |
| LaPha / sg@128Base model=Qwen2.5-7B, Tool=✓2026.02 | 46.7 | — | — | |
| ToRLBase model=Qwen2.5-Math-7B, Tool=✓2026.02 | 43.3 | — | — | |
| DAPOBase model=Qwen2.5-Math-7B, Tool=✗2026.02 | 36.7 | — | — | |
| SimpleRLBase model=Qwen2.5-Math-7B, Tool=✗2026.02 | 33.3 | — | — | |
| LaPhaBase model=Qwen2.5-Math-1.5B, Tool=✓2026.02 | 30 | — | — | |
| TreePO / maj@16Base model=Qwen2.5-7B, Tool=✗2026.02 | 28.9 | — | — | |
| ToRLBase model=Qwen2.5-Math-1.5B, Tool=✓2026.02 | 26.7 | — | — | |
| SFTBase model=Qwen2.5-Math-7B, Tool=✓2026.02 | 26.7 | — | — | |
| PrimeBase model=Qwen2.5-Math-7B, Tool=✗2026.02 | 26.7 | — | — | |
| SFTBase model=Qwen2.5-Math-1.5B, Tool=✓2026.02 | 23.3 | — | — | |
| LaPha / sg@128Base model=Qwen2.5-1.5B, Tool=✓2026.02 | 20 | — | — | |
| DAPOBase model=Qwen2.5-Math-1.5B, Tool=✗2026.02 | 20 | — | — | |
| DAPOBase model=Qwen2.5-7B, Tool=✗2026.02 | 16.7 | — | — | |
| LaPhaBase model=Qwen2.5-1.5B, Tool=✓2026.02 | 12.7 | — | — | |
| SFTBase model=Qwen2.5-Math-7B, Tool=✗2026.02 | 10 | — | — | |
| GPT-4oTool=✗2026.02 | 9.3 | — | — | |
| SFTBase model=Qwen2.5-7B, Tool=✗2026.02 | 9.1 | — | — | |
| DAPOBase model=Qwen2.5-1.5B, Tool=✗2026.02 | 6.7 | — | — | |
| SFTBase model=Qwen2.5-Math-1.5B, Tool=✗2026.02 | 3.3 | — | — | |
| SFTBase model=Qwen2.5-1.5B, Tool=✗2026.02 | 0.9 | — | — |