Mathematical Reasoning on AIME 2026
95pass@1Nemotron-Cascade-2 30B-A3B
Evaluation Results
| Method | Links | |
|---|---|---|
| Nemotron-Cascade-2 30B-A3BTool-Integrated Reasoning (TIR)=true2026.03 | 95 | |
| Qwen3.5 35B-A3BOfficial/Recommended Settings=true2026.03 | 91.1 | |
| Nemotron-Cascade-2 30B-A3B2026.03 | 90.9 | |
| Nemotron-3-Nano 30B-A3BOfficial/Recommended Settings=true2026.03 | 89.9 | |
| Nemotron-3-Super 120B-A12BOfficial/Recommended Settings=true2026.03 | 89.8 | |
| Qwen3-4B-Thinking-2507 + CHIMERA# Params=4B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 82.7 | |
| Qwen3-4B-Thinking-2507# Params=4B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 80.8 | |
| DeepSeek-R1-0528-Qwen3-8B# Params=8B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 78 | |
| Qwen3-32B# Params=32B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 74.3 | |
| DeepSeek-R1-Distill-Llama-70B# Params=70B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 59.4 | |
| Qwen3-4B-Thinking-2507 + OpenScience# Params=4B, Scale Category=Small to Medium Scale (≤ 70B)2026.03 | 53 | |
| NFPOBase Model=Qwen3-8B-Base, Algorithm=NFPO2026.05 | 28.6 | |
| DPPOBase Model=Qwen3-8B-Base, Algorithm=DPPO2026.05 | 26.5 | |
| GRPOBase Model=Qwen3-8B-Base, Algorithm=GRPO2026.05 | 23 | |
| NFPOBase Model=Qwen3-1.7B-Base, Algorithm=NFPO2026.05 | 8.1 | |
| DPPOBase Model=Qwen3-1.7B-Base, Algorithm=DPPO2026.05 | 5.6 | |
| Qwen3-8B-BaseBase Model=Qwen3-8B-Base, Algorithm=Base2026.05 | 5.6 | |
| GRPOBase Model=Qwen3-1.7B-Base, Algorithm=GRPO2026.05 | 4.7 | |
| Qwen3-1.7B-BaseBase Model=Qwen3-1.7B-Base, Algorithm=Base2026.05 | 2.1 |