Mathematical Reasoning on IMO-AnswerBench (Pass@1)
83.3Pass@1Gemini-3.0 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3.0 Pro2025.12 | 83.3 | |
| Nemotron-Cascade-2 30B-A3B2026.03 | 79.3 | |
| Kimi-K2thinking mode=true2025.12 | 78.6 | |
| DeepSeek-V3.2thinking mode=true2025.12 | 78.3 | |
| Nemotron-3-Super 120B-A12BOfficial/Recommended Settings=true2026.03 | 77.2 | |
| GPT-5 High2025.12 | 76 | |
| Qwen3.5 35B-A3BOfficial/Recommended Settings=true2026.03 | 74.8 | |
| Nemotron-3-Nano 30B-A3BOfficial/Recommended Settings=true2026.03 | 70.4 | |
| RSEBackbone=Qwen3-30B-A3B-Thinking-2507, Inference Strategy=Recycling Search Experience, Iteration=It32026.01 | 60.3 | |
| BaseBackbone=Qwen3-30B-A3B-Thinking-2507, Inference Strategy=Standard Sampling, Iteration=It02026.01 | 50.5 | |
| RSEBackbone=Phi-4-Reasoning, Inference Strategy=Recycling Search Experience, Iteration=It32026.01 | 42.3 | |
| BaseBackbone=Phi-4-Reasoning, Inference Strategy=Standard Sampling, Iteration=It02026.01 | 34.5 |