Mathematical Reasoning on IMO-AnswerBench
93.71AccuracyKimi-K2.6
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Kimi-K2.6Tools usage=with tools, Parameter Count=1T-A32B2026.06 | 93.71 | — | |
| DS-v4-ProTools usage=no tools, Parameter Count=1.6T-A49B2026.06 | 93 | — | |
| Nemotron 3 UltraTools usage=with tools, Parameter Count=550B-A55B2026.06 | 92.3 | — | |
| GLM-5.1Tools usage=with tools, Parameter Count=744B-A40B2026.06 | 91.1 | — | |
| Kimi-K2.6Tools usage=no tools, Parameter Count=1T-A32B2026.06 | 91.1 | — | |
| DS-v4-FlashTools usage=no tools, Parameter Count=284B-A13B2026.06 | 91.1 | — | |
| DS-v4-FlashTools usage=with tools, Parameter Count=284B-A13B2026.06 | 89.6 | — | |
| Nemotron 3 UltraTools usage=no tools, Parameter Count=550B-A55B2026.06 | 88.6 | — | |
| GLM-5.1Tools usage=no tools, Parameter Count=744B-A40B2026.06 | 86.8 | — | |
| DS-v4-ProTools usage=with tools, Parameter Count=1.6T-A49B2026.06 | 85.4 | — | |
| Qwen-3.5Tools usage=with tools, Parameter Count=397B-17B2026.06 | 84.51 | — | |
| DeepSeek-V3.2variant=Speciale, protocol=Pass@12025.12 | 84.5 | 45 | |
| Gemini-3.0variant=Pro, protocol=Pass@12025.12 | 83.3 | 18 | |
| Gemini 3 ProTraining/Evaluation Protocol=proprietary2026.04 | 83.2 | — | |
| Gemini-3.0version=Pro2026.02 | 83.1 | 18,000 | |
| Qwen-3.5Tools usage=no tools, Parameter Count=397B-17B2026.06 | 83.1 | — | |
| GLM-4.72026.07 | 82 | — | |
| Kimi K2.52026.02 | 81.8 | 36,000 | |
| Kimi K2mode=Thinking2026.02 | 78.6 | 37,000 | |
| Kimi-K2variant=Thinking, protocol=Pass@12025.12 | 78.6 | 37 | |
| DeepSeek-V3.2mode=Thinking2026.02 | 78.3 | 27,000 | |
| DeepSeek-V3.2variant=Thinking, protocol=Pass@12025.12 | 78.3 | 27 | |
| QED-Nano (+ RSA test-time scaffold)Model Size=4B, Training/Evaluation Protocol=RSA test-time scaffold2026.04 | 76.5 | — | |
| GPT-5variant=High, protocol=Pass@12025.12 | 76 | 31 | |
| GPT-5 High2026.07 | 76 | — | |
| DeepSeek-Math-V2Model Size=685B2026.04 | 75.8 | — | |
| MiniMax-2.7Tools usage=with tools, Parameter Count=230B-A10B2026.06 | 75.1 | — | |
| Qwen3-30B-A3B SAOBackbone=Qwen3-30B-A3B, Training/Optimization method=SAO2026.07 | 74 | — | |
| Qwen3-30B-A3B - SAO (w/ DIS only)Backbone=Qwen3-30B-A3B, Training/Optimization method=SAO, DIS mechanism=w/ DIS only2026.07 | 71.3 | — | |
| Qwen3-235B-A22B-Thinking-2507Model Size=235B-A22B2026.04 | 70.5 | — | |
| GPT-OSS-120BModel Size=120B2026.04 | 70.5 | — | |
| Qwen3-30B-A3B - GRPO (+ DIS)Backbone=Qwen3-30B-A3B, Training/Optimization method=GRPO, DIS mechanism=+ DIS2026.07 | 70 | — | |
| MiniMax-2.7Tools usage=no tools, Parameter Count=230B-A10B2026.06 | 68.3 | — | |
| QED-NanoModel Size=4B2026.04 | 67.5 | — | |
| Qwen3-30B-A3B-Thinking-2507Model Size=30B-A3B2026.04 | 67 | — | |
| Claude-Sonnet-4.52026.07 | 65.8 | — | |
| GPT-OSS-20BModel Size=20B2026.04 | 61.5 | — | |
| QED-Nano (SFT initialization only)Model Size=4B, Training/Evaluation Protocol=SFT initialization only2026.04 | 57.5 | — | |
| Qwen3-4B-Thinking-2507Model Size=4B2026.04 | 55.8 | — | |
| Qwen3-30B-A3B GRPO (w/ python)Backbone=Qwen3-30B-A3B, Training/Optimization method=GRPO, python usage=w/ python2026.07 | 55.8 | — | |
| Qwen3-30B-A3B w/o pythonBackbone=Qwen3-30B-A3B, python usage=w/o python2026.07 | 55.3 | — | |
| Qwen3-30B-A3B SFT (w/ python)Backbone=Qwen3-30B-A3B, Training/Optimization method=SFT, python usage=w/ python2026.07 | 53.3 | — | |
| Nomos-12026.04 | 49 | — | |
| Qwen3-30B-A3B SFT (w/o python)Backbone=Qwen3-30B-A3B, Training/Optimization method=SFT, python usage=w/o python2026.07 | 42 | — | |
| Qwen3-30B-A3B w/ pythonBackbone=Qwen3-30B-A3B, python usage=w/ python2026.07 | 7.8 | — |