Reasoning & General on IMO-AnswerBench
86.3ScoreGPT-5.2 (xhigh)
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5.2 (xhigh)2026.02 | 86.3 | |
| Gemini 3 Pro2026.02 | 83.3 | |
| GLM-52026.02 | 82.5 | |
| GLM-4.72026.02 | 82 | |
| Kimi K2.52026.02 | 81.8 | |
| Audex 30B-A3BContext Length=1M2026.07 | 81.1 | |
| Nemotron-3-Super 120B-A12BEvaluation Setting=non-reasoning2026.06 | 79.53 | |
| Claude Opus 4.52026.02 | 78.5 | |
| DeepSeek-V3.22026.02 | 78.3 | |
| Ling-2.6-1T2026.06 | 65.81 | |
| Qwen3-Omni 30B-A3B ThinkingContext Length=64K2026.07 | 59.9 | |
| Ling-2.6-flash2026.06 | 54.28 | |
| Kimi-K2.5Mode=Instant2026.06 | 52.56 | |
| Qwen3.5-Omni Flash 35B-A3BContext Length=256K2026.07 | 51.5 | |
| DeepSeek-V3.2Thinking Mode=nothink2026.06 | 46.66 | |
| GPT-5.4Reasoning Mode=non-reasoning2026.06 | 44.75 | |
| GLM-5Thinking Mode=non-thinking2026.06 | 39.34 | |
| GPT-OSS-120BEvaluation Setting=low2026.06 | 38.59 | |
| GPT-5.4-miniEvaluation Setting=non-reasoning2026.06 | 21.16 |