Mathematical Reasoning on AMO-Bench (VeRA Metrics)
0.56Seed (Avg@5)GPT-5.1-high
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT-5.1-high2026.01 | 0.56 | 0.646 | 0.086 | 0.588 | 0.028 | |
| Gemini-3-Pro-Preview2026.01 | 0.56 | 0.611 | 0.051 | 0.56 | 0 | |
| Seed-1.6-1015-high2026.01 | 0.48 | 0.431 | -0.049 | 0.364 | -0.116 | |
| DeepSeek-V3.1-thinking2026.01 | 0.48 | 0.521 | 0.041 | 0.528 | 0.048 | |
| GLM-4.62026.01 | 0.4 | 0.529 | 0.129 | 0.48 | 0.08 | |
| Seed-1.6-Thinking-07152026.01 | 0.4 | 0.416 | 0.016 | 0.38 | -0.02 | |
| GPT-5-high2026.01 | 0.4 | 0.584 | 0.184 | 0.544 | 0.144 | |
| Seed-1.6-Lite-1015-high2026.01 | 0.36 | 0.397 | 0.037 | 0.364 | 0.004 | |
| Gemini-2.5-Pro2026.01 | 0.28 | 0.435 | 0.155 | 0.376 | 0.096 | |
| DeepSeek-V3.2-thinking2026.01 | 0.28 | 0.516 | 0.236 | 0.46 | 0.18 | |
| Kimi-K2-thinking2026.01 | 0.22 | 0.265 | 0.045 | 0.172 | -0.048 | |
| Minimax-M22026.01 | 0.2 | 0.366 | 0.166 | 0.284 | 0.084 | |
| Claude-Sonnet-4.5-thinking2026.01 | 0.18 | 0.36 | 0.18 | 0.324 | 0.144 | |
| qwen3-max-09232026.01 | 0.14 | 0.392 | 0.252 | 0.352 | 0.212 | |
| Kimi-K2-09052026.01 | 0.08 | 0.199 | 0.119 | 0.188 | 0.108 | |
| GPT-5.1-chat-latest2026.01 | 0.06 | 0.207 | 0.147 | 0.176 | 0.116 |