Mathematical Reasoning on GSM8K (Seed, VeRA-E, Delta Metrics)
95.6SeedGPT-5.1-high
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-5.1-highEvaluation Protocol=Avg@5 (%)2026.01 | 95.6 | 96.9 | 1.4 | |
| Seed-1.6-1015-highEvaluation Protocol=Avg@5 (%)2026.01 | 95.6 | 96.3 | 0.7 | |
| GPT-5-highEvaluation Protocol=Avg@5 (%)2026.01 | 95.6 | 97.1 | 1.6 | |
| qwen3-max-0923Evaluation Protocol=Avg@5 (%)2026.01 | 95.5 | 96.9 | 1.3 | |
| Claude-Sonnet-4.5-thinkingEvaluation Protocol=Avg@5 (%)2026.01 | 95.3 | 96.2 | 0.9 | |
| Seed-1.6-Thinking-0715Evaluation Protocol=Avg@5 (%)2026.01 | 95.1 | 95.9 | 0.8 | |
| Seed-1.6-Lite-1015-highEvaluation Protocol=Avg@5 (%)2026.01 | 95.1 | 95.8 | 0.7 | |
| Kimi-K2-thinkingEvaluation Protocol=Avg@5 (%)2026.01 | 95 | 86.6 | -8.4 | |
| Kimi-K2-0905Evaluation Protocol=Avg@5 (%)2026.01 | 95 | 95.6 | 0.6 | |
| Minimax-M2Evaluation Protocol=Avg@5 (%)2026.01 | 94.7 | 94.9 | 0.2 | |
| Gemini-3-Pro-PreviewEvaluation Protocol=Avg@5 (%)2026.01 | 94.5 | 95 | 0.5 | |
| DeepSeek-V3.1-thinkingEvaluation Protocol=Avg@5 (%)2026.01 | 94.5 | 94.8 | 0.3 | |
| GPT-5.1-chat-latestEvaluation Protocol=Avg@5 (%)2026.01 | 94.3 | 96 | 1.7 | |
| GLM-4.6Evaluation Protocol=Avg@5 (%)2026.01 | 94.1 | 94.8 | 0.7 | |
| DeepSeek-V3.2-thinkingEvaluation Protocol=Avg@5 (%)2026.01 | 94 | 95.4 | 1.4 | |
| Gemini-2.5-ProEvaluation Protocol=Avg@5 (%)2026.01 | 93.9 | 95 | 1.2 |