Mathematical Reasoning on AIME 2025 (Accuracy, Output Token Count)
96AccuracyDeepSeek-V3.2
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-V3.2variant=Speciale, protocol=Pass@12025.12 | 96 | 23,000 | |
| Gemini-3.0variant=Pro, protocol=Pass@12025.12 | 95 | 15,000 | |
| GPT-5variant=High, protocol=Pass@12025.12 | 94.6 | 13,000 | |
| Kimi-K2variant=Thinking, protocol=Pass@12025.12 | 94.5 | 24,000 | |
| DeepSeek-V3.2variant=Thinking, protocol=Pass@12025.12 | 93.1 | 16,000 |