Advanced Reasoning on ruAIME 2025
80AccuracyDeepSeek-R1
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-R1Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 80 | |
| o4-mini (medium)Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 77.1 | |
| T-pro 2.0Model Category=Open Source Models (27B-32B class)2025.12 | 64.6 | |
| Qwen3-32BModel Category=Open Source Models (27B-32B class)2025.12 | 62.5 | |
| RuadaptQwen3-32B-InstructModel Category=Open Source Models (27B-32B class)2025.12 | 45 | |
| DeepSeek-R1-Distill-Qwen-32BModel Category=Open Source Models (27B-32B class)2025.12 | 40.2 | |
| DeepSeek-V3Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 28.5 | |
| Gemma 3 27BModel Category=Open Source Models (27B-32B class)2025.12 | 23.1 | |
| GPT-4oModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 6.9 | |
| GigaChat 2 MaxModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 6.2 | |
| YandexGPT5-ProModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 4.6 |