Advanced Reasoning on ruGPQA Diamond
0.773Accuracyo4-mini (medium)
Evaluation Results
| Method | Links | |
|---|---|---|
| o4-mini (medium)Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.773 | |
| DeepSeek-R1Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.763 | |
| DeepSeek-V3Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.657 | |
| DeepSeek-R1-Distill-Qwen-32BModel Category=Open Source Models (27B-32B class)2025.12 | 0.631 | |
| Qwen3-32BModel Category=Open Source Models (27B-32B class)2025.12 | 0.606 | |
| T-pro 2.0Model Category=Open Source Models (27B-32B class)2025.12 | 0.591 | |
| RuadaptQwen3-32B-InstructModel Category=Open Source Models (27B-32B class)2025.12 | 0.591 | |
| GPT-4oModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.51 | |
| GigaChat 2 MaxModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.475 | |
| Gemma 3 27BModel Category=Open Source Models (27B-32B class)2025.12 | 0.439 | |
| YandexGPT5-ProModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.354 |