Coding Reasoning on ruLCB
0.705Accuracyo4-mini (medium)
Evaluation Results
| Method | Links | |
|---|---|---|
| o4-mini (medium)Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.705 | |
| DeepSeek-R1Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.69 | |
| T-pro 2.0Model Category=Open Source Models (27B-32B class)2025.12 | 0.563 | |
| Qwen3-32BModel Category=Open Source Models (27B-32B class)2025.12 | 0.537 | |
| RuadaptQwen3-32B-InstructModel Category=Open Source Models (27B-32B class)2025.12 | 0.5 | |
| DeepSeek-R1-Distill-Qwen-32BModel Category=Open Source Models (27B-32B class)2025.12 | 0.493 | |
| DeepSeek-V3Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.444 | |
| GigaChat 2 MaxModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.272 | |
| YandexGPT5-ProModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.265 | |
| Gemma 3 27BModel Category=Open Source Models (27B-32B class)2025.12 | 0.261 | |
| GPT-4oModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 0.131 |