Advanced Reasoning on T-Math
63.4Accuracyo4-mini (medium)
Evaluation Results
| Method | Links | |
|---|---|---|
| o4-mini (medium)Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 63.4 | |
| DeepSeek-R1Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 61.9 | |
| T-pro 2.0Model Category=Open Source Models (27B-32B class)2025.12 | 54.1 | |
| Qwen3-32BModel Category=Open Source Models (27B-32B class)2025.12 | 52.9 | |
| RuadaptQwen3-32B-InstructModel Category=Open Source Models (27B-32B class)2025.12 | 44.4 | |
| DeepSeek-V3Model Category=Open Source Larger Scale & Proprietary Models2025.12 | 27.8 | |
| DeepSeek-R1-Distill-Qwen-32BModel Category=Open Source Models (27B-32B class)2025.12 | 25.4 | |
| Gemma 3 27BModel Category=Open Source Models (27B-32B class)2025.12 | 20.8 | |
| GigaChat 2 MaxModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 14.2 | |
| YandexGPT5-ProModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 13 | |
| GPT-4oModel Category=Open Source Larger Scale & Proprietary Models2025.12 | 10.6 |