Mathematical Reasoning on Mathematics (Accuracy, SUSTAINSCORE)
85.9AccuracyQwen3-32B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-32B2026.01 | 85.9 | 91.3 | |
| GLM-Z1-32B2026.01 | 85.7 | 89.9 | |
| Qwen3-235B-A22B-Instruct2026.01 | 84.9 | 97.8 | |
| Claude-Sonnet-4-52026.01 | 84.8 | 96.7 | |
| QwQ-32B2026.01 | 83.7 | 90.6 | |
| Deepseek-V3.12026.01 | 83.2 | 94.8 | |
| l-MoEaccActive Params=8.02B, Total Params=72.8B2026.01 | 80.19 | — | |
| GPT-4.1-MINI2026.01 | 78.9 | 94.3 | |
| BaselineActive Params=8.09B, Total Params=72.6B2026.01 | 78.32 | — | |
| l-MoEeffActive Params=5.91B, Total Params=72.8B2026.01 | 77.01 | — | |
| DeepSeek-R1-Distill-Qwen-14B2026.01 | 76.7 | 88.8 | |
| DeepSeek-R1-Distill-Qwen-32B2026.01 | 76.1 | 92.2 | |
| Qwen2.5-72B-Instruct2026.01 | 74.5 | 90.9 | |
| Qwen2.5-32B-Instruct2026.01 | 73.5 | 92.8 | |
| Qwen2.5-14B-Instruct2026.01 | 73.3 | 90.3 | |
| Gemini-2.5-Flash2026.01 | 72.7 | 84.5 | |
| OpenReasoning-Nemotron-14B2026.01 | 72.3 | 91.1 | |
| Meta-Llama-3.1-70B-Instruct2026.01 | 70 | 92.9 | |
| DeepSeek-R1-Distill-Qwen-7B2026.01 | 68.8 | 80.5 | |
| Grok-4-Fast2026.01 | 68.5 | 80.7 | |
| Qwen2.5-7B-Instruct2026.01 | 67.3 | 86.5 | |
| DeepSeek-R1-Distill-Llama-8B2026.01 | 64.8 | 75.2 | |
| Meta-Llama-3.1-8B-Instruct2026.01 | 56.1 | 87.2 | |
| Qwen2.5-1.5B-Instruct2026.01 | 33.1 | 65.3 |