Multilingual Math Reasoning on MT-AIME
85.67Mean@3DeepSeek-V3.2
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-V3.2Reasoning Mode=Think2026.04 | 85.67 | |
| o4-mini-highReasoning Mode=Think2026.04 | 75.33 | |
| DeepSeek-V3.2Reasoning Mode=Non-Think2026.04 | 53 | |
| x1-Qwen3-32BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 34.5 | |
| Qwen3-32BReasoning Mode=Think2026.04 | 33.89 | |
| x1-Qwen3-14BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 33.11 | |
| Qwen3-14BReasoning Mode=Think2026.04 | 29.22 | |
| x1-DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 27 | |
| DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Think2026.04 | 25.83 | |
| x1-Qwen3-4BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 22.83 | |
| x1-Qwen3-32BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 22.11 | |
| Qwen3-32BReasoning Mode=Non-Think2026.04 | 21.83 | |
| Qwen3-4BReasoning Mode=Think2026.04 | 21.78 | |
| x1-Qwen3-14BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 19.44 | |
| Qwen3-14BReasoning Mode=Non-Think2026.04 | 19.33 | |
| x1-DeepSeek-R1-Distill-Llama-8BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 17 | |
| DeepSeek-R1-Distill-Llama-8BReasoning Mode=Think2026.04 | 14.44 | |
| x1-Qwen3-4BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 13.56 | |
| Qwen3-4BReasoning Mode=Non-Think2026.04 | 12.89 | |
| x1-DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 9 | |
| DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Non-Think2026.04 | 8.33 | |
| x1-DeepSeek-R1-Distill-Llama-8BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 2.89 | |
| DeepSeek-R1-Distill-Llama-8BReasoning Mode=Non-Think2026.04 | 2.67 |