Multilingual Math Reasoning on MGSM (Mean@3)
84.43Mean@3x1-Qwen3-32B
Evaluation Results
| Method | Links | |
|---|---|---|
| x1-Qwen3-32BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 84.43 | |
| Qwen3-32BReasoning Mode=Think2026.04 | 83.98 | |
| x1-Qwen3-14BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 83.64 | |
| Qwen3-14BReasoning Mode=Think2026.04 | 82.56 | |
| o4-mini-highReasoning Mode=Think2026.04 | 82.32 | |
| Qwen3-32BReasoning Mode=Non-Think2026.04 | 80.52 | |
| x1-Qwen3-32BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 80.12 | |
| DeepSeek-V3.2Reasoning Mode=Non-Think2026.04 | 79.24 | |
| x1-Qwen3-4BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 77.69 | |
| Qwen3-14BReasoning Mode=Non-Think2026.04 | 77.64 | |
| x1-Qwen3-14BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 77.38 | |
| Qwen3-4BReasoning Mode=Think2026.04 | 76.59 | |
| DeepSeek-V3.2Reasoning Mode=Think2026.04 | 76.32 | |
| x1-Qwen3-4BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 70.3 | |
| Qwen3-4BReasoning Mode=Non-Think2026.04 | 70.21 | |
| x1-DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 63.24 | |
| DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Think2026.04 | 60.05 | |
| DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Non-Think2026.04 | 54.76 | |
| x1-DeepSeek-R1-Distill-Qwen-7BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 54.52 | |
| x1-DeepSeek-R1-Distill-Llama-8BReasoning Mode=Think, Task LoRA=+ Math2026.04 | 52.17 | |
| DeepSeek-R1-Distill-Llama-8BReasoning Mode=Think2026.04 | 40.36 | |
| DeepSeek-R1-Distill-Llama-8BReasoning Mode=Non-Think2026.04 | 38.17 | |
| x1-DeepSeek-R1-Distill-Llama-8BReasoning Mode=Non-Think, Task LoRA=+ Math2026.04 | 38.01 |