Multilingual Mathematical Reasoning on MGSM Thai (test)
41.6AccuracyOne Step + KMM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| One Step + KMMk (auxiliary subset size)=4, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 41.6 | — | |
| TV + KMMk (auxiliary subset size)=1, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 40.4 | — | |
| TVk (auxiliary subset size)=2, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 39.6 | — | |
| One Stepk (auxiliary subset size)=3, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 39.2 | — | |
| Randomk (auxiliary subset size)=4, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 38.4 | — | |
| TV + KMMk (auxiliary subset size)=2, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 38 | — | |
| One Step + KMMk (auxiliary subset size)=1, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 37.6 | — | |
| One Step + KMMk (auxiliary subset size)=2, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 37.6 | — | |
| TVk (auxiliary subset size)=3, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 37.6 | — | |
| One Step + KMMk (auxiliary subset size)=5, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 37.2 | — | |
| TV + KMMk (auxiliary subset size)=3, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 37.2 | — | |
| TV + KMMk (auxiliary subset size)=4, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 37.2 | — | |
| TVk (auxiliary subset size)=1, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 36.8 | — | |
| Randomk (auxiliary subset size)=5, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 36.8 | — | |
| One Stepk (auxiliary subset size)=1, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 35.6 | — | |
| One Step + KMMk (auxiliary subset size)=3, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 35.6 | — | |
| One Stepk (auxiliary subset size)=4, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 34.8 | — | |
| TVk (auxiliary subset size)=5, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 34.8 | — | |
| Randomk (auxiliary subset size)=2, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 34.8 | — | |
| TVk (auxiliary subset size)=4, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 34.4 | — | |
| TV + KMMk (auxiliary subset size)=5, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 34.4 | — | |
| One Stepk (auxiliary subset size)=2, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 32.8 | — | |
| Randomk (auxiliary subset size)=3, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 32 | — | |
| One Stepk (auxiliary subset size)=5, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 31.2 | — | |
| Randomk (auxiliary subset size)=1, Pre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | 30.8 | — | |
| One StepPre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | — | 6 | |
| One Step + KMMPre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | — | 14 | |
| RandomPre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | — | 7 | |
| TVPre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | — | 11 | |
| TV + KMMPre-trained model=Qwen2.5-1.5B-Instruct, Algorithm=GRPO2026.05 | — | 12 |