Mathematical Reasoning on MGSM Bangla
0.88Accuracy (Original)Qwen 3
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Qwen 3Model Category=Reasoning Models, Parameter Count=8B2026.01 | 0.88 | 0.705 | 0.175 | 0.714 | 2,074 | 3,662 | 3,128 | |
| Qwen 3Model Category=Reasoning Models, Parameter Count=4B2026.01 | 0.828 | 0.629 | 0.199 | 0.671 | 2,068 | 3,842 | 3,140 | |
| †DAGGERBackbone=Gemma 3, Parameter Count=12B, Training Stage=GRPO2026.01 | 0.784 | 0.64 | 0.144 | 0.694 | 411 | 519 | 359 | |
| Gemma 3Model Category=LLM (5-shot CoT), Parameter Count=12B, Protocol=5-shot CoT2026.01 | 0.768 | 0.543 | 0.225 | 0.557 | 582 | 659 | 599 | |
| †DAGGERBackbone=Gemma 3, Parameter Count=12B, Training Stage=SFT2026.01 | 0.7 | 0.568 | 0.132 | 0.667 | 407 | 500 | 334 | |
| †DAGGERBackbone=Gemma 3, Parameter Count=12B, Training Stage=w/o train.2026.01 | 0.604 | 0.477 | 0.127 | 0.617 | 376 | 484 | 316 | |
| Gemma 3Model Category=LLM (5-shot CoT), Parameter Count=4B, Protocol=5-shot CoT2026.01 | 0.548 | 0.263 | 0.285 | 0.362 | 706 | 680 | 700 | |
| †DAGGERBackbone=Gemma 3, Parameter Count=4B, Training Stage=GRPO2026.01 | 0.548 | 0.314 | 0.234 | 0.473 | 371 | 458 | 349 | |
| LLaMA 3Model Category=LLM (5-shot CoT), Parameter Count=8B, Protocol=5-shot CoT2026.01 | 0.492 | 0.222 | 0.27 | 0.33 | 295 | 413 | 386 | |
| †DAGGERBackbone=Gemma 3, Parameter Count=4B, Training Stage=SFT2026.01 | 0.404 | 0.251 | 0.153 | 0.443 | 392 | 538 | 370 | |
| Qwen 2.5Model Category=LLM (5-shot CoT), Parameter Count=3B, Protocol=5-shot CoT2026.01 | 0.384 | 0.137 | 0.247 | 0.226 | 536 | 784 | 636 | |
| Qwen 2.5Model Category=LLM (5-shot CoT), Parameter Count=7B, Protocol=5-shot CoT2026.01 | 0.264 | 0.232 | 0.032 | 0.371 | 560 | 707 | 613 | |
| †DAGGERBackbone=Gemma 3, Parameter Count=4B, Training Stage=w/o train.2026.01 | 0.248 | 0.099 | 0.149 | 0.333 | 620 | 911 | 399 |