Mathematical Reasoning on GSMPlus
68.38AccuracyGemma-3-4B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemma-3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=4T2025.11 | 68.38 | — | |
| LFM2-8B-A1BTotal Parameters=8.3B, Active Parameters=1.5B, Trained Tokens=13T2025.11 | 64.76 | — | |
| LFM2-2.6BTotal Parameters=2.6B, Active Parameters=2.6B, Trained Tokens=11T2025.11 | 60.75 | — | |
| Granite-4.0-HTotal Parameters=7B, Active Parameters=1B, Trained Tokens=15T2025.11 | 59.14 | — | |
| SmolLM3-3BTotal Parameters=3.1B, Active Parameters=3.1B, Trained Tokens=11T2025.11 | 58.91 | — | |
| Qwen3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=36T2025.11 | 56.16 | — | |
| Llama-3.2-3BTotal Parameters=3.2B, Active Parameters=3.2B, Trained Tokens=9T2025.11 | 38.68 | — | |
| MonoSoupSetting=M-2 (Cosine)2026.02 | 31.9 | — | |
| MonoSoup R = 0.8Setting=M-2 (Cosine)2026.02 | 31.7 | — | |
| MonoSoupSetting=M-3 (Ext.)2026.02 | 31.7 | — | |
| ModelStock (M-2, M-3)Setting=Pairwise2026.02 | 31.6 | — | |
| ModelStock (M-1, M-2)Setting=Pairwise2026.02 | 31.5 | — | |
| LINESSetting=M-2 (Cosine)2026.02 | 31.4 | — | |
| MonoSoup R = 0.8Setting=M-3 (Ext.)2026.02 | 31.4 | — | |
| ModelStock (M-1, M-3)Setting=Pairwise2026.02 | 31.3 | — | |
| StandardSetting=M-2 (Cosine)2026.02 | 30.8 | — | |
| LINESSetting=M-3 (Ext.)2026.02 | 30.8 | — | |
| StandardSetting=M-3 (Ext.)2026.02 | 30.6 | — | |
| MonoSoup R = 0.8Setting=M-1 (Linear)2026.02 | 30.3 | — | |
| MonoSoupSetting=M-1 (Linear)2026.02 | 30.2 | — | |
| LINESSetting=M-1 (Linear)2026.02 | 30.1 | — | |
| StandardSetting=M-1 (Linear)2026.02 | 29.5 | — | |
| QWEN3-0.6B-BASESetting=Reference2026.02 | 22.5 | — | |
| Gemma-3-1B# Total Params=1B, # Trained Tokens=2T2025.11 | — | 40.16 | |
| LFM2-1.2B# Total Params=1.2B, # Trained Tokens=11T2025.11 | — | 36.09 | |
| LFM2-350M# Total Params=0.35B, # Trained Tokens=11T2025.11 | — | 22.16 | |
| LFM2-700M# Total Params=0.70B, # Trained Tokens=11T2025.11 | — | 29.99 | |
| Llama-3.2-1B# Total Params=1.2B, # Trained Tokens=9T2025.11 | — | 18.57 | |
| Original (T2T)Backbone=LLaDA2.1-mini, Inference Strategy=Text-to-Text2026.04 | — | 67.54 | |
| Qwen3-0.6B# Total Params=0.6B, # Trained Tokens=36T2025.11 | — | 9.1 | |
| Qwen3-1.7B# Total Params=1.7B, # Trained Tokens=36T2025.11 | — | 29.62 | |
| T2MBackbone=LLaDA2.1-mini, Inference Strategy=Token-to-Mask, Remasking Strategy=LOWPROB, τ=0.3, Cmax=1, ρmax=0.252026.04 | — | 67.21 |