Reasoning on BIG-Bench Hard and GSM8K
45.2BBH ScoreQwen3.5 2B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3.5 2BNact/Ntotal=1.9B, Model Type=Instruct-tuned2026.05 | 45.2 | 61.3 | |
| MobileMoE-LNact/Ntotal=922M/5.3B, Model Type=Instruct-tuned2026.05 | 40.1 | 77.6 | |
| MobileMoE-MNact/Ntotal=528M/2.8B, Model Type=Instruct-tuned2026.05 | 39 | 67.5 | |
| Qwen3.5 0.8BNact/Ntotal=749M, Model Type=Instruct-tuned2026.05 | 37.8 | 45.7 | |
| OLMoE-1B-7BNact/Ntotal=1.3B/6.9B, Model Type=Instruct-tuned2026.05 | 37.1 | 49.1 | |
| Gemma 3 1BNact/Ntotal=1.0B, Model Type=Instruct-tuned2026.05 | 35.8 | 38.9 | |
| SmolLM2 1.7BNact/Ntotal=1.7B, Model Type=Instruct-tuned2026.05 | 35.3 | 46.1 | |
| OLMo 2 1BNact/Ntotal=1.5B, Model Type=Instruct-tuned2026.05 | 35 | 46.9 | |
| Llama 3.2 1BNact/Ntotal=1.2B, Model Type=Instruct-tuned2026.05 | 33.9 | 46 | |
| MobileMoE-SNact/Ntotal=272M/1.3B, Model Type=Instruct-tuned2026.05 | 32.2 | 52.2 | |
| Gemma 3 270MNact/Ntotal=270M, Model Type=Instruct-tuned2026.05 | 31.8 | 5.8 | |
| SmolLM2 360MNact/Ntotal=362M, Model Type=Instruct-tuned2026.05 | 30.5 | 10 | |
| MobileLLM-ProNact/Ntotal=1.1B, Model Type=Instruct-tuned2026.05 | 29.1 | 31.8 |