Mathematical Reasoning on GSM8k (Average Accuracy)
66.42Average AccuracyQwen-4B (teacher)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen-4B (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 66.42 | 84.61 | |
| Phi-mini (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 63.72 | 82.71 | |
| Model SwarmsSetting=single-task, Number of runs=52026.06 | 48.9 | — | |
| Llama-3B (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 46.93 | 24.94 | |
| LoraHubSetting=single-task, Number of runs=52026.06 | 46.47 | — | |
| MERGEvolveSetting=single-task, Number of runs=52026.06 | 44 | — | |
| Best Single ExpertSetting=single-task, Number of runs=52026.06 | 40.8 | — | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Llama-3B, Few-shot protocol=3-shot2026.05 | 40.48 | 20.39 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Qwen-4B + Llama-3B, Few-shot protocol=3-shot2026.05 | 40.15 | 19.18 | |
| Expert FusionSetting=single-task, Number of runs=52026.06 | 39.5 | — | |
| X-Token (H-KL)Setting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 39.18 | 19.11 | |
| X-Token (P-KL)Setting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 38.85 | 15.54 | |
| GOLDSetting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 38.66 | 16.5 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Qwen-4B, Few-shot protocol=3-shot2026.05 | 38.49 | 14.63 | |
| Llama-3B → 1BSetting=Same tokenizer, Teacher=Llama-3B, Few-shot protocol=3-shot2026.05 | 38.4 | 12.89 | |
| ULDSetting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 38.31 | 17.97 | |
| EMMSetting=single-task, Number of runs=52026.06 | 38.31 | — | |
| Pack of LLMsSetting=single-task, Number of runs=52026.06 | 37.6 | — | |
| ULDSetting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 36.77 | 14.56 | |
| Continued pre-trainingSetting=No distillation, Few-shot protocol=3-shot2026.05 | 36.63 | 10.25 | |
| GOLDSetting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 35.03 | 2.56 | |
| Llama-1B (base)Setting=No distillation, Few-shot protocol=3-shot2026.05 | 33.96 | 5.69 | |
| Data MergeSetting=single-task, Number of runs=52026.06 | 33 | — | |
| TIESSetting=single-task, Number of runs=52026.06 | 28 | — |