Mathematical Reasoning on MATH (MATH, Avg.)
61.2MATH ScoreGrid
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GridBackbone=Qwen3-4B, Data Scale=# 20K, Training Mixture Context=AM-Thinking-v1-Distilled-Code&Math2026.07 | 61.2 | 74.27 | |
| CAUSALMIXBackbone=Qwen3-4B, Data Scale=# 20K, Training Mixture Context=AM-Thinking-v1-Distilled-Code&Math2026.07 | 60.58 | 74.72 | |
| EqualBackbone=Qwen3-4B, Data Scale=# 20K, Training Mixture Context=AM-Thinking-v1-Distilled-Code&Math2026.07 | 56.78 | 73.62 | |
| DMOBackbone=Qwen3-4B, Data Scale=# 20K, Training Mixture Context=AM-Thinking-v1-Distilled-Code&Math2026.07 | 54.38 | 72 | |
| DoReMiBackbone=Qwen3-4B, Data Scale=# 20K, Training Mixture Context=AM-Thinking-v1-Distilled-Code&Math2026.07 | 42.22 | 65.39 | |
| ODMBackbone=Qwen3-4B, Data Scale=# 20K, Training Mixture Context=AM-Thinking-v1-Distilled-Code&Math2026.07 | 41.16 | 64.74 | |
| RegMixBackbone=Qwen3-4B, Data Scale=# 20K, Training Mixture Context=AM-Thinking-v1-Distilled-Code&Math2026.07 | 40.8 | 65.21 | |
| Qwen-4B (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 27.76 | 66.42 | |
| Phi-mini (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 19.3 | 63.72 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Llama-3B, Few-shot protocol=3-shot2026.05 | 9.02 | 40.48 | |
| Llama-3B (teacher)Setting=Teachers (reference), Few-shot protocol=3-shot2026.05 | 8.82 | 46.93 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Qwen-4B + Llama-3B, Few-shot protocol=3-shot2026.05 | 8.56 | 40.15 | |
| X-Token (H-KL)Setting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 8.32 | 39.18 | |
| Llama-3B → 1BSetting=Same tokenizer, Teacher=Llama-3B, Few-shot protocol=3-shot2026.05 | 8.16 | 38.4 | |
| X-TokenSetting=Multi-teacher, Teachers=Phi-mini + Qwen-4B, Few-shot protocol=3-shot2026.05 | 8.1 | 38.49 | |
| X-Token (P-KL)Setting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 7.96 | 38.85 | |
| GOLDSetting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 7.8 | 38.66 | |
| Continued pre-trainingSetting=No distillation, Few-shot protocol=3-shot2026.05 | 6.9 | 36.63 | |
| ULDSetting=Cross tokenizer (single teacher), Teacher=Phi-mini, Few-shot protocol=3-shot2026.05 | 6.24 | 38.31 | |
| Llama-1B (base)Setting=No distillation, Few-shot protocol=3-shot2026.05 | 5.48 | 33.96 | |
| GOLDSetting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 4.5 | 35.03 | |
| ULDSetting=Cross tokenizer (single teacher), Teacher=Qwen-4B, Few-shot protocol=3-shot2026.05 | 4.04 | 36.77 |