Mathematical Reasoning on AddSub (Accuracy Metrics)
85.6AccuracyExpertCondenser
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ExpertCondenserModel=GPT-OSS, Model Size=20B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 85.6 | — | — | |
| ExpertCondenserModel=DeepSeek-V2-Lite, Model Size=16B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 79.5 | — | — | |
| ExpertCondenserModel=Qwen1.5-MoE, Model Size=14B, Post-train Type=ExpertCondenser, Post-training Dataset=math14k, Evaluation Protocol=zero-shot2026.04 | 72.8 | — | — | |
| ExpertCondenserModel=OLMoE, Model Size=7B, Post-train Type=ExpertCondenser, Post-training Dataset=math14k, Evaluation Protocol=zero-shot2026.04 | 68.8 | — | — | |
| ExpertCondenserModel=OLMoE, Model Size=7B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 63.4 | — | — | |
| ExpertCondenserModel=Qwen1.5-MoE, Model Size=14B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 61.8 | — | — | |
| ExpertCondenserModel=DeepSeek-V2-Lite, Model Size=16B, Post-train Type=ExpertCondenser, Post-training Dataset=math14k, Evaluation Protocol=zero-shot2026.04 | 60.8 | — | — | |
| GLM-4-Voice2026.04 | — | 59.4 | — | |
| MoshiRAG2026.04 | — | 76.6 | 61.7 | |
| MoshiRAG_GPT-4.oreference generator=GPT-4.o2026.04 | — | 87.9 | 64.8 | |
| MoshiRAGsumm.summarized reference=true2026.04 | — | — | 62 | |
| STITCH-S2026.04 | — | 81.7 | — | |
| Vanilla Moshi2026.04 | — | — | 8.3 |