Mathematical Reasoning on SingleEq (Accuracy)
93.2Accuracy (SingleEq)ExpertCondenser
Evaluation Results
| Method | Links | |
|---|---|---|
| ExpertCondenserModel=GPT-OSS, Model Size=20B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 93.2 | |
| ExpertCondenserModel=DeepSeek-V2-Lite, Model Size=16B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 92.5 | |
| ExpertCondenserModel=OLMoE, Model Size=7B, Post-train Type=ExpertCondenser, Post-training Dataset=math14k, Evaluation Protocol=zero-shot2026.04 | 81.6 | |
| ExpertCondenserModel=Qwen1.5-MoE, Model Size=14B, Post-train Type=ExpertCondenser, Post-training Dataset=math14k, Evaluation Protocol=zero-shot2026.04 | 81.2 | |
| ExpertCondenserModel=DeepSeek-V2-Lite, Model Size=16B, Post-train Type=ExpertCondenser, Post-training Dataset=math14k, Evaluation Protocol=zero-shot2026.04 | 81.2 | |
| ExpertCondenserModel=OLMoE, Model Size=7B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 79.8 | |
| ExpertCondenserModel=Qwen1.5-MoE, Model Size=14B, Post-train Type=ExpertCondenser, Post-training Dataset=math7k, Evaluation Protocol=zero-shot2026.04 | 73.6 |