Math Reasoning on AIME 2025 (Pass@4)
50Pass@4Qwen3 (ExpertCondenser)
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3 (ExpertCondenser)Model Size=30B, Distill Type=ExpertCondenser, Fine-tuning with Stanford-S1=true2026.04 | 50 | |
| Qwen3 (SFT)Model Size=30B, Distill Type=SFT, Fine-tuning with Stanford-S1=true2026.04 | 48.3 | |
| Qwen3 (DenseMixer)Model Size=30B, Distill Type=DenseMixer, Fine-tuning with Stanford-S1=true2026.04 | 46.7 | |
| Qwen3 (ESFT)Model Size=30B, Distill Type=ESFT, Fine-tuning with Stanford-S1=true2026.04 | 44.2 | |
| Qwen1.5-MoE (ExpertCondenser)Model Size=14B, Distill Type=ExpertCondenser, Fine-tuning with Stanford-S1=true2026.04 | 2.5 | |
| DeepSeek-Coder-V2-Lite (SFT)Model Size=16B, Distill Type=SFT, Fine-tuning with Stanford-S1=true2026.04 | 2.5 | |
| DeepSeek-Coder-V2-Lite (ESFT)Model Size=16B, Distill Type=ESFT, Fine-tuning with Stanford-S1=true2026.04 | 2.5 | |
| DeepSeek-Coder-V2-Lite (DenseMixer)Model Size=16B, Distill Type=DenseMixer, Fine-tuning with Stanford-S1=true2026.04 | 2.5 | |
| DeepSeek-Coder-V2-Lite (ExpertCondenser)Model Size=16B, Distill Type=ExpertCondenser, Fine-tuning with Stanford-S1=true2026.04 | 2.5 | |
| Qwen1.5-MoE (DenseMixer)Model Size=14B, Distill Type=DenseMixer, Fine-tuning with Stanford-S1=true2026.04 | 1.7 | |
| Qwen1.5-MoE (ESFT)Model Size=14B, Distill Type=ESFT, Fine-tuning with Stanford-S1=true2026.04 | 0.8 | |
| Qwen1.5-MoE (SFT)Model Size=14B, Distill Type=SFT, Fine-tuning with Stanford-S1=true2026.04 | 0 |