Math Reasoning on GPQA Diamond (Pass@4)
65.8Pass@4Qwen3 (ExpertCondenser)
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3 (ExpertCondenser)Model Size=30B, Distill Type=ExpertCondenser, Fine-tuning with Stanford-S1=true2026.04 | 65.8 | |
| Qwen3 (DenseMixer)Model Size=30B, Distill Type=DenseMixer, Fine-tuning with Stanford-S1=true2026.04 | 61 | |
| Qwen3 (SFT)Model Size=30B, Distill Type=SFT, Fine-tuning with Stanford-S1=true2026.04 | 58.6 | |
| Qwen3 (ESFT)Model Size=30B, Distill Type=ESFT, Fine-tuning with Stanford-S1=true2026.04 | 52.7 | |
| DeepSeek-Coder-V2-Lite (ExpertCondenser)Model Size=16B, Distill Type=ExpertCondenser, Fine-tuning with Stanford-S1=true2026.04 | 38.7 | |
| DeepSeek-Coder-V2-Lite (DenseMixer)Model Size=16B, Distill Type=DenseMixer, Fine-tuning with Stanford-S1=true2026.04 | 34.8 | |
| Qwen1.5-MoE (ExpertCondenser)Model Size=14B, Distill Type=ExpertCondenser, Fine-tuning with Stanford-S1=true2026.04 | 34.6 | |
| DeepSeek-Coder-V2-Lite (SFT)Model Size=16B, Distill Type=SFT, Fine-tuning with Stanford-S1=true2026.04 | 34.2 | |
| DeepSeek-Coder-V2-Lite (ESFT)Model Size=16B, Distill Type=ESFT, Fine-tuning with Stanford-S1=true2026.04 | 32.2 | |
| Qwen1.5-MoE (DenseMixer)Model Size=14B, Distill Type=DenseMixer, Fine-tuning with Stanford-S1=true2026.04 | 31.6 | |
| Qwen1.5-MoE (SFT)Model Size=14B, Distill Type=SFT, Fine-tuning with Stanford-S1=true2026.04 | 27.8 | |
| Qwen1.5-MoE (ESFT)Model Size=14B, Distill Type=ESFT, Fine-tuning with Stanford-S1=true2026.04 | 26.4 |