Coding on MBPP (pass@1 accuracy)
95.33Pass@1 AccuracySwiR
Evaluation Results
| Method | Links | |
|---|---|---|
| SwiRBackbone=Qwen3-8B2025.10 | 95.33 | |
| CoTBackbone=Qwen3-8B2025.10 | 94.16 | |
| Soft ThinkingBackbone=Qwen3-8B2025.10 | 94.16 | |
| Claude-Sonnet-4.5Category=Candidate Models2026.03 | 93.3 | |
| CoT (Greedy)Backbone=Qwen3-8B, decoding=Greedy2025.10 | 91.44 | |
| FineRouterCategory=Routers2026.03 | 90 | |
| Llama-4-MaverickCategory=Candidate Models2026.03 | 86.7 | |
| Claude-Haiku-4.5Category=Candidate Models2026.03 | 86.7 | |
| DeepSeek-R1Category=Candidate Models2026.03 | 86.7 | |
| Qwen3-235B-A22BCategory=Candidate Models2026.03 | 83.3 | |
| IPRCategory=Routers2026.03 | 83.3 | |
| MT (Teacher)Teacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 82.3 | |
| GPT-OSS-120BCategory=Candidate Models2026.03 | 80 | |
| MLPCategory=Routers2026.03 | 80 | |
| GraphRouterCategory=Routers2026.03 | 80 | |
| CoT2-MetaStrategy=Ours (CoT2-Meta), Inference Budget=C=162026.03 | 79.5 | |
| ReST-MCTS*Strategy=ReST-MCTS*, Inference Budget=C=162026.03 | 76.8 | |
| HPDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 75.4 | |
| Self-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 75.4 | |
| Vanilla ToTStrategy=Vanilla ToT, Inference Budget=C=162026.03 | 75.1 | |
| MT (Teacher)Teacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 74.9 | |
| RKLDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 74.9 | |
| JSDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 74.6 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 74.3 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 74.3 | |
| BaseBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Pre-trained2026.05 | 74.3 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 74.1 | |
| DeepSeek-v3Category=Candidate Models2026.03 | 73.3 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 73.3 | |
| Best-of-16Strategy=Best-of-16, Inference Budget=C=162026.03 | 72.5 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 71.4 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 1: iGSM2026.05 | 71.4 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 70.4 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 70.4 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 1: iGSM2026.05 | 70.4 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 3: iGSM -> MedCalc -> IFEval2026.05 | 70.4 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 69.3 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 69.3 | |
| Greedy CoTStrategy=Greedy CoT, Inference Budget=C=162026.03 | 68.5 | |
| MS (Student)Teacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 68.5 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 68 | |
| SFTTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 67.7 | |
| Anchored LearningBackbone=Llama-3.2-3B-Instruct2026.05 | 67.7 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 2: iGSM -> MedCalc2026.05 | 67.7 | |
| KDTeacher-Student Model Configuration=Qwen-Coder (7B -> 1.5B)2026.04 | 67.5 | |
| kNNCategory=Routers2026.03 | 66.7 | |
| Self-sftBackbone=Qwen2.5-3B-Instruct2026.05 | 66.1 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 3: iGSM -> MedCalc -> IFEval2026.05 | 65.9 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 2: iGSM -> MedCalc2026.05 | 64.3 | |
| KDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 64 | |
| DFTBackbone=Qwen2.5-3B-Instruct2026.05 | 63.5 | |
| HPDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 63.2 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.05 | 62.7 | |
| SFTTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 61.9 | |
| Iter-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 61.9 | |
| RKLDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 61.6 | |
| MS (Student)Teacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 61.1 | |
| JSDTeacher-Student Model Configuration=DS-Coder (6.7B -> 1.3B)2026.04 | 61.1 | |
| Low-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 61.1 | |
| Self-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 60.1 | |
| STMBackbone=Llama-3.2-3B-Instruct2026.05 | 57.9 | |
| DFTBackbone=Llama-3.2-3B-Instruct2026.05 | 57.4 | |
| Xuanwu VL-2BNumber of Parameters=2B, Input Modality=Text-Only2026.03 | 50.6 | |
| InternVL 3.5 2BNumber of Parameters=2B, Input Modality=Text-Only2026.03 | 49.8 | |
| DFTBackbone=Qwen2.5-3B-Instruct2026.05 | 48.1 | |
| RouterDCCategory=Routers2026.03 | 46.7 | |
| InternVL 3.0 2BNumber of Parameters=2B, Input Modality=Text-Only2026.03 | 45.6 | |
| RouteLLMCategory=Routers2026.03 | 43.3 | |
| Llama-3.3-70BCategory=Candidate Models2026.03 | 40 | |
| Qwen3-32BCategory=Candidate Models2026.03 | 33.3 | |
| Mistral-SmallCategory=Candidate Models2026.03 | 26.7 | |
| Mistral-LargeCategory=Candidate Models2026.03 | 23.3 | |
| KL-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 18.5 | |
| KL-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 14.3 | |
| KL-SFT 0.2Backbone=Llama-3.2-3B-Instruct2026.05 | 10.1 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 2.6 | |
| SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 1.9 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 1.2 |