Code Generation on LiveCodeBench (Score)
73.1ScoreDeepseek-R1
Evaluation Results
| Method | Links | |
|---|---|---|
| Deepseek-R1Cost=12.3272026.05 | 73.1 | |
| SFT-based Classification RouterRouting Strategy=Auto-routing, Cost=8.6672026.05 | 70.5 | |
| MoMA RouterRouting Strategy=Performance-priority, Cost=10.042026.05 | 66.5 | |
| Qwen3-235B-A22BCost=14.652026.05 | 65.9 | |
| Contrastive learning based RouterRouting Strategy=Performance-priority, Cost=12.4982026.05 | 61.3 | |
| Qwen3-32BCost=14.652026.05 | 60.7 | |
| DreamReasoner-8BBlock Size=82026.06 | 53.9 | |
| DreamReasoner-8BBlock Size=162026.06 | 53.6 | |
| Qwen3-8B-Thinking2026.06 | 52.8 | |
| AceReason-Nemotron-1.1-7B2026.06 | 52.1 | |
| DreamReasoner-8BBlock Size=42026.06 | 51.3 | |
| MiMo-7B-RL2026.06 | 50.7 | |
| DreamReasoner-8BBlock Size=322026.06 | 50.4 | |
| MoMA RouterRouting Strategy=Auto-routing, Cost=6.3062026.05 | 45.3 | |
| LLaDA-2.0-Flash (100B-A6B)Block Size=322026.06 | 41 | |
| Contrastive learning based RouterRouting Strategy=Auto-routing, Cost=6.9402026.05 | 40.1 | |
| DeepSeek-R1-Distill-Qwen-7B2026.06 | 37.8 | |
| SDAR-30B-A3B-SciBlock Size=42026.06 | 29 | |
| Contrastive learning based RouterRouting Strategy=Cost-priority, Cost=1.6672026.05 | 27.6 | |
| Deepseek-V3Cost=9.4982026.05 | 27.2 | |
| JT-Code-8BCost=1.6672026.05 | 26.3 | |
| MoMA RouterRouting Strategy=Cost-priority, Cost=1.3572026.05 | 24.6 | |
| SDAR-8B-ChatBlock Size=42026.06 | 16.4 | |
| TraDo-8B-InstructBlock Size=42026.06 | 16.1 | |
| TraDo-8B-ThinkingBlock Size=82026.06 | 14.2 | |
| Dream-7B-Instruct2026.06 | 9.7 | |
| SDAR-30B-A3B-SciBlock Size=82026.06 | 6.4 | |
| SDAR-30B-A3B-SciBlock Size=162026.06 | 3.2 | |
| SDAR-30B-A3B-SciBlock Size=322026.06 | 2.6 |