Code Generation on MBPP+ (Score)
94.2ScoreGLM-4.6
Evaluation Results
| Method | Links | |
|---|---|---|
| GLM-4.6Parameters=357B-A32B, Non-Thinking Mode=true2026.01 | 94.2 | |
| DeepSeek-V3.1Parameters=685B-A37B, Non-Thinking Mode=true2026.01 | 92.6 | |
| Trinity Large Base2026.02 | 88.62 | |
| A.X K1Parameters=519B-A33B, Non-Thinking Mode=true2026.01 | 85.7 | |
| Qwen3-Coder-30B-A3B-Instruct (Base)Model=Qwen3-Coder-30B-A3B-Instruct, Strategy=Base2026.02 | 75.13 | |
| DeepSeek V3.1Model Variant=Base, # Shots=0-shot, # Activated Params=37B, # Total Params=671B2026.02 | 72.2 | |
| MiMo-V2 FlashModel Variant=Base, # Shots=0-shot, # Activated Params=15B, # Total Params=309B2026.02 | 71.4 | |
| Step 3.5 FlashModel Variant=Base, # Shots=0-shot, # Activated Params=11B, # Total Params=196B2026.02 | 70.6 | |
| Qwen 3 32BModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=false2025.12 | 70.6 | |
| DS-R1 32BModel Family=DeepSeek-R1, Parameter Count=32B, Thinking Capability=true2025.12 | 70.1 | |
| DeepSeek V3.2Model Variant=Exp Base, # Shots=0-shot, # Activated Params=37B, # Total Params=671B2026.02 | 69.8 | |
| Qwen 3 8B2025.12 | 69.1 | |
| Olmo 3.1 Think 32BTraining Stage=Final Think 3.1, Model Family=Olmo 3.1, Parameter Count=32B, Thinking Capability=true2025.12 | 68.3 | |
| Olmo 3 Think (Final 3.0)Training Stage=Final Think 3.0, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 68 | |
| Olmo 3 Think (DPO)Training Stage=DPO, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 67.2 | |
| Olmo 3 Think (SFT)Training Stage=SFT, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 66.7 | |
| Qwen 3 VL 8B Inststage=Instruct2025.12 | 66.3 | |
| Qwen 3 VL 32B ThinkModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=true2025.12 | 66.2 | |
| Nemotron Nano 9B v22025.12 | 66.1 | |
| K2-V2 70B InstructModel Family=K2, Parameter Count=70B, Thinking Capability=false2025.12 | 66 | |
| Olmo 3 7B ThinkStage=Final Think2025.12 | 64.7 | |
| Qwen 3 8Bstage=Instruct2025.12 | 64.4 | |
| DS-R1 Qwen 7B2025.12 | 63.5 | |
| Olmo 3 7B ThinkStage=SFT2025.12 | 63.2 | |
| Olmo 3 7B ThinkStage=DPO2025.12 | 63 | |
| Qwen 3 VL 8B Think2025.12 | 63 | |
| Qwen 2.5 7Bstage=Instruct2025.12 | 62.6 | |
| OpenThinker3 7B2025.12 | 61.4 | |
| OR Nemotron 7B2025.12 | 61.2 | |
| Olmo 3 7B Instructstage=Final Instruct2025.12 | 60.2 | |
| Qwen3-4B-Instruct-2507 (C/E Weighted)Model=Qwen3-4B-Instruct-2507, Strategy=C/E Weighted2026.02 | 58.73 | |
| Qwen3-4B-Instruct-2507 (Base)Model=Qwen3-4B-Instruct-2507, Strategy=Base2026.02 | 56.61 | |
| Qwen3-4B-Instruct-2507 (Edge Only)Model=Qwen3-4B-Instruct-2507, Strategy=Edge Only2026.02 | 56.61 | |
| Olmo 3 7B Instructstage=SFT2025.12 | 56.5 | |
| Qwen3-4B-Instruct-2507 (Basic Only)Model=Qwen3-4B-Instruct-2507, Strategy=Basic Only2026.02 | 56.34 | |
| Olmo 3 7B Instructstage=DPO2025.12 | 55.9 | |
| Granite 3.3 8B Inststage=Instruct2025.12 | 54 | |
| Qwen3-4B-Instruct-2507 (Uniform)Model=Qwen3-4B-Instruct-2507, Strategy=Uniform2026.02 | 53.7 | |
| Qwen3-4B-Instruct-2507 (B/I Weighted)Model=Qwen3-4B-Instruct-2507, Strategy=B/I Weighted2026.02 | 52.38 | |
| Qwen3-4B-Instruct-2507 (Complex Only)Model=Qwen3-4B-Instruct-2507, Strategy=Complex Only2026.02 | 51.85 | |
| Apertus 8B Inststage=Instruct2025.12 | 42.1 | |
| OLMo 2 7B Inststage=Instruct2025.12 | 40.7 | |
| Qwen3-4B-Instruct-2507 (C/E Weighted (Rev))Model=Qwen3-4B-Instruct-2507, Strategy=C/E Weighted (Rev)2026.02 | 35.98 |