Coding on MBPP+
97.88Pass@1K-EXAONE-236B-A23B
Evaluation Results
| Method | Links | |
|---|---|---|
| K-EXAONE-236B-A23BReasoning=On2026.03 | 97.88 | |
| Qwen-3-30B-A3BReasoning=On2026.03 | 90.48 | |
| HyperCLOVAX-SEED-Think-32BReasoning=Off2026.03 | 89.9 | |
| Mi:dm K 2.5 Pro (March ‘26)Reasoning=On2026.03 | 89.68 | |
| Gemini 3-Pro2026.02 | 86.21 | |
| GPT-5 (High)2026.02 | 83.1 | |
| HyperCLOVAX-SEED-Think-32BReasoning=On2026.03 | 83.07 | |
| ERNIE 5.02026.02 | 82.54 | |
| Llama-3.1-8B-instParameters=8B, Type=Instruction-tuned2026.01 | 81.8 | |
| Youtu-LLMSize=2B, Type=Base, Prompt Setting=3-shot2025.12 | 81.8 | |
| DS V3.2-Thinking2026.02 | 81.48 | |
| Qwen3Size=4B, Type=Base, Prompt Setting=3-shot2025.12 | 80.8 | |
| RDPOPass@k protocol=@12026.05 | 79.63 | |
| Exaone-3.5-7.8B-instParameters=7.8B, Type=Instruction-tuned2026.01 | 79.4 | |
| RL initialization modelPass@k protocol=@12026.05 | 78.84 | |
| GRPOPass@k protocol=@12026.05 | 78.57 | |
| Qwen-3-30B-A3BReasoning=Off2026.03 | 77.7 | |
| Mi:dm 2.0 Base-instType=Instruction-tuned2026.01 | 77.5 | |
| Mi:dm K 2.5 Pro (March ‘26)Reasoning=Off2026.03 | 76.7 | |
| P-GRPO (Code+RM)Size=7B-Inst, Training Setting=Training-based2026.05 | 76.2 | |
| Solar-Open-100BReasoning=On2026.03 | 75.13 | |
| GRPO (Code)Size=7B-Inst, Training Setting=Training-based2026.05 | 75.1 | |
| HASP-Evolve + RSSize=7B-Inst, Training Setting=Training-based, Backbone=Qwen2.5-7B-Instruct2026.05 | 74 | |
| Gemini 2.5-Pro2026.02 | 73.8 | |
| Qwen3-14BParameters=14B2026.01 | 73.4 | |
| HASP-Intervention (w. Teacher)Size=7B-Inst, Training Setting=Training-free / inference-time, Backbone=Qwen2.5-7B-Instruct2026.05 | 73 | |
| KodCode-RL Step 256Size=7B-Inst, Training Setting=Training-based2026.05 | 72.8 | |
| AceCoderRMSize=7B-Inst, Training Setting=Training-based2026.05 | 71.2 | |
| Qwen3Size=1.7B, Type=Base, Prompt Setting=3-shot2025.12 | 71 | |
| GPT-4oSize=~200B, Training Setting=Training-free / inference-time2026.05 | 71 | |
| KodCode-RL Step 128Size=7B-Inst, Training Setting=Training-based2026.05 | 70.9 | |
| Qwen 3 VL 32B InstructParameters=32B2025.12 | 69 | |
| AceCoderRuleSize=7B-Inst, Training Setting=Training-based2026.05 | 68.3 | |
| Qwen 3 32BThinking=No, Parameters=32B2025.12 | 67.9 | |
| Qwen2.5-7B-InstructSize=7B-Inst, Training Setting=Training-free / inference-time2026.05 | 67.7 | |
| GPT-4o-miniSize=~8B, Training Setting=Training-free / inference-time2026.05 | 67 | |
| K-EXAONE-236B-A23BReasoning=Off2026.03 | 66.9 | |
| Qwen 2.5 32BParameters=32B2025.12 | 66.6 | |
| SmolLM3Size=3B, Type=Base, Prompt Setting=3-shot2025.12 | 66.1 | |
| Prompt-Only SkillsSize=7B-Inst, Training Setting=Training-free / inference-time2026.05 | 66 | |
| Gemma 3 27BParameters=27B2025.12 | 65.7 | |
| Self-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 65.6 | |
| Olmo 3.1 32B InstructStage=Final Instruct 3.12025.12 | 65.1 | |
| HASP-Intervention (PF-only)Size=7B-Inst, Training Setting=Training-free / inference-time, Backbone=Qwen2.5-7B-Instruct2026.05 | 65 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 63.8 | |
| Olmo 3.1 32B InstructStage=DPO2025.12 | 63.6 | |
| DFTBackbone=Qwen2.5-3B-Instruct2026.05 | 63 | |
| SFT (vanilla)Size=7B-Inst, Training Setting=Training-based2026.05 | 63 | |
| Llama3.1Size=8B, Type=Base, Prompt Setting=3-shot2025.12 | 62.7 | |
| Qwen3-4BParameters=4B2026.01 | 62.4 | |
| Qwen3-8BParameters=8B2026.01 | 62.2 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 62.2 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 62.2 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 62.2 | |
| BaseBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Pre-trained2026.05 | 62.2 | |
| RA-Agent (multi-loop)Size=7B-Inst, Training Setting=Training-free / inference-time2026.05 | 62 | |
| Gemma3Size=4B, Type=Base, Prompt Setting=3-shot2025.12 | 61.9 | |
| Olmo 3.1 32B InstructStage=SFT2025.12 | 61.5 | |
| Gemma 2 27BParameters=27B2025.12 | 61.2 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 61.1 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 61.1 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 61.1 | |
| Mi:dm 2.0 Mini-instType=Instruction-tuned2026.01 | 60.9 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 3: iGSM -> MedCalc -> IFEval2026.05 | 60.3 | |
| OLMo3-7B-InstructParameters=7B, Type=Instruct2026.01 | 60.2 | |
| Exaone-3.5-2.4B-instParameters=2.4B, Type=Instruction-tuned2026.01 | 59.8 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 59.8 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 59.8 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 59.8 | |
| Qwen3-4BParameters=4B2026.01 | 59.5 | |
| KL-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 59.5 | |
| Self-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 59.5 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 59.5 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 59.3 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 1: iGSM2026.05 | 59.3 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 59 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 1: iGSM2026.05 | 59 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 58.2 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 3: iGSM -> MedCalc -> IFEval2026.05 | 57.7 | |
| Molmo2-8BParameters=8B2026.01 | 57.5 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 56.9 | |
| Molmo2-4BParameters=4B2026.01 | 56.2 | |
| Molmo2-O-7BParameters=7B2026.01 | 55.7 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 55.6 | |
| Self-sftBackbone=Qwen2.5-3B-Instruct2026.05 | 55.6 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 2: iGSM -> MedCalc2026.05 | 55.3 | |
| DFTBackbone=Qwen2.5-3B-Instruct2026.05 | 55 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct, Incremental Learning Stage=Stage 2: iGSM -> MedCalc2026.05 | 54.5 | |
| Iter-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 53.5 | |
| UnSafeChainbase_model=DeepSeek-R1-Distill-Qwen-7B2026.03 | 53.2 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.05 | 51.6 | |
| SafeChainbase_model=DeepSeek-R1-Distill-Qwen-7B2026.03 | 50.8 | |
| Low-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 50 | |
| Self-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 50 | |
| OLMo 2 32BParameters=32B2025.12 | 49 | |
| PreSafebase_model=DeepSeek-R1-Distill-Qwen-7B2026.03 | 48.4 | |
| STMBackbone=Llama-3.2-3B-Instruct2026.05 | 48.4 | |
| DFTBackbone=Llama-3.2-3B-Instruct2026.05 | 48.1 | |
| Apertus 70BParameters=70B2025.12 | 45.8 | |
| Anchored LearningBackbone=Llama-3.2-3B-Instruct2026.05 | 45.5 |