Code Generation on MBPP+ (pass@1)
84.39Pass@1Gemini-2.5-Pro
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-2.5-ProInstitution=Google2026.03 | 84.39 | — | |
| Claude-Opus-4.5Model Category=Closed-APIs2026.03 | 83.9 | — | |
| GLM 4.6Evaluation Mode=Chat2025.12 | 83.6 | — | |
| Kimi-K2-ThinkingModel Scale=20B+2026.03 | 82.3 | — | |
| Claude-Sonnet-4.5Model Category=Closed-APIs2026.03 | 82.3 | — | |
| Qwen3-235B-A22B-Thinking-2507Model Scale=20B+2026.03 | 81.5 | — | |
| Qwen3-Coder-480B-A35B-InstructModel Scale=20B+2026.03 | 80.2 | — | |
| DeepSeek V3.2Evaluation Mode=Chat2025.12 | 79.9 | — | |
| LongCat-Flash ChatEvaluation Mode=Chat2025.12 | 79.6 | — | |
| LongCat-Flash Exp-ChatEvaluation Mode=Chat2025.12 | 79.1 | — | |
| GPT-5.1Institution=OpenAI2026.03 | 79.1 | — | |
| GPT-4.1Institution=OpenAI2026.03 | 79.1 | — | |
| ReflexiCoder-8B (Multiple)Institution=Ours, setup=full iterative reasoning-reflection setup2026.03 | 79.1 | — | |
| Gemini-3-Flash-previewModel Category=Closed-APIs2026.03 | 79.1 | — | |
| ExpGraphBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 79 | — | |
| ReflexiCoder-8B (Single)Institution=Ours, setup=single-attempt without system prompt2026.03 | 78.57 | — | |
| Qwen2.5-Coder-32B-InstructModel Scale=20B+2026.03 | 77.8 | — | |
| Qwen3-235B-A22B-Instruct-2507Model Scale=20B+2026.03 | 77.8 | — | |
| IQuest-Coder-V1-40B-InstructModel Scale=20B+2026.03 | 77.8 | — | |
| Qwen2.5-Coder-14B-InstructModel Scale=13B+2026.03 | 77.2 | — | |
| Qwen3-Coder-30B-A3B-InstructModel Scale=13B+2026.03 | 77.2 | — | |
| Deepseek-V3.2Model Scale=20B+2026.03 | 77.2 | — | |
| IQuest-Coder-V1-40B-Loop-InstructModel Scale=20B+2026.03 | 77.2 | — | |
| KAT-DevModel Scale=20B+2026.03 | 76.2 | — | |
| IQuest-Coder-V1-40B-Loop-ThinkingModel Scale=20B+2026.03 | 76.2 | — | |
| UCoder-32BSize Category=32B+ LLMs2025.12 | 75.7 | — | |
| GLM-4.7Model Scale=20B+2026.03 | 75.7 | — | |
| Claude-Sonnet-4.5Institution=Anthropic2026.03 | 75.4 | — | |
| DS-Coder-V2-InstructSize Category=32B+ LLMs2025.12 | 75.1 | — | |
| Qwen2.5-Coder-32B-InstructSize Category=32B+ LLMs2025.12 | 75.1 | — | |
| IQuest-Coder-V1-40B-ThinkingModel Scale=20B+2026.03 | 75.1 | — | |
| Claude-3.5-Sonnet-20241022Size Category=Closed-APIs2025.12 | 74.6 | — | |
| UCoder-14BSize Category=13B+ LLMs2025.12 | 74.3 | — | |
| Kimi-K2-Instruct-0905Model Scale=20B+2026.03 | 74.1 | — | |
| Kimi-K2 Base# Shots=3-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 73.8 | — | |
| Seed-Coder-8B-InstructModel Scale=6B+2026.03 | 73.3 | — | |
| Qwen2.5-Coder-14B-InstructSize Category=13B+ LLMs2025.12 | 72.8 | — | |
| GPT-4o-2024-08-06Size Category=Closed-APIs2025.12 | 72.5 | — | |
| Seed-Coder-8B-InstructInstitution=ByteDance2026.03 | 72.49 | — | |
| DeepSeek-V3.1 Base# Shots=3-shot, # Activated Params=37B, # Total Params=671B2026.01 | 72.2 | — | |
| UCoder-7BSize Category=6B+ LLMs2025.12 | 72.2 | — | |
| Qwen2.5-Coder-7B-InstructModel Scale=6B+2026.03 | 72.2 | — | |
| GPT-5.1Model Category=Closed-APIs2026.03 | 72.2 | — | |
| IQuest-Coder-V1-14B-ThinkingModel Scale=13B+2026.03 | 72 | — | |
| Qwen2.5-Coder-7B-InstructSize Category=6B+ LLMs2025.12 | 71.7 | — | |
| AGRPOReinforcement Learning Strategy=+ AGRPO2026.03 | 71.7 | — | |
| GASP + Real-data RLBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 71.6 | — | |
| DreamModel size category=Reference Models (≥7B), Number of parameters=7B2025.09 | 71.5 | 72.5 | |
| MiMo-V2-Flash Base# Shots=3-shot, # Activated Params=15B, # Total Params=309B2026.01 | 71.4 | — | |
| LFPO (All Loss)Reinforcement Learning Strategy=+ LFPO (All Loss)2026.03 | 71.3 | — | |
| GPT-4-TurboParams=-, Base Model=-2024.05 | 70.7 | — | |
| GASPBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 70.63 | — | |
| DeepSeek-Coder-V2-Lite-InstructModel Scale=6B+2026.03 | 70.6 | — | |
| DS-Coder-V2-Lite-InstructSize Category=13B+ LLMs2025.12 | 70.4 | — | |
| OpenCoderModel Category=AR Coding Models2025.10 | 70.4 | — | |
| Qwen2.5-Coder-7B-InstRole=Teacher Model2026.04 | 70.4 | — | |
| Qwen3-8BInstitution=Alibaba2026.03 | 70.37 | — | |
| DS-Coder-33B-InstructSize Category=32B+ LLMs2025.12 | 70.1 | — | |
| Qwen2.5-Coder-7B-InstructInstitution=Alibaba2026.03 | 69.84 | — | |
| DeepSeek-R1-Distill-Qwen-32BModel=DeepSeek-R1-Distill-Qwen-32B, #Bits=W16A162026.03 | 69.84 | — | |
| DeepSeek-V3.2 Exp Base# Shots=3-shot, # Activated Params=37B, # Total Params=671B2026.01 | 69.8 | — | |
| AZRBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 69.58 | — | |
| Qwen2.5-Coder-7BDecoding=Greedy2026.03 | 69.58 | — | |
| GPT-3.5-TurboParams=-, Base Model=-2024.05 | 69.4 | — | |
| Real-data RLBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 69.31 | — | |
| KAT-Dev-72B-ExpModel Scale=20B+2026.03 | 69.3 | — | |
| SPGReinforcement Learning Strategy=+ SPG2026.03 | 69.2 | — | |
| SliderQuantModel=DeepSeek-R1-Distill-Qwen-32B, #Bits=W4A162026.03 | 69.05 | — | |
| OpenCoder-8B-InstructSize Category=6B+ LLMs2025.12 | 69 | — | |
| UniGRPOReinforcement Learning Strategy=+ UniGRPO2026.03 | 68.8 | — | |
| Kimi-Dev-72BModel Scale=20B+2026.03 | 68.8 | — | |
| LFPO (Neg. Only)Reinforcement Learning Strategy=+ LFPO (Neg. Only)2026.03 | 68.5 | — | |
| IQuest-Coder-V1-14B-InstructModel Scale=13B+2026.03 | 68.5 | — | |
| Qwen2.5-Coder-32BSize Category=32B+ LLMs2025.12 | 68.2 | — | |
| LFPO (Pos. Only)Reinforcement Learning Strategy=+ LFPO (Pos. Only)2026.03 | 68.2 | — | |
| Qwen-7BFT data=-, LoRA=-2026.03 | 68 | — | |
| coupled-GRPOReinforcement Learning Strategy=+ coupled-GRPO2026.03 | 67.5 | — | |
| SecCoderXBackbone=Qwen2.5-Coder-7B2026.02 | 67.49 | 83.86 | |
| DeepCoder-14B-PreviewInstitution=rLLM2026.03 | 67.46 | — | |
| diffu-GRPOReinforcement Learning Strategy=+ diffu-GRPO2026.03 | 67.3 | — | |
| Qwen2.5-Coder-14BSize Category=13B+ LLMs2025.12 | 66.7 | — | |
| Ouro 2.6BModel Category=Looped Latent Reasoning Models2025.10 | 66.6 | — | |
| SWaRLWatermarking Status=Watermarked2026.01 | 66.2 | 83.03 | |
| OmniQuantModel=DeepSeek-R1-Distill-Qwen-32B, #Bits=W4A162026.03 | 65.61 | — | |
| DS-Coder-6.7B-InstructSize Category=6B+ LLMs2025.12 | 65.6 | — | |
| EXP-editWatermarking Status=Watermarked2026.01 | 65.38 | 82.9 | |
| Base+SFTWatermarking Status=No Watermark2026.01 | 65.26 | 84.18 | |
| Starcoder2-15B-Instruct-v0.1Size Category=13B+ LLMs2025.12 | 65.1 | — | |
| BaseBackbone=Qwen2.5-Coder-7B2026.02 | 64.95 | 84.66 | |
| PrivCodeBackbone=Qwen2.5-Coder-7B, Privacy Epsilon=42025.12 | 64.8 | — | |
| Gemini-3-Pro-previewModel Category=Closed-APIs2026.03 | 64.8 | — | |
| ProSecBackbone=Qwen2.5-Coder-7B, Reproduced=true2026.02 | 64.1 | 76.19 | |
| DeepSeek-Coder-7B-InstructInstitution=DeepSeek2026.03 | 63.76 | — | |
| IQuest-Coder-V1-7B-InstructModel Scale=6B+2026.03 | 63.5 | — | |
| IQuest-Coder-V1-7B-ThinkingModel Scale=6B+2026.03 | 63.5 | — | |
| GKDTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 63.2 | — | |
| DistillLM-2Teacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 63.2 | — | |
| Qwen2.5-Coder-7BSize Category=6B+ LLMs2025.12 | 62.9 | — | |
| DistilLLMTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 62.7 | — | |
| VCRDTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 62.7 | — |