Code Generation on HumanEval+
100Pass@1Gemini-3-Pro-preview
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini-3-Pro-previewModel Category=Closed-APIs2026.03 | 100 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-235B-A22B-Thinking-2507Model Scale=20B+2026.03 | 98.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-Opus-4.5Model Category=Closed-APIs2026.03 | 98.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-Sonnet-4.5Model Category=Closed-APIs2026.03 | 98.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2-ThinkingModel Scale=20B+2026.03 | 98.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-Coder-480B-A35B-InstructModel Scale=20B+2026.03 | 97.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-Loop-InstructModel Scale=20B+2026.03 | 97.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-Loop-ThinkingModel Scale=20B+2026.03 | 97.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-5.1Model Category=Closed-APIs2026.03 | 97 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-235B-A22B-Instruct-2507Model Scale=20B+2026.03 | 96.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-InstructModel Scale=20B+2026.03 | 96.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2-Instruct-0905Model Scale=20B+2026.03 | 94.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-3-Pro-previewModel Category=Closed-APIs2026.03 | 94.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-Coder-30B-A3B-InstructModel Scale=13B+2026.03 | 93.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Deepseek-V3.2Model Scale=20B+2026.03 | 93.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-ThinkingModel Scale=20B+2026.03 | 93.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-32B-InstructModel Scale=20B+2026.03 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-Dev-72BModel Scale=20B+2026.03 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-235B-A22B-Thinking-2507Model Scale=20B+2026.03 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-Opus-4.5Model Category=Closed-APIs2026.03 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-Sonnet-4.5Model Category=Closed-APIs2026.03 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLM 4.6Evaluation Mode=Chat2025.12 | 92.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-14B-ThinkingModel Scale=13B+2026.03 | 92.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-Coder-480B-A35B-InstructModel Scale=20B+2026.03 | 92.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2-ThinkingModel Scale=20B+2026.03 | 92.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-2.5-ProInstitution=Google2026.03 | 92.07 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-235B-A22B-Instruct-2507Model Scale=20B+2026.03 | 91.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-Loop-InstructModel Scale=20B+2026.03 | 91.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KAT-DevModel Scale=20B+2026.03 | 90.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-InstructModel Scale=20B+2026.03 | 90.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-5.1Model Category=Closed-APIs2026.03 | 90 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2-Instruct-0905Model Scale=20B+2026.03 | 89.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-Loop-ThinkingModel Scale=20B+2026.03 | 89.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek V3.2Evaluation Mode=Chat2025.12 | 89 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LongCat-Flash ChatEvaluation Mode=Chat2025.12 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KAT-Dev-72B-ExpModel Scale=20B+2026.03 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-3-Flash-previewModel Category=Closed-APIs2026.03 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Deepseek-V3.2Model Scale=20B+2026.03 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReflexiCoder-8B (Multiple)Institution=Ours, setup=full iterative reasoning-reflection setup2026.03 | 87.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-40B-ThinkingModel Scale=20B+2026.03 | 87.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LongCat-Flash Exp-ChatEvaluation Mode=Chat2025.12 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-14B-InstructSize Category=13B+ LLMs2025.12 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-32B-InstructSize Category=32B+ LLMs2025.12 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-5.1Institution=OpenAI2026.03 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReflexiCoder-8B (Single)Institution=Ours, setup=single-attempt without system prompt2026.03 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-InstructModel Scale=6B+2026.03 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLM-4.7Model Scale=20B+2026.03 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-Coder-30B-A3B-InstructModel Scale=13B+2026.03 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-32B-InstructModel Scale=20B+2026.03 | 86.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KAT-DevModel Scale=20B+2026.03 | 86.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o2024.07 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o-2024-08-06Size Category=Closed-APIs2025.12 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-3.5-Sonnet-20241022Size Category=Closed-APIs2025.12 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-14B-ThinkingModel Scale=13B+2026.03 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-Dev-72BModel Scale=20B+2026.03 | 86 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oStrategy=Direct2024.10 | 84.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2 Base# Shots=1-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 84.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-3-Flash-previewModel Category=Closed-APIs2026.03 | 84.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LaDiRModel Category=Reasoning Methods2025.10 | 84.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-InstructSize Category=6B+ LLMs2025.12 | 84.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AZRBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 83.54 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-14B-InstructModel Scale=13B+2026.03 | 83.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oStrategy=MBTI2024.10 | 82.9 | — | -1.9 | — | — | — | — | — | — | — | — | — | — | — | |
| UCoder-32BSize Category=32B+ LLMs2025.12 | 82.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SecCoderXBackbone=Qwen2.5-Coder-7B2026.02 | 82.74 | 90.85 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama 3 405BParameters=405B2024.07 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude 3.5 Sonnet2024.07 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o miniStrategy=MBTI2024.10 | 82.3 | — | 1.8 | — | — | — | — | — | — | — | — | — | — | — | |
| DS-Coder-V2-InstructSize Category=32B+ LLMs2025.12 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-32BModel=DeepSeek-R1-Distill-Qwen-32B, #Bits=W16A162026.03 | 81.71 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4-TurboParams=-, Base Model=-2024.05 | 81.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-InstructModel Scale=6B+2026.03 | 81.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KAT-Dev-72B-ExpModel Scale=20B+2026.03 | 81.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UCoder-14BSize Category=13B+ LLMs2025.12 | 81.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek-Coder-V2-Lite-InstructModel Scale=6B+2026.03 | 81.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seed-Coder-8B-InstructModel Scale=6B+2026.03 | 81.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o miniStrategy=Direct2024.10 | 80.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek-Coder V2Strategy=Direct2024.10 | 80.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-8BInstitution=Alibaba2026.03 | 80.49 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seed-Coder-8B-InstructInstitution=ByteDance2026.03 | 80.49 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GASP + Real-data RLBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 80.49 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SliderQuantModel=DeepSeek-R1-Distill-Qwen-32B, #Bits=W4A162026.03 | 80.49 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-7B-InstructModel Scale=6B+2026.03 | 79.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLM-4.7Model Scale=20B+2026.03 | 79.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BaseBackbone=Qwen2.5-Coder-7B2026.02 | 79.88 | 92.68 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-InstructInstitution=Alibaba2026.03 | 79.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7BDecoding=Greedy2026.03 | 79.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GASPBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 79.67 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TaH+Model Category=Reasoning Methods2025.10 | 79.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4.1Institution=OpenAI2026.03 | 78.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-LongStrategy=MBTI2024.10 | 78.7 | — | 1.9 | — | — | — | — | — | — | — | — | — | — | — | |
| OpenCoder-8B-InstructSize Category=6B+ LLMs2025.12 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IQuest-Coder-V1-14B-InstructModel Scale=13B+2026.03 | 78.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Real-data RLBase Model=Qwen2.5-Coder-7B, Decoding=Greedy2026.03 | 78.66 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SecCoderXBackbone=Qwen2.5-Coder-3B2026.02 | 77.74 | 92.68 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-Sonnet-4.5Institution=Anthropic2026.03 | 77.44 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-42024.07 | 77.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-LongStrategy=Direct2024.10 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CodestralStrategy=MBTI2024.10 | 76.8 | — | 1.2 | — | — | — | — | — | — | — | — | — | — | — | |
| UCoder-7BSize Category=6B+ LLMs2025.12 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | — | — |