Code Generation on EvalPlus
89Pass@1EVOLVECODER-4B (r3)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| EVOLVECODER-4B (r3)Refinement round=r3, Backbone=Qwen3-4B (Thinking)2026.03 | 89 | — | — | |
| EVOLVECODER-4B (r1)Refinement round=r1, Backbone=Qwen3-4B (Thinking)2026.03 | 88.8 | — | — | |
| EVOLVECODER-4B (r2)Refinement round=r2, Backbone=Qwen3-4B (Thinking)2026.03 | 88.8 | — | — | |
| GPT-o1Backbone=GPT-o12025.09 | 88.6 | — | — | |
| o12026.03 | 88.6 | — | — | |
| EVOLVECODER-4B (r0)Refinement round=r0, Backbone=Qwen3-4B (Thinking)2026.03 | 88.3 | — | — | |
| EAPO2025.09 | 88.04 | — | — | |
| SRGenModel=Qwen3-32B2025.10 | 87.8 | — | — | |
| CRITIQUE-CODERBackbone=Qwen3-8B, Thinking Mode=true, Training Strategy=CRL2025.09 | 87.7 | — | — | |
| Self-RefineModel=Qwen3-32B2025.10 | 87.2 | — | — | |
| CoTModel=Qwen3-32B2025.10 | 87.1 | — | — | |
| GLM-5 BaseArchitecture=MoE, Activated Params=40B, Total Params=744B2026.02 | 87 | — | — | |
| CRITIQUE-CODERBackbone=Qwen3-4B, Thinking Mode=true, Training Strategy=CRL2025.09 | 86.5 | — | — | |
| Critique-Coder-4B2026.03 | 86.5 | — | — | |
| Qwen3-8B-RLBackbone=Qwen3-8B, Thinking Mode=true, Training Strategy=Standard RL2025.09 | 86.2 | — | — | |
| Baseline (Qwen3-8B)Backbone=Qwen3-8B, Thinking Mode=true, Training Strategy=Baseline2025.09 | 85.8 | — | — | |
| DeepCoder-14BBackbone=DeepCoder-14B2025.09 | 85.3 | — | — | |
| DeepCoder-14B2026.03 | 85.3 | — | — | |
| Baseline (Qwen3-4B)Backbone=Qwen3-4B, Thinking Mode=true, Training Strategy=Baseline2025.09 | 85.2 | — | — | |
| Qwen3-4B (Thinking)Thinking-mode sampling=true2026.03 | 85.2 | — | — | |
| Self-Exploratory RL2025.09 | 85.16 | — | — | |
| Qwen3-4B-RLBackbone=Qwen3-4B, Thinking Mode=true, Training Strategy=Standard RL2025.09 | 84.9 | — | — | |
| DeepSeek-V2.5-238BBackbone=DeepSeek-V2.5-238B2025.09 | 83.8 | — | — | |
| DeepSeek-V2.5-238B2026.03 | 83.8 | — | — | |
| Base Model2025.09 | 83.72 | — | — | |
| MapCoderLLM=GPT-42024.05 | 83.5 | — | — | |
| AceCoder-7BBackbone=AceCoder-7B2025.09 | 82.7 | — | — | |
| AceCoder-7B2026.03 | 82.7 | — | — | |
| DeepSeek-R1-Distill-14BBackbone=DeepSeek-R1-Distill-14B2025.09 | 82.4 | — | — | |
| DeepSeek-R1-Distill-14B2026.03 | 82.4 | — | — | |
| DirectLLM=GPT-42024.05 | 81.7 | — | — | |
| ReflexionLLM=GPT-42024.05 | 81.7 | — | — | |
| EXG-SE-RevModel=Qwen3-Coder-Flash2026.05 | 81.7 | — | 91.5 | |
| EXGModel=Qwen3-Coder-Flash2026.05 | 81.1 | — | 86.6 | |
| Kimi-K2 BaseArchitecture=MoE, Activated Params=32B, Total Params=1043B2026.02 | 80.3 | — | — | |
| EXG-ReflexionModel=Qwen3-Coder-Flash2026.05 | 79.3 | — | 86.6 | |
| EXG-SEModel=Qwen3-Coder-Flash2026.05 | 79.3 | — | 84.8 | |
| GLM-4.5 BaseArchitecture=MoE, Activated Params=32B, Total Params=355B2026.02 | 78.1 | — | — | |
| MAPRBase Model=Qwen3-14B-Base2025.09 | 77.66 | — | — | |
| GRPOBase Model=Qwen3-14B-Base2025.09 | 77.32 | — | — | |
| CodeSimLLM=LLaMa3.1-70B2025.02 | 76.2 | — | — | |
| SRGenModel=DS-R1-Qwen-7B2025.10 | 73.7 | — | — | |
| CoTModel=DS-R1-Qwen-7B2025.10 | 73.1 | — | — | |
| Self-RefineModel=DS-R1-Llama-8B2025.10 | 73.1 | — | — | |
| CodeSimLLM=Gemma2-9B2025.02 | 72.6 | — | — | |
| Self-RefineModel=DS-R1-Qwen-7B2025.10 | 72.6 | — | — | |
| SRGenModel=DS-R1-Llama-8B2025.10 | 72 | — | — | |
| MapCoderLLM=ChatGPT2024.05 | 71.3 | — | — | |
| CoTModel=DS-R1-Llama-8B2025.10 | 71.3 | — | — | |
| UniGRPOReinforcement Learning Strategy=+ UniGRPO2026.03 | 70.3 | — | — | |
| CoTLLM=LLaMa3.1-70B2025.02 | 70.1 | — | — | |
| LFPO (All Loss)Reinforcement Learning Strategy=+ LFPO (All Loss)2026.03 | 69.5 | — | — | |
| AGRPOReinforcement Learning Strategy=+ AGRPO2026.03 | 69.3 | — | — | |
| SPGReinforcement Learning Strategy=+ SPG2026.03 | 68.9 | — | — | |
| LFPO (Neg. Only)Reinforcement Learning Strategy=+ LFPO (Neg. Only)2026.03 | 68.9 | — | — | |
| LFPO (Pos. Only)Reinforcement Learning Strategy=+ LFPO (Pos. Only)2026.03 | 68.3 | — | — | |
| ReflexionLLM=LLaMa3.1-70B2025.02 | 68.3 | — | — | |
| diffu-GRPOReinforcement Learning Strategy=+ diffu-GRPO2026.03 | 67.9 | — | — | |
| coupled-GRPOReinforcement Learning Strategy=+ coupled-GRPO2026.03 | 67.9 | — | — | |
| DirectLLM=ChatGPT2024.05 | 66.5 | — | — | |
| DeepSeek-V3 BaseArchitecture=MoE, Activated Params=37B, Total Params=671B2026.02 | 65.6 | — | — | |
| CoTLLM=ChatGPT2024.05 | 65.2 | — | — | |
| Qwen-3.5Model Type=Open-weight, Number of Parameters=4B2026.03 | 65 | — | — | |
| DiffuCoderReinforcement Learning Strategy=Base2026.03 | 63.6 | — | — | |
| YulanModel Type=Fully-open, Number of Parameters=2.4B2026.03 | 62.25 | — | — | |
| ReflexionLLM=ChatGPT2024.05 | 62.2 | — | — | |
| AnalogicalLLM=GPT-42024.05 | 62.2 | — | — | |
| MSA-PTTraining Strategy=from-scratch sparse pretraining2026.06 | 61.8 | — | — | |
| CodeSimLLM=LLaMa3.1-8B2025.02 | 61.2 | — | — | |
| CodeSimLLM=Mixtral8x7B2025.02 | 61 | — | — | |
| SE-AgentModel=Qwen3-Coder-Flash2026.05 | 61 | — | 76.8 | |
| MSA-CPTTraining Strategy=sparse continued pretraining2026.06 | 60 | — | — | |
| Qwen-3Model Type=Open-weight, Number of Parameters=4B2026.03 | 59.45 | — | — | |
| FullTraining Strategy=Full-Attention baseline2026.06 | 59.4 | — | — | |
| AnalogicalLLM=ChatGPT2024.05 | 59.1 | — | — | |
| SE-Agent-RevModel=Qwen3-Coder-Flash2026.05 | 59.1 | — | 84.8 | |
| ReflexionModel=Qwen3-Coder-Flash2026.05 | 58.5 | — | 87.2 | |
| daVinciModel Type=Fully-open, Number of Parameters=3B2026.03 | 57.32 | — | — | |
| Qwen2-57B-A14BArchitecture=MoE, # Act Params=14B, # Params=57B2024.07 | 57.2 | — | — | |
| DirectLLM=Gemma2-9B2025.02 | 56.1 | — | — | |
| ReflexionLLM=Gemma2-9B2025.02 | 55.5 | — | — | |
| Qwen2-7BParameters=7B2024.07 | 54.2 | — | — | |
| OLMO-3Model Type=Fully-open, Number of Parameters=7B2026.03 | 53.62 | — | — | |
| Qwen-2.5Model Type=Open-weight, Number of Parameters=3B2026.03 | 53.23 | — | — | |
| DirectLLM=LLaMa3.1-70B2025.02 | 52.4 | — | — | |
| EXG-SEModel=Qwen3-8B2026.05 | 52.4 | — | 53.7 | |
| Yi-1.5-34BArchitecture=Dense, # Act Params=32B, # Params=32B2024.07 | 51.9 | — | — | |
| Qwen1.5-32BArchitecture=Dense, # Act Params=34B, # Params=34B2024.07 | 50.4 | — | — | |
| SRGenModel=Qwen2.5-Math-7B2025.10 | 47.6 | — | — | |
| Mixtral-8x7BArchitecture=MoE, # Act Params=12B, # Params=47B2024.07 | 46.4 | — | — | |
| CoTModel=Qwen2.5-Math-7B2025.10 | 45.7 | — | — | |
| Self-RefineModel=Qwen2.5-Math-7B2025.10 | 43.9 | — | — | |
| CoTLLM=LLaMa3.1-8B2025.02 | 43.3 | — | — | |
| EXGModel=Qwen3-8B2026.05 | 40.9 | — | 53 | |
| EXG-SEModel=Qwen3-1.7B2026.05 | 40.8 | — | 41.5 | |
| Llama-3-8BParameters=8B2024.07 | 40.3 | — | — | |
| Qwen1.5-7BParameters=7B2024.07 | 40 | — | — | |
| Gemma-7BParameters=7B2024.07 | 39.6 | — | — | |
| CoTLLM=Mixtral8x7B2025.02 | 39 | — | — | |
| DirectLLM=LLaMa3.1-8B2025.02 | 39 | — | — |