Code Generation on BigCodeBench Hard
35.1Pass@1Claude-Opus-4.5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Claude-Opus-4.5Model Category=Closed-APIs2026.03 | 35.1 | — | |
| IQuest-Coder-V1-40B-InstructModel Scale=20B+2026.03 | 33.1 | — | |
| Kimi-Dev-72BModel Scale=20B+2026.03 | 31.8 | — | |
| DS2Base model=Seed-Coder-8B, Eff. Tokens (%)=4.62026.06 | 31 | 66 | |
| Kimi-K2-Instruct-0905Model Scale=20B+2026.03 | 30.4 | — | |
| IQuest-Coder-V1-40B-Loop-ThinkingModel Scale=20B+2026.03 | 29.7 | — | |
| CLAMBase model=Seed-Coder-8B, Eff. Tokens (%)=3.62026.06 | 29.7 | 64.9 | |
| IQuest-Coder-V1-40B-ThinkingModel Scale=20B+2026.03 | 29.1 | — | |
| Claude-Sonnet-4.5Model Category=Closed-APIs2026.03 | 29.1 | — | |
| GPT-5.1Model Category=Closed-APIs2026.03 | 29.1 | — | |
| CODEBLOCKBase model=Seed-Coder-8B, Eff. Tokens (%)=1.92026.06 | 29.1 | 65.5 | |
| Kimi-K2-ThinkingModel Scale=20B+2026.03 | 28.4 | — | |
| FULL TOKENSBase model=Seed-Coder-8B, Eff. Tokens (%)=100.02026.06 | 28.3 | 65.1 | |
| Qwen3-Coder-30B-A3B-InstructModel Scale=13B+2026.03 | 27.7 | — | |
| Qwen3-Coder-480B-A35B-InstructModel Scale=20B+2026.03 | 27.7 | — | |
| IQuest-Coder-V1-40B-Loop-InstructModel Scale=20B+2026.03 | 27.7 | — | |
| Deepseek-V3.2Model Scale=20B+2026.03 | 27 | — | |
| IQuest-Coder-V1-14B-InstructModel Scale=13B+2026.03 | 26.4 | — | |
| KAT-Dev-72B-ExpModel Scale=20B+2026.03 | 26.4 | — | |
| GLM-4.7Model Scale=20B+2026.03 | 26.4 | — | |
| Qwen3-235B-A22B-Instruct-2507Model Scale=20B+2026.03 | 25.7 | — | |
| KAT-DevModel Scale=20B+2026.03 | 25.7 | — | |
| BASEBase model=Seed-Coder-8B2026.06 | 25.7 | 62.5 | |
| RANDOM SELECTIONBase model=Seed-Coder-8B, Eff. Tokens (%)=1.92026.06 | 25.7 | 63.9 | |
| Gemini-3-Flash-previewModel Category=Closed-APIs2026.03 | 25.6 | — | |
| Gemini-3-Pro-previewModel Category=Closed-APIs2026.03 | 25 | — | |
| TOKEN CLEANINGBase model=Seed-Coder-8B, Eff. Tokens (%)=1.72026.06 | 25 | 62.8 | |
| Qwen2.5-Coder-32B-InstructModel Scale=20B+2026.03 | 24.3 | — | |
| IQuest-Coder-V1-14B-ThinkingModel Scale=13B+2026.03 | 23.7 | — | |
| Seed-Coder-8B-InstructModel Scale=6B+2026.03 | 23.6 | — | |
| IQuest-Coder-V1-7B-InstructModel Scale=6B+2026.03 | 23 | — | |
| Qwen3-235B-A22B-Thinking-2507Model Scale=20B+2026.03 | 23 | — | |
| IQuest-Coder-V1-7B-ThinkingModel Scale=6B+2026.03 | 19.6 | — | |
| BASEBase model=Qwen2.5-Coder-7B-Instruct2026.06 | 19.4 | — | |
| DeepSeek-Coder-V2-Lite-InstructModel Scale=6B+2026.03 | 18.9 | — | |
| FULL TOKENSBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=100.02026.06 | 18.9 | — | |
| DS2Base model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=4.62026.06 | 18.9 | — | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.92026.06 | 18.2 | — | |
| CODEBLOCKBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.92026.06 | 18.2 | — | |
| CLAMBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=3.52026.06 | 16.2 | — | |
| FULL TOKENSBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=14.92026.06 | 14.9 | — | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.72026.06 | 14.9 | — | |
| CLAMBase model=OpenCoder-8B-Base, Eff. Tokens (%)=3.52026.06 | 14.2 | 54.6 | |
| DS2Base model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=4.62026.06 | 14.2 | — | |
| CLAMBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=3.52026.06 | 14.2 | — | |
| Qwen2.5-Coder-7B-InstructModel Scale=6B+2026.03 | 13.5 | — | |
| BASEBase model=Qwen2.5-Coder-3B-Instruct2026.06 | 13.5 | — | |
| AGRPOReinforcement Learning Strategy=+ AGRPO2026.03 | 13.1 | — | |
| UniGRPOReinforcement Learning Strategy=+ UniGRPO2026.03 | 12.9 | — | |
| SPGReinforcement Learning Strategy=+ SPG2026.03 | 12.9 | — | |
| LFPO (All Loss)Reinforcement Learning Strategy=+ LFPO (All Loss)2026.03 | 12.9 | — | |
| FULL TOKENSBase model=OpenCoder-8B-Base, Eff. Tokens (%)=100.02026.06 | 12.8 | 54.5 | |
| TOKEN CLEANINGBase model=OpenCoder-8B-Base, Eff. Tokens (%)=1.72026.06 | 12.8 | 51.3 | |
| CODEBLOCKBase model=OpenCoder-8B-Base, Eff. Tokens (%)=1.92026.06 | 12.8 | 57.1 | |
| diffu-GRPOReinforcement Learning Strategy=+ diffu-GRPO2026.03 | 12.5 | — | |
| LFPO (Neg. Only)Reinforcement Learning Strategy=+ LFPO (Neg. Only)2026.03 | 12.5 | — | |
| LFPO (Pos. Only)Reinforcement Learning Strategy=+ LFPO (Pos. Only)2026.03 | 12.4 | — | |
| DiffuCoderReinforcement Learning Strategy=Base2026.03 | 12.2 | — | |
| DS2Base model=OpenCoder-8B-Base, Eff. Tokens (%)=4.62026.06 | 12.2 | 54.6 | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.72026.06 | 12.2 | — | |
| CODEBLOCKBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.92026.06 | 12.2 | — | |
| coupled-GRPOReinforcement Learning Strategy=+ coupled-GRPO2026.03 | 10.8 | — | |
| RANDOM SELECTIONBase model=OpenCoder-8B-Base, Eff. Tokens (%)=1.92026.06 | 10.8 | 54.6 | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.92026.06 | 10.1 | — | |
| BASEBase model=OpenCoder-8B-Base2026.06 | 9.5 | 54.4 | |
| CLAMBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=3.52026.06 | 8.1 | 51.8 | |
| DS2Base model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=4.62026.06 | 7.4 | 53.6 | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=1.92026.06 | 6.8 | 53.6 | |
| Qwen2.5-Coder-14B-InstructModel Scale=13B+2026.03 | 6.1 | — | |
| CODEBLOCKBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=1.92026.06 | 6.1 | 54.6 | |
| FULL TOKENSBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=100.02026.06 | 5.4 | 52.7 | |
| BASEBase model=Qwen2.5-Coder-1.5B-Instruct2026.06 | 4.1 | 49.4 | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=1.82026.06 | 4.1 | 49.1 |