Code Generation on BigCodeBench Full
54.2Pass@1IQuest-Coder-V1-40B-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| IQuest-Coder-V1-40B-InstructModel Scale=20B+2026.03 | 54.2 | — | |
| TOKEN CLEANINGBase model=Seed-Coder-8B, Eff. Tokens (%)=1.72026.06 | 53.6 | 62.8 | |
| Claude-Opus-4.5Model Category=Closed-APIs2026.03 | 53.3 | — | |
| Claude-Opus-4.52026.06 | 53.3 | — | |
| CODEBLOCKBase model=Seed-Coder-8B, Eff. Tokens (%)=1.92026.06 | 53.1 | 65.5 | |
| FULL TOKENSBase model=Seed-Coder-8B, Eff. Tokens (%)=100.02026.06 | 52.8 | 65.1 | |
| DS2Base model=Seed-Coder-8B, Eff. Tokens (%)=4.62026.06 | 52.7 | 66 | |
| CLAMBase model=Seed-Coder-8B, Eff. Tokens (%)=3.62026.06 | 52.3 | 64.9 | |
| RANDOM SELECTIONBase model=Seed-Coder-8B, Eff. Tokens (%)=1.92026.06 | 51.6 | 63.9 | |
| Claude-Sonnet-4.5Model Category=Closed-APIs2026.03 | 51.4 | — | |
| BASEBase model=Seed-Coder-8B2026.06 | 51.3 | 62.5 | |
| IQuest-Coder-V1-40B-ThinkingModel Scale=20B+2026.03 | 51.1 | — | |
| IQuest-Coder-V1-40B-Loop-ThinkingModel Scale=20B+2026.03 | 50.6 | — | |
| IQuest-Coder-V1-40B-Loop-InstructModel Scale=20B+2026.03 | 49.9 | — | |
| Kimi-K2-Instruct-0905Model Scale=20B+2026.03 | 49.8 | — | |
| Kimi-K2-Instruct-09052026.06 | 49.8 | — | |
| Qwen3-Coder-480B-A35B-InstructModel Scale=20B+2026.03 | 49.4 | — | |
| Qwen3-Coder-480B-A35B-Instruct2026.06 | 49.4 | — | |
| KAT-Dev-72B-ExpModel Scale=20B+2026.03 | 48.3 | — | |
| Deepseek-V3.2Model Scale=20B+2026.03 | 48.1 | — | |
| DeepSeek-V3.22026.06 | 48.1 | — | |
| Qwen2.5-Coder-32B-InstructModel Scale=20B+2026.03 | 48 | — | |
| Qwen2.5-Coder-32B-Instruct2026.06 | 48 | — | |
| IQuest-Coder-V1-14B-ThinkingModel Scale=13B+2026.03 | 47.7 | — | |
| Qwen3-235B-A22B-Instruct-2507Model Scale=20B+2026.03 | 47.4 | — | |
| Qwen3-235B-A22B-Instruct-25072026.06 | 47.4 | — | |
| Gemini-3-Pro-previewModel Category=Closed-APIs2026.03 | 47.1 | — | |
| Gemini-3-Pro2026.06 | 47.1 | — | |
| Qwen2.5-Coder-14B-InstructModel Scale=13B+2026.03 | 47 | — | |
| Qwen2.5-Coder-14B-Instruct2026.06 | 47 | — | |
| Qwen3-Coder-30B-A3B-InstructModel Scale=13B+2026.03 | 46.9 | — | |
| Kimi-K2-ThinkingModel Scale=20B+2026.03 | 46.8 | — | |
| GPT-5.1Model Category=Closed-APIs2026.03 | 46.8 | — | |
| GPT-5.12026.06 | 46.8 | — | |
| IQuest-Coder-V1-14B-InstructModel Scale=13B+2026.03 | 46.3 | — | |
| KAT-DevModel Scale=20B+2026.03 | 46.2 | — | |
| LoopCoder-v2R (number of refinement loops)=22026.06 | 46.1 | — | |
| GLM-4.7Model Scale=20B+2026.03 | 45.7 | — | |
| GLM-4.72026.06 | 45.7 | — | |
| Kimi-Dev-72BModel Scale=20B+2026.03 | 45.4 | — | |
| Kimi-Dev-72B2026.06 | 45.4 | — | |
| LFPO (All Loss)Reinforcement Learning Strategy=+ LFPO (All Loss)2026.03 | 44.8 | — | |
| AGRPOReinforcement Learning Strategy=+ AGRPO2026.03 | 44.6 | — | |
| Seed-Coder-8B-InstructModel Scale=6B+2026.03 | 44.6 | — | |
| Seed-Coder-8B-Instruct2026.06 | 44.6 | — | |
| Gemini-3-Flash-previewModel Category=Closed-APIs2026.03 | 44.5 | — | |
| Qwen3-235B-A22B-Thinking-2507Model Scale=20B+2026.03 | 44.1 | — | |
| LFPO (Neg. Only)Reinforcement Learning Strategy=+ LFPO (Neg. Only)2026.03 | 43.4 | — | |
| LoopCoder-v2R (number of refinement loops)=32026.06 | 43.3 | — | |
| LFPO (Pos. Only)Reinforcement Learning Strategy=+ LFPO (Pos. Only)2026.03 | 43.1 | — | |
| FULL TOKENSBase model=OpenCoder-8B-Base, Eff. Tokens (%)=100.02026.06 | 42.8 | 54.5 | |
| UniGRPOReinforcement Learning Strategy=+ UniGRPO2026.03 | 42.6 | — | |
| SPGReinforcement Learning Strategy=+ SPG2026.03 | 42.1 | — | |
| DS2Base model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=4.62026.06 | 42 | — | |
| DS2Base model=OpenCoder-8B-Base, Eff. Tokens (%)=4.62026.06 | 41.4 | 54.6 | |
| CODEBLOCKBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.92026.06 | 41.4 | — | |
| RANDOM SELECTIONBase model=OpenCoder-8B-Base, Eff. Tokens (%)=1.92026.06 | 40.9 | 54.6 | |
| LoopCoder-v2R (number of refinement loops)=42026.06 | 40.8 | — | |
| IQuest-Coder-V1-7B-ThinkingModel Scale=6B+2026.03 | 40.5 | — | |
| BASEBase model=OpenCoder-8B-Base2026.06 | 40.5 | 54.4 | |
| coupled-GRPOReinforcement Learning Strategy=+ coupled-GRPO2026.03 | 40.4 | — | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.92026.06 | 40.4 | — | |
| LoopCoder-v2R (number of refinement loops)=12026.06 | 40.1 | — | |
| FULL TOKENSBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=100.02026.06 | 39.9 | — | |
| CLAMBase model=OpenCoder-8B-Base, Eff. Tokens (%)=3.52026.06 | 39.8 | 54.6 | |
| CLAMBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=3.52026.06 | 39.8 | — | |
| BASEBase model=Qwen2.5-Coder-7B-Instruct2026.06 | 39.7 | — | |
| CODEBLOCKBase model=OpenCoder-8B-Base, Eff. Tokens (%)=1.92026.06 | 39.3 | 57.1 | |
| diffu-GRPOReinforcement Learning Strategy=+ diffu-GRPO2026.03 | 39.2 | — | |
| IQuest-Coder-V1-7B-InstructModel Scale=6B+2026.03 | 38.9 | — | |
| DeepSeek-Coder-V2-Lite-InstructModel Scale=6B+2026.03 | 37.8 | — | |
| Qwen2.5-Coder-7B-InstructModel Scale=6B+2026.03 | 37.8 | — | |
| DeepSeek-Coder-V2-Lite-Instruct2026.06 | 37.8 | — | |
| Qwen2.5-Coder-7B-Instruct2026.06 | 37.8 | — | |
| FULL TOKENSBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=100.02026.06 | 37.2 | — | |
| DS2Base model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=4.62026.06 | 36.7 | — | |
| CLAMBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=3.52026.06 | 36.7 | — | |
| CODEBLOCKBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.92026.06 | 36.7 | — | |
| TOKEN CLEANINGBase model=OpenCoder-8B-Base, Eff. Tokens (%)=1.72026.06 | 35.9 | 51.3 | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.72026.06 | 35.9 | — | |
| DiffuCoderReinforcement Learning Strategy=Base2026.03 | 35.7 | — | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.92026.06 | 35.7 | — | |
| BASEBase model=Qwen2.5-Coder-3B-Instruct2026.06 | 34.8 | — | |
| CODEBLOCKBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=1.92026.06 | 31.1 | 54.6 | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.72026.06 | 30.9 | — | |
| DS2Base model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=4.62026.06 | 30.5 | 53.6 | |
| CLAMBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=3.52026.06 | 30.4 | 51.8 | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=1.92026.06 | 29.9 | 53.6 | |
| FULL TOKENSBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=100.02026.06 | 29.8 | 52.7 | |
| BASEBase model=Qwen2.5-Coder-1.5B-Instruct2026.06 | 24.8 | 49.4 | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-1.5B-Instruct, Eff. Tokens (%)=1.82026.06 | 24.5 | 49.1 |