Code Generation on BigCodeBench (pass@1)
88.5pass@1Claude Opus 4.6
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude Opus 4.6Prompting Strategy=Zero-shot prompting2026.03 | 88.5 | |
| GPT-5.2Prompting Strategy=Zero-shot prompting2026.03 | 81.1 | |
| ARYAPrompting Strategy=Zero-shot prompting, Parameters=02026.03 | 80.5 | |
| o3-miniContext=Best Published Baseline2026.03 | 61.4 | |
| Claude-Sonnet-4.5Institution=Anthropic2026.03 | 45 | |
| GPT-4.1Institution=OpenAI2026.03 | 41.32 | |
| Gemini-2.5-ProInstitution=Google2026.03 | 41.32 | |
| GPT-5.1Institution=OpenAI2026.03 | 39.56 | |
| ReflexiCoder-8B (Multiple)Institution=Ours, setup=full iterative reasoning-reflection setup2026.03 | 36.84 | |
| Seed-Coder-8B-InstructInstitution=ByteDance2026.03 | 36.05 | |
| ReflexiCoder-8B (Single)Institution=Ours, setup=single-attempt without system prompt2026.03 | 35 | |
| Qwen2.5-Coder-7B-InstructInstitution=Alibaba2026.03 | 33.33 | |
| Qwen3-8BInstitution=Alibaba2026.03 | 32.63 | |
| DeepCoder-14B-PreviewInstitution=rLLM2026.03 | 28.33 | |
| DeepSeek-Coder-7B-InstructInstitution=DeepSeek2026.03 | 27.02 | |
| CodeGemma-7B-ITInstitution=Google2026.03 | 25.44 | |
| CodeLlama-7b-InstructInstitution=Meta2026.03 | 16.58 | |
| DeepCoder-1.5B-PreviewInstitution=rLLM2026.03 | 6.84 | |
| EmbedLLMModel scale=1.5B, Evaluation mode=0-shot2026.06 | 6.8 | |
| RouterDCModel scale=1.5B, Evaluation mode=0-shot2026.06 | 6.1 | |
| Qwen2.5-Coder-1.5B-InstructModel scale=1.5B, Evaluation mode=0-shot2026.06 | 5.4 | |
| LinearModel scale=1.5B, Evaluation mode=0-shot2026.06 | 4.7 | |
| SLERPModel scale=1.5B, Evaluation mode=0-shot2026.06 | 4.1 | |
| Entropy WeightingModel scale=1.5B, Evaluation mode=0-shot2026.06 | 4.1 | |
| Pack of LLMsModel scale=1.5B, Evaluation mode=0-shot2026.06 | 4.1 | |
| DLLGModel scale=1.5B, Evaluation mode=0-shot2026.06 | 4.1 | |
| Qwen2.5-1.5B-InstructModel scale=1.5B, Evaluation mode=0-shot2026.06 | 3.4 | |
| Task ArithmeticModel scale=1.5B, Evaluation mode=0-shot2026.06 | 3.4 | |
| Token Maj-VotingModel scale=1.5B, Evaluation mode=0-shot2026.06 | 3.4 | |
| GaCModel scale=1.5B, Evaluation mode=0-shot2026.06 | 2.7 | |
| UniTeModel scale=1.5B, Evaluation mode=0-shot2026.06 | 1.4 | |
| Qwen2.5-Math-1.5B-InstructModel scale=1.5B, Evaluation mode=0-shot2026.06 | 0 |