Code Evaluation on BigCodeBench
82.02AccuracyGPT-o1-preview
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-o1-previewApproach=System-22025.02 | 82.02 | |
| MCTS-JudgeApproach=System-2, Base Model=GPT-4o-mini2025.02 | 79.12 | |
| ICE-ScoreApproach=System-1, Base Model=GPT-4o-mini2025.02 | 77.37 | |
| GPT-o1-miniApproach=System-22025.02 | 75.7 | |
| VanillaApproach=System-1, Base Model=GPT-4o-mini2025.02 | 72.37 | |
| MCTS-JudgeApproach=System-2, Base Model=Llama-3.1-8B-Instruct2025.02 | 71.84 | |
| MCTS-JudgeApproach=System-2, Base Model=Qwen2.5-Coder-14B-Instruct2025.02 | 71.23 | |
| CodeJudgeApproach=System-1, Base Model=GPT-4o-mini2025.02 | 70.7 | |
| ICE-ScoreApproach=System-1, Base Model=Qwen2.5-Coder-14B-Instruct2025.02 | 70.44 | |
| MCTS-JudgeApproach=System-2, Base Model=Mistralai-Codestral-22B2025.02 | 68.77 | |
| CodeJudgeApproach=System-1, Base Model=Llama-3.1-8B-Instruct2025.02 | 63.86 | |
| VanillaApproach=System-1, Base Model=Qwen2.5-Coder-14B-Instruct2025.02 | 63.33 | |
| CodeJudgeApproach=System-1, Base Model=Qwen2.5-Coder-14B-Instruct2025.02 | 63.33 | |
| MCTS-JudgeApproach=System-2, Base Model=DeepSeek-Coder-V2-16B-Instruct2025.02 | 62.46 | |
| ICE-ScoreApproach=System-1, Base Model=DeepSeek-Coder-V2-16B-Instruct2025.02 | 57.89 | |
| CodeJudgeApproach=System-1, Base Model=DeepSeek-Coder-V2-16B-Instruct2025.02 | 52.45 | |
| ICE-ScoreApproach=System-1, Base Model=Mistralai-Codestral-22B2025.02 | 51.93 | |
| VanillaApproach=System-1, Base Model=DeepSeek-Coder-V2-16B-Instruct2025.02 | 51.75 | |
| Qwen-QwQ-32BApproach=System-22025.02 | 50.96 | |
| CodeJudgeApproach=System-1, Base Model=Mistralai-Codestral-22B2025.02 | 49.04 | |
| ICE-ScoreApproach=System-1, Base Model=Llama-3.1-8B-Instruct2025.02 | 45.88 | |
| VanillaApproach=System-1, Base Model=Llama-3.1-8B-Instruct2025.02 | 43.16 | |
| VanillaApproach=System-1, Base Model=Mistralai-Codestral-22B2025.02 | 42.81 |