Code Generation on MBPP Sanitized
85.7AccuracyDeepSeek-Coder V2
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-Coder V2Strategy=MBTI2024.10 | 85.7 | 10.8 | |
| GPT-4oStrategy=MBTI2024.10 | 84.3 | 6.1 | |
| GPT-4o miniStrategy=MBTI2024.10 | 82.2 | 12.9 | |
| Llama3.1Strategy=MBTI2024.10 | 81 | 11.2 | |
| Qwen-LongStrategy=MBTI2024.10 | 80.8 | 12.4 | |
| GPT-4oStrategy=Direct2024.10 | 78.2 | — | |
| DeepSeek-Coder V2Strategy=Direct2024.10 | 74.9 | — | |
| Qwen3-4BParams=4B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 74.32 | — | |
| Qwen3-4BParams=4B2025.12 | 74.32 | — | |
| CodestralStrategy=MBTI2024.10 | 73.8 | 9.6 | |
| Llama3.1Strategy=Direct2024.10 | 69.8 | — | |
| GPT-4o miniStrategy=Direct2024.10 | 69.3 | — | |
| Qwen-LongStrategy=Direct2024.10 | 68.4 | — | |
| CodeTBase Model=Codex002, Number of samples=1002022.11 | 67.7 | — | |
| Qwen2.5-3BParams=3B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 66.54 | — | |
| Qwen2.5-3BParams=3B2025.12 | 66.54 | — | |
| Coder-ReviewerBase Model=Codex002, Number of samples=1002022.11 | 64.7 | — | |
| CodestralStrategy=Direct2024.10 | 64.2 | — | |
| Qwen3-1.7BParams=1.7B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 64.2 | — | |
| Qwen3-1.7BParams=1.7B2025.12 | 64.2 | — | |
| YuLan-Mini-2.4BParams=2.4B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 62.26 | — | |
| SmolLM3-3BParams=3B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 62.26 | — | |
| YuLan-Mini-2.4BParams=2.4B2025.12 | 62.26 | — | |
| SmolLM3-3BParams=3B2025.12 | 62.26 | — | |
| N. Coder-ReviewerBase Model=Codex002, Number of samples=1002022.11 | 61 | — | |
| Qwen2.5-1.5BParams=1.5B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 58.37 | — | |
| Qwen2.5-1.5BParams=1.5B2025.12 | 58.37 | — | |
| PCMind-2.1-Kaiyuan-2BParams=2B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 56.42 | — | |
| PCMind-2.1-Kaiyuan-2BParams=2B2025.12 | 56.42 | — | |
| Qwen3-0.6BParams=0.6B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 51.75 | — | |
| Qwen3-0.6BParams=0.6B2025.12 | 51.75 | — | |
| Qwen2-1.5BParams=1.5B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 50.58 | — | |
| Qwen2-1.5BParams=1.5B2025.12 | 50.58 | — | |
| Coder-ReviewerBase Model=CodeGen16B, Number of samples=1002022.11 | 50.3 | — | |
| CodeTBase Model=CodeGen16B, Number of samples=1002022.11 | 49.5 | — | |
| Llama-3.2-3BParams=3B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 49.42 | — | |
| SmolLM2-1.7BParams=1.7B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 49.42 | — | |
| llama-3.2-3BParams=3B2025.12 | 49.42 | — | |
| SmolLM2-1.7BParams=1.7B2025.12 | 49.42 | — | |
| CodeLlamaStrategy=MBTI2024.10 | 46.8 | 3.5 | |
| N. Coder-ReviewerBase Model=CodeGen16B, Number of samples=1002022.11 | 46.1 | — | |
| CodeLlamaStrategy=Direct2024.10 | 43.3 | — | |
| Gemma2-2BParams=2B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 38.91 | — | |
| gemma2-2BParams=2B2025.12 | 38.91 | — | |
| Coder-ReviewerBase Model=InCoder6B, Number of samples=1002022.11 | 35.8 | — | |
| Llama-3.2-1BParams=1B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 34.63 | — | |
| llama-3.2-1BParams=1B2025.12 | 34.63 | — | |
| CodeTBase Model=InCoder6B, Number of samples=1002022.11 | 34.4 | — | |
| N. Coder-ReviewerBase Model=InCoder6B, Number of samples=1002022.11 | 30.2 | — | |
| OLMo-2-0425-1BParams=1B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 15.56 | — | |
| OLMo-2-0425-1BParams=1B2025.12 | 15.56 | — |