Functional Code Generation on HumanEval 8 instruct cot (test)
94.7Pass@1gpt-oss-20b
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| gpt-oss-20bPrompting=Chain-of-thought, Instruction-following=true2026.05 | 94.7 | 99.4 | |
| Qwen3-4BPrompting=Chain-of-thought, Instruction-following=true2026.05 | 93.3 | 99.4 | |
| GPT-5 nanoPrompting=Chain-of-thought, Instruction-following=true2026.05 | 84.8 | 98.8 | |
| Qwen3-8BPrompting=Chain-of-thought, Instruction-following=true2026.05 | 80.1 | 97.6 | |
| gemma-3-12b-itPrompting=Chain-of-thought, Instruction-following=true2026.05 | 76.2 | 87.8 | |
| gemma-3-4b-itPrompting=Chain-of-thought, Instruction-following=true2026.05 | 64.1 | 81.1 | |
| Moonlight-16B-A3B-InstructPrompting=Chain-of-thought, Instruction-following=true2026.05 | 61.1 | 87.2 | |
| EngGPT2-16B-A3BPrompting=Chain-of-thought, Instruction-following=true2026.05 | 45.2 | 92.1 | |
| Llama-3.2-3B-InstructPrompting=Chain-of-thought, Instruction-following=true2026.05 | 33.9 | 59.2 | |
| Llama-3.1-8B-InstructPrompting=Chain-of-thought, Instruction-following=true2026.05 | 28.7 | 54.9 | |
| LLaMAntino-3-ANITA-8B-Inst-DPO-ITAPrompting=Chain-of-thought, Instruction-following=true2026.05 | 26 | 57.9 | |
| Ministral-3-8B-Instruct-2512-BF16Prompting=Chain-of-thought, Instruction-following=true2026.05 | 25.1 | 80.5 | |
| FastwebMIIA-7BPrompting=Chain-of-thought, Instruction-following=true2026.05 | 18.5 | 56.7 | |
| deepseek-moe-16b-chatPrompting=Chain-of-thought, Instruction-following=true2026.05 | 17.8 | 54.9 | |
| Velvet-14BPrompting=Chain-of-thought, Instruction-following=true2026.05 | 12.7 | 36 | |
| Minerva-7B-instruct-v1.0Prompting=Chain-of-thought, Instruction-following=true2026.05 | 5 | 16.5 |