Code Generation on HumanEval 2021 (test)
98.78AccuracyPoT
Evaluation Results
| Method | Links | |
|---|---|---|
| PoTBase Model=Qwen3-4B-Instruct-25072026.01 | 98.78 | |
| Claude-Opus-4Inference Strategy=Zero-shot2026.01 | 95.8 | |
| Gemini-2.5-flashInference Strategy=Zero-shot2026.01 | 94.01 | |
| Gpt-4oInference Strategy=Zero-shot2026.01 | 92.7 | |
| Deepseek-V3Inference Strategy=Zero-shot2026.01 | 91.5 | |
| Best of N (N=20)Base Model=Qwen3-4B-Instruct-2507, Method Category=Standard Inference2026.01 | 90.85 | |
| AB-MCTSBase Model=Qwen3-4B-Instruct-2507, Method Category=Search-Based Reasoning2026.01 | 90.85 | |
| ToTBase Model=Qwen3-4B-Instruct-2507, Method Category=Search-Based Reasoning2026.01 | 90.24 | |
| LLM (32B)Model Group=Single Models, Size=32B2026.04 | 89.02 | |
| PG-TDBase Model=Qwen3-4B-Instruct-2507, Method Category=Search-Based Reasoning2026.01 | 89.02 | |
| RethinkMCTSBase Model=Qwen3-4B-Instruct-2507, Method Category=Search-Based Reasoning2026.01 | 89.02 | |
| CODETBase Model=Qwen3-4B-Instruct-2507, Method Category=Self-Refinement / Debugging2026.01 | 88.41 | |
| ReflexionBase Model=Qwen3-4B-Instruct-2507, Method Category=Self-Refinement / Debugging2026.01 | 87.8 | |
| LATSBase Model=Qwen3-4B-Instruct-2507, Method Category=Search-Based Reasoning2026.01 | 87.8 | |
| RAPBase Model=Qwen3-4B-Instruct-2507, Method Category=Self-Refinement / Debugging2026.01 | 87.19 | |
| Few-shot + CoTBase Model=Qwen3-4B-Instruct-2507, Method Category=Standard Inference2026.01 | 85.98 | |
| TandemModel Group=Collaboration, SLM=7B, LLM=32B, Guidance Level=Adaptive2026.04 | 85.37 | |
| Zero-shotBase Model=Qwen3-4B-Instruct-2507, Method Category=Standard Inference2026.01 | 84.76 | |
| Self-RefineBase Model=Qwen3-4B-Instruct-2507, Method Category=Self-Refinement / Debugging2026.01 | 84.15 | |
| 7B+32B (high)Model Group=Collaboration, SLM=7B, LLM=32B, Guidance Level=high2026.04 | 83.54 | |
| Qwen3-235B-A22BInference Strategy=Zero-shot2026.01 | 79.88 | |
| 7B+32B (medium)Model Group=Collaboration, SLM=7B, LLM=32B, Guidance Level=medium2026.04 | 79.27 | |
| Merged (α=0.2)Model variant=Merged (α=0.2), Interpolation Weight=0.22026.04 | 76.1 | |
| 7B+32B (low)Model Group=Collaboration, SLM=7B, LLM=32B, Guidance Level=low2026.04 | 73.17 | |
| SFT No-TestModel variant=SFT No-Test2026.04 | 70.1 | |
| SLM (7B)Model Group=Single Models, Size=7B2026.04 | 65.24 | |
| Qwen3-4B-Instruct-2507Model variant=Base Model2026.04 | 63.7 |