Code Repair on HumanEvalFix (test)
59.1Success Rate (Python)WaveCoder-Pro-6.7B
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| WaveCoder-Pro-6.7BDecoding strategy=Greedy decoding2023.12 | 59.1 | 56.7 | 54.2 | 45.1 | 45.7 | 34.1 | 49.2 | — | — | |
| WaveCoder-Ultra-6.7BDecoding strategy=Greedy decoding2023.12 | 58.5 | 57.3 | 61 | 53 | 50 | 37.2 | 52.8 | — | — | |
| WaveCoder-DS-6.7BDecoding strategy=Greedy decoding2023.12 | 57.9 | 52.4 | 57.3 | 47.5 | 45.1 | 36 | 49.4 | — | — | |
| Deepseek-instruct-6.7BDecoding strategy=Self-reported2023.12 | 56.1 | 58.5 | 57.3 | 49.4 | 45.1 | 36.6 | 50.5 | — | — | |
| Magicoder-S-DSDecoding strategy=Self-reported2023.12 | 56.1 | 55.4 | 58.5 | 51.2 | 45.7 | 35.3 | 50.3 | — | — | |
| DeepseekCoder-CodeAlpaca-6.7BDecoding strategy=Greedy decoding2023.12 | 49.4 | 51.8 | 45.1 | 48.8 | 44.5 | 31.7 | 45.2 | — | — | |
| WaveCoder-CL-13BDecoding strategy=Greedy decoding2023.12 | 48.8 | 48.2 | 50.6 | 51.8 | 45.1 | 40.2 | 47.4 | — | — | |
| GPT-4Decoding strategy=Self-reported2023.12 | 47 | 48.2 | 50 | 50.6 | 47.6 | 43.3 | 47.8 | — | — | |
| CodeLLaMa-CodeAlpaca-13BDecoding strategy=Self-reported2023.12 | 42.7 | 43.9 | 50 | 45.7 | 39.6 | 37.2 | 43.2 | — | — | |
| Magicoder-DSDecoding strategy=Self-reported2023.12 | 42 | 43.3 | 50.6 | 41.4 | 38.4 | 29.2 | 40.8 | — | — | |
| WaveCoder-CL-7BDecoding strategy=Greedy decoding2023.12 | 41.4 | 41.4 | 42 | 47.1 | 42.7 | 34.7 | 41.5 | — | — | |
| WaveCoder-SC-15BDecoding strategy=Greedy decoding2023.12 | 39.3 | 35.1 | 34.8 | 36.2 | 30.2 | 22.5 | 33 | — | — | |
| CodeLLaMa-CodeAlpaca-7BDecoding strategy=Self-reported2023.12 | 37.8 | 39 | 42 | 37.8 | 37.2 | 29.2 | 37.1 | — | — | |
| WizardCoderDecoding strategy=Self-reported2023.12 | 31.8 | 29.5 | 30.7 | 30.4 | 18.7 | 13 | 25.7 | — | — | |
| OctoCoderDecoding strategy=Self-reported2023.12 | 30.4 | 28.4 | 30.6 | 30.2 | 26.1 | 16.5 | 27 | — | — | |
| DeepseekCoder-6.7BDecoding strategy=Self-reported2023.12 | 29.9 | 29.2 | 39 | 29.2 | 25 | 21.9 | 29 | — | — | |
| CodeLLaMa-instruct-13BDecoding strategy=Self-reported2023.12 | 29.2 | 19.5 | 32.3 | 24.4 | 12.8 | 0.1 | 19.7 | — | — | |
| CodeLLaMa-instruct-7BDecoding strategy=Self-reported2023.12 | 28 | 23.2 | 23.2 | 18.3 | 0.1 | 0.1 | 15.5 | — | — | |
| StarCoderDecoding strategy=Self-reported2023.12 | 8.7 | 15.7 | 13.3 | 20.1 | 15.6 | 6.7 | 13.4 | — | — | |
| InternLM-StepProver + CoTBackbone=InternLM-StepProver, Decoding Strategy=CoT2026.06 | — | — | — | — | — | — | — | 55.8 | 1,987 | |
| InternLM-StepProver + CoT-SCBackbone=InternLM-StepProver, Decoding Strategy=CoT-SC(k=8)2026.06 | — | — | — | — | — | — | — | 60.3 | 15,896 | |
| Llama-3.1-70B + CoTBackbone=Llama-3.1-70B, Decoding Strategy=CoT2026.06 | — | — | — | — | — | — | — | 57.3 | 1,913 | |
| Llama-3.1-70B + ToTBackbone=Llama-3.1-70B, Decoding Strategy=ToT(b=5)2026.06 | — | — | — | — | — | — | — | 60.1 | 9,561 | |
| Qwen2.5-72B + CoTBackbone=Qwen2.5-72B, Decoding Strategy=CoT2026.06 | — | — | — | — | — | — | — | 61.4 | 1,842 | |
| Qwen2.5-72B + CoT-SCBackbone=Qwen2.5-72B, Decoding Strategy=CoT-SC2026.06 | — | — | — | — | — | — | — | 66.8 | 29,472 | |
| TRI (Full)Fine-tuning Protocol=SFT + DPO + Repair2026.06 | — | — | — | — | — | — | — | 74.9 | 1,268 | |
| TRI (SFT + DPO)Fine-tuning Protocol=SFT + DPO2026.06 | — | — | — | — | — | — | — | 73.1 | 1,267 | |
| TRI (SFT only)Fine-tuning Protocol=SFT only2026.06 | — | — | — | — | — | — | — | 68.3 | 1,534 |