Coding on HumanEval+ (test)
67.7Pass@1Base
Evaluation Results
| Method | Links | |
|---|---|---|
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 67.7 | |
| KL-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 67.1 | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 67.1 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 65.9 | |
| DFTBackbone=Qwen2.5-3B-Instruct2026.05 | 65.9 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 65.9 | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 64.6 | |
| Self-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 64 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 62.2 | |
| FReDA-4BNumber of shots=0, Evaluation protocol=Best-of-N2026.06 | 59.76 | |
| FReDA-4BNumber of shots=0, Evaluation protocol=Self-refine2026.06 | 58.54 | |
| Qwen3-4B-BaseNumber of shots=02026.06 | 57.93 | |
| TiDAR-8BNumber of shots=0, Evaluation protocol=Trust Diff2026.06 | 55.49 | |
| TiDAR-8BNumber of shots=0, Evaluation protocol=Trust AR2026.06 | 52.44 | |
| BlockDiff-4BNumber of shots=02026.06 | 51.83 | |
| Dream-7B-BaseNumber of shots=02026.06 | 50 | |
| LLaDA-MoE-7B-A1B-BaseNumber of shots=02026.06 | 42.07 | |
| Qwen2.5-3B-BaseNumber of shots=02026.06 | 36 | |
| LLaDA-8B-BaseNumber of shots=02026.06 | 31.1 |