Code Generation on HumanEval (r* pass and r_self)
73.4r* Pass RatePhi / Base
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Phi / BaseSampling=Power, Evaluation Protocol=Self-reward Best-of-N2026.05 | 73.4 | -0.294 | |
| Phi / DistilledSampling=Temperature, Evaluation Protocol=All completions2026.05 | 71.5 | -0.627 | |
| Phi / BaseSampling=Power, Evaluation Protocol=All completions2026.05 | 71.2 | -0.33 | |
| Phi / DistilledSampling=Temperature, Evaluation Protocol=Self-reward Best-of-N2026.05 | 67.5 | -0.447 | |
| Phi / DistilledSampling=Standard, Evaluation Protocol=All completions2026.05 | 63.4 | -0.73 | |
| Phi / DistilledSampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 60.2 | -0.473 | |
| Qwen / DistilledSampling=Temperature, Evaluation Protocol=Self-reward Best-of-N2026.05 | 60 | -0.235 | |
| Qwen / BaseSampling=Power, Evaluation Protocol=All completions2026.05 | 57.3 | -0.13 | |
| Qwen / BaseSampling=Power, Evaluation Protocol=Self-reward Best-of-N2026.05 | 56.8 | -0.096 | |
| Qwen-Math / DistilledSampling=Temperature, Evaluation Protocol=Self-reward Best-of-N2026.05 | 56.6 | -0.208 | |
| Qwen-Math / BaseSampling=Power, Evaluation Protocol=Self-reward Best-of-N2026.05 | 56.2 | -0.106 | |
| Phi / BaseSampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 56.2 | -0.589 | |
| Phi / BaseSampling=Standard, Evaluation Protocol=All completions2026.05 | 54.9 | -0.913 | |
| Qwen-Math / DistilledSampling=Temperature, Evaluation Protocol=All completions2026.05 | 54.1 | -0.304 | |
| Qwen / DistilledSampling=Temperature, Evaluation Protocol=All completions2026.05 | 54.1 | -0.479 | |
| Qwen-Math / BaseSampling=Power, Evaluation Protocol=All completions2026.05 | 53.8 | -0.144 | |
| Qwen / DistilledSampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 47 | -0.325 | |
| Qwen-Math / DistilledSampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 45.2 | -0.334 | |
| Qwen / DistilledSampling=Standard, Evaluation Protocol=All completions2026.05 | 42.5 | -0.849 | |
| Qwen-Math / DistilledSampling=Standard, Evaluation Protocol=All completions2026.05 | 41.6 | -0.563 | |
| Qwen-Math / BaseSampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 38.3 | -0.427 | |
| Qwen / BaseSampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 37.6 | -0.426 | |
| Qwen / BaseSampling=Standard, Evaluation Protocol=All completions2026.05 | 32.6 | -0.966 | |
| Qwen-Math / BaseSampling=Standard, Evaluation Protocol=All completions2026.05 | 32 | -0.741 |