Science Reasoning on GPQA (r* Accuracy and r_self)
30.9r* AccuracyPhi
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PhiVariant=Base, Sampling=Power, Evaluation Protocol=All completions2026.05 | 30.9 | 0.215 | |
| PhiVariant=Base, Sampling=Power, Evaluation Protocol=Self-reward Best-of-N2026.05 | 30.9 | 0.214 | |
| PhiVariant=Distilled, Sampling=Temperature, Evaluation Protocol=Self-reward Best-of-N2026.05 | 29.8 | 0.262 | |
| PhiVariant=Distilled, Sampling=Temperature, Evaluation Protocol=All completions2026.05 | 29.2 | 0.284 | |
| QwenVariant=Distilled, Sampling=Temperature, Evaluation Protocol=Self-reward Best-of-N2026.05 | 29.1 | 0.201 | |
| QwenVariant=Base, Sampling=Power, Evaluation Protocol=Self-reward Best-of-N2026.05 | 28.7 | 0.118 | |
| QwenVariant=Distilled, Sampling=Temperature, Evaluation Protocol=All completions2026.05 | 28.5 | 0.21 | |
| QwenVariant=Base, Sampling=Power, Evaluation Protocol=All completions2026.05 | 28.3 | 0.118 | |
| Qwen-MathVariant=Distilled, Sampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 28.1 | 0.113 | |
| QwenVariant=Distilled, Sampling=Standard, Evaluation Protocol=All completions2026.05 | 28 | 0.437 | |
| Qwen-MathVariant=Base, Sampling=Power, Evaluation Protocol=Self-reward Best-of-N2026.05 | 27.9 | 0.087 | |
| QwenVariant=Distilled, Sampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 27.8 | 0.426 | |
| Qwen-MathVariant=Base, Sampling=Power, Evaluation Protocol=All completions2026.05 | 27.7 | 0.088 | |
| Qwen-MathVariant=Distilled, Sampling=Temperature, Evaluation Protocol=All completions2026.05 | 27.7 | 0.149 | |
| Qwen-MathVariant=Distilled, Sampling=Temperature, Evaluation Protocol=Self-reward Best-of-N2026.05 | 27.7 | 0.109 | |
| Qwen-MathVariant=Distilled, Sampling=Standard, Evaluation Protocol=All completions2026.05 | 27.5 | 0.165 | |
| PhiVariant=Distilled, Sampling=Standard, Evaluation Protocol=All completions2026.05 | 26.8 | 0.321 | |
| PhiVariant=Distilled, Sampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 26.7 | 0.319 | |
| QwenVariant=Base, Sampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 24.5 | 1.527 | |
| QwenVariant=Base, Sampling=Standard, Evaluation Protocol=All completions2026.05 | 24.4 | 1.531 | |
| PhiVariant=Base, Sampling=Standard, Evaluation Protocol=All completions2026.05 | 22.3 | 0.802 | |
| PhiVariant=Base, Sampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 22.3 | 0.8 | |
| Qwen-MathVariant=Base, Sampling=Standard, Evaluation Protocol=Self-reward Best-of-N2026.05 | 10.3 | 0.675 | |
| Qwen-MathVariant=Base, Sampling=Standard, Evaluation Protocol=All completions2026.05 | 10 | 0.675 |