Reasoning on GPQA Main (accuracy)
33.93Accuracy (GPQA Main)Rubric-grounded GRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| Rubric-grounded GRPOCheckpoint=Best by held-out rubric reward, Backbone=Llama-3.1-8B-Instruct2026.05 | 33.93 | |
| Llama-3.1-8B-InstructCheckpoint=Base2026.05 | 25.22 |