Process Reward Modeling on ProcessBench
65.7GSM8K AccuracyLCA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LCATraining Base=Qwen2.5-Math-7B-Instruct, alpha=12026.06 | 65.7 | 48.7 | 26.2 | 19.9 | 40.2 | |
| Qwen2.5-Math-7B-Math-Shepherd-PRMModel Category=Open Source Process Reward Models2026.06 | 62.5 | 31.6 | 13.7 | 7.7 | 28.9 | |
| Qwen2.5-Math-7B-Math-ShepherdBackbone=Qwen2.5-Math, Parameters=7B2025.08 | 62.5 | 31.6 | 13.7 | 7.7 | — | |
| SCANTraining Base=Qwen2.5-Math-7B-Instruct2026.06 | 60.6 | 33.9 | 15 | 6.8 | 29.1 | |
| MathShepherdTraining Base=Qwen2.5-Math-7B-Instruct2026.06 | 59.2 | 31.1 | 12.7 | 8.1 | 27.8 | |
| PQMTraining Base=Qwen2.5-Math-7B-Instruct2026.06 | 55.7 | 43.9 | 28.5 | 24.5 | 38.1 | |
| ImplicitPRMTraining Base=Llama3.2-3B-Instruct2026.06 | 54.8 | 34.1 | 22.2 | 12.4 | 30.9 | |
| RLHFlow-PRM-Mistral-8BModel Category=Open Source Process Reward Models2026.06 | 50.4 | 33.4 | 13.8 | 15.8 | 28.4 | |
| LCATraining Base=Llama3.2-3B-Instruct, alpha=12026.06 | 48.4 | 42.1 | 30.3 | 29.4 | 37.6 | |
| Math-Shepherd-PRM-7BModel Category=Open Source Process Reward Models2026.06 | 47.9 | 29.5 | 24.8 | 23.8 | 31.5 | |
| Math-Shepherd-PRM-7BModel Family=Math-Shepherd, Parameters=7B2025.08 | 47.9 | 29.5 | 24.8 | 23.8 | — | |
| SCANTraining Base=Llama3.2-3B-Instruct2026.06 | 47.6 | 27.6 | 16.2 | 12.2 | 25.9 | |
| EurusPRM-Stage2Model Category=Open Source Process Reward Models2026.06 | 47.3 | 35.7 | 21.2 | 20.9 | 31.3 | |
| BiPRMBackbone=Qwen2.5-Math-1.5B2025.08 | 46.8 | 37.9 | 32.9 | 28 | — | |
| ImplicitPRMTraining Base=Qwen2.5-Math-7B-Instruct2026.06 | 45.2 | 36.3 | 23.5 | 25.1 | 32.5 | |
| OmegaPRMTraining Base=Qwen2.5-Math-7B-Instruct2026.06 | 44.8 | 37.1 | 22.9 | 18.4 | 30.8 | |
| EurusPRM-Stage1Model Category=Open Source Process Reward Models2026.06 | 44.3 | 35.6 | 21.7 | 23.1 | 31.2 | |
| Qwen2.5-Math-RM-72BBackbone=Qwen2.5-Math-RM, Parameters=72B2025.08 | 43.5 | 47.2 | 37.6 | 27.4 | — | |
| RLHFlow-PRM-Deepseek-8BModel Category=Open Source Process Reward Models2026.06 | 38.8 | 33.8 | 16.9 | 16.9 | 26.6 | |
| MathShepherdTraining Base=Llama3.2-3B-Instruct2026.06 | 38.8 | 22.7 | 14.5 | 11.4 | 21.8 | |
| RLHFlow-PRM-Deepseek-8BBackbone=Deepseek, Parameters=8B2025.08 | 38.8 | 33.8 | 16.9 | 16.9 | — | |
| OmegaPRMTraining Base=Llama3.2-3B-Instruct2026.06 | 38.7 | 30.1 | 18.2 | 14.1 | 25.3 | |
| Qwen2.5-7B-InstructModel Category=Large Language Models as Critic2026.06 | 36.5 | 36.6 | 29.7 | 27.4 | 32.6 | |
| PQMTraining Base=Llama3.2-3B-Instruct2026.06 | 34.4 | 38.3 | 20.2 | 27.9 | 30.2 | |
| Qwen2.5-Math-7B-InstructModel Category=Large Language Models as Critic2026.06 | 26.8 | 25.7 | 14.2 | 12.7 | 19.9 | |
| Qwen2.5-Coder-7B-InstructModel Category=Large Language Models as Critic2026.06 | 14.3 | 6.5 | 4.1 | 1.8 | 6.7 | |
| Llama-3-8B-InstructModel Category=Large Language Models as Critic2026.06 | 13.1 | 13.8 | 4.8 | 12.6 | 11.1 | |
| Llama-3.1-8B-InstructModel Category=Large Language Models as Critic2026.06 | 10.9 | 5.1 | 2.8 | 1.6 | 5.1 |