Best-of-N Reranking on Average of 7 Benchmarks (AIME24, LeetCode) (test)
52Average AccuracyDeepSeek-R1-Distill-Qwen-32B
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-R1-Distill-Qwen-32BEvaluator Type=Reasoning Process + Outcome Evaluators, N=82025.03 | 52 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluator Type=Reasoning Outcome Evaluators, N=82025.03 | 51.1 | |
| Qwen2.5-Math-PRM-72BEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 50.6 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluator Type=Reasoning Process Evaluators, N=82025.03 | 50.3 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Reasoning Process + Outcome Evaluators, N=82025.03 | 50.1 | |
| Skywork-o1-Open-PRM-Qwen-2.5-7BEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 49.9 | |
| Qwen2.5-Math-PRM-7BEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 48.7 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Reasoning Outcome Evaluators, N=82025.03 | 48.7 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluator Type=Reasoning Process + Outcome Evaluators, N=42025.03 | 48.5 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Reasoning Process Evaluators, N=82025.03 | 48.3 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluator Type=Single-step Reasoning Process Evaluators, N=82025.03 | 48.1 | |
| Skywork-o1-Open-PRM-Qwen-2.5-1.5BEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 46.7 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Reasoning Process + Outcome Evaluators, N=42025.03 | 46.6 | |
| Skywork-Reward-Gemma-2-27B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=322025.03 | 45.6 | |
| Qwen2.5-72B-InstructEvaluator Type=Non-Reasoning Generative Evaluators, N=642025.03 | 45.6 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Single-step Reasoning Process Evaluators, N=82025.03 | 45.6 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=162025.03 | 45.5 | |
| Skywork-Reward-Gemma-2-27B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=162025.03 | 45.5 | |
| Skywork-Reward-Gemma-2-27B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=642025.03 | 45.4 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=322025.03 | 45.2 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=642025.03 | 44.8 | |
| Skywork-Reward-Gemma-2-27B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=82025.03 | 44.8 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=82025.03 | 44.6 | |
| Qwen2.5-Math-7B-PRM800KEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 44.6 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluator Type=Reasoning Process + Outcome Evaluators, N=22025.03 | 44.4 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=42025.03 | 43.4 | |
| Skywork-Reward-Gemma-2-27B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=42025.03 | 43.4 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Reasoning Process + Outcome Evaluators, N=22025.03 | 43.3 | |
| math-shepherd-mistral-7b-prmEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 43.2 | |
| Llama3-8B-CLoud-RMEvaluator Type=Non-Reasoning Generative Evaluators, N=642025.03 | 42.5 | |
| Skywork-Reward-Gemma-2-27B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=22025.03 | 41.8 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=22025.03 | 41.4 | |
| prometheus-8x7b-v2.0Evaluator Type=Non-Reasoning Generative Evaluators, N=642025.03 | 40.8 | |
| prometheus-7b-v2.0Evaluator Type=Non-Reasoning Generative Evaluators, N=642025.03 | 39.7 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=12025.03 | 38.2 | |
| Skywork-Reward-Gemma-2-27B-v0.2Evaluator Type=Direct Outcome Evaluators (ORMs), N=12025.03 | 38.2 | |
| math-shepherd-mistral-7b-prmEvaluator Type=Direct Process Evaluators (PRMs), N=12025.03 | 38.2 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Reasoning Outcome Evaluators, N=12025.03 | 38.2 | |
| DeepSeek-R1-Distill-Qwen-7BEvaluator Type=Reasoning Process + Outcome Evaluators, N=12025.03 | 38.2 | |
| DeepSeek-R1-Distill-Qwen-32BEvaluator Type=Reasoning Process + Outcome Evaluators, N=12025.03 | 38.2 | |
| RLHFlow/Llama3.1-8B-PRM-DeepseekEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 37.8 | |
| RLHFlow/Llama3.1-8B-PRM-MistralEvaluator Type=Direct Process Evaluators (PRMs), N=642025.03 | 35.5 |