Best-of-N evaluation on RewardBench v2
58.69AccuracyPC2-based LLM-as-a-Judge
Evaluation Results
| Method | Links | |
|---|---|---|
| PC2-based LLM-as-a-JudgeEvaluation Method=ours2025.05 | 58.69 | |
| Naive Pointwise EvaluationEvaluation Method=naive2025.05 | 55.51 |
| Method | Links | |
|---|---|---|
| PC2-based LLM-as-a-JudgeEvaluation Method=ours2025.05 | 58.69 | |
| Naive Pointwise EvaluationEvaluation Method=naive2025.05 | 55.51 |