Reward Modeling Evaluation on Reward Bench Safety 2
72.3Pairwise AccuracyDistribution-Calibrated Aggregation
Evaluation Results
| Method | Links | |
|---|---|---|
| Distribution-Calibrated Aggregationn=12, Judge LLM=gemini-2.5-flash2025.12 | 72.3 | |
| Distribution-Calibrated Aggregationn=4, Judge LLM=gemini-2.5-flash2025.12 | 69.1 | |
| SCn=4, Judge LLM=gemini-2.5-flash2025.12 | 65 | |
| Soft-SCn=4, Judge LLM=gemini-2.5-flash2025.12 | 63.5 | |
| Soft-SCn=12, Judge LLM=gemini-2.5-flash2025.12 | 63.3 | |
| SCn=12, Judge LLM=gemini-2.5-flash2025.12 | 63 | |
| CI-SCn=12, Judge LLM=gemini-2.5-flash2025.12 | 62.9 | |
| CI-SCn=4, Judge LLM=gemini-2.5-flash2025.12 | 62.6 | |
| GSCn=4, Judge LLM=gemini-2.5-flash2025.12 | 62.5 | |
| USCn=12, Judge LLM=gemini-2.5-flash2025.12 | 62.3 | |
| USCn=4, Judge LLM=gemini-2.5-flash2025.12 | 61.9 | |
| GSCn=12, Judge LLM=gemini-2.5-flash2025.12 | 61.8 |