Reward Modeling Evaluation on Reward Bench Math 2
72.3Pairwise AccuracyDistribution-Calibrated Aggregation
Evaluation Results
| Method | Links | |
|---|---|---|
| Distribution-Calibrated Aggregationn=12, Judge LLM=gemini-2.5-flash2025.12 | 72.3 | |
| Distribution-Calibrated Aggregationn=4, Judge LLM=gemini-2.5-flash2025.12 | 70.9 | |
| SCn=4, Judge LLM=gemini-2.5-flash2025.12 | 65.8 | |
| Soft-SCn=12, Judge LLM=gemini-2.5-flash2025.12 | 65.4 | |
| SCn=12, Judge LLM=gemini-2.5-flash2025.12 | 63.5 | |
| CI-SCn=12, Judge LLM=gemini-2.5-flash2025.12 | 63.4 | |
| CI-SCn=4, Judge LLM=gemini-2.5-flash2025.12 | 63.2 | |
| Soft-SCn=4, Judge LLM=gemini-2.5-flash2025.12 | 62.6 | |
| USCn=12, Judge LLM=gemini-2.5-flash2025.12 | 61.9 | |
| USCn=4, Judge LLM=gemini-2.5-flash2025.12 | 61.6 | |
| GSCn=4, Judge LLM=gemini-2.5-flash2025.12 | 60.9 | |
| GSCn=12, Judge LLM=gemini-2.5-flash2025.12 | 60.5 |