Reward Modeling on RewardBench 2 (test)
76.3RWBench2 ScoreCE-RM-4B
Evaluation Results
| Method | Links | |
|---|---|---|
| CE-RM-4BRM Type=Pointwise, Training Data=5.7K, Scaling=42026.01 | 76.3 | |
| Gemini-2.5-FlashRM Type=Pointwise, Training Data=-2026.01 | 75.9 | |
| CE-RM-4BRM Type=Pointwise, Training Data=5.7K, Scaling=22026.01 | 75.9 | |
| CE-RM-4BRM Type=Pointwise, Training Data=5.7K, Scaling=12026.01 | 74.6 | |
| TIR-Judge-Zero-8BRM Type=Pointwise, Training Data=26K2026.01 | 73.4 | |
| TIR-Judge-Distill-8BRM Type=Pointwise, Training Data=26K2026.01 | 71.6 | |
| TIR-Judge-Zero-4BRM Type=Pointwise, Training Data=26K2026.01 | 68.3 | |
| TIR-Judge-Distill-4BRM Type=Pointwise, Training Data=26K2026.01 | 67.3 | |
| CompassJudger1-32BRM Type=Pointwise, Training Data=900K2026.01 | 56.5 |