Reward Modeling on PPE Correctness (test)
75PPE CorrCE-RM-4B
Evaluation Results
| Method | Links | |
|---|---|---|
| CE-RM-4BRM Type=Pointwise, Scaling=42026.01 | 75 | |
| CE-RM-4BRM Type=Pointwise, Scaling=22026.01 | 73.1 | |
| J1-Llama-70BRM Type=Pairwise2026.01 | 72.9 | |
| TIR-Judge-Distill-8BRM Type=Pointwise2026.01 | 71 | |
| TIR-Judge-Zero-8BRM Type=Pointwise2026.01 | 70.3 | |
| TIR-Judge-Zero-4BRM Type=Pointwise2026.01 | 69.8 | |
| CE-RM-4BRM Type=Pointwise, Scaling=12026.01 | 69.7 | |
| RRM-32BRM Type=Pairwise2026.01 | 67.9 | |
| TIR-Judge-Distill-4BRM Type=Pointwise2026.01 | 65.9 | |
| Llama-3.3-70B-InstructRM Type=Pairwise2026.01 | 65.7 | |
| J1-Llama-70BRM Type=Pointwise2026.01 | 65 | |
| CLoud-Gemma-2-27BRM Type=Pointwise2026.01 | 62.4 | |
| Gemini-2.5-FlashRM Type=Pointwise2026.01 | 61.9 | |
| RISE-Judge-7BRM Type=Pairwise2026.01 | 60.4 | |
| CompassJudger2-7BRM Type=Pairwise2026.01 | 60.2 | |
| RRM-7BRM Type=Pairwise2026.01 | 60.1 | |
| DeepSeek-GRM-27BRM Type=Pointwise2026.01 | 59.8 | |
| RM-R1-Qwen-32BRM Type=Pairwise2026.01 | 59.3 | |
| J1-Llama-8BRM Type=Pairwise2026.01 | 59.2 | |
| RM-R1-Qwen-7BRM Type=Pairwise2026.01 | 57.9 | |
| GPT-4oRM Type=Pairwise2026.01 | 57.6 | |
| CompassJudger2-32BRM Type=Pairwise2026.01 | 56.6 | |
| RISE-Judge-32BRM Type=Pairwise2026.01 | 55 | |
| J1-Llama-8BRM Type=Pointwise2026.01 | 53.8 | |
| CompassJudger1-32BRM Type=Pointwise2026.01 | 48 | |
| JudgeLRM-7BRM Type=Pairwise2026.01 | 42.6 |