Reward Modeling on Human Evaluation
2.73Self-contain AccuracyOPENREWARD
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OPENREWARDBackbone=Qwen-2.5-7B-Instruct2025.10 | 2.73 | 2.83 | 2.78 | — | |
| LLM-as-JudgeBackbone=DeepSeek-V3.12025.10 | 2.67 | 2.47 | 2.57 | 7.55 | |
| RM-R1Backbone=Qwen-2.5-7B-Instruct2025.10 | 2.13 | 1.73 | 1.93 | 30.58 |