Reward Modeling on RewardBench Chat
96.4AccuracyClaude-3.5-Sonnet
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude-3.5-Sonnet2026.02 | 96.4 | |
| Claude-3.5-SonnetModel Category=Black-box LLMs2026.05 | 96.4 | |
| RM-R1-32B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Size=32B2026.02 | 95.3 | |
| RM-R1-32B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Size=32B2026.02 | 95.3 | |
| RM-R1-32B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Model Category=Larger White-box LLMs2026.05 | 95.3 | |
| RM-R1-32B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Model Category=Larger White-box LLMs2026.05 | 95.3 | |
| Gemini-2.5-Flash2026.02 | 95 | |
| Gemini-2.5-FlashModel Category=Black-box LLMs2026.05 | 95 | |
| RRM-32BSize=32B2026.02 | 94.7 | |
| RRM-32BModel Category=Larger White-box LLMs2026.05 | 94.7 | |
| JudgeLRM-7BSize=7B2026.02 | 92.1 | |
| JudgeLRM-7BModel Category=White-box Judge/Reward LLMs2026.05 | 92.1 | |
| Rubric-ARROW-voting@5Voting Strategy=voting@52026.05 | 90.8 | |
| RM-R1-14B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Size=14B2026.02 | 90.3 | |
| Rubric-ARM-voting@5voting=52026.02 | 90.3 | |
| RM-R1-14B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Model Category=Larger White-box LLMs2026.05 | 90.3 | |
| Rubric-RM-voting@5voting=52026.02 | 89.9 | |
| Rubric-RM-voting@5Voting Strategy=voting@52026.05 | 89.9 | |
| RIFL-voting@5Voting Strategy=voting@52026.05 | 89.7 | |
| API (direct Judge)Judge Model=Gemini-2.5-Flash Lite2026.02 | 89.6 | |
| API Prompting (direct Judge)Judge=Gemini-2.5-Flash-Lite, Evaluation Protocol=direct Judge2026.05 | 89.6 | |
| RIFLModel Category=Rubric-based Methods2026.05 | 89.6 | |
| Rubric-ARM2026.02 | 89.4 | |
| Rubric-ARROWModel Category=Rubric-based Methods2026.05 | 89.1 | |
| Rubric-RM2026.02 | 88.2 | |
| Rubric-RMModel Category=Rubric-based Methods2026.05 | 88.2 | |
| Rubric-ARROW w/o RL-voting@5Training Protocol=w/o RL, Voting Strategy=voting@52026.05 | 87.7 | |
| Rubric-ARROW w/o RLTraining Protocol=w/o RL2026.05 | 87.4 | |
| RM-R1-7B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Size=7B2026.02 | 85.3 | |
| RM-R1-7B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Model Category=White-box Judge/Reward LLMs2026.05 | 85.3 | |
| RM-R1-7B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Size=7B2026.02 | 83 | |
| RM-R1-7B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Model Category=White-box Judge/Reward LLMs2026.05 | 83 | |
| API (Rubric+Judge)Rubric Model=GPT-4.1-Mini, Judge Model=Gemini-2.5-Flash Lite2026.02 | 79.6 | |
| API Prompting (Rubric+Judge pairwise)Rubric Generator=GPT-4.1-Mini, Judge=Gemini-2.5-Flash-Lite, Evaluation Protocol=Rubric+Judge pairwise2026.05 | 79.6 | |
| RRM-7BSize=7B2026.02 | 77.7 | |
| RRM-7BModel Category=White-box Judge/Reward LLMs2026.05 | 77.7 | |
| Qwen-3-8B (Rubric+Judge)Backbone=Qwen-3-8B2026.02 | 73.9 | |
| Qwen-3-8B (Rubric+Judge pairwise)Backbone=Qwen-3-8B, Evaluation Protocol=Rubric+Judge pairwise2026.05 | 73.9 | |
| RM-R1-14B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Size=14B2026.02 | 73.5 | |
| RM-R1-14B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Model Category=Larger White-box LLMs2026.05 | 73.5 | |
| Qwen-3-8B (Rubric+Judge pointwise)Backbone=Qwen-3-8B, Evaluation Protocol=Rubric+Judge pointwise2026.05 | 69.3 | |
| API Prompting (Rubric+Judge pointwise)Rubric Generator=GPT-4.1-Mini, Judge=Gemini-2.5-Flash-Lite, Evaluation Protocol=Rubric+Judge pointwise2026.05 | 51.7 |