Reward Modeling on JudgeBench (Category Scores)
74.6KnowledgeFlexible Principles GenRM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Flexible Principles GenRMRM Type (Scalar vs Generative)=Generative, Latency (seconds/task)=>10 seconds/task2025.09 | 74.6 | 85.7 | 85.7 | 90.5 | 81.4 | |
| Flexible Principles ScalarRMRM Type (Scalar vs Generative)=Scalar, Latency (seconds/task)=<0.1 second/task2025.09 | 74 | 74.5 | 82.1 | 81 | 76.3 | |
| Llama-3.3-Nemotron-Super-49B-GenRMRM Type (Scalar vs Generative)=Generative, Latency (seconds/task)=>10 seconds/task2025.09 | 71.4 | 73.5 | 87.5 | 76.2 | 75.1 | |
| Llama-3.3-Nemotron-70B-RewardRM Type (Scalar vs Generative)=Scalar, Latency (seconds/task)=<0.1 second/task2025.09 | 70.8 | 76.5 | 82.1 | 66.7 | 73.7 | |
| Bradley-TerryRM Type (Scalar vs Generative)=Scalar, Latency (seconds/task)=<0.1 second/task2025.09 | 63 | 69.4 | 82.1 | 71.4 | 68.9 | |
| URM-Llama-3.1-8BBackbone=Llama-3.1-8B2026.01 | 62.3 | 67.4 | 76.8 | 47.6 | 64.3 | |
| Llama-3.1-Nemotron-70B-RewardRM Type (Scalar vs Generative)=Scalar, Latency (seconds/task)=<0.1 second/task2025.09 | 62.3 | 72.5 | 76.8 | 57.1 | 66.9 | |
| Qwen3-4B + AdaJudgeBackbone=Qwen3-4B, Aggregation Strategy=AdaJudge2026.01 | 61.7 | 58.2 | 80.4 | 81 | 66 | |
| Qwen3-8B + last token poolingBackbone=Qwen3-8B, Aggregation Strategy=last token pooling2026.01 | 61.7 | 54.1 | 78.6 | 69 | 63.1 | |
| RewardAnything-8B-v1RM Type (Scalar vs Generative)=Generative, Latency (seconds/task)=>10 seconds/task2025.09 | 61 | 57.1 | 73.2 | 66.7 | 62.6 | |
| Skywork-Reward-Gemma-2-27BBackbone=Gemma-2-27B2026.01 | 59.7 | 66.3 | 83.9 | 50 | 64.3 | |
| Skywork-Reward-Llama-3.1-8BBackbone=Llama-3.1-8B2026.01 | 59.1 | 64.3 | 76.8 | 50 | 62.3 | |
| Qwen3-8B + AdaJudgeBackbone=Qwen3-8B, Aggregation Strategy=AdaJudge2026.01 | 58.4 | 64.3 | 80.4 | 78.6 | 66 | |
| Ray2333/GRM-llama3-8B-distillBackbone=Llama-3-8B-distill2026.01 | 57.1 | 66.3 | 78.6 | 54.8 | 62 | |
| Qwen3-4B + last token poolingBackbone=Qwen3-4B, Aggregation Strategy=last token pooling2026.01 | 57.1 | 62.2 | 76.8 | 61.9 | 62.3 | |
| Phi-3.5-mini-instruct + AdaJudgeBackbone=Phi-3.5-mini-instruct, Aggregation Strategy=AdaJudge2026.01 | 56.5 | 57.1 | 64.3 | 52.4 | 57.4 | |
| Qwen3-4B + mean poolingBackbone=Qwen3-4B, Aggregation Strategy=mean pooling2026.01 | 56.5 | 65.3 | 69.6 | 81 | 64 | |
| RM-R1-DeepSeek-Distilled-Qwen-32BRM Type (Scalar vs Generative)=Generative, Latency (seconds/task)=>10 seconds/task2025.09 | 56.5 | 66.3 | 85.7 | 73.8 | 66 | |
| Qwen3-8B + mean poolingBackbone=Qwen3-8B, Aggregation Strategy=mean pooling2026.01 | 55.2 | 64.3 | 76.8 | 73.8 | 63.4 | |
| Phi-3.5-mini-instruct + mean poolingBackbone=Phi-3.5-mini-instruct, Aggregation Strategy=mean pooling2026.01 | 53.9 | 51 | 62.5 | 42.9 | 53.1 | |
| R3-QWEN3-14B-LORA-4KRM Type (Scalar vs Generative)=Generative, Latency (seconds/task)=>10 seconds/task2025.09 | 50 | 64.3 | 76.8 | 71.4 | 60.9 | |
| Phi-3.5-mini-instruct + last token poolingBackbone=Phi-3.5-mini-instruct, Aggregation Strategy=last token pooling2026.01 | 45.5 | 46.9 | 28.6 | 45.2 | 43.1 |