Reward Modeling on HelpSteer 3
83.15Accuracyo3-2025-04-16
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| o3-2025-04-16Model Category=Base Models (LLM-as-a-Judge)2026.01 | 83.15 | — | — | — | |
| RM-NLHF-Qwen-32BModel Category=Our Generative Reward Models, Backbone=Qwen-32B2026.01 | 83.15 | — | — | — | |
| gpt-5-2025-08-07Model Category=Base Models (LLM-as-a-Judge)2026.01 | 82.45 | — | — | — | |
| gemini-2.5-proModel Category=Base Models (LLM-as-a-Judge)2026.01 | 82.07 | — | — | — | |
| qwen3-maxModel Category=Base Models (LLM-as-a-Judge)2026.01 | 81.02 | — | — | — | |
| INF-ORM-Llama3.1-70BModel Category=Scalar Reward Models, Backbone=Llama-3.1-70B2026.01 | 80.75 | — | — | — | |
| claude-3-7-sonnet-20250219Model Category=Base Models (LLM-as-a-Judge)2026.01 | 80.42 | — | — | — | |
| qwen-plus-latestModel Category=Base Models (LLM-as-a-Judge)2026.01 | 80.38 | — | — | — | |
| URM-LLaMa-3.1-8BModel Category=Scalar Reward Models, Backbone=Llama-3.1-8B2026.01 | 80.12 | — | — | — | |
| Skywork-Reward-Llama-3.1-8B-v0.2Model Category=Scalar Reward Models, Backbone=Llama-3.1-8B2026.01 | 79.5 | — | — | — | |
| RRM-Qwen-32BModel Category=Specialized Generative Reward Models, Backbone=Qwen-32B2026.01 | 79.42 | — | — | — | |
| gpt-4o-latestModel Category=Base Models (LLM-as-a-Judge)2026.01 | 78.96 | — | — | — | |
| deepseek-r1-0528Model Category=Base Models (LLM-as-a-Judge)2026.01 | 78.27 | — | — | — | |
| RM-R1-Qwen-32BModel Category=Specialized Generative Reward Models, Backbone=Qwen-32B2026.01 | 78.18 | — | — | — | |
| ArmoRM-Llama3-8B-v0.1Model Category=Scalar Reward Models, Backbone=Llama3-8B2026.01 | 76.4 | — | — | — | |
| RM-R1-32B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Size=32B2026.02 | 75.6 | — | — | — | |
| RM-R1-32B (DeepSeek-Dist)Backbone=DeepSeek-Dist2026.05 | 75.6 | — | — | — | |
| RRM-32BSize=32B2026.02 | 75.4 | — | — | — | |
| RRM-32BModel Category=Larger White-box LLMs2026.05 | 75.4 | — | — | — | |
| deepseek-v3Model Category=Base Models (LLM-as-a-Judge)2026.01 | 75.22 | — | — | — | |
| RM-R1-14B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Size=14B2026.02 | 74.8 | — | — | — | |
| RM-R1-14B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst2026.05 | 74.8 | — | — | — | |
| RM-R1-14B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Size=14B2026.02 | 74.6 | — | — | — | |
| RM-R1-14B (DeepSeek-Dist)Backbone=DeepSeek-Dist2026.05 | 74.6 | — | — | — | |
| RM-NLHF-Qwen-7BModel Category=Our Generative Reward Models, Backbone=Qwen-7B2026.01 | 73.81 | — | — | — | |
| R1-Distill-Qwen-32BModel Category=Base Models (LLM-as-a-Judge), Backbone=Qwen-32B2026.01 | 73.76 | — | — | — | |
| RM-R1-32B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Size=32B2026.02 | 72.9 | — | — | — | |
| RM-R1-32B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst2026.05 | 72.9 | — | — | — | |
| Rubric-ARROW-voting@5Voting Strategy=voting@52026.05 | 72 | — | — | — | |
| API (Rubric+Judge)Rubric Model=GPT-4.1-Mini, Judge Model=Gemini-2.5-Flash Lite2026.02 | 71.4 | — | — | — | |
| API Prompting (Rubric+Judge pairwise)Evaluation Protocol=Rubric+Judge pairwise2026.05 | 71.4 | — | — | — | |
| Rubric-ARM-voting@5voting=52026.02 | 71.1 | — | — | — | |
| Gemini-2.5-Flash2026.02 | 70.6 | — | — | — | |
| Gemini-2.5-FlashModel Category=Black-box LLMs2026.05 | 70.6 | — | — | — | |
| API (direct Judge)Judge Model=Gemini-2.5-Flash Lite2026.02 | 70.3 | — | — | — | |
| API Prompting (direct Judge)Evaluation Protocol=direct Judge2026.05 | 70.3 | — | — | — | |
| Rubric-ARROWModel Category=Rubric-based Methods2026.05 | 70.2 | — | — | — | |
| Rubric-ARM2026.02 | 69.8 | — | — | — | |
| RRM-Qwen-7BModel Category=Specialized Generative Reward Models, Backbone=Qwen-7B2026.01 | 67.94 | — | — | — | |
| RIFLModel Category=Rubric-based Methods2026.05 | 67.7 | — | — | — | |
| RIFL-voting@5Voting Strategy=voting@52026.05 | 67.7 | — | — | — | |
| Rubric-RM-voting@5voting=52026.02 | 67.5 | — | — | — | |
| Rubric-RM-voting@5Voting Strategy=voting@52026.05 | 67.5 | — | — | — | |
| Rubric-ARROW w/o RLTraining Protocol=w/o RL2026.05 | 67.5 | — | — | — | |
| Rubric-ARROW w/o RL-voting@5Voting Strategy=voting@5, Training Protocol=w/o RL2026.05 | 67.5 | — | — | — | |
| Rubric-RM2026.02 | 67 | — | — | — | |
| Rubric-RMModel Category=Rubric-based Methods2026.05 | 67 | — | — | — | |
| RM-R1-7B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst, Size=7B2026.02 | 65.2 | — | — | — | |
| RM-R1-7B (Qwen-2.5-Inst)Backbone=Qwen-2.5-Inst2026.05 | 65.2 | — | — | — | |
| RM-R1-Qwen-7BModel Category=Specialized Generative Reward Models, Backbone=Qwen-7B2026.01 | 64.99 | — | — | — | |
| R1-Distill-Llama-8BModel Category=Base Models (LLM-as-a-Judge), Backbone=Llama-8B2026.01 | 64 | — | — | — | |
| RM-R1-7B (DeepSeek-Dist)Backbone=DeepSeek-Dist, Size=7B2026.02 | 62.6 | — | — | — | |
| RM-R1-7B (DeepSeek-Dist)Backbone=DeepSeek-Dist2026.05 | 62.6 | — | — | — | |
| RRM-7BSize=7B2026.02 | 62.4 | — | — | — | |
| RRM-7BModel Category=White-box Judge/Reward LLMs2026.05 | 62.4 | — | — | — | |
| Qwen-3-8B (Rubric+Judge)Backbone=Qwen-3-8B2026.02 | 61.8 | — | — | — | |
| Qwen-3-8B (Rubric+Judge pairwise)Backbone=Qwen-3-8B, Evaluation Protocol=Rubric+Judge pairwise2026.05 | 61.8 | — | — | — | |
| JudgeLRM-7BSize=7B2026.02 | 60.2 | — | — | — | |
| JudgeLRM-7BModel Category=White-box Judge/Reward LLMs2026.05 | 60.2 | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7BModel Category=Base Models (LLM-as-a-Judge), Backbone=Qwen-7B2026.01 | 59.99 | — | — | — | |
| Qwen-3-8B (Rubric+Judge pointwise)Backbone=Qwen-3-8B, Evaluation Protocol=Rubric+Judge pointwise2026.05 | 55.5 | — | — | — | |
| API Prompting (Rubric+Judge pointwise)Evaluation Protocol=Rubric+Judge pointwise2026.05 | 39.2 | — | — | — | |
| Claude-Sonnet-4.5Category=LLM-as-Judge2026.02 | — | 78.4 | 15.7 | 66 | |
| Claude-Sonnet-4.5-thinkingCategory=LLM-as-Judge, Mode=thinking2026.02 | — | 79.9 | 12.5 | 69.9 | |
| DeepSeek-V3.2-chatCategory=LLM-as-Judge, Mode=chat2026.02 | — | 75.9 | 24.9 | 57 | |
| DeepSeek-V3.2-thinkingCategory=LLM-as-Judge, Mode=thinking2026.02 | — | 77.5 | 29.5 | 54 | |
| Gemini-2.5-ProCategory=LLM-as-Judge, Tier=Pro2026.02 | — | 78.4 | 12.4 | 68.6 | |
| Gemini-3-ProCategory=LLM-as-Judge, Tier=Pro2026.02 | — | 78.1 | 5.6 | 73.7 | |
| GenRM-R-Align-14BCategory=Our Methods, Parameters=14B, Training=R-Align2026.02 | — | 76.3 | 29.2 | 54 | |
| GenRM-R-Align-8BCategory=Our Methods, Parameters=8B, Training=R-Align2026.02 | — | 73.1 | 34.6 | 47.8 | |
| GenRM-RLVR-14BCategory=Our Methods (Baseline), Parameters=14B, Training=RLVR2026.02 | — | 75.5 | 46.9 | 40.1 | |
| GenRM-RLVR-8BCategory=Our Methods (Baseline), Parameters=8B, Training=RLVR2026.02 | — | 72.9 | 44.6 | 40.4 | |
| GPT-5-chatCategory=LLM-as-Judge, Mode=chat2026.02 | — | 77.5 | 20.5 | 61.5 | |
| GPT-5-thinkingCategory=LLM-as-Judge, Mode=thinking2026.02 | — | 76 | 11.7 | 67.1 | |
| GPT-OSS-120BCategory=LLM-as-Judge, Parameters=120B2026.02 | — | 76.5 | 37.9 | 47.4 | |
| Qwen3-14BCategory=Our Methods (Baseline), Parameters=14B2026.02 | — | 74.1 | 40 | 44.5 | |
| Qwen3-4B-Instruct-2507Category=LLM-as-Judge, Parameters=4B, Mode=Instruct2026.02 | — | 72.8 | 44.8 | 40.2 | |
| Qwen3-4B-Thinking-2507Category=LLM-as-Judge, Parameters=4B, Mode=Thinking2026.02 | — | 72.5 | 34.9 | 47.2 | |
| Qwen3-8BCategory=Our Methods (Baseline), Parameters=8B2026.02 | — | 72.9 | 45.8 | 39.5 | |
| RM-R1-DS-32BCategory=Specialized Generative Reward Models, Parameters=32B2026.02 | — | 73.2 | 50.7 | 36.1 | |
| RM-R1-Qwen-32BCategory=Specialized Generative Reward Models, Parameters=32B2026.02 | — | 75.1 | 32.9 | 50.4 | |
| RRM-32BCategory=Specialized Generative Reward Models, Parameters=32B2026.02 | — | 74.7 | 59 | 30.6 |