Reward Modeling on PPE Preference ZH
82.3AccuracyOpenRS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OpenRSModel Category=LLM-as-a-Judge & Rubric System, Base Model=Qwen3-235B-A22B-Instruct-25072026.02 | 82.3 | 89.4 | |
| Skywork-Reward-V2-Llama-3.1-8BModel Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 79.6 | 84.3 | |
| OpenRSModel Category=LLM-as-a-Judge & Rubric System, Base Model=DeepSeek-V3.22026.02 | 78.9 | 86.7 | |
| OpenRSModel Category=LLM-as-a-Judge & Rubric System, Base Model=Qwen3-30B-A3B-Instruct-25072026.02 | 78.4 | 84.9 | |
| OpenRSModel Category=LLM-as-a-Judge & Rubric System, Base Model=DeepSeek-V3.12026.02 | 78.1 | 86.2 | |
| Auto-RubricModel Category=LLM-as-a-Judge & Generative Reward Models & Rubric, Base Model=Qwen3-235B-A22B-Instruct-2507, Re-implementation=true2026.02 | 78 | — | |
| OpenRSModel Category=LLM-as-a-Judge & Rubric System, Base Model=gpt-oss-120b2026.02 | 76.6 | 86.8 | |
| RM-R1-Qwen-Instruct-32BModel Category=LLM-as-a-Judge & Generative Reward Models & Rubric, Re-implementation=true2026.02 | 76.3 | — | |
| RM-R1-DeepSeek-Distill-Qwen-32BModel Category=LLM-as-a-Judge & Generative Reward Models & Rubric, Re-implementation=true2026.02 | 75.3 | — | |
| DeepSeek-GRM-27B (w/ MetaRM)Model Category=LLM-as-a-Judge & Generative Reward Models & Rubric, Re-implementation=true2026.02 | 72.2 | — | |
| DeepSeek-GRM-27BModel Category=LLM-as-a-Judge & Generative Reward Models & Rubric, Re-implementation=true2026.02 | 71.1 | — | |
| Internlm2-20b-rewardModel Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 69.4 | 64.8 | |
| Llama-3.1-Nemotron-70BModel Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 68.7 | 70.5 | |
| INF-ORM-Llama3.1-70BModel Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 65.7 | 71.7 | |
| Skywork-Reward-Llama-3.1-8B-v0.2Model Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 62.2 | 67.1 | |
| Llama-3-OffsetBias-RM-8BModel Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 60.6 | 65 | |
| ArmoRM-Llama3-8B-v0.1Model Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 58.3 | 63.4 | |
| Skywork-Reward-Gemma-2-27B-v0.2Model Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 57.6 | 67.3 | |
| LDL-Reward-Gemma-2-27B-v0.1Model Category=Open Scalar Reward Model, Re-implementation=true2026.02 | 43.3 | 67.4 |