Reward Modeling on OPENRM-LIB v1 (test)
93Accuracy (Wikipedia)OPENREWARD-Qwen-2.5-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OPENREWARD-Qwen-2.5-7BEvaluation Protocol=Agentic Reward Modeling2025.10 | 93 | 90 | 91 | 91.33 | |
| OPENREWARD-Qwen-3-4BEvaluation Protocol=Agentic Reward Modeling2025.10 | 80.98 | 82.4 | 82.6 | 81.99 | |
| Deepseek-V3.1Evaluation Protocol=Agentic Reward Modeling2025.10 | 77 | 48 | 34 | 53 | |
| GPT-4oEvaluation Protocol=Agentic Reward Modeling2025.10 | 76.4 | 58.6 | 53.4 | 62.8 | |
| Claude-Opus-4-1-20250805Evaluation Protocol=Agentic Reward Modeling2025.10 | 75.6 | 56.1 | 58 | 63.23 | |
| Deepseek-V3.1Evaluation Protocol=LLM-as-a-Judge2025.10 | 75 | 46 | 33 | 51.33 | |
| Claude-Opus-4-1-20250805Evaluation Protocol=LLM-as-a-Judge2025.10 | 74.6 | 49.2 | 51.1 | 58.3 | |
| Gemini-2.5-proEvaluation Protocol=Agentic Reward Modeling2025.10 | 72.5 | 54.6 | 42.4 | 56.5 | |
| Gemini-2.5-ProEvaluation Protocol=LLM-as-a-Judge2025.10 | 72.2 | 46.6 | 36 | 51.6 | |
| GPT-4oEvaluation Protocol=LLM-as-a-Judge2025.10 | 70 | 48.2 | 44 | 54.07 | |
| RM-R1Evaluation Protocol=LLM-as-a-Judge, Training Data Configuration=OPENRM-LIB-27K2025.10 | 66 | 73 | 65 | 68 | |
| RRM-7BEvaluation Protocol=LLM-as-a-Judge2025.10 | 56.9 | 52.95 | 53.1 | 54.32 | |
| RM-R1-Qwen2.5-Instruct-7BEvaluation Protocol=LLM-as-a-Judge2025.10 | 55.4 | 54.8 | 52.3 | 54.17 | |
| JudgeLRM-7BEvaluation Protocol=LLM-as-a-Judge2025.10 | 50.8 | 50.6 | 48.44 | 49.94 | |
| Skywork-Reward-Gemma-2-27BEvaluation Protocol=LLM-as-a-Judge2025.10 | 45.2 | 55.4 | 47.74 | 49.45 |