Reward Modeling Evaluation on Reward-Bench
84.79AgreementFairJudge-8B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| FairJudge-8BParameter Scale=8B2026.02 | 84.79 | 58.14 | 56.34 | 56.94 | |
| DeepSeek-V3-671BParameter Scale=671B2026.02 | 83.34 | 57.29 | 55.48 | 56.33 | |
| Qwen3-VL-8BParameter Scale=8B2026.02 | 82.08 | 55.99 | 54.71 | 55.34 | |
| Qwen3-8BParameter Scale=8B2026.02 | 77.25 | 53.17 | 51.67 | 52.11 | |
| InternVL3-8BParameter Scale=8B2026.02 | 76.38 | 54.91 | 50.96 | 52.8 | |
| FlexVL-7B*Parameter Scale=7B2026.02 | 75.55 | 50.49 | 50.41 | 50.41 | |
| InternVL3-14BParameter Scale=14B2026.02 | 75.23 | 58.26 | 50.03 | 53.7 | |
| GLM-4-9BParameter Scale=9B2026.02 | 68.24 | 46.68 | 45.54 | 46.05 | |
| PandaLM-7B*Parameter Scale=7B2026.02 | 53 | 54.53 | 53.77 | 51.3 | |
| JudgeLM-7B*Parameter Scale=7B2026.02 | 46.57 | 41.43 | 30.89 | 35.27 | |
| LLaVA-1.5-7BParameter Scale=7B2026.02 | 42.07 | 39.57 | 30.29 | 25.73 | |
| LLaVA-1.5-13BParameter Scale=13B2026.02 | 40.28 | 36.12 | 25.71 | 28.36 |