Reward Modeling on Aggregated Benchmarks Macro
74.74Average Score (excl. MM-RB, VL-RB)Claude 4.6 Opus
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Claude 4.6 OpusModel Category=Proprietary Frontier VLMs2026.04 | 74.74 | 76.17 | 76.01 | |
| Qwen3-VL-32B-InstructModel Category=Open-source Generalist VLMs, Prompting Strategy=Rubric2026.04 | 73.25 | 75.05 | 74.81 | |
| Qwen3-VL-32B-InstructModel Category=Open-source Generalist VLMs, Prompting Strategy=Vanilla2026.04 | 73.14 | 73.9 | 74.09 | |
| Qwen3-VL-32B-InstructModel Category=Open-source Generalist VLMs, Prompting Strategy=CoT2026.04 | 72.97 | 74.45 | 75.06 | |
| Qwen3-VL-30B-A3B-InstructModel Category=Open-source Generalist VLMs, Prompting Strategy=Rubric2026.04 | 72.85 | 74.83 | 74.91 | |
| Qwen3-VL-30B-A3B-InstructModel Category=Open-source Generalist VLMs, Prompting Strategy=Vanilla2026.04 | 72.44 | 71.73 | 70.69 | |
| Gemini 3.0 FlashModel Category=Proprietary Frontier VLMs2026.04 | 72.23 | — | — | |
| Gemini 3.1 ProModel Category=Proprietary Frontier VLMs2026.04 | 72.05 | — | — | |
| GPT-5.4Model Category=Proprietary Frontier VLMs2026.04 | 71.45 | 74.91 | 76.84 | |
| Qwen3-VL-30B-A3B-InstructModel Category=Open-source Generalist VLMs, Prompting Strategy=CoT2026.04 | 71.07 | 69.18 | 67.64 | |
| Gemini 3.1 Flash-LiteModel Category=Proprietary Frontier VLMs2026.04 | 68.55 | — | — | |
| Skywork-VL RewardModel Category=Open-source Specialist RMs2026.04 | 65.56 | 68.89 | 69.24 | |
| IXC-2.5-RewardModel Category=Open-source Specialist RMs2026.04 | — | — | 68.3 |