Video Understanding Reward Modeling on VURB
65.6General Video UnderstandingVideoDRM
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| VideoDRMSetting=Pointwise2026.05 | 65.6 | 62.8 | 62.2 | 63.8 | |
| GPT 5.2Setting=Pairwise2026.05 | 62.6 | 63.1 | 63 | 62.9 | |
| Seed1.6-VL-ThinkingSetting=Pairwise2026.05 | 61.6 | 60.3 | 69 | 62.6 | |
| SkyWork-VL-RewardSetting=Pointwise2026.05 | 58.8 | 58.4 | 59.8 | 58.9 | |
| MiMo-VL-7B-RL-2508Setting=Pairwise2026.05 | 57.1 | 52.9 | 62 | 56.4 | |
| Qwen3VL-PlusSetting=Pairwise2026.05 | 56.5 | 61.5 | 62.7 | 59.9 | |
| UnifiedReward-3.0-Qwen(8B)[Pointwise]Setting=Pointwise2026.05 | 56.5 | 52 | 55.7 | 54.5 | |
| VideoGRMSetting=Pairwise2026.05 | 54.7 | 62 | 61.8 | 59.3 | |
| Qwen3VL-32B-ThinkingSetting=Pairwise2026.05 | 54.5 | 53.5 | 56.6 | 54.5 | |
| UnifiedReward-3.0-Qwen(8B)[Pairwise]Setting=Pairwise2026.05 | 52.9 | 55.1 | 54.1 | 54 | |
| InternVL-3.5-38BSetting=Pairwise2026.05 | 52.8 | 53 | 56.7 | 53.7 | |
| UnifiedReward-Thinking-3.0-Qwen(8B)Setting=Pairwise2026.05 | 52.7 | 53 | 59.5 | 54.3 | |
| Qwen3VL-8B-InstructSetting=Pairwise2026.05 | 51.3 | 51 | 57.1 | 52.4 | |
| Qwen3VL-8B-ThinkingSetting=Pairwise2026.05 | 50.4 | 52.8 | 57.5 | 53 | |
| InternVL-3.5-8BSetting=Pairwise2026.05 | 50.1 | 50.7 | 55.6 | 51.5 | |
| R1-Reward(8B)Setting=Pairwise2026.05 | 48.8 | 49.5 | 51.9 | 49.8 | |
| LLaVA-Critic-R1(7B)Setting=Pairwise2026.05 | 46.3 | 47.5 | 51.5 | 47.9 | |
| Flex-Judge(8B)Setting=Pairwise2026.05 | 40.6 | 32.6 | 40.4 | 37.3 |