Video Reward Modeling on VideoRewardBench
76.2Perception (long)VideoGRM
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| VideoGRMModel Type=Slow-Thinking Generative MRMs (with critic training), #Param=8B2026.05 | 76.2 | 54.1 | 58.2 | 60.3 | 73.5 | 64.6 | 63.9 | |
| VideoDRMModel Type=Discriminative Multimodal Reward Models, #Param=8B2026.05 | 73.9 | 56.2 | 54.2 | 56.8 | 80.3 | 64.7 | 63.3 | |
| IXC-2.5-RewardModel Type=Discriminative Multimodal Reward Models, #Param=7B2026.05 | 73.5 | 51.3 | 56.3 | 52.2 | 38.7 | 53.4 | 54.4 | |
| LLaVA-Critic-72B (LLaVA-OV-72B)Model Type=Fast-Thinking Generative MRMs (with critic training), #Param=72B2026.05 | 72.4 | 43.8 | 55.9 | 56.5 | 88 | 63 | 63.3 | |
| Gemini-2.5-ProModel Type=Proprietary Models (w/o critic training), #Param=-2026.05 | 70.7 | 55.9 | 65.5 | 67.3 | 62.7 | 63.6 | 64.4 | |
| InternVL3-78BModel Type=Open-Source Models (w/o critic training), #Param=78B2026.05 | 70 | 49.2 | 57.1 | 50 | 65.8 | 58 | 58.4 | |
| Qwen2.5-VL-72BModel Type=Open-Source Models (w/o critic training), #Param=72B2026.05 | 68.9 | 48.4 | 56.7 | 52.5 | 44.7 | 53.3 | 54.3 | |
| UnifiedReward (LLaVA-OV-7B)Model Type=Fast-Thinking Generative MRMs (with critic training), #Param=7B2026.05 | 67.1 | 48.2 | 50.4 | 45.3 | 71.2 | 56.6 | 56.5 | |
| Skywork-VL RewardModel Type=Discriminative Multimodal Reward Models, #Param=7B2026.05 | 65.7 | 49.2 | 52.9 | 54 | 80.1 | 60.5 | 60.4 | |
| Claude-3.7-Sonnet (2025-02-19)Model Type=Proprietary Models (w/o critic training), #Param=-2026.05 | 65 | 48.4 | 63.4 | 58.3 | 82.9 | 63.2 | 63.6 | |
| LLaVA-OneVision-72BModel Type=Open-Source Models (w/o critic training), #Param=72B2026.05 | 64.7 | 40.9 | 59.7 | 53.6 | 73.5 | 57.6 | 58.5 | |
| GPT-4oModel Type=Proprietary Models (w/o critic training), #Param=-2026.05 | 63.3 | 50.8 | 58.8 | 57.9 | 57.3 | 57 | 57.6 | |
| UnifiedReward-Think (LLaVA-OV-7B)Model Type=Slow-Thinking Generative MRMs (with critic training), #Param=7B2026.05 | 59.7 | 53.3 | 50.5 | 52.9 | 55.6 | 54.4 | 54.3 | |
| MM-RLHF-Reward(LLaVA-OV-7B)Model Type=Semni-Scalar Multimodal Reward Models, #Param=7B2026.05 | 59.4 | 37 | 44.1 | 52.2 | 65.2 | 51.2 | 51.6 | |
| R1-RewardModel Type=Slow-Thinking Generative MRMs (with critic training), #Param=7B2026.05 | 36 | 40 | 37.8 | 30.6 | 47.9 | 39 | 38.4 | |
| Flex_JudgeModel Type=Slow-Thinking Generative MRMs (with critic training), #Param=7B2026.05 | 35 | 35.1 | 37 | 37.1 | 30.2 | 34.6 | 34.9 |