Multimodal Reward Modeling on VL-RewardBench
83.5AccuracyDT2IT-MRM (Qwen3-VL)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| DT2IT-MRM (Qwen3-VL)#Param=8B, Category=Ours2026.04 | 83.5 | 77 | 63 | 91.3 | 76.7 | |
| DT2IT-MRM2026.04 | 83.5 | — | — | — | — | |
| DT2IT-MRM (Qwen2.5-VL)#Param=7B, Category=Ours2026.04 | 82.5 | 75.8 | 63.5 | 91.3 | 72.6 | |
| BaseReward (Qwen2.5-VL)#Param=7B, Category=Discriminative Multimodal Reward Models2026.04 | 82.2 | 80.9 | 68.6 | 92.2 | 81.8 | |
| BaseReward2026.04 | 82.2 | — | — | — | — | |
| MSRL + voting@16Params.=8B, Backbone=InternVL3.5-8B, Voting Strategy=voting@162026.03 | 77.5 | — | 76.5 | 79.3 | 76.8 | |
| EGT# Param=7B2026.02 | 77.15 | — | — | — | — | |
| MSRLParams.=8B, Backbone=InternVL3.5-8B2026.03 | 75.9 | — | 75.4 | 78.2 | 74.2 | |
| MR. Judge-7B-SFT-RL#Param=7B, Category=Slow-Thinking Generative MRMs (with critic training)2026.04 | 75.5 | 71.1 | 68.7 | 83.2 | 61.4 | |
| Proxy-GRM-RLProxy Agent=Proxy-SFT, Data Size=50k2026.03 | 75.22 | 73.93 | — | — | — | |
| Gemini-2.5-ProAccessibility=Proprietary2026.02 | 74.9 | 72.9 | 59.1 | 85.2 | 74.4 | |
| Unified-Reward-ThinkProxy Agent=None, Data Size=>200k2026.03 | 73.8 | 72.3 | — | — | — | |
| Gemini-3-ProParams.=-, Category=MLLM-as-a-Judge2026.03 | 73.4 | — | 54.6 | 78.9 | 86.7 | |
| Proxy-GRM-RLProxy Agent=Proxy-RL, Data Size=60k2026.03 | 73.38 | 72.38 | — | — | — | |
| Skywork-VL Reward#Param=7B, Category=Discriminative Multimodal Reward Models2026.04 | 73.1 | 69 | 66 | 80 | 61 | |
| Skywork-VL Reward2026.04 | 73.1 | — | — | — | — | |
| VL-MDR#Param=7B2026.04 | 73.06 | 71.84 | 71.27 | 75.17 | 69.09 | |
| UnifiedReward-Think*#Param=7B, Category=Slow-Thinking Generative MRMs (with critic training)2026.04 | 73 | 65.4 | 51.4 | 82.9 | 61.8 | |
| Skywork-VL-Reward#Param=7B2026.04 | 72.98 | 68.82 | 65.75 | 79.84 | 60.88 | |
| R1-Reward# Param=7B2026.02 | 72.89 | — | — | — | — | |
| UnifiedReward-think (w/ CoT)Params.=7B, Category=Open-Source Multimodal Reward Models2026.03 | 72.2 | — | 77.9 | 72.7 | 66 | |
| Proxy-GRM-RLProxy Agent=None, Data Size=45k2026.03 | 72.17 | 71.18 | — | — | — | |
| Omni-RewardModel-BT*#Param=8B, Category=Discriminative Multimodal Reward Models2026.04 | 72 | 64.8 | 51.4 | 81.4 | 61.5 | |
| R1-RewardProxy Agent=None, Data Size=>200k2026.03 | 71.92 | 71.44 | — | — | — | |
| Gemini-2.5-FlashAccessibility=Proprietary2026.02 | 71.9 | 72 | 58 | 77 | 81 | |
| UnifiedReward-Think#Param=7B2026.04 | 71.45 | 71.82 | 77.35 | 72.5 | 65.62 | |
| R1-RewardParams.=7B, Category=Open-Source Multimodal Reward Models2026.03 | 71.4 | — | 63.8 | 85.7 | 64.8 | |
| MSRL + voting@16Params.=4B, Backbone=InternVL3.5-4B, Voting Strategy=voting@162026.03 | 71.1 | — | 63.5 | 82.2 | 67.6 | |
| GPT-5-miniParams.=-, Category=MLLM-as-a-Judge2026.03 | 70.9 | — | 60.2 | 76.5 | 76 | |
| UnifiedRewardParams.=7B, Category=Open-Source Multimodal Reward Models2026.03 | 70.8 | — | 76.5 | 70.5 | 65.4 | |
| GPT-5.2*Category=Proprietary Models (w/o critic training)2026.04 | 70.2 | 67 | 52.5 | 71.8 | 76.7 | |
| GPT-5.22026.04 | 70.2 | — | — | — | — | |
| MSRLParams.=4B, Backbone=InternVL3.5-4B2026.03 | 69.7 | — | 62.4 | 80.5 | 66.3 | |
| Proxy-GRM-SFTProxy Agent=None, Data Size=10k2026.03 | 69.53 | 69.18 | — | — | — | |
| Gemini-1.5-Pro#Param=-2026.04 | 67.2 | 62.5 | 50.8 | 72.5 | 64.2 | |
| Gemini-1.5-Pro (2024-09-24)Category=Proprietary Models (w/o critic training)2026.04 | 67.2 | 62.5 | 50.8 | 72.5 | 64.2 | |
| Generative MRMParams.=8B, Backbone=InternVL3.5-8B2026.03 | 66.6 | — | 57.2 | 74.5 | 68.2 | |
| Claude-3.7-SonnetParams.=-, Category=MLLM-as-a-Judge2026.03 | 66.5 | — | 68.1 | 70.7 | 60.8 | |
| Claude-3.7-Sonnet (2025-02-24)2026.02 | 66.31 | — | — | — | — | |
| Claude-3.7-SonnetProxy Agent=None, Data Size=-2026.03 | 66.31 | 66.53 | — | — | — | |
| Claude 3.7 Sonnet#Param=-2026.04 | 66.31 | 66.53 | 68.08 | 70.7 | 60.81 | |
| Claude-3.7-SonnetCategory=Proprietary Models (w/o critic training)2026.04 | 66.3 | 66.5 | 68.1 | 70.7 | 60.8 | |
| IXC-2.5-Reward#Param=7B2026.04 | 66.16 | 68.55 | 80.11 | 65.29 | 60.25 | |
| Unified-Reward-SFTProxy Agent=None, Data Size=>200k2026.03 | 66.1 | 66.5 | — | — | — | |
| GPT-4o (2024-08-06)2026.02 | 65.8 | — | — | — | — | |
| IXC-2.5-Reward# Param=7B2026.02 | 65.8 | — | — | — | — | |
| GPT-4oAccessibility=Proprietary2026.02 | 65.8 | 62.4 | 49.6 | 67.6 | 70.5 | |
| GPT-4o-(2024-08-06)Proxy Agent=None, Data Size=-2026.03 | 65.8 | 62.4 | — | — | — | |
| IXC-2.5-RewardProxy Agent=None, Data Size=>200k2026.03 | 65.8 | 70 | — | — | — | |
| GPT-4o#Param=-2026.04 | 65.8 | 62.4 | 49.1 | 67.6 | 70.5 | |
| GPT-4o (2024-08-06)Category=Proprietary Models (w/o critic training)2026.04 | 65.8 | 62.4 | 49.1 | 67.6 | 70.5 | |
| IXC-2.5-Reward#Param=7B, Category=Discriminative Multimodal Reward Models2026.04 | 65.8 | 70 | 84.7 | 62.5 | 62.9 | |
| Discriminative MRMParams.=8B, Backbone=InternVL3.5-8B2026.03 | 64.3 | — | 54.8 | 75.6 | 62.4 | |
| InternVL3-78B#Param=78B, Category=Open-Source Models (w/o critic training)2026.04 | 63.3 | 61.6 | 67.8 | 52.5 | 64.5 | |
| UnifiedReward*#Param=7B, Category=Fast-Thinking Generative MRMs (with critic training)2026.04 | 63 | 58.8 | 45.9 | 66.6 | 64 | |
| UnifiedReward#Param=7B2026.04 | 62.79 | 66.61 | 76.24 | 58.61 | 64.98 | |
| Gemini-1.5-ProParams.=-, Category=MLLM-as-a-Judge2026.03 | 62.5 | — | 50.8 | 72.5 | 64.2 | |
| GPT-4oParams.=-, Category=MLLM-as-a-Judge2026.03 | 62.4 | — | 49.1 | 67.6 | 70.5 | |
| Discriminative MRMParams.=4B, Backbone=InternVL3.5-4B2026.03 | 61.2 | — | 54.4 | 66.8 | 62.4 | |
| Generative MRMParams.=4B, Backbone=InternVL3.5-4B2026.03 | 60.5 | — | 56.6 | 65.4 | 59.4 | |
| InternVL3-78B#Param=78B2026.04 | 57.98 | 62.15 | 69.61 | 52.47 | 64.35 | |
| Gemini-1.5-Flash#Param=-2026.04 | 57.6 | 55.3 | 47.8 | 59.6 | 58.4 | |
| PhyCritic-7BAccessibility=Open-Source, Parameters=7B, Base Model=Qwen2.5-VL-7B-Instruct2026.02 | 57.3 | 54.9 | 45.3 | 58.6 | 60.9 | |
| InternVL3-8B#Param=8B, Category=Open-Source Models (w/o critic training)2026.04 | 57 | 55.6 | 60.6 | 44 | 62.3 | |
| Llama-3.2-90BProxy Agent=None, Data Size=-2026.03 | 56.2 | 53.9 | — | — | — | |
| Llama-3.2-90B#Param=90B2026.04 | 56.2 | 53.9 | 42.6 | 57.3 | 61.7 | |
| Llama-3.2-90B#Param=90B, Category=Open-Source Models (w/o critic training)2026.04 | 56.2 | 53.9 | 42.6 | 57.3 | 61.7 | |
| Claude-3.5-Sonnet-(2024-06-22)Proxy Agent=None, Data Size=-2026.03 | 55.3 | 53.6 | — | — | — | |
| Claude 3.5 Sonnet#Param=-2026.04 | 55.3 | 53.6 | 43.4 | 55 | 62.3 | |
| Claude-3.5-Sonnet (2024-06-22)Category=Proprietary Models (w/o critic training)2026.04 | 55.3 | 53.6 | 43.4 | 55 | 62.3 | |
| InternVL3.5-14BParams.=14B, Category=MLLM-as-a-Judge2026.03 | 54.5 | — | 56.8 | 65.4 | 41.4 | |
| Qwen2.5-VL-7BAccessibility=Open-Source, Parameters=7B2026.02 | 53.2 | 50.9 | 40.9 | 54.3 | 57.4 | |
| InternVL3.5-8BParams.=8B, Category=MLLM-as-a-Judge2026.03 | 52.7 | — | 53.3 | 68 | 36.8 | |
| Qwen3-VL-8B-InstructParams.=8B, Category=MLLM-as-a-Judge2026.03 | 51.7 | — | 39 | 52 | 64.2 | |
| Qwen2.5-VL-72B-Instruct#Param=72B, Category=Open-Source Models (w/o critic training)2026.04 | 51.6 | 52.7 | 47.8 | 46.8 | 63.5 | |
| Qwen2.5-VL-72B#Param=72B2026.04 | 51.16 | 52.73 | 48.07 | 46.73 | 63.41 | |
| MM-RLHF-RewardParams.=7B, Category=Open-Source Multimodal Reward Models2026.03 | 51 | — | 45 | 50.5 | 57.6 | |
| InternVL3-8B#Param=8B2026.04 | 51 | 55.54 | 60.22 | 43.93 | 62.46 | |
| R1-Reward*#Param=7B, Category=Slow-Thinking Generative MRMs (with critic training)2026.04 | 50.8 | 46.6 | 43.1 | 57.5 | 39.1 | |
| LLaVA-CriticParams.=7B, Category=Open-Source Multimodal Reward Models2026.03 | 50.7 | — | 54.6 | 38.3 | 59.1 | |
| Eagle-2.5-8BAccessibility=Open-Source, Parameters=8B2026.02 | 50.2 | 49.7 | 41.4 | 48.6 | 59.3 | |
| MM-RLHF-Reward#Param=7B, Category=Semi-Scalar Multimodal Reward Models2026.04 | 50.2 | 51 | 45 | 50.5 | 57.6 | |
| MM-RLHF-Reward# Param=7B2026.02 | 50.15 | — | — | — | — | |
| MM-RLHF-Reward#Param=7B2026.04 | 50.15 | 51.01 | 45.04 | 50.45 | 57.55 | |
| Qwen2.5-VL-7B-Instruct#Param=7B, Category=Open-Source Models (w/o critic training)2026.04 | 48 | 49.5 | 43.4 | 42 | 63 | |
| Cosmos-R1-7BAccessibility=Open-Source, Parameters=7B2026.02 | 44.8 | 44.8 | 33.1 | 41.3 | 59.9 | |
| Llama-3.2-11B#Param=11B2026.04 | 42.9 | 42.8 | 33.3 | 38.4 | 56.6 | |
| Robobrain2.0-7BAccessibility=Open-Source, Parameters=7B2026.02 | 42.4 | 44.2 | 39.2 | 55.8 | 37.5 | |
| GPT-4o-mini#Param=-2026.04 | 41.5 | 44.8 | 41.7 | 34.5 | 58.2 | |
| LLaVA-Critic#Param=7B2026.04 | 41.2 | 44 | 54.6 | 38.3 | 59.1 | |
| LLaVA-Critic#Param=8B, Category=Fast-Thinking Generative MRMs (with critic training)2026.04 | 41.2 | 44 | 54.6 | 38.3 | 59.1 | |
| NVLM-D-72BProxy Agent=None, Data Size=-2026.03 | 40.1 | 44.1 | — | — | — | |
| Qwen2-VL-72B# Param=72B2026.02 | 39.5 | — | — | — | — | |
| Qwen2-VL-72BProxy Agent=None, Data Size=-2026.03 | 39.5 | 43 | — | — | — | |
| Qwen2-VL-72B#Param=72B2026.04 | 39.5 | 43 | 38.1 | 32.8 | 58 | |
| Qwen2.5-VL-7B#Param=7B2026.04 | 31.92 | 36.86 | 34.25 | 21.76 | 54.57 | |
| LLaVA-OneVision-7B#Param=7B2026.04 | 29.6 | 36.5 | 32.2 | 20.1 | 57.1 | |
| Qwen2-VL-7B#Param=7B2026.04 | 28.3 | 33.9 | 31.6 | 19.1 | 51.1 | |
| SliME# Param=7B2026.02 | 19.04 | — | — | — | — | |
| SliMEProxy Agent=None, Data Size=-2026.03 | 19.04 | 17.64 | — | — | — |