Multi-modal Preference Evaluation on VL-Reward
79.6AccuracyGemini-2.5-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-2.5-Pro2026.01 | 79.6 | |
| Doubao-1.5-Vision-Pro2026.01 | 77.3 | |
| Gemini-2.0-Flash2026.01 | 73.4 | |
| Omni-RRMModel Scale=7B, Training Stage=sft+rl, Rationale Supervision=yes2026.01 | 67.1 | |
| UnifiedReward-think-7B2026.01 | 66.6 | |
| R1-Reward-7B2026.01 | 65.8 | |
| Omni-RMModel Scale=7B, Training Stage=sft+rl, Rationale Supervision=no2026.01 | 64 | |
| Qwen2.5-VLModel Scale=72B2026.01 | 62.3 | |
| Skywork-VL-Reward-7B2026.01 | 60.4 | |
| Omni-RRMModel Scale=7B, Training Stage=sft, Rationale Supervision=yes2026.01 | 60.4 | |
| GPT-4o-mini2026.01 | 59.8 | |
| Omni-RRMModel Scale=3B, Training Stage=sft+rl, Rationale Supervision=yes2026.01 | 58.5 | |
| Qwen2.5-VLModel Scale=7B2026.01 | 58.2 | |
| Qwen2.5-OmniModel Scale=7B2026.01 | 57.8 | |
| Omni-RMModel Scale=3B, Training Stage=sft+rl, Rationale Supervision=no2026.01 | 56.9 | |
| Omni-RRMModel Scale=3B, Training Stage=sft, Rationale Supervision=yes2026.01 | 56.8 | |
| LLaVA-Critic-7B2026.01 | 54.1 | |
| Qwen2.5-OmniModel Scale=3B2026.01 | 53.7 | |
| Qwen2.5-VLModel Scale=3B2026.01 | 53.2 |