Video Preference Evaluation on ShareGPT-Video
80.2AccuracyOmni-RRM
Evaluation Results
| Method | Links | |
|---|---|---|
| Omni-RRMModel Scale=7B, Training Stage=sft+rl, Rationale Supervision=yes2026.01 | 80.2 | |
| Gemini-2.5-Pro2026.01 | 78.8 | |
| UnifiedReward-think-7B2026.01 | 77.8 | |
| Doubao-1.5-Vision-Pro2026.01 | 77 | |
| Gemini-2.0-Flash2026.01 | 74.6 | |
| Qwen2.5-VLModel Scale=72B2026.01 | 72.9 | |
| Qwen2.5-VLModel Scale=7B2026.01 | 70.5 | |
| Omni-RRMModel Scale=7B, Training Stage=sft, Rationale Supervision=yes2026.01 | 70.5 | |
| Omni-RMModel Scale=7B, Training Stage=sft+rl, Rationale Supervision=no2026.01 | 68.7 | |
| Omni-RRMModel Scale=3B, Training Stage=sft+rl, Rationale Supervision=yes2026.01 | 67.4 | |
| Qwen2.5-OmniModel Scale=7B2026.01 | 66.3 | |
| Omni-RRMModel Scale=3B, Training Stage=sft, Rationale Supervision=yes2026.01 | 64.9 | |
| Qwen2.5-VLModel Scale=3B2026.01 | 61.2 | |
| Omni-RMModel Scale=3B, Training Stage=sft+rl, Rationale Supervision=no2026.01 | 60 | |
| Skywork-VL-Reward-7B2026.01 | 59.9 | |
| R1-Reward-7B2026.01 | 58.7 | |
| Qwen2.5-OmniModel Scale=3B2026.01 | 58.1 | |
| GPT-4o-mini2026.01 | 53.9 |