Human Preference Prediction on HPD v3
78.3AccuracyGemini 3.1 Pro + ARR
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 3.1 Pro + ARRModel Category=ARR (Ours), Evaluation Protocol=ARR2026.05 | 78.3 | |
| HPSv3Model Category=Trained Reward Model2026.05 | 76.9 | |
| Gemini 3.1 ProModel Category=VLM-as-Judge (Direct), Evaluation Protocol=Direct2026.05 | 76.6 | |
| GPT-5 + ARRModel Category=ARR (Ours), Evaluation Protocol=ARR2026.05 | 76.1 | |
| GPT-5Model Category=VLM-as-Judge (Direct), Evaluation Protocol=Direct2026.05 | 72.4 | |
| Qwen3vl-8B + ARRModel Category=ARR (Ours), Evaluation Protocol=ARR2026.05 | 70.2 | |
| UnifiedReward-ThinkingModel Category=Trained Reward Model2026.05 | 68.1 | |
| Qwen3-VL-8BModel Category=VLM-as-Judge (Direct), Evaluation Protocol=Direct2026.05 | 67.2 | |
| UnifiedRewardModel Category=Trained Reward Model2026.05 | 66 | |
| PickScoreModel Category=Trained Reward Model2026.05 | 65.6 | |
| HPSv2Model Category=Trained Reward Model2026.05 | 65.3 | |
| CMMSPreference Model=CMMS2026.03 | 61.3 | |
| ImageRewardModel Category=Trained Reward Model2026.05 | 58.6 | |
| DEQAPreference Model=DEQA2026.03 | 52.7 | |
| QUALIPreference Model=QUALI2026.03 | 51.2 | |
| MDIQAPreference Model=MDIQA2026.03 | 51.1 | |
| CLIP-IQAPreference Model=CLIP-IQA2026.03 | 48.9 | |
| Q-AlignPreference Model=Q-Align2026.03 | 48.9 | |
| CLIP-ScorePreference Model=CLIP-Score2026.03 | 48.6 | |
| CLIPScoreModel Category=Trained Reward Model2026.05 | 48.6 | |
| MUSIQPreference Model=MUSIQ2026.03 | 39.4 |