Text-to-image preference evaluation on MM-RewardBench T2I 2
78.9AccuracyGemini 3.1 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 3.1 ProARR=true2026.05 | 78.9 | |
| Gemini 3.1 ProARR=false2026.05 | 75.1 | |
| GPT-5ARR=true2026.05 | 74.7 | |
| GPT-5ARR=false2026.05 | 70.5 | |
| UnifiedReward-Thinking2026.05 | 66 | |
| Qwen3-VL-8BARR=true2026.05 | 62.7 | |
| HPSv32026.05 | 60.2 | |
| UnifiedReward2026.05 | 59.8 | |
| PickScore2026.05 | 58.6 | |
| Qwen3-VL-8BARR=false2026.05 | 57.6 | |
| ImageReward2026.05 | 54 |