Preference evaluation on ImageReward
67.5AccuracyMPS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MPS2026.06 | 67.5 | — | |
| DiT-Reward2026.06 | 67 | — | |
| HPSv32026.06 | 66.8 | — | |
| HPSv22026.06 | 65.7 | — | |
| ImageReward2026.06 | 65.1 | — | |
| PickScore2026.06 | 61.6 | — | |
| HPS2026.06 | 61.2 | — | |
| Q-Real-ScoreTraining Dataset=subset of Q-Real derived from Q-Eval-100K2025.11 | 60.7 | — | |
| Q-Eval-ScoreTraining Dataset=subset of Q-Real derived from Q-Eval-100K2025.11 | 60.1 | — | |
| Aesthetic Score Predictor2026.06 | 57.4 | — | |
| CLIP ViT-H/14Backbone=ViT-H/142026.06 | 57.1 | — | |
| Q-AlignTraining Dataset=subset of Q-Real derived from Q-Eval-100K2025.11 | 55.6 | — | |
| IPCETraining Dataset=subset of Q-Real derived from Q-Eval-100K2025.11 | 49.5 | — | |
| CLIP-IQATraining Dataset=subset of Q-Real derived from Q-Eval-100K2025.11 | 49.3 | — | |
| APO-imageJudge Model=Llama-4-Scout-17B-16E-instruct2026.02 | 37 | 31 | |
| BLPOJudge Model=Llama-4-Scout-17B-16E-instruct2026.02 | 36 | 34 | |
| BLPOJudge Model=Llama-4-Maverick-17B-128E-instruct2026.02 | 35 | 32 | |
| BLPOJudge Model=Qwen2.5-VL-32B-instruct2026.02 | 34 | 29 | |
| OPROJudge Model=Llama-4-Scout-17B-16E-instruct2026.02 | 34 | 32 | |
| OPROJudge Model=Llama-4-Maverick-17B-128E-instruct2026.02 | 33 | 31 | |
| No Optim.Judge Model=Llama-4-Scout-17B-16E-instruct2026.02 | 29 | 21 | |
| TextGradJudge Model=Llama-4-Maverick-17B-128E-instruct2026.02 | 29 | 27 | |
| TextGradJudge Model=Qwen2.5-VL-32B-instruct2026.02 | 28 | 26 | |
| OPROJudge Model=Qwen2.5-VL-32B-instruct2026.02 | 27 | 24 | |
| APO-imageJudge Model=Qwen2.5-VL-32B-instruct2026.02 | 27 | 23 | |
| APO-imageJudge Model=Llama-4-Maverick-17B-128E-instruct2026.02 | 26 | 25 | |
| No Optim.Judge Model=Qwen2.5-VL-32B-instruct2026.02 | 25 | 19 | |
| TextGradJudge Model=Llama-4-Scout-17B-16E-instruct2026.02 | 24 | 27 | |
| No Optim.Judge Model=Llama-4-Maverick-17B-128E-instruct2026.02 | 22 | 14 |