Aesthetic Quality Assessment on AVA v1 (test)
0.883Kendall's TauRARL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RARLBackbone=Qwen2.5-VL-7B, Training strategy=RARL, Number of images=4-imgs2026.01 | 0.883 | — | |
| RARLBackbone=Qwen2.5-VL-7B, Training strategy=RARL, Number of images=2-imgs2026.01 | 0.875 | — | |
| RARLBackbone=Qwen2.5-VL-7B, Training strategy=RARL, Number of images=8-imgs2026.01 | 0.871 | — | |
| RARLBackbone=Qwen2.5-VL-3B, Training strategy=RARL, Number of images=2-imgs2026.01 | 0.84 | — | |
| RARLBackbone=Qwen2.5-VL-3B, Training strategy=RARL, Number of images=4-imgs2026.01 | 0.832 | — | |
| Qwen2.5-VL-7B + SFTBackbone=Qwen2.5-VL-7B, Training strategy=SFT, Number of images=8-imgs2026.01 | 0.823 | — | |
| RARLBackbone=Qwen2.5-VL-3B, Training strategy=RARL, Number of images=8-imgs2026.01 | 0.821 | — | |
| Qwen2.5-VL-7B + SFTBackbone=Qwen2.5-VL-7B, Training strategy=SFT, Number of images=2-imgs2026.01 | 0.815 | — | |
| Qwen2.5-VL-7B + SFTBackbone=Qwen2.5-VL-7B, Training strategy=SFT, Number of images=4-imgs2026.01 | 0.813 | — | |
| Qwen2.5-VL-3B + SFTBackbone=Qwen2.5-VL-3B, Training strategy=SFT, Number of images=8-imgs2026.01 | 0.806 | — | |
| Qwen2.5-VL-3B + SFTBackbone=Qwen2.5-VL-3B, Training strategy=SFT, Number of images=4-imgs2026.01 | 0.801 | — | |
| Qwen2.5-VL-3B + SFTBackbone=Qwen2.5-VL-3B, Training strategy=SFT, Number of images=2-imgs2026.01 | 0.77 | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, Training strategy=Baseline, Number of images=8-imgs2026.01 | 0.731 | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, Training strategy=Baseline, Number of images=4-imgs2026.01 | 0.723 | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, Training strategy=Baseline, Number of images=2-imgs2026.01 | 0.714 | — | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B, Training strategy=Baseline, Number of images=4-imgs2026.01 | 0.677 | — | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B, Training strategy=Baseline, Number of images=2-imgs2026.01 | 0.656 | — | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B, Training strategy=Baseline, Number of images=8-imgs2026.01 | 0.634 | — | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B, Training strategy=Baseline, Number of images=1-img2026.01 | — | 0.527 | |
| Qwen2.5-VL-3B + SFTBackbone=Qwen2.5-VL-3B, Training strategy=SFT, Number of images=1-img2026.01 | — | 0.755 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, Training strategy=Baseline, Number of images=1-img2026.01 | — | 0.643 | |
| Qwen2.5-VL-7B + SFTBackbone=Qwen2.5-VL-7B, Training strategy=SFT, Number of images=1-img2026.01 | — | 0.771 | |
| RankingCLIPNumber of images=1-img2026.01 | — | 0.747 | |
| RARLBackbone=Qwen2.5-VL-3B, Training strategy=RARL, Number of images=1-img2026.01 | — | 0.783 | |
| RARLBackbone=Qwen2.5-VL-7B, Training strategy=RARL, Number of images=1-img2026.01 | — | 0.803 | |
| VILA-RNumber of images=1-img2026.01 | — | 0.774 |