3D Appearance Preference Evaluation on Human Preference Evaluation Dataset Appearance
90AccuracyGen3DEval (CLIP)
Evaluation Results
| Method | Links | |
|---|---|---|
| Gen3DEval (CLIP)backbone=CLIP2025.04 | 90 | |
| Gen3DEval (CLIP + Fit3D)backbone=CLIP + Fit3D2025.04 | 90 | |
| Gen3DEval (CLIP)backbone=CLIP2025.04 | 89 | |
| Gen3DEval (CLIP + DinoV2)backbone=CLIP + DinoV22025.04 | 86 | |
| Gen3DEval w/ Fit3Dbackbone=Fit3D2025.04 | 81 | |
| Gen3DEval (CLIP + Fit3D)backbone=CLIP + Fit3D2025.04 | 78 | |
| Gen3DEval (CLIP + DinoV2)backbone=CLIP + DinoV22025.04 | 78 | |
| Gen3DEval w/ DinoV2backbone=DinoV22025.04 | 77 | |
| Image Reward Score2025.04 | 73 | |
| GPT-4omulti-image support=false2025.04 | 69 | |
| Image Reward Score2025.04 | 66 | |
| GPT-4omulti-image support=false, input_format=8 input images in sequence2025.04 | 59 | |
| Gen3DEval w/ Fit3Dbackbone=Fit3D2025.04 | 55 | |
| LLaVA-Qwen-7B2025.04 | 54 | |
| Phi-3.5-Vision2025.04 | 54 | |
| LLaVA-Qwen-7B2025.04 | 54 | |
| Gen3DEval w/ DinoV2backbone=DinoV22025.04 | 54 | |
| Phi-3.5-Vision2025.04 | 53 | |
| LLaVA-Llama3-8b2025.04 | 50 | |
| LLaVA-Llama3-8b2025.04 | 47 | |
| PickScore2025.04 | 37 | |
| PickScore2025.04 | 34 | |
| CLIP Score2025.04 | 30 | |
| PaliGemmamulti-image support=false2025.04 | 21 | |
| BLIPmulti-image support=false2025.04 | 20 | |
| CLIP Score2025.04 | 17 | |
| Llama3.2-Vision-11Bmulti-image support=false, input_format=4x2 grids or sequence2025.04 | 6 | |
| BLIPmulti-image support=false, input_format=4x2 grids or sequence2025.04 | 5 | |
| Llama3.2-Vision-11Bmulti-image support=false2025.04 | 5 | |
| PaliGemmamulti-image support=false, input_format=4x2 grids or sequence2025.04 | 2 |