UGC Quality Assessment on CASTER-Bench 1.0 (test)
60.3HQ PrecisionMEDEA
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| MEDEAModel Category=Ours2026.06 | 60.3 | 70.5 | 65 | 85 | 84.5 | 84.7 | 72.7 | 77.5 | 74.9 | |
| GPT-5.2Model Category=Flagship Models with Social-CoT Simulation, Reasoning Protocol=social-CoT2026.06 | 44.2 | 30.4 | 36 | 76.2 | 85.3 | 80.5 | 60.2 | 57.8 | 58.2 | |
| GPT-5.2Model Category=Reasoning-Enhanced LMMs, Reasoning Protocol=Long-CoT2026.06 | 40.1 | 90.3 | 55.5 | 92.8 | 48.3 | 63.5 | 66.5 | 69.3 | 59.5 | |
| Q-AlignModel Category=Traditional VQA Methods2026.06 | 38.2 | 40.4 | 39.2 | 76.6 | 74.9 | 75.8 | 57.4 | 57.7 | 57.5 | |
| Qwen3-VL-PlusModel Category=Flagship Models with Social-CoT Simulation, Reasoning Protocol=social-CoT2026.06 | 38 | 76.6 | 50.8 | 85.3 | 52.1 | 64.7 | 61.7 | 64.4 | 57.8 | |
| Claude-4.5-opusModel Category=Flagship Models with Social-CoT Simulation, Reasoning Protocol=social-CoT2026.06 | 37.1 | 81 | 51 | 86.7 | 47.4 | 61.3 | 61.9 | 64.2 | 56.1 | |
| Qwen3-VL-PlusModel Category=Standard LMMs, Reasoning Protocol=Standard2026.06 | 36.6 | 89.3 | 51.9 | 91 | 41.1 | 56.6 | 63.8 | 65.2 | 54.2 | |
| Claude-4.5-opusModel Category=Reasoning-Enhanced LMMs, Reasoning Protocol=Long-CoT2026.06 | 36.4 | 96.4 | 52.8 | 96.2 | 35.3 | 51.7 | 66.3 | 65.8 | 52.2 | |
| VQA2Model Category=Traditional VQA Methods2026.06 | 35.8 | 45.4 | 40 | 76.6 | 68.8 | 72.5 | 56.2 | 57.1 | 56.2 | |
| Gemini-2.5-FlashModel Category=Flagship Models with Social-CoT Simulation, Reasoning Protocol=social-CoT2026.06 | 35.3 | 62.9 | 45.2 | 77.9 | 61.5 | 68.7 | 56.6 | 62.2 | 57 | |
| FastVQAModel Category=Traditional VQA Methods2026.06 | 34.7 | 44 | 38.8 | 76.1 | 68.2 | 71.9 | 55.4 | 56.1 | 55.4 | |
| GPT-5.2Model Category=Standard LMMs, Reasoning Protocol=Standard2026.06 | 34.7 | 93.3 | 50.6 | 92.9 | 33.2 | 48.9 | 63.8 | 63.3 | 49.8 | |
| MaxVQAModel Category=Traditional VQA Methods2026.06 | 34.5 | 51.8 | 41.4 | 77.2 | 62.3 | 69 | 55.8 | 57.1 | 55.2 | |
| FineVQModel Category=Traditional VQA Methods2026.06 | 32.3 | 34.3 | 33.3 | 74.2 | 72.4 | 73.3 | 53.2 | 53.4 | 53.3 | |
| Qwen3-VL-PlusModel Category=Reasoning-Enhanced LMMs, Reasoning Protocol=Long-CoT2026.06 | 31.6 | 90.5 | 46.8 | 87.2 | 24.7 | 38.5 | 59.4 | 57.6 | 42.7 | |
| Gemini-3.0-ProModel Category=Reasoning-Enhanced LMMs, Reasoning Protocol=Long-CoT2026.06 | 31.3 | 97.8 | 47.4 | 95.4 | 17.6 | 29.7 | 63.4 | 57.7 | 38.5 | |
| Claude-4.5-opusModel Category=Standard LMMs, Reasoning Protocol=Standard2026.06 | 30.9 | 99.5 | 47.2 | 98.8 | 14.8 | 25.7 | 64.8 | 57.1 | 36.4 | |
| DOVERModel Category=Traditional VQA Methods2026.06 | 30.8 | 37.7 | 33.9 | 73.9 | 67.6 | 70.6 | 52.4 | 52.6 | 52.3 | |
| Qwen3-VL-8B-ThinkModel Category=Reasoning-Enhanced LMMs, Reasoning Protocol=Long-CoT, Backbone=True2026.06 | 26.5 | 11.5 | 16 | 72.1 | 89.2 | 79.7 | 49.3 | 50.4 | 47.9 |