Video Quality Assessment on YouTube-UGC
0.91SROCCSoft Ranking (Stage 1)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Soft Ranking (Stage 1)Training Stage=Stage 1, Evaluation Protocol=10-fold cross-validation2025.05 | 0.91 | 0.911 | — | |
| KVQPre-training Dataset=LSVQ [56]2025.03 | 0.903 | 0.905 | — | |
| KVQprotocol=supervised2026.04 | 0.903 | 0.905 | — | |
| KSVQE2024.02 | 0.9 | 0.912 | — | |
| DOVERprotocol=supervised2026.04 | 0.89 | 0.891 | — | |
| MinimalisticVQAEvaluation Protocol=10-fold cross-validation2025.05 | 0.89 | 0.891 | — | |
| qstRectifiers=Spatial and Temporal2024.02 | 0.876 | 0.877 | — | |
| DOVER2024.02 | 0.875 | 0.874 | — | |
| DPC-VQAprotocol=few-shot2026.04 | 0.868 | 0.871 | — | |
| FastVQA2024.02 | 0.863 | 0.859 | — | |
| FasterVQAPre-training Dataset=LSVQ [56]2025.03 | 0.863 | 0.859 | — | |
| Soft Ranking (Base)Training Stage=Base, Evaluation Protocol=10-fold cross-validation2025.05 | 0.858 | 0.848 | — | |
| qsRectifiers=Spatial2024.02 | 0.857 | 0.859 | — | |
| LMM-PVQA (Soft Ranking Stage 2)Training data=PLVD-P1/ P2 (600k), Human Label=No, Setting=Zero-shot2025.05 | 0.856 | 0.859 | — | |
| FastVQA2024.02 | 0.855 | 0.852 | — | |
| Fast-VQAPre-training Dataset=LSVQ [56]2025.03 | 0.855 | 0.852 | — | |
| FAST-VQAEvaluation Protocol=10-fold cross-validation2025.05 | 0.855 | 0.852 | — | |
| qtRectifiers=Temporal2024.02 | 0.854 | 0.858 | — | |
| LMM-PVQA (Soft Ranking Stage 3)Training data=PLVD-P1/ P2/ P3 (700k), Human Label=No, Setting=Zero-shot2025.05 | 0.854 | 0.861 | — | |
| FastVQAprotocol=supervised2026.04 | 0.852 | 0.855 | — | |
| SimpleVQAType=VQA2022.04 | 0.847 | 0.856 | — | |
| SimpleVQA2024.02 | 0.847 | 0.856 | — | |
| SimpleVQAprotocol=supervised2026.04 | 0.847 | 0.856 | — | |
| SimpleVQAEvaluation Protocol=10-fold cross-validation2025.05 | 0.847 | 0.856 | — | |
| LMM-PVQA (Soft Ranking Stage 1)Training data=PLVD-P1 (500k), Human Label=No, Setting=Zero-shot2025.05 | 0.845 | 0.849 | — | |
| Doveraesthetic branch=false2024.02 | 0.841 | 0.851 | — | |
| qbRectifiers=None2024.02 | 0.841 | 0.847 | — | |
| DOVEREvaluation Protocol=10-fold cross-validation2025.05 | 0.841 | 0.851 | — | |
| LMM-PVQA (Hard Ranking)Training data=PLVD-P1 (500k), Human Label=No, Setting=Zero-shot2025.05 | 0.839 | 0.844 | — | |
| VQTPre-training Dataset=ImageNet [7], Kinetics-400 [18]2025.03 | 0.836 | 0.851 | — | |
| Q-AlignTraining data=fused [11, 17, 28, 40, 70], Human Label=Yes, Setting=Zero-shot2025.05 | 0.834 | 0.846 | — | |
| CONVIQTprotocol=few-shot2026.04 | 0.832 | 0.822 | — | |
| Li et al.Type=VQA2022.04 | 0.831 | 0.819 | — | |
| Q-Alignprotocol=supervised2026.04 | 0.831 | 0.847 | — | |
| StarVQA+Pre-training Dataset=ImageNet [7], LSVQ [56], ...2025.03 | 0.826 | 0.82 | — | |
| MinimalisticVQA(IX)Training data=LSVQ [70], Human Label=Yes, Setting=Zero-shot2025.05 | 0.826 | 0.821 | — | |
| Li222024.02 | 0.825 | 0.818 | — | |
| CONTRIQUEprotocol=few-shot2026.04 | 0.825 | 0.813 | — | |
| SimpleVQA2024.02 | 0.819 | 0.817 | — | |
| Li et al.Pre-training Dataset=BID [6], LIVE [9], KonIQ-10k [15], SPAQ [8]2025.03 | 0.818 | 0.826 | — | |
| BVQAprotocol=supervised2026.04 | 0.818 | 0.826 | — | |
| COINVQPre-training Dataset=self-collected2025.03 | 0.816 | 0.802 | — | |
| Li22Inference Time (Sec)=27.632, Testing Protocol=Cross-dataset Testing2024.02 | 0.802 | 0.792 | — | |
| SimpleVQAInference Time (Sec)=0.714, Testing Protocol=Cross-dataset Testing2024.02 | 0.802 | 0.806 | — | |
| qstInference Time (Sec)=0.159, Testing Protocol=Cross-dataset Testing2024.02 | 0.788 | 0.804 | — | |
| VSFA2024.02 | 0.787 | 0.789 | — | |
| VSFAprotocol=supervised2026.04 | 0.787 | 0.789 | — | |
| VSFAEvaluation Protocol=10-fold cross-validation2025.05 | 0.787 | 0.789 | — | |
| qbInference Time (Sec)=0.159, Testing Protocol=Cross-dataset Testing2024.02 | 0.782 | 0.801 | — | |
| PVQTraining=KoNViD-1k2021.08 | 0.7799 | 0.7803 | — | |
| VIDEVALType=VQA2022.04 | 0.779 | 0.773 | — | |
| VIDEVALPre-training Dataset=NA (pure handcraft)2025.03 | 0.779 | 0.773 | — | |
| DOVERInference Time (Sec)=0.047, Testing Protocol=Cross-dataset Testing2024.02 | 0.777 | 0.792 | — | |
| MinimalisticVQA(VII)Training data=LSVQ [70], Human Label=Yes, Setting=Zero-shot2025.05 | 0.775 | 0.779 | — | |
| qtInference Time (Sec)=0.159, Testing Protocol=Cross-dataset Testing2024.02 | 0.774 | 0.791 | — | |
| RAPIQUEprotocol=supervised2026.04 | 0.774 | 0.781 | — | |
| 2BiVQAPre-training Dataset=ImageNet [7], KonIQ-10k [15]2025.03 | 0.771 | 0.79 | — | |
| DOVERTraining data=LSVQ [70], Human Label=Yes, Setting=Zero-shot2025.05 | 0.771 | 0.781 | — | |
| CSPTprotocol=few-shot2026.04 | 0.762 | 0.77 | — | |
| RAPIQUEType=VQA2022.04 | 0.759 | 0.768 | — | |
| RAPIQUE2024.02 | 0.759 | 0.768 | — | |
| UCDAprotocol=few-shot2026.04 | 0.756 | 0.761 | — | |
| SSL-VQAprotocol=few-shot2026.04 | 0.75 | 0.757 | — | |
| FastVQAInference Time (Sec)=0.045, Testing Protocol=Cross-dataset Testing2024.02 | 0.73 | 0.747 | — | |
| FAST-VQATraining data=LSVQ [70], Human Label=Yes, Setting=Zero-shot2025.05 | 0.73 | 0.747 | — | |
| VSFAType=VQA2022.04 | 0.724 | 0.743 | — | |
| VSFA2024.02 | 0.724 | 0.743 | — | |
| VSFAPre-training Dataset=None2025.03 | 0.724 | 0.743 | — | |
| qsInference Time (Sec)=0.159, Testing Protocol=Cross-dataset Testing2024.02 | 0.723 | 0.743 | — | |
| TLVQMTraining=KoNViD-1k2021.08 | 0.7206 | 0.753 | — | |
| ResNet50Type=IQA2022.04 | 0.718 | 0.71 | — | |
| VSFAInference Time (Sec)=11.109, Testing Protocol=Cross-dataset Testing2024.02 | 0.718 | 0.721 | — | |
| VSFATraining=KoNViD-1k2021.08 | 0.7175 | 0.7597 | — | |
| VISIONprotocol=few-shot2026.04 | 0.706 | 0.717 | — | |
| VGG19Type=IQA2022.04 | 0.703 | 0.7 | — | |
| VIDEVALTraining=KoNViD-1k2021.08 | 0.7022 | 0.7153 | — | |
| VIDEALEvaluation Protocol=10-fold cross-validation2025.05 | 0.687 | 0.709 | — | |
| TLVQMEvaluation Protocol=10-fold cross-validation2025.05 | 0.685 | 0.692 | — | |
| TLVQMType=VQA2022.04 | 0.669 | 0.659 | — | |
| TLVQM2024.02 | 0.669 | 0.659 | — | |
| VIDEVAL2024.02 | 0.669 | 0.659 | — | |
| TLVQMPre-training Dataset=NA (pure handcraft)2025.03 | 0.669 | 0.659 | — | |
| TLVQMprotocol=supervised2026.04 | 0.669 | 0.659 | — | |
| PVQTraining=LIVE-VQC2021.08 | 0.6025 | 0.6029 | — | |
| KonCept512Type=IQA2022.04 | 0.587 | 0.594 | — | |
| BRISQUETraining=KoNViD-1k2021.08 | 0.5682 | 0.6052 | — | |
| V-BLIINDSType=VQA2022.04 | 0.559 | 0.555 | — | |
| BUONA-VISTATraining data=None, Human Label=No, Setting=Zero-shot2025.05 | 0.525 | 0.556 | — | |
| VIQE2024.02 | 0.513 | 0.496 | — | |
| BRISQUEprotocol=supervised2026.04 | 0.472 | 0.447 | — | |
| 2BiVQATraining Dataset=KoNViD-1K2022.08 | 0.428 | — | — | |
| VSFATraining=LIVE-VQC2021.08 | 0.4221 | 0.4525 | — | |
| 2BiVQATraining Dataset=LIVE-VQC2022.08 | 0.416 | — | — | |
| VIDEVALTraining Dataset=KoNViD-1K2022.08 | 0.392 | — | — | |
| BRISQUEType=IQA2022.04 | 0.382 | 0.395 | — | |
| FAST-VQATraining Dataset=KoNViD-1K2022.08 | 0.373 | — | — | |
| GM-LOGType=IQA2022.04 | 0.368 | 0.392 | — | |
| FAST-VQATraining Dataset=LIVE-VQC2022.08 | 0.365 | — | — | |
| RAPIQUETraining Dataset=LIVE-VQC2022.08 | 0.352 | — | — | |
| RAPIQUETraining Dataset=KoNViD-1K2022.08 | 0.318 | — | — |