Video Question Answering on MSRVTT (test)
92.7AccuracyNorton
Evaluation Results
| Method | Links | |
|---|---|---|
| NortonEvaluation Protocol=Supervised2024.01 | 92.7 | |
| TempCLREvaluation Protocol=Supervised2024.01 | 92.2 | |
| VideoCLIPEvaluation Protocol=Supervised2024.01 | 92.1 | |
| MERLOTEvaluation Protocol=Supervised2024.01 | 90.9 | |
| ClipBERTEvaluation Protocol=Supervised2024.01 | 88.2 | |
| ActBERTEvaluation Protocol=Supervised2024.01 | 85.7 | |
| JSFusionEvaluation Protocol=Supervised2024.01 | 83.4 | |
| NortonEvaluation Protocol=Zero-shot2024.01 | 77.1 | |
| MLBEvaluation Protocol=Supervised2024.01 | 76.1 | |
| TempCLREvaluation Protocol=Zero-shot2024.01 | 74.4 | |
| VideoCLIPEvaluation Protocol=Zero-shot2024.01 | 73.9 | |
| EITanqueEvaluation Protocol=Supervised2024.01 | 65.5 | |
| X2-VLM (large)# Params=593M, mode=fine-tuning2022.11 | 45.5 | |
| X2-VLM (base)# Params=255M, mode=fine-tuning2022.11 | 45 | |
| All-in-one# Params=110M, mode=fine-tuning2022.11 | 44.3 | |
| OmniVL# Params=288M, mode=fine-tuning2022.11 | 44.1 | |
| OmniVLPreTrain VLData=18M2024.03 | 44.1 | |
| VIOLET# Params=163M, mode=fine-tuning2022.11 | 43.9 | |
| OmniViD2024.03 | 42.3 | |
| ALPRO# Params=513M, mode=fine-tuning2022.11 | 42.1 | |
| ALIPROPreTrain VLData=5.5M2024.03 | 42.1 | |
| JustAskPreTrain VLData=69M2024.03 | 41.5 | |
| JustAsk2024.03 | 39.6 | |
| CoMVTPreTrain VLData=100M2024.03 | 39.5 | |
| ClipBERTPreTrain VLData=5.6M2024.03 | 37.4 | |
| HCRN2024.03 | 35.6 |