Video Question Answering on TGIF-QA (test)
95.5AccuracyAll-in-one-B *
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| All-in-one-B *Nets=CE, Params=110M, PT Samples=9.72M, Frames=3, Resolution=224x224, Note=Pre-training with additional YT-Temporal 180M2022.03 | 95.5 | — | — | — | — | — | — | — | |
| All-in-one-B [384]Nets=CE, Params=110M, PT Samples=3.72M, Frames=3, Resolution=384x3842022.03 | 94.7 | — | — | — | — | — | — | — | |
| All-in-one-BNets=CE, Params=110M, PT Samples=3.72M, Frames=1, Resolution=224x2242022.03 | 92.9 | — | — | — | — | — | — | — | |
| All-in-one-BNets=CE, Params=110M, PT Samples=3.72M, Frames=3, Resolution=224x2242022.03 | 92.7 | — | — | — | — | — | — | — | |
| All-in-one-SNets=CE, Params=33M, PT Samples=3.72M, Frames=3, Resolution=224x2242022.03 | 91.2 | — | — | — | — | — | — | — | |
| VIOLETNets=T+V+CE, Params=198M, PT Samples=5.5M, Frames=16, Resolution=224x2242022.03 | 87.1 | — | — | — | — | — | — | — | |
| ClipBERTNets=T+V+CE, Params=137M, PT Samples=5.6M, Frames=1x1, Resolution=224x2242022.03 | 82.9 | — | — | — | — | — | — | — | |
| HCRNhierarchy levels=22020.02 | 81.4 | — | — | — | — | — | — | — | |
| All-in-one-TiNets=CE, Params=12M, PT Samples=3.72M, Frames=3, Resolution=224x2242022.03 | 80.6 | — | — | — | — | — | — | — | |
| PLLaVALLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 80.6 | — | — | — | — | 4.3 | — | — | |
| SF-LLaVA-34BLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 80.6 | — | — | — | — | 4.3 | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 79.1 | — | — | — | — | 4.2 | — | — | |
| IG-VLMLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 79.1 | — | — | — | — | 4.2 | — | — | |
| SF-LLaVA-7BLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 78.7 | — | — | — | — | 4.2 | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=13B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 78 | — | — | — | — | 4 | — | — | |
| HME2020.02 | 77.8 | — | — | — | — | — | — | — | |
| PLLaVALLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 77.5 | — | — | — | — | 4.1 | — | — | |
| PSAC2020.02 | 76.9 | — | — | — | — | — | — | — | |
| IG-VLM (CogAgent)Vision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 76.7 | — | — | — | — | 4 | — | — | |
| QueSTNets=T+V+LSTM, Frames=16, Resolution=224x2242022.03 | 75.9 | — | — | — | — | — | — | — | |
| HCRNhierarchy levels=22020.02 | 75 | — | — | — | — | — | — | — | |
| HCRNNets=T+V+LSTM, Frames=16, Resolution=224x2242022.03 | 75 | — | — | — | — | — | — | — | |
| VideoGPT+LLM Size=3.8B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 74.6 | — | — | — | — | 4.1 | — | — | |
| Co-mem2020.02 | 74.3 | — | — | — | — | — | — | — | |
| HME2020.02 | 73.9 | — | — | — | — | — | — | — | |
| HeterogeneousNets=T+V+LSTM, Frames=35, Resolution=224x2242022.03 | 73.9 | — | — | — | — | — | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 73 | — | — | — | — | 4 | — | — | |
| IG-VLMLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 73 | — | — | — | — | 4 | — | — | |
| PSAC2020.02 | 70.4 | — | — | — | — | — | — | — | |
| Video-LLaVAVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 70 | — | — | — | — | 4 | — | — | |
| Video-LLaVALLM Size=7B, Vision Encoder=ViT-L, Training Mode=Fine-tuned2024.07 | 70 | — | — | — | — | 4 | — | — | |
| Video-LLaVALLM size=7B2023.11 | 70 | — | — | — | — | 4 | — | — | |
| MERLOTPretrain Dataset=YT, CC, Pretrain Size=180M, Text=BERT2023.02 | 69.5 | — | — | — | — | — | — | — | |
| ST-TP2020.02 | 69.4 | — | — | — | — | — | — | — | |
| Co-mem2020.02 | 68.2 | — | — | — | — | — | — | — | |
| All-in-one-B *Nets=CE, Params=110M, PT Samples=9.72M, Frames=3, Resolution=224x224, Note=Pre-training with additional YT-Temporal 180M2022.03 | 66.3 | — | — | — | — | — | — | — | |
| Supervised SOTAEvaluation protocol=supervised2023.12 | 66.3 | — | — | — | — | — | — | — | |
| ProViQEvaluation protocol=zero-shot2023.12 | 66.1 | — | — | — | — | — | — | — | |
| All-in-one-B [384]Nets=CE, Params=110M, PT Samples=3.72M, Frames=3, Resolution=384x3842022.03 | 65.4 | — | — | — | — | — | — | — | |
| IG-VLM (GPT-4V)Vision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 65.3 | — | — | — | — | 3.7 | — | — | |
| All-in-one-BNets=CE, Params=110M, PT Samples=3.72M, Frames=3, Resolution=224x2242022.03 | 64.2 | — | — | — | — | — | — | — | |
| All-in-one-SNets=CE, Params=33M, PT Samples=3.72M, Frames=3, Resolution=224x2242022.03 | 64 | — | — | — | — | — | — | — | |
| ST-TP2020.02 | 62.9 | — | — | — | — | — | — | — | |
| All-in-one-BNets=CE, Params=110M, PT Samples=3.72M, Frames=1, Resolution=224x2242022.03 | 62.5 | — | — | — | — | — | — | — | |
| VGT (PT)Pretrain Dataset=WV, Pretrain Size=0.18M, Text=BERT2023.02 | 61.7 | — | — | — | — | — | — | — | |
| CoVGT (PT)Pretrain Dataset=WV, Pretrain Size=0.18M, Text=RoBERTa2023.02 | 61.7 | — | — | — | — | — | — | — | |
| VGTText=BERT2023.02 | 61.6 | — | — | — | — | — | — | — | |
| CoVGTText=RoBERTa2023.02 | 61.6 | — | — | — | — | — | — | — | |
| HQGAText=BERT2023.02 | 61.3 | — | — | — | — | — | — | — | |
| PGATText=GloVe2023.02 | 61.1 | — | — | — | — | — | — | — | |
| ClipBERTPretrain Dataset=VG, COCO, Text=BERT2023.02 | 60.3 | — | — | — | — | — | — | — | |
| Chat-UniViVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 60.3 | — | — | — | — | 3.4 | — | — | |
| Chat-UniViLLM size=7B2023.11 | 60.3 | — | — | — | — | 3.4 | — | — | |
| HAIRText=GloVe2023.02 | 60.2 | — | — | — | — | — | — | — | |
| SiaSReaPretrain Dataset=VG, COCO, Text=BERT2023.02 | 60.2 | — | — | — | — | — | — | — | |
| QueSTNets=T+V+LSTM, Frames=16, Resolution=224x2242022.03 | 59.7 | — | — | — | — | — | — | — | |
| MASNText=GloVe2023.02 | 59.5 | — | — | — | — | — | — | — | |
| ClipBERTNets=T+V+CE, Params=137M, PT Samples=5.6M, Frames=1x1, Resolution=224x2242022.03 | 59.4 | — | — | — | — | — | — | — | |
| HOSTRText=GloVe2023.02 | 58 | — | — | — | — | — | — | — | |
| HCRNhierarchy levels=22020.02 | 55.9 | — | — | — | — | — | — | — | |
| HCRNNets=T+V+LSTM, Frames=16, Resolution=224x2242022.03 | 55.9 | — | — | — | — | — | — | — | |
| PSAC2020.02 | 55.7 | — | — | — | — | — | — | — | |
| All-in-one-TiNets=CE, Params=12M, PT Samples=3.72M, Frames=3, Resolution=224x2242022.03 | 53.9 | — | — | — | — | — | — | — | |
| HME2020.02 | 53.8 | — | — | — | — | — | — | — | |
| HeterogeneousNets=T+V+LSTM, Frames=35, Resolution=224x2242022.03 | 53.8 | — | — | — | — | — | — | — | |
| R2AAdaptation type=Language based Adaptation, Language model=DeBERTa-v2-XL, Language model #params=890M, Vision model=ViT-L/14, Vision model #params=300M, Evaluation protocol=Zero-shot2023.06 | 52.2 | — | — | — | — | — | — | — | |
| Co-mem2020.02 | 51.5 | — | — | — | — | — | — | — | |
| Video-ChatGPTVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 51.4 | — | — | — | — | 3 | — | — | |
| Video-ChatGPTLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 51.4 | — | — | — | — | 3 | — | — | |
| Video-ChatGPTLLM size=7B2023.11 | 51.4 | — | — | — | — | 3 | — | — | |
| ST-TP2020.02 | 49.5 | — | — | — | — | — | — | — | |
| Video-LLaVAVision Size=425M, Modality=V+I, Pretrain Data=1.26M, Finetune Data=765K2024.12 | 47 | — | — | — | — | 3.3 | — | — | |
| Video-PandaVision Size=45M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 42.9 | — | — | — | — | 3.2 | — | — | |
| FrozenBiLMAdaptation type=Training based Adaptation, Language model=DeBERTa-v2-XL, Language model #params=890M, Vision model=ViT-L/14, Vision model #params=300M, Evaluation protocol=Zero-shot2023.06 | 41.9 | — | — | — | — | — | — | — | |
| FrozenBiLMVision Encoder=ViT-L, LLM Size=1.3B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 41.9 | — | — | — | — | — | — | — | |
| FrozenBiLMEvaluation protocol=zero-shot2023.12 | 41.9 | — | — | — | — | — | — | — | |
| Video-LLaVAVision Size=425M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 41.7 | — | — | — | — | — | — | — | |
| FrozenBiLMLLM size=1B2023.11 | 41 | — | — | — | — | — | — | — | |
| Video-ChatGPTVision Size=307M, Modality=V, Pretrain Data=100K, Finetune Data=100K2024.12 | 40.7 | — | — | — | — | 3.1 | — | — | |
| EVE*Vision Size=30M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 39.2 | — | — | — | — | 2.9 | — | — | |
| ChatUniViVision Size=307M, Modality=V+I, Pretrain Data=1.6M, Finetune Data=649K2024.12 | 38.2 | — | — | — | — | 3 | — | — | |
| VideoChatVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 34.4 | — | — | — | — | 2.3 | — | — | |
| VideoChatLLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 34.4 | — | — | — | — | 2.3 | — | — | |
| VideoChatLLM size=7B2023.11 | 34.4 | — | — | — | — | 2.3 | — | — | |
| VideoChatVision Size=1.2B, Modality=V, Pretrain Data=25M, Finetune Data=18K2024.12 | 21.3 | — | — | — | — | 1.9 | — | — | |
| Video-LLaMALLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 12.4 | — | — | — | — | 1.1 | — | — | |
| CLIP VIT-L/14Adaptation type=Training based Adaptation, Language model=Custom, Language model #params=123M, Vision model=ViT-L/14, Vision model #params=300M, Evaluation protocol=Zero-shot2023.06 | 3.6 | — | — | — | — | — | — | — | |
| CLIP-VIT-L/14Evaluation protocol=zero-shot2023.12 | 3.6 | — | — | — | — | — | — | — | |
| RandomEvaluation protocol=zero-shot2023.12 | 0.1 | — | — | — | — | — | — | — | |
| Bridge to Answer2021.04 | — | 3.71 | 75.9 | 82.6 | 57.5 | — | — | — | |
| Bridge2Answer2022.08 | — | — | 75.9 | 82.6 | 57.5 | — | — | — | |
| ClipBERTmode=Pre-trained2021.11 | — | — | 82.8 | 87.8 | 60.3 | — | — | — | |
| ClipBERTNtest=12022.08 | — | — | 82.9 | 87.5 | 59.4 | — | — | — | |
| ClipBERTNtest=standard2022.08 | — | — | 82.8 | 87.8 | 60.3 | — | — | — | |
| CLIPBERTNtrain x T (Training Input Sampling Method)=1x1, Ntest (Number of test clips/frames)=12021.02 | — | — | 82.9 | 87.5 | 59.4 | — | — | — | |
| CLIPBERTNtrain x T (Training Input Sampling Method)=1x12021.02 | — | — | 82.8 | 87.8 | 60.3 | — | — | — | |
| Co-mem2021.04 | — | 4.1 | 68.2 | 74.3 | 51.5 | — | — | — | |
| Co-MemBackbone=ResNet+C3D2019.04 | — | 4.1 | 68.2 | 74.3 | 51.5 | — | — | — | |
| Co-Memory2021.02 | — | — | 68.2 | 74.3 | 51.5 | — | — | — | |
| Co-Memory2021.11 | — | — | 68.2 | 74.3 | 51.5 | — | — | — |