Video Question Answering on TGIF-QA
97.2AccuracyHiTeA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| HiTeA#PT Data=17M2022.12 | 97.2 | — | — | — | — | |
| HiTeA#PT Data=5M2022.12 | 96.8 | — | — | — | — | |
| LAVENDER#PT Data=5M2022.12 | 96.6 | — | — | — | — | |
| ATMPre-trained with large-scale external data=false2023.09 | 96 | — | — | — | — | |
| All-in-one#PT Data=283M2022.12 | 95.5 | — | — | — | — | |
| VGTPre-trained with large-scale external data=false2023.09 | 95 | — | — | — | — | |
| Clover#PT Data=5M2022.12 | 94.9 | — | — | — | — | |
| MERLOT#PT Data=180M2022.12 | 94 | — | — | — | — | |
| MERLOTPre-trained with large-scale external data=true2023.09 | 94 | — | — | — | — | |
| VIOLET#PT Data=183M2022.12 | 92.5 | — | — | — | — | |
| MASNPre-trained with large-scale external data=false2023.09 | 84.4 | — | — | — | — | |
| MHNPre-trained with large-scale external data=false2023.09 | 83.5 | — | — | — | — | |
| ClipBERT#PT Data=0.2M2022.12 | 82.8 | — | — | — | — | |
| ClipBERTPre-trained with large-scale external data=true2023.09 | 82.8 | — | — | — | — | |
| TS-LLAVALLM Size=34B, Vision Encoder=CLIP-L, Training-free=true2024.11 | 81 | — | — | — | 4.2 | |
| PGATPre-trained with large-scale external data=false2023.09 | 80.6 | — | — | — | — | |
| PLLaVA 34BVision Encoder=ViT-L, LLM Size=34B2024.04 | 80.6 | — | — | — | 4.3 | |
| PLLaVALLM Size=34B, Vision Encoder=CLIP-L2024.07 | 80.6 | — | — | — | 4.3 | |
| SF-LLaVA-34BLLM Size=34B, Vision Encoder=CLIP-L2024.07 | 80.6 | — | — | — | 4.3 | |
| PLLAVALLM Size=34B, Vision Encoder=CLIP-L, Training-free=false2024.11 | 80.6 | — | — | — | 4.3 | |
| SF-LLAVALLM Size=34B, Vision Encoder=CLIP-L, Training-free=true2024.11 | 80.6 | — | — | — | 4.3 | |
| SiaSReaPre-trained with large-scale external data=true2023.09 | 79.7 | — | — | — | — | |
| COSA (1.2B)Example=415M2023.06 | 79.5 | — | — | — | — | |
| VAST(1.3B)Sample=442M, extra_modalities=audio or subtitle2023.05 | 79.1 | — | — | — | — | |
| IG-VLM LLaVA 34BVision Encoder=ViT-L, LLM Size=34B2024.04 | 79.1 | — | — | — | 4.2 | |
| IG-VLM (LLaVA-v1.6)LLM Size=34B, Vision Encoder=CLIP-L2024.07 | 79.1 | — | — | — | 4.2 | |
| IG-VLMLLM Size=34B, Vision Encoder=CLIP-L, Training-free=true2024.11 | 79.1 | — | — | — | 4.2 | |
| VALOR-LExample=433.5M2023.06 | 78.7 | — | — | — | — | |
| VALOR-LSample=433.5M, extra_modalities=audio or subtitle2023.05 | 78.7 | — | — | — | — | |
| SF-LLaVALLM Size=7B, Vision Encoder=CLIP-L, Training-free=true2024.11 | 78.7 | — | — | — | 4.2 | |
| IG-VLM LLaVA 13BVision Encoder=ViT-L, LLM Size=13B2024.04 | 78 | — | — | — | 4 | |
| HAIRPre-trained with large-scale external data=false2023.09 | 77.8 | — | — | — | — | |
| PLLaVA 13BVision Encoder=ViT-L, LLM Size=13B2024.04 | 77.8 | — | — | — | 4.2 | |
| TS-LLaVALLM Size=7B, Vision Encoder=CLIP-L, Training-free=true2024.11 | 77.7 | — | — | — | 4.1 | |
| COSA-LExample=417M2023.06 | 77.6 | — | — | — | — | |
| PLLaVA 7BVision Encoder=ViT-L, LLM Size=7B2024.04 | 77.5 | — | — | — | 4.1 | |
| PLLAVALLM Size=7B, Vision Encoder=CLIP-L, Training-free=false2024.11 | 77.5 | — | — | — | 4.1 | |
| IG-VLM CogAgentVision Encoder=CLIP-E, LLM Size=7B2024.04 | 76.7 | — | — | — | 4 | |
| B2APre-trained with large-scale external data=false2023.09 | 75.9 | — | — | — | — | |
| mPLUG-2Example=417M2023.06 | 75.4 | — | — | — | — | |
| mPLUG-2Sample=417M2023.05 | 75.4 | — | — | — | — | |
| HGAPre-trained with large-scale external data=false2023.09 | 75.4 | — | — | — | — | |
| Video-LLaVA + PADL+ (16×)LLM size=7B, Compression=16×, adapter fine-tuning=true2026.06 | 75.3 | — | — | — | 43 | |
| COSA-BExample=17M2023.06 | 75 | — | — | — | — | |
| HCRNPre-trained with large-scale external data=false2023.09 | 75 | — | — | — | — | |
| HOSTRPre-trained with large-scale external data=false2023.09 | 75 | — | — | — | — | |
| GIT2 (5.1B)Example=12.9B2023.06 | 74.9 | — | — | — | — | |
| GIT2 (5.1B)Sample=12.9B2023.05 | 74.9 | — | — | — | — | |
| VideoGPT+Protocol=Zero-shot2024.06 | 74.6 | — | — | — | 4.1 | |
| VideoGPT+LLM Size=3.8B, Vision Encoder=CLIP-L, Training-free=false2024.11 | 74.6 | — | — | — | 4.1 | |
| LGCNPre-trained with large-scale external data=false2023.09 | 74.3 | — | — | — | — | |
| LAVENDERExample=30M2023.06 | 73.5 | — | — | — | — | |
| LAVENDERSample=30M2023.05 | 73.5 | — | — | — | — | |
| COSA-BExample=5M2023.06 | 73.4 | — | — | — | — | |
| HiTeA#PT Data=17M2022.12 | 73.2 | — | — | — | — | |
| HiTeAExample=17M2023.06 | 73.2 | — | — | — | — | |
| HiTeASample=17M2023.05 | 73.2 | — | — | — | — | |
| IG-VLM LLaVA 7BVision Encoder=ViT-L, LLM Size=7B2024.04 | 73 | — | — | — | 4 | |
| IG-VLMLLM Size=7B, Vision Encoder=CLIP-L, Training-free=true2024.11 | 73 | — | — | — | 4 | |
| VIOLETV2Example=5M2023.06 | 72.8 | — | — | — | — | |
| GITExample=1.7B2023.06 | 72.8 | — | — | — | — | |
| VIOLETv2Sample=5M2023.05 | 72.8 | — | — | — | — | |
| GITSample=1.7B2023.05 | 72.8 | — | — | — | — | |
| VIOLETv22023.10 | 72.8 | — | — | — | — | |
| HiTeA#PT Data=5M2022.12 | 72.5 | — | — | — | — | |
| LAVENDER#PT Data=5M2022.12 | 72.2 | — | — | — | — | |
| Intern VideoExample=646M2023.06 | 72.2 | — | — | — | — | |
| InternVideoSample=646M2023.05 | 72.2 | — | — | — | — | |
| Clover#PT Data=5M2022.12 | 71.4 | — | — | — | — | |
| CloverExample=5M2023.06 | 71.4 | — | — | — | — | |
| CloverSample=5M2023.05 | 71.4 | — | — | — | — | |
| SimVTP2022.12 | 70.2 | — | — | — | — | |
| Video-LLaVAVision Encoder=ViT-L, LLM Size=7B2024.04 | 70 | — | — | — | 4 | |
| Video-LLaVAProtocol=Zero-shot2024.06 | 70 | — | — | — | 4 | |
| Video-LLaVALLM Size=7B, Vision Encoder=ViT-L, Training-free=false2024.11 | 70 | — | — | — | 4 | |
| Video-LLaVALLM Backbone=Vicuna-7B, Data Scale=700K, Zero-shot=true2026.01 | 70 | — | — | — | 4 | |
| Video-LLaVALLM size=7B2026.06 | 70 | — | — | — | 40 | |
| Video-LLaVA + PADL (16×)LLM size=7B, Compression=16×, adapter fine-tuning=false2026.06 | 69.9 | — | — | — | 41 | |
| MERLOT2022.12 | 69.5 | — | — | — | — | |
| MERLOT#PT Data=180M2022.12 | 69.5 | — | — | — | — | |
| MERLOTExample=180M2023.06 | 69.5 | — | — | — | — | |
| MERLOTSample=180M2023.05 | 69.5 | — | — | — | — | |
| MERLOT2023.10 | 69.5 | — | — | — | — | |
| InternVideo2023.10 | 69.3 | — | — | — | — | |
| VIOLET2022.12 | 68.9 | — | — | — | — | |
| VIOLET#PT Data=183M2022.12 | 68.9 | — | — | — | — | |
| VIOLET2023.10 | 68.9 | — | — | — | — | |
| FrozenBiLMExample=410M2023.06 | 68.6 | — | — | — | — | |
| FrozenBiLMSample=410M2023.05 | 68.6 | — | — | — | — | |
| FrozenBiLMSupervision=100% (fully-supervised)2022.06 | 68.6 | — | — | — | — | |
| InternVideovariant=w/o decoder2023.10 | 67.2 | — | — | — | — | |
| mPLUG-Owl2Evaluation Category=GPT-Assisted, Zero-shot=true2023.11 | 67.1 | — | — | — | 3.7 | |
| InternVideobackbone=ViT2023.10 | 67 | — | — | — | — | |
| Elysium2024.03 | 66.6 | — | — | — | 3.6 | |
| All-in-one#PT Data=283M2022.12 | 66.3 | — | — | — | — | |
| All-in-oneExample=228.5M2023.06 | 66.3 | — | — | — | — | |
| All-in-oneSample=228.5M2023.05 | 66.3 | — | — | — | — | |
| IG-VLM GPT-4VVision Encoder=Unk, LLM Size=GPT-42024.04 | 65.3 | — | — | — | 3.7 | |
| All-in-one2023.10 | 64.2 | — | — | — | — | |
| Co-Tokenization2023.10 | 62.5 | — | — | — | — |