Video Question Answering on NExT-QA (val)
88.4Overall AccHuman
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Human2022.12 | 88.4 | 87.6 | 88.6 | 90.4 | |
| LLaVAOneVision-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 79.4 | — | — | — | |
| TarsierEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=34B2024.03 | 79.2 | — | — | — | |
| LLaVA-NeXT-Interleave-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 78.2 | — | — | — | |
| MM1.5-Video-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 76.9 | — | — | — | |
| MM1.5-Video-7BTraining Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 76.1 | — | — | — | |
| MM1.5-Video-3BTraining Mode=SFT, Video Data=true, Model Size=3B2024.09 | 74.7 | — | — | — | |
| LLoViLM=GPT-4, Params=1.5T, Zero-shot=true2023.12 | 73.8 | 73.7 | 70.2 | 81.9 | |
| SeViLAEvaluation Protocol=finetuned, Video Pretrain=true, Params=4B2024.03 | 73.8 | 74.2 | 69.4 | 81.3 | |
| LLoVi (GPT-4)Params=N/A, Zero-shot=true2024.12 | 73.8 | 73.7 | 70.2 | 81.9 | |
| TCRtrain params=76M2023.12 | 73.5 | 73.5 | 69.8 | 82.2 | |
| VideoTreeEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 73.5 | 75.2 | 67 | 81.3 | |
| VideoTree (GPT-4)Params=N/A, Zero-shot=true2024.12 | 73.5 | 75.2 | 67 | 81.3 | |
| SeViLAtrain params=346M2023.12 | 73.4 | 73.4 | 68.8 | 83.5 | |
| LVNetEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=<1.8T2024.03 | 72.9 | 75 | 65.5 | 81.5 | |
| MM1.5-Video-3BTraining Mode=Training-free, Video Data=false, Model Size=3B2024.09 | 72.8 | — | — | — | |
| VamosParams=7B, Protocol=finetuned2024.05 | 72.5 | 72.6 | 69.6 | 78 | |
| VamosEvaluation Protocol=finetuned, Video Pretrain=true, Params=7B2024.03 | 72.5 | 72.6 | 69.6 | 78 | |
| LLaMA-VQAParams=7B, Protocol=finetuned2024.05 | 72 | 72.7 | 69.2 | 75.8 | |
| LLama-VQAEvaluation Protocol=finetuned, Video Pretrain=true, Params=7B2024.03 | 72 | 72.7 | 69.2 | 75.8 | |
| MM1.5-Video-1BTraining Mode=SFT, Video Data=true, Model Size=1B2024.09 | 71.8 | — | — | — | |
| TarsierEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=7B2024.03 | 71.6 | — | — | — | |
| VideoAgentEvaluation Protocol=Zero-shot2024.03 | 71.3 | 72.7 | 64.5 | 81.1 | |
| VideoAgentEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 71.3 | 72.7 | 64.5 | 81.1 | |
| VideoAgent (GPT-4)Params=N/A, Zero-shot=true2024.12 | 71.3 | 72.7 | 64.5 | 81.1 | |
| VidCtxParams=7B, Zero-shot=true2024.12 | 70.7 | 71.7 | 65.1 | 79.2 | |
| BLIP-2Params=4B, Protocol=finetuned2024.05 | 70.1 | 70.1 | 65.2 | 80.1 | |
| BLIP2train params=188M, re-implementation=by [40]2023.12 | 70.1 | 72.9 | 65.2 | 80.1 | |
| BLIP-2Evaluation Protocol=finetuned, Video Pretrain=true, Params=4B2024.03 | 70.1 | 70.1 | 65.2 | 80.1 | |
| MM1.5-Video-1BTraining Mode=Training-free, Video Data=false, Model Size=1B2024.09 | 70 | — | — | — | |
| MoReVQAEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=340B2024.03 | 69.2 | 70.2 | 64.6 | — | |
| VideoChat2-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 68.6 | — | — | — | |
| IG-VLMEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 68.6 | 69.8 | 63.6 | 74.7 | |
| TraveLERTraining Protocol=Zero-shot2024.04 | 68.2 | 70 | 60.5 | 78.2 | |
| TraveLEREvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 68.2 | 70 | 60.5 | 78.2 | |
| LLoViEvaluation Protocol=Zero-shot2024.03 | 67.7 | 69.5 | 61 | 75.6 | |
| LLoViTraining Protocol=Zero-shot2024.04 | 67.7 | 69.5 | 61 | 75.6 | |
| LLoViEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 67.7 | 69.5 | 61 | 75.6 | |
| LLoViLM=GPT-3.5, Params=175B, Zero-shot=true2023.12 | 66.3 | 67.1 | 60.1 | 76.5 | |
| LLoVi (GPT-3.5)Params=N/A, Zero-shot=true2024.12 | 66.3 | 67.1 | 60.1 | 76.5 | |
| Q-ViDParams=12B, Zero-shot=true2024.12 | 66.3 | 67.6 | 61.6 | 72.2 | |
| VideoStreamingParams=7B+1.3B, Protocol=zero-shot2024.05 | 66.2 | 65.1 | 62.2 | 78.1 | |
| MC-ViT-LTraining Protocol=Uses fine-tuned components2024.04 | 65 | — | — | — | |
| MC-ViT-LEvaluation Protocol=finetuned, Video Pretrain=true, Params=424M2024.03 | 65 | — | — | — | |
| ProViQEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=175B2024.03 | 64.6 | — | — | — | |
| SlowFast-LLaVA-7BTraining Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 64.2 | — | — | — | |
| ProViQTraining Protocol=Zero-shot2024.04 | 63.8 | — | — | — | |
| SeViLAEvaluation Protocol=Zero-shot2024.03 | 63.6 | 61.3 | 61.5 | 75.6 | |
| SeViLAParams=4B, Protocol=zero-shot2024.05 | 63.6 | 61.3 | 61.5 | 75.6 | |
| SeViLALM=Flan-T5, Params=4B, Zero-shot=true2023.12 | 63.6 | 61.3 | 61.5 | 75.6 | |
| SeViLATraining Protocol=Uses fine-tuned components2024.04 | 63.6 | 61.3 | 61.5 | 75.6 | |
| SeViLAEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=4B2024.03 | 63.6 | 61.3 | 61.5 | 75.6 | |
| SeViLAParams=4B, Zero-shot=true, Pre-trained on video=true2024.12 | 63.6 | 61.5 | 61.3 | 75.6 | |
| BLIP2train params=188M2023.12 | 63.5 | 64.9 | 59.7 | 77.8 | |
| InternVideoEvaluation Protocol=finetuned, Video Pretrain=true, Params=478M2024.03 | 63.2 | 62.5 | 58.5 | 75.8 | |
| HiTeA# PT Data=5M2022.12 | 63.1 | 62.4 | 58.3 | 75.6 | |
| HiTeAEvaluation Protocol=Supervised2024.03 | 63.1 | 62.4 | 58.3 | 75.6 | |
| HiTeA2023.12 | 63.1 | 62.4 | 58.3 | 75.6 | |
| IG-VLM-7B (LLaVA-v1.6)Training Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 63.1 | — | — | — | |
| HiTeAEvaluation Protocol=finetuned, Video Pretrain=true, Params=297M2024.03 | 63.1 | 62.4 | 58.3 | 75.6 | |
| STAR-MINIBase Model=GPT-3.5-turbo, Frames=0.6 fps / 22.6, #LLM calls=5.42025.12 | 62 | 62.8 | 55.1 | 73.7 | |
| VideoChat2Params=7B, Zero-shot=true, Pre-trained on video=true2024.12 | 61.7 | 61.9 | 57.4 | 69.9 | |
| DeepStack-L-7BTraining Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 61 | — | — | — | |
| LangRepoParams=8x7B, Protocol=zero-shot2024.05 | 60.9 | 64.4 | 51.4 | 69.1 | |
| LangRepoEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=false, Params=12B2024.03 | 60.9 | 64.4 | 51.4 | 69.1 | |
| LangRepoParams=12B, Zero-shot=true2024.12 | 60.9 | 64.4 | 51.4 | 69.1 | |
| CoVGT (PT)Pretrain Dataset=WebVid(WV), Pretrain Size=0.18M, Video=R, F, Text=ROBERTa2023.02 | 60.73 | 59.69 | 58 | 69.88 | |
| CoVGTEvaluation Protocol=Supervised2024.03 | 60.7 | 59.7 | 58 | 69.9 | |
| Vista-LLaMA-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 60.7 | — | — | — | |
| SeViTFiDEvaluation Protocol=finetuned, Video Pretrain=true, Params=215M2024.03 | 60.6 | — | — | — | |
| CoVGTVideo=R, F, Text=ROBERTa2023.02 | 60.01 | 58.8 | 57.44 | 69.37 | |
| ViperGPTEvaluation Protocol=Zero-shot2024.03 | 60 | — | — | — | |
| ViperGPTLM=GPT-3, Params=175B, Zero-shot=true2023.12 | 60 | — | — | — | |
| ViperGPTTraining Protocol=Zero-shot2024.04 | 60 | — | — | — | |
| CoVGTEvaluation Protocol=finetuned, Video Pretrain=true, Params=149M2024.03 | 60 | 58.8 | 57.4 | 69.3 | |
| ViperGPTEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=175B2024.03 | 60 | — | — | — | |
| LLaVA-NeXT-Interleave-0.5BTraining Mode=SFT, Video Data=true, Model Size=0.5B2024.09 | 59.5 | — | — | — | |
| GF(uns)Backbone=CLIP2024.01 | 58.83 | 56.93 | 57.07 | 70.53 | |
| GFEvaluation Protocol=Supervised2024.03 | 58.8 | 56.9 | 57.1 | 70.5 | |
| AssistGPTEvaluation Protocol=Zero-shot2024.03 | 58.4 | 60 | 51.4 | 67.3 | |
| ATM2023.09 | 58.27 | 56.04 | 58.44 | 65.38 | |
| LLoViEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=false, Params=12B2024.03 | 58.2 | 60.2 | 51.2 | 66 | |
| MISTEvaluation Protocol=Supervised2024.03 | 57.2 | 54.6 | 56.6 | 66.9 | |
| LLaVAOneVision-0.5BTraining Mode=SFT, Video Data=true, Model Size=0.5B2024.09 | 57.2 | — | — | — | |
| MISTbackbone=CLIP2022.12 | 57.18 | 54.62 | 56.64 | 66.92 | |
| MISTBackbone=CLIP2024.01 | 57.18 | 54.62 | 56.64 | 66.92 | |
| PAXIONPatcher Training Loss=VTC+DVDM2023.05 | 57 | 56 | 53 | 68.5 | |
| VGT# PT Data=0.18M2022.12 | 56.9 | 53.4 | 56.4 | 69.5 | |
| VGT (PT)Pretrain Dataset=WebVid(WV), Pretrain Size=0.18M, Video=R, F, Text=BERT2023.02 | 56.89 | 53.43 | 56.39 | 69.5 | |
| VGTPre-training=WebVid-2M2023.09 | 56.89 | 53.43 | 56.39 | 69.5 | |
| SeViT_FiDFrame Sampling=Frame Retrieval2023.01 | 56.7 | 54 | 54.1 | 71.3 | |
| SeViTEvaluation Protocol=Supervised2024.03 | 56.7 | 54 | 54.1 | 71.3 | |
| GF(uns)Backbone=C3D2024.01 | 56.33 | 53.59 | 55.71 | 66.8 | |
| SeViT_FiDFrame Sampling=Uniform Sampling2023.01 | 56.3 | 53 | 54.1 | 71.9 | |
| Side-TuningPatcher Training Loss=VTC+DVDM2023.05 | 56.3 | 54.9 | 52 | 69.8 | |
| SeViT_MARFrame Sampling=Frame Retrieval2023.01 | 56.1 | 53.5 | 54 | 69.2 | |
| DoraemonGPTBase Model=GPT-3.5-turbo, Frames=28.7 fps / 1144.4, #LLM calls=8.52025.12 | 55.7 | 54.7 | 50.4 | 70.3 | |
| SeViT_MARFrame Sampling=Uniform Sampling2023.01 | 55.2 | 52.3 | 52.3 | 71.2 | |
| VGT2022.07 | 55.02 | 52.28 | 55.09 | 64.09 | |
| VGT2022.12 | 55.02 | 52.28 | 55.09 | 64.09 |