Video Question Answering on NExT-GQA (test)
39.6Acc@GQAMoReVQA
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MoReVQAVideo Pretrain=false, Params=340B, Protocol=zero-shot (proprietary LLMs)2024.03 | 39.6 | 37.8 | 37.6 | 19.7 | 15.4 | |
| MoReVQAFine-tuned=false2024.04 | 31 | 44.1 | 43 | 20.7 | 18.3 | |
| VITED (Temporal Evidence Distillation)Params=7B, Base Model=TimeChat Base2025.03 | 27.61 | — | — | — | — | |
| VITED (Temporal Evidence Distillation)Params=7B, Base Model=LLaVA-Video Base2025.03 | 25.19 | — | — | — | — | |
| LLoViVideo Pretrain=false, Params=1.8T, Protocol=zero-shot (proprietary LLMs)2024.03 | 24.3 | 37.3 | 36.9 | 20 | 15.3 | |
| FrozenBiLMVideo Pretrain=true, Params=1B, Protocol=weakly-supervised2024.03 | 17.5 | 24.2 | 23.7 | 9.6 | 6.1 | |
| FrozenBiLMFine-tuned=true2024.04 | 17.5 | 24.2 | 23.7 | 9.6 | 6.1 | |
| LangRepoVideo Pretrain=false, Params=12B, Protocol=zero-shot (open-source LLMs)2024.03 | 17.1 | 31.3 | 28.7 | 18.5 | 12.2 | |
| SeViLAVideo Pretrain=true, Params=4B, Protocol=weakly-supervised2024.03 | 16.6 | 29.5 | 22.9 | 21.7 | 13.8 | |
| SeViLAParams=4B2025.03 | 16.6 | — | — | — | — | |
| SeViLAFine-tuned=true, Explicit localization supervision during pretraining=true2024.04 | 16.6 | 29.5 | 22.9 | 21.7 | 13.8 | |
| LLoViVideo Pretrain=false, Params=12B, Protocol=zero-shot (open-source LLMs)2024.03 | 16.2 | 31.4 | 28.8 | 18.4 | 12 | |
| Temp[CLIP]Video Pretrain=true, Params=130M, Protocol=weakly-supervised2024.03 | 16 | 25.7 | 25.5 | 12.1 | 8.9 | |
| Temp[CLIP]Fine-tuned=true2024.04 | 16 | 25.7 | 25.5 | 12.1 | 8.9 | |
| LLaMA-3.2VParams=11B2025.03 | 11.64 | — | — | — | — | |
| LLoViVideo Pretrain=false, Params=7B, Protocol=zero-shot (open-source LLMs)2024.03 | 11.2 | 20.7 | 20.5 | 8.7 | 6 | |
| LangRepoVideo Pretrain=false, Params=7B, Protocol=zero-shot (open-source LLMs)2024.03 | 11.2 | 20.3 | 20 | 8.7 | 6 | |
| IGVVideo Pretrain=true, Params=110M, Protocol=weakly-supervised2024.03 | 10.2 | 21.4 | 18.9 | 14 | 9.6 | |
| IGVFine-tuned=true2024.04 | 10.2 | 21.4 | 18.9 | 14 | 9.6 | |
| MistralVideo Pretrain=false, Params=7B, Protocol=zero-shot (open-source LLMs)2024.03 | 9.2 | 20.4 | 20.2 | 8.7 | 5.9 | |
| LLaVA-VideoParams=7B, Base Model=LLaVA-Video Base2025.03 | 0.04 | — | — | — | — | |
| LLaVA-Video (Chain-of-Thought)Params=7B, Base Model=LLaVA-Video Base2025.03 | 0.03 | — | — | — | — | |
| LLaVA-OneVisionParams=7B2025.03 | 0 | — | — | — | — | |
| TimeChatParams=7B, Base Model=TimeChat Base2025.03 | 0 | — | — | — | — | |
| TimeChat (Chain-of-Thought)Params=7B, Base Model=TimeChat Base2025.03 | 0 | — | — | — | — | |
| TimeChat (Video Instruction Tuning)Params=7B, Base Model=TimeChat Base2025.03 | 0 | — | — | — | — | |
| VITED (Dense Caption Distillation)Params=7B, Base Model=TimeChat Base2025.03 | 0 | — | — | — | — | |
| LLaVA-Video (Video Instruction Tuning)Params=7B, Base Model=LLaVA-Video Base2025.03 | 0 | — | — | — | — | |
| VITED (Dense Caption Distillation)Params=7B, Base Model=LLaVA-Video Base2025.03 | 0 | — | — | — | — |