Video Question Answering on EgoSchema 500-question subset
71.2AccuracyMMCTAgent
Evaluation Results
| Method | Links | |
|---|---|---|
| MMCTAgentCritic=Included2024.05 | 71.2 | |
| MMCTAgentCritic=Excluded2024.05 | 68.8 | |
| TarsierEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=34B2024.03 | 68.6 | |
| LVNetEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=<1.8T2024.03 | 68.2 | |
| LifelongMemoryEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 68 | |
| VideoTreeEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T2024.03 | 66.2 | |
| LangRepoEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=false, Params=12B2024.03 | 66.2 | |
| VideoChat2Evaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=7B2024.03 | 63.6 | |
| GPT-4V2024.03 | 63.5 | |
| GPT-4V2024.03 | 63.5 | |
| GPT-4V2024.05 | 63.5 | |
| VideoAgent-M2024.05 | 62.8 | |
| VideoAgentEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T, citation=[Fan et al., 2024]2024.03 | 62.8 | |
| MC-ViT-LFrames=128+2024.03 | 62.6 | |
| MC-ViT-L2024.05 | 62.6 | |
| MC-ViT-LEvaluation Protocol=finetuned, Video Pretrain=true, Params=424M2024.03 | 62.6 | |
| Gemini 1.0 Pro2024.05 | 61.5 | |
| LangRepoEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=false, Params=7B2024.03 | 60.8 | |
| VideoAgentFrames=8.42024.03 | 60.2 | |
| VideoAgent2024.03 | 60.2 | |
| VideoAgent2024.05 | 60.2 | |
| VideoAgentEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=1.8T, citation=[Wang et al., 2024b]2024.03 | 60.2 | |
| LLoViFrames=1802024.03 | 57.6 | |
| LLoVi2024.05 | 57.6 | |
| LLoViEvaluation Protocol=zero-shot (with proprietary LLMs), Video Pretrain=false, Params=175B2024.03 | 57.6 | |
| LongViViTFrames=2562024.03 | 56.8 | |
| LongViViTEvaluation Protocol=finetuned, Video Pretrain=true, Params=1B2024.03 | 56.8 | |
| TarsierEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=7B2024.03 | 56 | |
| LLoViEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=false, Params=7B2024.03 | 50.8 | |
| ShortViViT_locFrames=322024.03 | 49.6 | |
| MistralEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=false, Params=7B2024.03 | 48.8 | |
| ShortViViTEvaluation Protocol=finetuned, Video Pretrain=true, Params=1B2024.03 | 47.9 | |
| Bard + PALI2024.03 | 44.8 | |
| Bard + PALI2024.03 | 44.8 | |
| Bard + ShortViViT2024.03 | 42 | |
| Bard + ShortViViT2024.03 | 42 | |
| ImageViTFrames=162024.03 | 40.8 | |
| ImageViTEvaluation Protocol=finetuned, Video Pretrain=true, Params=1B2024.03 | 40.8 | |
| Video-LLaVa2024.05 | 36.8 | |
| Bard + ImageViT2024.03 | 35 | |
| Bard + ImageViT2024.03 | 35 | |
| GPT-4 Turbo (blind)Blind=true2024.03 | 31 | |
| GPT-4 Turbomode=blind2024.03 | 31 | |
| Bard only (blind)Blind=true2024.03 | 27 | |
| Bardmode=blind2024.03 | 27 | |
| SeViLAFrames=322024.03 | 25.7 | |
| SeViLAEvaluation Protocol=zero-shot (with open-source LLMs), Video Pretrain=true, Params=4B2024.03 | 25.7 | |
| Random Chance2024.03 | 20 | |
| Random Chance2024.03 | 20 | |
| ViperGPT2024.05 | 15.8 |