Video Question Answering on VideoMME (long split)
67.4AccuracyGemini 1.5 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 1.5 ProModel category=Proprietary MLLM2024.05 | 67.4 | |
| GPT-4oModel category=Proprietary MLLM2024.05 | 65.3 | |
| Qwen2-VL-72BModel category=Open-Source MLLM2024.05 | 62.2 | |
| LLaVA-NeXT-Video-72BModel category=Open-Source MLLM2024.05 | 61.5 | |
| Oryx-1.5-34BModel category=Open-Source MLLM2024.05 | 59.3 | |
| VIDEOTREEModel category=Training-free Approach2024.05 | 54.2 | |
| LVNet (GPT-4o)LLM Param/Type=<1.8T/PP, VT Free=true, TS Cap.=242024.06 | 53.9 | |
| VILA-1.5-40BModel category=Open-Source MLLM2024.05 | 53.8 | |
| GPT-4VModel category=Proprietary MLLM2024.05 | 53.5 | |
| LVNet (DS-V3)LLM Param/Type=37B/OS, VT Free=true, TS Cap.=242024.06 | 53.1 | |
| InternVL2-34BModel category=Open-Source MLLM2024.05 | 52.6 | |
| VideoTree+GPT-4oLLM Param/Type=<1.8T/PP, VT Free=true, TS Cap.=242024.06 | 52.3 | |
| VideoAgent+GPT-4oLLM Param/Type=<1.8T/PP, VT Free=true, TS Cap.=242024.06 | 51.3 | |
| Frame-VoyagerLLM Param/Type=34B/OS, VT Free=false, TS Cap.=N/A2024.06 | 51.2 | |
| LLoViModel category=Training-free Approach2024.05 | 48.8 | |
| VITAModel category=Open-Source MLLM2024.05 | 48.6 | |
| LongVAModel category=Open-Source MLLM2024.05 | 46.2 | |
| VideoChat-TLLM Param/Type=7B/OS, VT Free=false, TS Cap.=N/A2024.06 | 43.8 |