Multiple-Choice Video QA on EgoSchema latest (test)
72.2AccuracyGPT4-O
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT4-O2024.06 | 72.2 | |
| VideoLLaMA2Model Scale=72B, #Frames=162024.06 | 63.9 | |
| Gemini 1.5 ProModel Scale=Pro2024.06 | 63.2 | |
| LLaVA-OneVisionModel Scale=72B, #Frames=322024.06 | 62 | |
| Gemini 1.0 UltraModel Scale=Ultra2024.06 | 61.5 | |
| LLaVA-NeXT-VideoModel Scale=32B, #Frames=322024.06 | 60.9 | |
| VILA 1.5Model Scale=34B2024.06 | 58 | |
| Gemini 1.0 ProModel Scale=Pro2024.06 | 55.7 | |
| GPT4-V2024.06 | 55.6 | |
| VideoLLaMA2Model Scale=8x7B, #Frames=82024.06 | 53.3 | |
| VideoLLaMA2.1Model Scale=7B, #Frames=162024.06 | 53.1 | |
| VideoLLaMA2Model Scale=7B, #Frames=162024.06 | 51.7 | |
| VideoLLaMA2Model Scale=7B, #Frames=82024.06 | 50.5 | |
| LLaVA-NeXT-VideoModel Scale=7B, #Frames=322024.06 | 43.9 | |
| VideoChat2Model Scale=7B, #Frames=162024.06 | 42.2 | |
| LLAMA-VIDModel Scale=7B, #Frames=1 fps2024.06 | 38.5 | |
| Video-LLaVAModel Scale=7B, #Frames=82024.06 | 38.4 |