Open Ended Question Answering on ActivityNet
62.3AccuracyLLaVA-OneVision
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaVA-OneVisionModel Size=72B, # Frames=322024.06 | 62.3 | — | |
| GPT4-O2024.06 | 61.9 | — | |
| PLLAVAModel Size=34B, # Frames=162024.06 | 60.9 | — | |
| Pegasus-12024.06 | 59.9 | — | |
| GPT4-V2024.06 | 59.5 | — | |
| Gemini 1.5 Pro2024.06 | 56.7 | — | |
| VideoLLaMA 2Model Size=72B, # Frames=82024.06 | 55.2 | 3.4 | |
| LLaVA-NeXT-VideoModel Size=32B, # Frames=322024.06 | 54.3 | — | |
| LLaVA-NeXT-VideoModel Size=7B, # Frames=322024.06 | 53.5 | 3.2 | |
| VideoLLaMA 2.1Model Size=7B, # Frames=162024.06 | 53 | 3.4 | |
| Gemini 1.0 Ultra2024.06 | 52.2 | — | |
| VideoLLaMA 2Model Size=8x7B, # Frames=82024.06 | 50.3 | 3.4 | |
| VideoLLaMA 2Model Size=7B, # Frames=162024.06 | 50.2 | 3.3 | |
| VideoLLaMA 2Model Size=7B, # Frames=82024.06 | 49.9 | 3.3 | |
| Gemini 1.0 Pro2024.06 | 49.8 | — | |
| VideoChat2Model Size=7B, # Frames=162024.06 | 49.1 | 3.3 | |
| LLaMA-VID-7BUsing Subtitles=false, Zero-shot=true2024.04 | 47.4 | 3.3 | |
| LLaMA-VIDModel Size=7B, # Frames=1 fps2024.06 | 47.4 | 3.3 | |
| Chat-UniViModel Size=7B, # Frames=82024.06 | 46.1 | 3.3 | |
| MiniGPT4-VideoBackbone=Llama 2-7B, Using Subtitles=false, Zero-shot=true2024.04 | 45.85 | 3.23 | |
| Video-LLaVAModel Size=7B, # Frames=82024.06 | 45.3 | 3.3 | |
| MiniGPT4-VideoBackbone=Mistral-7B, Using Subtitles=false, Zero-shot=true2024.04 | 44.25 | 3.35 | |
| Video-ChatGPTUsing Subtitles=false, Zero-shot=true2024.04 | 35.2 | 2.7 | |
| Video-ChatGPTModel Size=7B, # Frames=82024.06 | 35.2 | 2.7 | |
| Video ChatUsing Subtitles=false, Zero-shot=true2024.04 | 26.5 | — | |
| VideoChatModel Size=7B, # Frames=82024.06 | 26.5 | 2.2 | |
| FrozenBiLMUsing Subtitles=false, Zero-shot=true2024.04 | 24.7 | — | |
| Video LLaMAUsing Subtitles=false, Zero-shot=true2024.04 | 12.4 | 1.1 | |
| VideoLLaMAModel Size=7B, # Frames=82024.06 | 12.4 | 1.1 |