Video Question Answering on ActivityNet (test)
62.3AccuracyLLaVA-OneVision
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaVA-OneVisionLLM size=72B, Video-Gen=false2024.12 | 62.3 | — | |
| GPT4-OVideo-Gen=false2024.12 | 61.9 | — | |
| PLLAVALLM size=34B, Video-Gen=false2024.12 | 60.9 | — | |
| GPT4-VVideo-Gen=false2024.12 | 59.5 | — | |
| Tarsier# Frames=16, # Tokens=23042025.03 | 59.5 | 3.6 | |
| HIComModel Scale=7B, # Frames=128, # Tokens=26242025.03 | 59.5 | 3.7 | |
| HIComModel Scale=7B, # Frames=64, # Tokens=13282025.03 | 59.4 | 3.7 | |
| InternVL2.5Model Size=7-8B, Zero-shot evaluation=true2025.12 | 58.9 | — | |
| HIComModel Scale=7B, # Frames=32, # Tokens=6802025.03 | 58.3 | 3.7 | |
| JavisGPTModel Size=7-8B, #Samples=1.5M, Zero-shot evaluation=true2025.12 | 58.1 | — | |
| Qwen2-VLModel Size=7-8B, Zero-shot evaluation=true2025.12 | 57.4 | — | |
| Qwen2.5-OmniModel Size=7-8B, Zero-shot evaluation=true2025.12 | 57.2 | — | |
| Gemini 1.5 ProVideo-Gen=false2024.12 | 56.7 | — | |
| LLaVA-OneVision# Frames=32, # Tokens=62722025.03 | 56.6 | 3.6 | |
| Video-LLaVAModel Size=7-8B, Zero-shot evaluation=true2025.12 | 56.5 | — | |
| PLLAVA# Frames=16, # Tokens=23042025.03 | 56.3 | 3.5 | |
| SF-LLAVA# Frames=50, # Tokens=36802025.03 | 56.3 | 3.4 | |
| Divot-LLMLLM size=7B, Video-Gen=true2024.12 | 55.8 | — | |
| LLaVA-OVModel Size=7-8B, Zero-shot evaluation=true2025.12 | 55.6 | — | |
| VideoLLaMA2LLM size=72B, Video-Gen=false2024.12 | 55.2 | — | |
| LLaVA-NeXT-VideoLLM size=32B, Video-Gen=false2024.12 | 54.3 | — | |
| LLaVA-NeXT-VideoLLM size=7B, Video-Gen=false2024.12 | 53.5 | — | |
| LLaVA-Next-Video# Frames=32, # Tokens=46082025.03 | 53.5 | 3.2 | |
| LLaVA-NeXTModel Size=7-8B, Zero-shot evaluation=true2025.12 | 53.5 | — | |
| HIComModel Scale=1.5B, # Frames=32, # Tokens=6802025.03 | 53 | 3.5 | |
| VideoLLaMA2.1Model Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 53 | — | |
| VILA-ULLM size=7B, Video-Gen=true2024.12 | 52.7 | — | |
| ST-LLM# Frames=16, # Tokens=5122025.03 | 50.9 | 3.3 | |
| VideoLLaMA2LLM size=7B, Video-Gen=false2024.12 | 50.2 | — | |
| VideoLLaMA2# Frames=16, # Tokens=11522025.03 | 50.2 | 3.3 | |
| VideoLLaMA2Model Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 50.2 | — | |
| Video-LaVITLLM size=7B, Video-Gen=true2024.12 | 50.1 | — | |
| Gemini 1.0 ProVideo-Gen=false2024.12 | 49.8 | — | |
| VideoChat2LLM size=7B, Video-Gen=false2024.12 | 49.1 | — | |
| VideoChat2# Frames=16, # Tokens=15362025.03 | 49.1 | 3.3 | |
| LongVLM# Frames=100, # Tokens=3052025.03 | 47.6 | 3.3 | |
| LLaMA-VIDLLM size=7B, Video-Gen=false2024.12 | 47.4 | — | |
| LLaMA-VID# Frames=1fps, # Tokens=2tps2025.03 | 47.4 | 3.3 | |
| AV-LLMModel Size=7-8B, Zero-shot evaluation=true2025.12 | 47.2 | — | |
| Chat-Univi# Frames=64, # Tokens=4482025.03 | 45.8 | 3.2 | |
| Video-LLaVALLM size=7B, Video-Gen=false2024.12 | 45.3 | — | |
| Video-LLaVA# Frames=8, # Tokens=20482025.03 | 45.3 | 3.3 | |
| FrozenBiLMFv, Ft=CLIP [26], Extra MM Samples=10M, Delta GPU hours=160, ASR=false2022.06 | 43.2 | — | |
| FrozenBiLMFv, Ft=CLIP [26], Extra MM Samples=10M, Delta GPU hours=160, ASR=true2022.06 | 43.2 | — | |
| MERLOTExtra MM Samples=180M, ASR=false2022.06 | 41.4 | — | |
| Text Tokens + Text TransformerFv, Ft=CLIP [26], Extra MM Samples=0, Delta GPU hours=0, ASR=false2022.06 | 41.4 | — | |
| Text Tokens + Text TransformerFv, Ft=CLIP [26], Extra MM Samples=0, Delta GPU hours=0, ASR=true2022.06 | 41.4 | — | |
| SiaSamReaExtra MM Samples=5.6M + 80K, ASR=false2022.06 | 39.8 | — | |
| VQA-TFv, Ft=S3D [21], Extra MM Samples=69M + 3M, Delta GPU hours=350 + 30, ASR=false2022.06 | 39 | — | |
| Continuous Features + Multimodal TransformerFv, Ft=S3D [21], Extra MM Samples=69M, Delta GPU hours=400, ASR=false2022.06 | 38.9 | — | |
| Continuous Features + Multimodal TransformerFv, Ft=S3D [21], Extra MM Samples=69M, Delta GPU hours=400, ASR=true2022.06 | 38.9 | — | |
| Text Tokens + Text TransformerFv, Ft=S3D [21], Extra MM Samples=0, Delta GPU hours=0, ASR=true2022.06 | 38.8 | — | |
| Text Tokens + Text TransformerFv, Ft=S3D [21], Extra MM Samples=0, Delta GPU hours=0, ASR=false2022.06 | 38.7 | — | |
| Video-ChatGPTLLM size=7B, Video-Gen=false2024.12 | 35.2 | — | |
| UnifiedIO-2Model Size=7-8B, #Samples=9.2B, Zero-shot evaluation=true2025.12 | 23.2 | — | |
| NExT-GPTModel Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 21.5 | — | |
| VideoLLaMAModel Size=7-8B, Zero-shot evaluation=true2025.12 | 12.4 | — | |
| LongVA# Frames=128, # Tokens=184322025.03 | — | 2.8 |