Video Question Answering on MF2
59.9AccuracyVideo Panels (LLaVA-Video 72B)
Evaluation Results
| Method | Links | |
|---|---|---|
| Video Panels (LLaVA-Video 72B)#frames=64, Model Context Category=Medium-context VLMs2025.09 | 59.9 | |
| VideoLLaMA 3 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 58.9 | |
| Video Panels (LLaVA-OV 72B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 58.5 | |
| Video Panels (VideoLLaMA 3 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 58.3 | |
| LLaVA-Video 72B#frames=64, Model Context Category=Medium-context VLMs2025.09 | 58.2 | |
| Video Panels (Qwen-2VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 56.7 | |
| LLaVA-OV 72B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 56.6 | |
| Video Panels (Qwen-2.5VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 54.8 | |
| Video Panels (LLaVA-Video 7B)#frames=64, Model Context Category=Medium-context VLMs2025.09 | 54.4 | |
| Qwen-2VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 54.3 | |
| Qwen-2.5VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 54.2 | |
| Video Panels (Qwen-2.5VL 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 53.8 | |
| LLaVA-Video 7B#frames=64, Model Context Category=Medium-context VLMs2025.09 | 52.8 | |
| Qwen-2.5VL 7B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 52.6 | |
| Video Panels (LLaVA-OV 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 52.1 | |
| LLaVA-OV 7B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 51.5 | |
| Video-LLaVA 7B#frames=8, Model Context Category=Small-context VLMs2025.09 | 50.4 | |
| Video Panels (Video-LLaVA 7B)#frames=8, Model Context Category=Small-context VLMs2025.09 | 50.2 | |
| Video Panels (LLaVA-OV 0.5B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 50.2 | |
| LLaVA-OV 0.5B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 50.1 | |
| VideoChat2-HD#frames=16, Model Context Category=Small-context VLMs2025.09 | 50 | |
| Video Panels (VideoChat2-HD)#frames=16, Model Context Category=Small-context VLMs2025.09 | 50 |