Video Question Answering on STAR
78.6AccuracyInternVideo2
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVideo2Params=6B2025.03 | 78.6 | |
| VITED (Temporal Evidence Distillation)Params=7B, Base Model=TimeChat Base2025.03 | 70.99 | |
| Open-o3-Video-7BParameters=7B2025.10 | 70.5 | |
| Open-o3-Video-7BParameters=7B, Ablation=without Adaptive frame sampling2025.10 | 70.1 | |
| VITED (Dense Caption Distillation)Params=7B, Base Model=TimeChat Base2025.03 | 69.99 | |
| Open-o3-Video-7BParameters=7B, Ablation=without Gated mechanism2025.10 | 69.6 | |
| TimeChat (Video Instruction Tuning)Params=7B, Base Model=TimeChat Base2025.03 | 68.39 | |
| VITED (Temporal Evidence Distillation)Params=7B, Base Model=LLaVA-Video Base2025.03 | 68.23 | |
| LLaVA-Video (Video Instruction Tuning)Params=7B, Base Model=LLaVA-Video Base2025.03 | 67.74 | |
| VITED (Dense Caption Distillation)Params=7B, Base Model=LLaVA-Video Base2025.03 | 67.56 | |
| Qwen2.5-VL-7BParameters=7B2025.10 | 67.3 | |
| LLaVA-VideoParams=7B, Base Model=LLaVA-Video Base2025.03 | 66.88 | |
| LLaVA-OneVisionParams=7B2025.03 | 66.24 | |
| LLaVA-Video (Chain-of-Thought)Params=7B, Base Model=LLaVA-Video Base2025.03 | 65.34 | |
| SeViLAParams=4B2025.03 | 64.9 | |
| KTV-34B-denseLLM Size=34B, Vis Encoder=CLIP-L, Training Regime=Training-Free, Variant=dense2026.02 | 54.7 | |
| KTV-34B-normalLLM Size=34B, Vis Encoder=CLIP-L, Training Regime=Training-Free, Variant=normal2026.02 | 54.6 | |
| KTV-34B-sparseLLM Size=34B, Vis Encoder=CLIP-L, Training Regime=Training-Free, Variant=sparse2026.02 | 54.2 | |
| KTV-7B-denseLLM Size=7B, Vis Encoder=CLIP-L, Training Regime=Training-Free, Variant=dense2026.02 | 52.7 | |
| KTV-7B-normalLLM Size=7B, Vis Encoder=CLIP-L, Training Regime=Training-Free, Variant=normal2026.02 | 52.5 | |
| KTV-7B-sparseLLM Size=7B, Vis Encoder=CLIP-L, Training Regime=Training-Free, Variant=sparse2026.02 | 52.3 | |
| UIO-2XXLModel Size=XXL2023.12 | 52.2 | |
| UIO-2XLModel Size=XL2023.12 | 52 | |
| SF-LLaVA-7BLLM Size=34B, Vis Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 51.3 | |
| DYTOLLM Size=34B, Vis Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 51.1 | |
| UIO-2LModel Size=L2023.12 | 51 | |
| DYTOLLM Size=7B, Vis Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 50.7 | |
| IG-VLMLLM Size=34B, Vis Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 50.5 | |
| SF-LLaVA-7BLLM Size=7B, Vis Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 48.8 | |
| IG-VLMLLM Size=7B, Vis Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 48.6 | |
| AnyMAL-Image 70Bzero-shot=true, vision encoder=ViT-G, LLM size=70B, number of frames=42023.09 | 48.2 | |
| LLaMA-3.2VParams=11B2025.03 | 45.62 | |
| TraveLERTraining Protocol=Zero-shot2024.04 | 44.9 | |
| SeViLATraining Protocol=Uses fine-tuned components2024.04 | 44.6 | |
| AnyMAL-Image 13Bzero-shot=true, vision encoder=ViT-G, LLM size=13B, number of frames=42023.09 | 44.4 | |
| BLIPv2 ViTG FlanT5xxlzero-shot=true, vision encoder=ViTG, LLM=FlanT5xxl, number of frames=42023.09 | 42.2 | |
| BLIP-2concatTraining Protocol=Zero-shot2024.04 | 42.2 | |
| Flamingo-9Bzero-shot=true, LLM size=9B2023.09 | 41.8 | |
| Flamingo-9BTraining Protocol=Zero-shot2024.04 | 41.8 | |
| Internvideozero-shot=true, number of frames=82023.09 | 41.6 | |
| InternVideoTraining Protocol=Zero-shot2024.04 | 41.6 | |
| AnyMAL-Video 70Bzero-shot=true, vision encoder=Internvideo, LLM size=70B, number of frames=82023.09 | 41.3 | |
| FlamingoModel Size=9B, Few-shot=true2023.12 | 41.2 | |
| BLIP-2votingTraining Protocol=Zero-shot2024.04 | 40.3 | |
| Flamingo-80Bzero-shot=true, LLM size=80B2023.09 | 39.7 | |
| InstructBLIPZero-shot=true2023.12 | 38.3 | |
| AnyMAL-Video 13Bzero-shot=true, vision encoder=Internvideo, LLM size=13B, number of frames=82023.09 | 37.5 | |
| BLIP-2Zero-shot=true2023.12 | 36.7 | |
| TimeChatParams=7B, Base Model=TimeChat Base2025.03 | 21.03 | |
| TimeChat (Chain-of-Thought)Params=7B, Base Model=TimeChat Base2025.03 | 17.54 |