Multiple-choice Video Question Answering on TVQA (test)
57.8AccuracyGPT-4V
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4VVision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 57.8 | |
| LLaVA v1.6Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 51.1 | |
| IG-VLM LLaVA v1.6Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 44.5 | |
| LLaVA v1.6Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 42.1 | |
| VideoChat2Vision Encoder=UMT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=true2024.03 | 40.6 | |
| CogAgentVision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 38.6 | |
| SevillaVision Encoder=ViT-L, LLM Size=2.85B, Inference Vision=multiple, Inference LLM=multiple, Video Trained=true2024.03 | 38.2 | |
| InternVideoVision Encoder=ViT-L, LLM Size=1.3B, Inference Vision=multiple, Inference LLM=single, Video Trained=true2024.03 | 35.9 |