Video Text Generation on Video-ChatGPT benchmark v1 (test)
3.4Correctness of Information (CI)GPT-4V
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GPT-4VVision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=No2024.03 | 3.4 | 2.8 | 3.61 | 2.89 | 3.13 | 3.17 | |
| CogAgentVision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=No2024.03 | 3.26 | 2.76 | 3.57 | 2.34 | 3.28 | 3.04 | |
| IG-VLM LLaVA v1.6Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=No2024.03 | 3.21 | 2.87 | 3.54 | 2.51 | 3.34 | 3.09 | |
| LLaVA v1.6 (13B)Vision Encoder=ViT-L, LLM Size=13B, Inference Vision=single, Inference LLM=single, Video Trained=No2024.03 | 3.17 | 2.79 | 3.52 | 2.51 | 3.25 | 3.05 | |
| LLaVA v1.6 (7B)Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=No2024.03 | 3.11 | 2.78 | 3.51 | 2.44 | 3.29 | 3.03 | |
| VideoChat2Vision Encoder=UMT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 3.02 | 2.88 | 3.51 | 2.66 | 2.81 | 2.98 | |
| LLAMA-VIDVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 2.96 | 3 | 3.53 | 2.46 | 2.51 | 2.89 | |
| Chat-UniViVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 2.89 | 2.91 | 3.46 | 2.89 | 2.81 | 2.99 | |
| MovieChatVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 2.76 | 2.93 | 3.01 | 2.24 | 2.42 | 2.67 | |
| Video-ChatGPTVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 2.5 | 2.57 | 2.69 | 2.16 | 2.2 | 2.42 | |
| Vista-LLaMAVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 2.44 | 2.64 | 3.18 | 2.26 | 2.31 | 2.57 | |
| VideoChatVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 2.23 | 2.5 | 2.53 | 1.94 | 2.24 | 2.29 | |
| LLaMA-AdapterVision Encoder=ViT-B, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 2.03 | 2.32 | 2.3 | 1.98 | 2.15 | 2.16 | |
| Video-LLaMAVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=Yes2024.03 | 1.96 | 2.18 | 2.16 | 1.82 | 1.79 | 1.98 |