Video Generation Performance on VideoChatGPT (test)
2.34Temporal UnderstandingBT-Adapter-LLaVA (FT)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| BT-Adapter-LLaVA (FT)instruction tuning=with, backbone=Vicuna-1.0-7B2023.09 | 2.34 | 2.68 | 2.69 | 3.27 | 2.46 | 2.69 | |
| BT-Adapter-LLaVA (ZS)instruction tuning=without, backbone=Vicuna-1.0-7B2023.09 | 2.13 | 2.16 | 2.46 | 2.89 | 2.2 | 2.46 | |
| LLaMA-Adapter2023.09 | 1.98 | 2.03 | 2.32 | 2.3 | 2.15 | 2.16 | |
| VideoChatGPT2023.09 | 1.98 | 2.4 | 2.52 | 2.62 | 2.37 | 2.38 | |
| VideoChat2023.09 | 1.94 | 2.23 | 2.5 | 2.53 | 2.24 | 2.29 | |
| VideoLLaMA2023.09 | 1.82 | 1.96 | 2.18 | 2.16 | 1.79 | 1.98 |