Video Question Answering on VLEP
71Total AccuracyLLaMA-VQA
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaMA-VQALanguage Model=LLaMA, Number of Trainable Parameters=4.5M2023.10 | 71 | |
| ViLAFrames Numbers=42023.12 | 69.6 | |
| ViLAFrames Numbers=22023.12 | 69.2 | |
| SeViLAFrames Numbers=42023.12 | 68.9 | |
| MERLOTLanguage Model=ROBERTa, Number of Trainable Parameters=223M2023.10 | 68.4 | |
| BLIP-2Frames Numbers=42023.12 | 67 | |
| InternVideoLanguage Model=CLIP text encoder, Number of Trainable Parameters=1.3B2023.10 | 63.9 | |
| InternVideoFrames Numbers=82023.12 | 63.9 |