Video Question Answering on TVQA (val)
87.9QA AccuracyLlama 3-V 70B
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama 3-V 70BZero-shot=true, Parameters=70B2024.07 | 87.9 | |
| GPT-4VZero-shot=true2024.07 | 87.3 | |
| Llama 3-V 8BZero-shot=true, Parameters=8B2024.07 | 82.5 | |
| VX2TEXTSamples for Multimodal Pretext=02021.01 | 74.9 | |
| HEROSamples for Multimodal Pretext=7.6M2021.01 | 74.8 | |
| BERT QASamples for Multimodal Pretext=02021.01 | 72.4 | |
| MSANSamples for Multimodal Pretext=02021.01 | 71.6 | |
| HEROSamples for Multimodal Pretext=02021.01 | 70.7 | |
| STAGE backboneTemporal Supervision=true2019.04 | 70.5 | |
| STAGESamples for Multimodal Pretext=02021.01 | 70.5 | |
| STAGE backboneTemporal Supervision=false2019.04 | 68.56 | |
| TVQASamples for Multimodal Pretext=02021.01 | 67.7 | |
| STAGE backboneBackbone=GloVe, Temporal Supervision=true2019.04 | 66.92 | |
| STAGE backboneBackbone=GloVe, Temporal Supervision=false2019.04 | 66.46 | |
| PAMNTemporal Supervision=false2019.04 | 66.38 | |
| multi-taskTemporal Supervision=true2019.04 | 66.22 | |
| two-streamTemporal Supervision=false2019.04 | 65.85 |