Video Question Answering on iVQA (test)
60.9AccuracyMoReVQA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MoReVQAFT=false2024.04 | 60.9 | — | — | |
| JCEFFT=false2024.04 | 56.9 | — | — | |
| InstructBLIPFT=false, Variant=FlanT5XXL2024.04 | 53.8 | — | — | |
| InstructBLIPFT=false, Variant=FlanT5XL2024.04 | 51.1 | — | — | |
| ProViQEvaluation protocol=zero-shot2023.12 | 50.7 | — | — | |
| BLIP-2FT=false, Variant=FlanT5XXL2024.04 | 45.8 | — | — | |
| Supervised SOTAEvaluation protocol=supervised2023.12 | 40.9 | — | — | |
| Flamingo-80BAdaptation type=Training based Adaptation, Language model=Chinchilla-like, Language model #params=80B, Vision model=NFNet-F6, Vision model #params=629M, Evaluation protocol=Zero-shot2023.06 | 40.7 | — | — | |
| Text Tokens + Text TransformerFv, Ft=CLIP [26], Extra MM Samples=0, Delta GPU hours=0, ASR=true2022.06 | 40.2 | — | — | |
| Text+TextFT=false2024.04 | 40.2 | — | — | |
| FrozenBiLMFv, Ft=CLIP [26], Extra MM Samples=10M, Delta GPU hours=160, ASR=false2022.06 | 39.7 | — | — | |
| FrozenBiLMFT=true2024.04 | 39.7 | — | — | |
| FrozenBiLMFv, Ft=CLIP [26], Extra MM Samples=10M, Delta GPU hours=160, ASR=true2022.06 | 39.6 | — | — | |
| VideoCoCaFT=false2024.04 | 39 | — | — | |
| Text Tokens + Text TransformerFv, Ft=CLIP [26], Extra MM Samples=0, Delta GPU hours=0, ASR=false2022.06 | 36.9 | — | — | |
| Text Tokens + Text TransformerFv, Ft=S3D [21], Extra MM Samples=0, Delta GPU hours=0, ASR=true2022.06 | 36.8 | — | — | |
| Continuous Features + Multimodal TransformerFv, Ft=S3D [21], Extra MM Samples=69M, Delta GPU hours=400, ASR=true2022.06 | 36 | — | — | |
| Continuous Features + Multimodal TransformerFv, Ft=S3D [21], Extra MM Samples=69M, Delta GPU hours=400, ASR=false2022.06 | 35.4 | — | — | |
| VQA-TFv, Ft=S3D [21], Extra MM Samples=69M + 3M, Delta GPU hours=350 + 30, ASR=false2022.06 | 35.2 | — | — | |
| Flamingo-9BAdaptation type=Training based Adaptation, Language model=Chinchilla-like, Language model #params=8.7B, Vision model=NFNet-F6, Vision model #params=629M, Evaluation protocol=Zero-shot2023.06 | 35.2 | — | — | |
| Flamingo-3BAdaptation type=Training based Adaptation, Language model=Chinchilla-like, Language model #params=2.6B, Vision model=NFNet-F6, Vision model #params=629M, Evaluation protocol=Zero-shot2023.06 | 32.7 | — | — | |
| Text Tokens + Text TransformerFv, Ft=S3D [21], Extra MM Samples=0, Delta GPU hours=0, ASR=false2022.06 | 31.6 | — | — | |
| R2AAdaptation type=Language based Adaptation, Language model=DeBERTa-v2-XL, Language model #params=890M, Vision model=ViT-L/14, Vision model #params=300M, Evaluation protocol=Zero-shot2023.06 | 29.3 | — | — | |
| FrozenBiLMFT=false2024.04 | 27.3 | — | — | |
| FrozenBiLMEvaluation protocol=zero-shot2023.12 | 26.8 | — | — | |
| FrozenBiLMAdaptation type=Training based Adaptation, Language model=DeBERTa-v2-XL, Language model #params=890M, Vision model=ViT-L/14, Vision model #params=300M, Evaluation protocol=Zero-shot2023.06 | 26.2 | — | — | |
| Just AskAdaptation type=Training based Adaptation, Language model=DistilBERT, Language model #params=66M, Vision model=S3D, Vision model #params=12M, Evaluation protocol=Zero-shot2023.06 | 13.3 | — | — | |
| Just AskEvaluation protocol=zero-shot2023.12 | 13.3 | — | — | |
| CLIP VIT-L/14Adaptation type=Training based Adaptation, Language model=Custom, Language model #params=123M, Vision model=ViT-L/14, Vision model #params=300M, Evaluation protocol=Zero-shot2023.06 | 9.2 | — | — | |
| CLIP-VIT-L/14Evaluation protocol=zero-shot2023.12 | 9.2 | — | — | |
| RandomEvaluation protocol=zero-shot2023.12 | 0.1 | — | — | |
| QA-TPretraining Data=HowToVQA69M, Zero-shot=true2020.12 | — | 4.4 | 23.2 | |
| RandomPretraining Data=None, Zero-shot=true2020.12 | — | 0.09 | 0.9 | |
| VQA-TPretraining Data=HowTo100M, Zero-shot=true2020.12 | — | 1.9 | 11.9 | |
| VQA-TPretraining Data=HowToVQA69M, Zero-shot=true2020.12 | — | 12.2 | 43.3 |