Video Question Answering on STAR (test)
79.1Interaction ScoreIPRM
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| IPRMSetup=GT V2024.11 | 79.1 | 85.8 | 82 | 71.4 | — | 79.6 | — | — | — | — | — | |
| IPRMSetup=PR V2024.11 | 72 | 78.1 | 71.5 | 59.4 | — | 70.3 | — | — | — | — | — | |
| BLIP-2 + OursNumber of Frames=4, Text Augmentation=without2026.05 | 67.8 | 72.5 | 61.4 | 56.8 | — | — | — | — | — | — | — | |
| BLIP-2 + OursNumber of Frames=4, Text Augmentation=with2026.05 | 66.2 | 71.6 | 58.7 | 55.3 | — | — | — | — | — | — | — | |
| BLIP-2Number of Frames=4, Text Augmentation=without2026.05 | 65.4 | 69 | 59.7 | 54.2 | — | — | — | — | — | — | — | |
| InternVideo + OursNumber of Frames=8, Text Augmentation=without2026.05 | 63.8 | 67.7 | 58.9 | 55.2 | — | — | — | — | — | — | — | |
| InternVideoNumber of Frames=8, Text Augmentation=without2026.05 | 62.7 | 65.6 | 54.9 | 51.9 | — | — | — | — | — | — | — | |
| InternVideo + OursNumber of Frames=8, Text Augmentation=with2026.05 | 61.8 | 64.9 | 57.4 | 54.3 | — | — | — | — | — | — | — | |
| BLIP-2Number of Frames=4, Text Augmentation=with2026.05 | 60.9 | 66.3 | 54.3 | 50.1 | — | — | — | — | — | — | — | |
| mPLUGSetup=-2024.11 | 60.4 | 65.6 | 57.5 | 49.6 | — | 58.3 | — | — | — | — | — | |
| GLOBAL2026.01 | 59.36 | 57.52 | 57.21 | 46.68 | — | 52.85 | — | — | — | — | — | |
| VideoChat2Vision Encoder=UMT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=true2024.03 | 58.4 | 60.9 | 55.3 | 53.1 | — | 59 | — | — | — | — | — | |
| GF(sup)Supervision Type=supervised2024.01 | 56.1 | 61.27 | 52.65 | 45.74 | — | 53.94 | — | — | — | — | — | |
| GFSetup=-2024.11 | 56.1 | 61.3 | 52.7 | 45.7 | — | 53.9 | — | — | — | — | — | |
| InternVideoNumber of Frames=8, Text Augmentation=with2026.05 | 55.6 | 61 | 50.3 | 47.2 | — | — | — | — | — | — | — | |
| MISTBackbone=CLIP2026.01 | 55.59 | 54.23 | 54.24 | 44.48 | — | 51.13 | — | — | — | — | — | |
| LLaVA v1.6Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 53.4 | 53.9 | 49.5 | 48.4 | — | 53 | — | — | — | — | — | |
| MISTBackbone=AIO2026.01 | 53 | 52.37 | 49.52 | 43.87 | — | 49.69 | — | — | — | — | — | |
| GF(uns)Supervision Type=unsupervised2024.01 | 51.91 | 63.06 | 54.89 | 45.57 | — | 53.86 | — | — | — | — | — | |
| IG-VLM LLaVA v1.6Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 51.5 | 52 | 51 | 51.8 | — | 51.7 | — | — | — | — | — | |
| ATP2026.01 | 50.63 | 52.87 | 49.36 | 40.61 | — | 48.37 | — | — | — | — | — | |
| GPT-4VVision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 50.6 | 52.3 | 44.9 | 46.9 | — | 50.7 | — | — | — | — | — | |
| LLaVA v1.6Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 49.3 | 50.1 | 48.4 | 48.8 | — | 49.6 | — | — | — | — | — | |
| SevillaVision Encoder=ViT-L, LLM Size=2.85B, Inference Vision=multiple, Inference LLM=multiple, Video Trained=true2024.03 | 48.3 | 45 | 44.4 | 40.8 | — | 44.6 | — | — | — | — | — | |
| All-in-One + OursNumber of Frames=32, Text Augmentation=without2026.05 | 48.3 | 51.9 | 49.6 | 45.7 | — | — | — | — | — | — | — | |
| SHG-VQASetup=-2024.11 | 48 | 42 | 35.3 | 32.5 | — | 39.5 | — | — | — | — | — | |
| SHG-VQABackbone=SlowR50-K400 + ResNext101-K400, Obj.=false, Hyper.=true2023.04 | 47.98 | 42.03 | 35.34 | 32.52 | 39.47 | — | — | — | — | — | — | |
| All-in-One + OursNumber of Frames=32, Text Augmentation=with2026.05 | 47.9 | 51.3 | 48.7 | 44.3 | — | — | — | — | — | — | — | |
| AIO2026.01 | 47.53 | 50.81 | 47.75 | 44.08 | — | 47.54 | — | — | — | — | — | |
| All-in-OneNumber of Frames=32, Text Augmentation=without2026.05 | 47.5 | 50.8 | 47.7 | 44 | — | — | — | — | — | — | — | |
| SHG-VQABackbone=ResNext101-ImageNet1K, Obj.=false, Hyper.=true2023.04 | 45.8 | 42.77 | 34.64 | 29.91 | 38.28 | — | — | — | — | — | — | |
| RESERVE-B2024.01 | 44.8 | 42.4 | 38.8 | 36.2 | — | 40.5 | — | — | — | — | — | |
| RESERVEVariant=B2026.01 | 44.8 | 42.4 | 38.8 | 36.2 | — | 40.5 | — | — | — | — | — | |
| InternVideoVision Encoder=ViT-L, LLM Size=1.3B, Inference Vision=multiple, Inference LLM=single, Video Trained=true2024.03 | 43.8 | 43.2 | 42.3 | 37.4 | — | 41.6 | — | — | — | — | — | |
| All-in-OneNumber of Frames=32, Text Augmentation=with2026.05 | 42.9 | 48.5 | 44 | 40.2 | — | — | — | — | — | — | — | |
| NS-SRSetup=GT V2024.11 | 42.6 | 46.3 | 43.4 | 43.9 | — | 44.5 | — | — | — | — | — | |
| ClipBERTBackbone=ResNext101-K400, Obj.=false, Hyper.=false2023.04 | 39.81 | 43.59 | 32.24 | 31.42 | 36.7 | — | — | — | — | — | — | |
| ClipBERT2024.01 | 39.81 | 43.59 | 32.34 | 31.42 | — | 36.7 | — | — | — | — | — | |
| CogAgentVision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 39.8 | 47.4 | 40.5 | 43.6 | — | 44.4 | — | — | — | — | — | |
| CLIP-BERTSetup=-2024.11 | 39.8 | 43.6 | 32.2 | 31.4 | — | 36.7 | — | — | — | — | — | |
| CLIP2026.01 | 39.8 | 40.5 | 35.5 | 36 | — | 38 | — | — | — | — | — | |
| HRCNBackbone=ResNext101-K400, Obj.=true, Hyper.=false2023.04 | 39.1 | 38.17 | 28.75 | 27.27 | 33.32 | — | — | — | — | — | — | |
| LCGNBackbone=ResNext101-K400, Obj.=false, Hyper.=false2023.04 | 39.01 | 37.97 | 28.81 | 26.98 | 33.19 | — | — | — | — | — | — | |
| CLIP-BERTSetup=GT V2024.11 | 36.3 | 38.9 | 30.7 | 29.8 | — | 36.5 | — | — | — | — | — | |
| Vis-BERTSetup=GT V2024.11 | 34.7 | 35.9 | 31.2 | 31.4 | — | 34.7 | — | — | — | — | — | |
| Vis-BERTSetup=-2024.11 | 33.6 | 37.2 | 31 | 30.8 | — | 34.8 | — | — | — | — | — | |
| CNN-BERTBackbone=ResNext101-K400, Obj.=false, Hyper.=false2023.04 | 33.59 | 37.16 | 30.95 | 30.84 | 33.14 | — | — | — | — | — | — | |
| CNN-LSTMBackbone=ResNext101-K400, Obj.=false, Hyper.=false2023.04 | 33.25 | 32.67 | 30.69 | 30.43 | 31.76 | — | — | — | — | — | — | |
| Blind Model (BERT)Backbone=BERT, Obj.=false, Hyper.=false2023.04 | 32.68 | 34.21 | 29.98 | 29.26 | 31.53 | — | — | — | — | — | — | |
| Blind Model (LSTM)Backbone=GloVe, Obj.=false, Hyper.=false2023.04 | 32.24 | 32.17 | 28.56 | 28.41 | 30.34 | — | — | — | — | — | — | |
| NS-SRSetup=PR V2024.11 | 30.9 | 31.8 | 30.2 | 29.7 | — | 30.7 | — | — | — | — | — | |
| NS-SRBackbone=ResNext101-K400, Obj.=true, Hyper.=true2023.04 | 30.88 | 31.76 | 30.23 | 29.73 | 30.65 | — | — | — | — | — | — | |
| Q-type (Random)Obj.=false, Hyper.=false2023.04 | 25.06 | 24.93 | 24.79 | 24.81 | 24.89 | — | — | — | — | — | — | |
| Q-type (Frequent)Obj.=false, Hyper.=false2023.04 | 19.09 | 19.45 | 12.9 | 18.31 | 17.44 | — | — | — | — | — | — | |
| BOLDModel=Video-LLaMA, k=0.52024.10 | — | — | — | — | — | — | 37.19 | 33.02 | 22.29 | 14.55 | 14.58 | |
| BOLDModel=Video-LLaVA, k=0.52024.10 | — | — | — | — | — | — | 37.21 | 35.76 | 18.28 | 4.97 | 4.54 | |
| BOLDModel=SeViLA, k=0.52024.10 | — | — | — | — | — | — | 46.22 | 46.1 | 4.13 | 2.26 | 2.1 | |
| CoVGT2026.01 | — | — | — | — | — | 46.23 | — | — | — | — | — | |
| FlamingoParameters=9B2026.01 | — | — | — | — | — | 43.4 | — | — | — | — | — | |
| Flamingo-9B2024.01 | — | — | — | — | — | 43.4 | — | — | — | — | — | |
| HawkEyeFine-tuned=false2024.03 | — | — | — | — | — | — | 56.59 | — | — | — | — | |
| TimeChatFine-tuned=false2024.03 | — | — | — | — | — | — | 37.97 | — | — | — | — | |
| VideoChat2Fine-tuned=false2024.03 | — | — | — | — | — | — | 59 | — | — | — | — | |
| VideoChat2 (our impl.)Fine-tuned=false, Implementation=Reproduced2024.03 | — | — | — | — | — | — | 55.24 | — | — | — | — | |
| Weighted_BOLDModel=Video-LLaMA, k=0.52024.10 | — | — | — | — | — | — | 37.34 | 33.5 | 21.13 | 14.05 | 14.14 | |
| Weighted_BOLDModel=Video-LLaVA, k=0.52024.10 | — | — | — | — | — | — | 37.53 | 36.18 | 17.45 | 4.91 | 4.26 | |
| Weighted_BOLDModel=SeViLA, k=0.52024.10 | — | — | — | — | — | — | 46.2 | 46.08 | 4.01 | 2.2 | 2.07 |