Video Question Answering on PerceptionTest
78.6AccuracymPLUG-Owl3
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| mPLUG-Owl3zero-shot=true2025.03 | 78.6 | — | |
| VideoChat2-7Bzero-shot=true2025.03 | 78.6 | — | |
| LLaVAction-7Bzero-shot=true2025.03 | 70.2 | — | |
| CREMAModality=V, F, D, Fine-tuning=true, Trainable Params.=20M2024.02 | 68.7 | — | |
| BLIP-2Modality=V, F, Fine-tuning=true, Trainable Params.=216M2024.02 | 68.2 | — | |
| CREMAModality=V, F, Fine-tuning=true, Trainable Params.=8M2024.02 | 68.2 | — | |
| LLaVA-VideoParams=7B, Base Model=LLaVA-Video Base2025.03 | 67.9 | — | |
| BLIP-2Modality=V, F, D, Fine-tuning=true, Trainable Params.=324M2024.02 | 67.9 | — | |
| LLaVA-Video-7Bzero-shot=true2025.03 | 67.9 | — | |
| VITED (Temporal Evidence Distillation)Params=7B, Base Model=LLaVA-Video Base2025.03 | 67.5 | — | |
| BLIP-2Modality=V, Fine-tuning=true, Trainable Params.=108M2024.02 | 67.1 | — | |
| VITED (Dense Caption Distillation)Params=7B, Base Model=LLaVA-Video Base2025.03 | 67.07 | — | |
| LLaVA-Video (Video Instruction Tuning)Params=7B, Base Model=LLaVA-Video Base2025.03 | 67.03 | — | |
| LLaVA-Video (Chain-of-Thought)Params=7B, Base Model=LLaVA-Video Base2025.03 | 66.91 | — | |
| CREMAModality=V, Fine-tuning=true, Trainable Params.=4M2024.02 | 66.6 | — | |
| Seed 2.0 ProAccess=Proprietary2026.07 | 65.75 | — | |
| Gemini 3.0 ProAccess=Proprietary2026.07 | 63.73 | — | |
| VITED (Temporal Evidence Distillation)Params=7B, Base Model=TimeChat Base2025.03 | 63.66 | — | |
| CuReAccess=Open-source2026.07 | 63.56 | — | |
| Qwen3.6-35B-A3BAccess=Open-source2026.07 | 61.72 | — | |
| Qwen3-VL-235B-A22B-InstructAccess=Open-source2026.07 | 61.26 | — | |
| Qwen3-VL-30B-A3B-InstructAccess=Open-source2026.07 | 60.05 | — | |
| VITED (Dense Caption Distillation)Params=7B, Base Model=TimeChat Base2025.03 | 59.94 | — | |
| TimeChat (Video Instruction Tuning)Params=7B, Base Model=TimeChat Base2025.03 | 57.39 | — | |
| LLaVA-OneVisionParams=7B2025.03 | 57.1 | — | |
| LLaVA-OV-7Bzero-shot=true2025.03 | 57.1 | — | |
| LLaVA-OV-7BFLOPs (TB)=98.53, Total/Trainable Params=8.2B/-2025.03 | 57.1 | — | |
| PAVE-7B (w/ video feature)FLOPs (TB)=98.63, Total/Trainable Params=8.2B/170.5M2025.03 | 56 | — | |
| LLaMA-3.2VParams=11B2025.03 | 52.65 | — | |
| OwlCap-7BAccess=Open-source2026.07 | 52.58 | — | |
| VideoLLaMA2-7Bzero-shot=true2025.03 | 51.4 | — | |
| LLaVA-OV-0.5BFLOPs (TB)=8.01, Total/Trainable Params=0.9B/-2025.03 | 49.2 | — | |
| PAVE-0.5B (w/ video feature)FLOPs (TB)=8.08, Total/Trainable Params=0.9B/41.4M2025.03 | 48.8 | — | |
| SeViLAinference_mode=Zero-shot2024.10 | 45.3 | 45.11 | |
| Video-LLaMAinference_mode=Zero-shot2024.10 | 41.59 | 37.19 | |
| Video-LLaVAinference_mode=Zero-shot2024.10 | 40.73 | 35.69 | |
| Tarsier2-7BAccess=Open-source2026.07 | 37.51 | — | |
| TimeChat (Chain-of-Thought)Params=7B, Base Model=TimeChat Base2025.03 | 20.96 | — | |
| TimeChatParams=7B, Base Model=TimeChat Base2025.03 | 20.34 | — |