Video Perception on Perception (test)
70.5AccuracyQwen2.5-VL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-VLModel Size=7-8B, Zero-shot evaluation=true2025.12 | 70.5 | — | |
| Qwen2.5-OmniModel Size=7-8B, Zero-shot evaluation=true2025.12 | 70.4 | — | |
| Penguin-VLParameters=2B2026.03 | 70.4 | — | |
| JavisGPTModel Size=7-8B, #Samples=1.5M, Zero-shot evaluation=true2025.12 | 70.2 | — | |
| MACDBackbone=Qwen2.5-VL-7B2026.02 | 70 | — | |
| InternVL2.5Model Size=7-8B, Zero-shot evaluation=true2025.12 | 68.9 | — | |
| Video-LLaVAModel Size=7-8B, Zero-shot evaluation=true2025.12 | 67.9 | — | |
| MACDBackbone=Qwen2-VL-7B2026.02 | 67 | — | |
| APPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 66.9 | — | |
| DAPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 66.4 | — | |
| GRPOBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 66 | — | |
| APPOSize=7B, Training Data=34K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 64.7 | — | |
| InternVL3.5Parameters=2B2026.03 | 64.7 | — | |
| Qwen3-VLParameters=2B2026.03 | 64.5 | — | |
| BaselineBackbone=Qwen2-VL-7B2026.02 | 64 | — | |
| APPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 63.1 | — | |
| BaselineBackbone=Qwen2.5-VL-7B2026.02 | 63 | — | |
| DAPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 62.5 | — | |
| Qwen2-VLModel Size=7-8B, Zero-shot evaluation=true2025.12 | 62.3 | — | |
| GRPOBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 62.1 | — | |
| MACDBackbone=Qwen3-VL-2B2026.02 | 62 | — | |
| MACDBackbone=Qwen2.5-VL-3B2026.02 | 61 | — | |
| VideoRFTSize=7B, Training Data=310K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 60.4 | — | |
| SFTBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 60.2 | — | |
| SFTBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 58.2 | — | |
| VideoChat-R1Size=7B, Training Data=18K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 57.7 | — | |
| LLaVA-OVModel Size=7-8B, Zero-shot evaluation=true2025.12 | 57.1 | — | |
| GRPO-CARESize=7B, Training Data=260K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 57.1 | — | |
| Video-R1Size=7B, Training Data=260K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 57 | — | |
| Base ModelBackbone=Qwen2.5-VL-7B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 55.2 | — | |
| BaselineBackbone=Qwen3-VL-2B2026.02 | 55 | — | |
| VideoLLaMA2.1Model Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 54.9 | — | |
| TW-GRPOSize=7B, Training Data=1K, Resolution=224 x 224, Sampling Rate (fps)=1fps, Maximum Frames=30, Zero-shot evaluation=true2026.02 | 54.9 | — | |
| VCDBackbone=Qwen3-VL-2B2026.02 | 54 | — | |
| BaselineBackbone=Qwen2.5-VL-3B2026.02 | 52 | — | |
| SmolVLM2Parameters=2.2B2026.03 | 51.6 | — | |
| VideoLLaMA2Model Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 51.4 | — | |
| SIDBackbone=Qwen3-VL-2B2026.02 | 49 | — | |
| LLaVA-NeXTModel Size=7-8B, Zero-shot evaluation=true2025.12 | 48.8 | — | |
| Gemma3n E2B-itParameters=E2B-it, Inference template=Penguin's template, Frame budget=322026.03 | 48.6 | — | |
| MACDBackbone=InternVL3-8B2026.02 | 46 | — | |
| MACDBackbone=Qwen2-VL-2B2026.02 | 43 | — | |
| Base ModelBackbone=Qwen2.5-VL-3B, Resolution=224x224, Sampling Rate=1 fps, Max Frames=602026.02 | 42.9 | — | |
| BaselineBackbone=Qwen2-VL-2B2026.02 | 40 | — | |
| SIDBackbone=InternVL3-8B2026.02 | 38 | — | |
| BaselineBackbone=InternVL3-8B2026.02 | 37 | — | |
| SIDBackbone=Qwen2.5-VL-3B2026.02 | 35 | — | |
| VCDBackbone=Qwen2.5-VL-3B2026.02 | 35 | — | |
| SIDBackbone=Qwen2.5-VL-7B2026.02 | 35 | — | |
| SIDBackbone=Qwen2-VL-7B2026.02 | 35 | — | |
| VCDBackbone=Qwen2-VL-7B2026.02 | 35 | — | |
| VCDBackbone=InternVL3-8B2026.02 | 35 | — | |
| UnifiedIO-2Model Size=7-8B, #Samples=9.2B, Zero-shot evaluation=true2025.12 | 34.7 | — | |
| VCDBackbone=Qwen2.5-VL-7B2026.02 | 34 | — | |
| VCDBackbone=Qwen2-VL-2B2026.02 | 34 | — | |
| NExT-GPTModel Size=7-8B, #Samples=1.9M, Zero-shot evaluation=true2025.12 | 33.7 | — | |
| SIDBackbone=Qwen2-VL-2B2026.02 | 31 | — | |
| BIMBALLM backbone=Qwen2-7B2026.02 | — | 68.1 | |
| LLaMA-VIDLLM backbone=Vicuna-7B2026.02 | — | 44.6 | |
| LLaVA-NeXT-VideoLLM backbone=Qwen1.5-7B2026.02 | — | 48.8 | |
| LLaVA-OneVisionLLM backbone=Qwen2-7B2026.02 | — | 57.1 | |
| LLaVA-VideoLLM backbone=Qwen2-7B2026.02 | — | 67.9 | |
| ReMoRaLLM backbone=Qwen2-7B2026.02 | — | 67.7 | |
| Video-LaVITLLM backbone=LaVIT-7B2026.02 | — | 47.9 | |
| Video-LLaMA2LLM backbone=Mistral-7B2026.02 | — | 51.4 | |
| Video-LLaVALLM backbone=Vicuna-7B2026.02 | — | 44.3 | |
| VideoChat2LLM backbone=Vicuna-7B2026.02 | — | 47.3 |