Video Reasoning on VSI-Bench
61.9AccuracyLATENT-VC-9B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LATENT-VC-9BFrames=642026.07 | 61.9 | — | — | |
| LATENT-VC-9BFrames=322026.07 | 60.4 | — | — | |
| EFlow2026.07 | 59.1 | — | — | |
| Qwen3-VL-8BModel size=8B2026.07 | 57.9 | — | — | |
| LATENT-VC-9BFrames=162026.07 | 55.5 | — | — | |
| Qwen3.5-9BSFT+GRPOFrames=642026.07 | 52.6 | — | — | |
| Gemini-2.5-ProCategory=Proprietary Models, Size=-2026.06 | 51.5 | — | — | |
| Qwen3.5-9BSFT+GRPOFrames=322026.07 | 50 | — | — | |
| OmniAgentCategory=Ours, Size=7B, incorporate audio signals=true2026.06 | 48.4 | — | — | |
| Qwen3.5-9BSFT+GRPOFrames=162026.07 | 47.2 | — | — | |
| Gemini-1.5-ProCategory=Proprietary Models, Size=-2026.06 | 45.4 | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 22026.05 | 43.1 | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 32026.05 | 42.8 | — | — | |
| SFTDistillation Method=SFT, Teacher Model=Qwen3-235B2026.04 | 41.9 | — | — | |
| VITALCategory=Open-Source Agentic Models, Size=7B2026.06 | 41.8 | — | — | |
| VITAL-7BModel size=7B2026.07 | 41.8 | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Iter 12026.05 | 41.4 | — | — | |
| EvoVidModel=Qwen3-VL-4B-Instruct, Training Iteration=Frozen Questioner2026.05 | 40.6 | — | — | |
| Base ModelModel=Qwen3-VL-4B-Instruct, Training Iteration=w/o training2026.05 | 40.1 | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 32026.05 | 39.8 | — | — | |
| SFTDistillation Method=SFT, Teacher Model=Qwen3-32B2026.04 | 39.1 | — | — | |
| FFRBase Model=Video-R1-SFT, Training=RFT + FFR2026.04 | 38.9 | 22.33 | — | |
| FFR (GLM-4.5V)Distillation Method=FFR, Teacher Model=GLM-4.5V2026.04 | 38.9 | — | — | |
| FFRBase Model=VideoRFT-SFT, Training=RFT + FFR2026.04 | 38.6 | 21.77 | — | |
| FFRDistillation Method=FFR, Teacher Model=Qwen3-32B2026.04 | 38.5 | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 22026.05 | 38.3 | — | — | |
| FFRDistillation Method=FFR, Teacher Model=Qwen3-235B2026.04 | 38.1 | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Iter 12026.05 | 38 | — | — | |
| Base ModelModel=Qwen3-VL-8B-Instruct, Training Iteration=w/o training2026.05 | 37.4 | — | — | |
| Video-R1Category=Open-Source Thinking Models, Size=7B2026.06 | 37.1 | — | — | |
| Video-R1-7BModel size=7B2026.07 | 37.1 | — | — | |
| Video-R1-7BFrames=642026.07 | 37.1 | — | — | |
| VideoRFTTraining=RFT2026.04 | 36.8 | — | — | |
| VideoRFTCategory=Open-Source Thinking Models, Size=7B2026.06 | 36.8 | — | — | |
| VideoRFT-7BModel size=7B2026.07 | 36.8 | — | — | |
| Video-R1Training=RFT2026.04 | 35.8 | — | — | |
| Video-R1-7BFrames=322026.07 | 35.8 | — | — | |
| Video-R1-7BFrames/Resolution=32/1282025.10 | 35.6 | — | 39.2 | |
| Qwen2.5-OmniCategory=Open-Source Non-Thinking Models, Size=7B, incorporate audio signals=true2026.06 | 35.5 | — | — | |
| EvoVidModel=Qwen3-VL-8B-Instruct, Training Iteration=Frozen Questioner2026.05 | 35.3 | — | — | |
| VideoCFRFrames=642026.06 | 34.8 | — | — | |
| Video-R1-7BFrames=162026.07 | 34.6 | — | — | |
| LongVTCategory=Open-Source Agentic Models, Size=7B2026.06 | 34.4 | — | — | |
| GPT-4o2026.04 | 34 | — | — | |
| GPT-4o2025.10 | 34 | — | — | |
| GPT-4oCategory=Proprietary Models, Size=-2026.06 | 34 | — | — | |
| Video-R1-7BFrames/Resolution=16/2562025.10 | 33.8 | — | 37 | |
| Qwen2.5-VLCategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 33.5 | — | — | |
| Qwen3.5-9BCoTFrames=322026.07 | 33.4 | — | — | |
| VideoCFRFrames=322026.06 | 33.1 | — | — | |
| DeepVideo-R12026.06 | 33 | — | — | |
| LLaVA-OneV-7B2025.10 | 32.4 | — | — | |
| LLaVA-OneVision-7B2026.06 | 32.4 | — | — | |
| LLaVA-OneVisionCategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 32.4 | — | — | |
| LLaVA-OV-7BModel size=7B2026.07 | 32.4 | — | — | |
| LLaVA-OneVision-7BFrames=-2026.07 | 32.4 | — | — | |
| Video-R1-SFT2026.04 | 31.8 | — | — | |
| Video-R1-SFTDistillation Method=SFT, Teacher Model=Baseline2026.04 | 31.8 | — | — | |
| VideoCFRFrames=162026.06 | 31.8 | — | — | |
| Qwen2.5-VL-7BModel size=7B2026.07 | 31.8 | — | — | |
| VideoRFT-SFT2026.04 | 31.7 | — | — | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 32026.05 | 31.7 | — | — | |
| VILA-1.5-40B2026.06 | 31.2 | — | — | |
| VILA-1.5-40BFrames=-2026.07 | 31.2 | — | — | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 22026.05 | 30.9 | — | — | |
| V-Reason-7B (Lite)Frames/Resolution=32/1282025.10 | 30.5 | — | 23.7 | |
| V-Reason-7BFrames/Resolution=32/1282025.10 | 30.3 | — | 23.4 | |
| Video-R1-7BFrames=162026.06 | 30.3 | — | — | |
| Pixel-Reasoner2026.04 | 30.2 | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 12026.05 | 29.5 | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 32026.05 | 29.5 | — | — | |
| Qwen3.5-9BCoTFrames=162026.07 | 29.4 | — | — | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Frozen Questioner2026.05 | 29.3 | — | — | |
| LongVA-7B2025.10 | 29.2 | — | — | |
| LongVA-7B2026.06 | 29.2 | — | — | |
| LongVACategory=Open-Source Non-Thinking Models, Size=7B2026.06 | 29.2 | — | — | |
| LongVA-7BModel size=7B2026.07 | 29.2 | — | — | |
| LongVA-7BFrames=-2026.07 | 29.2 | — | — | |
| VILA-1.5-8B2025.10 | 28.9 | — | — | |
| VILA-1.5-8B2026.06 | 28.9 | — | — | |
| VideoChat-R1Frames=162026.06 | 28.9 | — | — | |
| VILA-1.5-8BModel size=8B2026.07 | 28.9 | — | — | |
| VILA-1.5-8BFrames=-2026.07 | 28.9 | — | — | |
| AoTD-7BModel size=7B2026.07 | 28.8 | — | — | |
| V-Reason-7BFrames/Resolution=16/2562025.10 | 28.5 | — | 22.6 | |
| EvoVidModel=Qwen2.5-VL-7B-Instruct, Training Iteration=Iter 12026.05 | 28.5 | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Iter 22026.05 | 28.4 | — | — | |
| Qwen2.5-VL-7BFrames/Resolution=32/1282025.10 | 28.1 | — | 22.3 | |
| V-Reason-7B (Lite)Frames/Resolution=16/2562025.10 | 27.9 | — | 21.6 | |
| Qwen3.5-9BCoTFrames=642026.07 | 27.9 | — | — | |
| Qwen2.5-VL2026.04 | 27.7 | — | — | |
| Base ModelModel=Qwen2.5-VL-7B-Instruct, Training Iteration=w/o training2026.05 | 27.7 | — | — | |
| MARC-3BFrames=12026.06 | 27.6 | — | — | |
| Qwen2.5-VL-7BFrames/Resolution=16/2562025.10 | 26.4 | — | 21.4 | |
| V-Reason-3B (Lite)Frames/Resolution=32/1282025.10 | 26.3 | — | 20.4 | |
| Video-Thinker2026.04 | 26.3 | — | — | |
| EvoVidModel=Qwen2.5-VL-3B-Instruct, Training Iteration=Frozen Questioner2026.05 | 26.3 | — | — | |
| Base ModelModel=Qwen2.5-VL-3B-Instruct, Training Iteration=w/o training2026.05 | 25.1 | — | — | |
| V-Reason-3BFrames/Resolution=32/1282025.10 | 24.7 | — | 17.5 | |
| Qwen2.5-VL-3BFrames/Resolution=32/1282025.10 | 24.3 | — | 17 |