Video Reasoning on VSI-Bench (test)
38.6AccuracyVideo-ToC
Evaluation Results
| Method | Links | |
|---|---|---|
| Video-ToCFrames=642026.04 | 38.6 | |
| Video-ToC-SFTFrames=642026.04 | 37.6 | |
| Video-R1Frames=642026.04 | 37.1 | |
| VideoRFTBase Model=Qwen2.5-VL-7B, Training Strategy=RFT2025.11 | 36.8 | |
| VIDEOP2RBase Model=Qwen2.5-VL-7B, Training Strategy=Process-aware Reinforcement Learning2025.11 | 36.8 | |
| Video-ToCFrames=322026.04 | 36.4 | |
| Video-R1Base Model=Qwen2.5-VL-7B, Training Strategy=RFT2025.11 | 35.8 | |
| Video-ToC-SFTFrames=322026.04 | 35.8 | |
| Video-R1Frames=322026.04 | 35.8 | |
| Video-ToCFrames=162026.04 | 35.3 | |
| Video-ToC-SFTFrames=162026.04 | 34.8 | |
| Video-R1-SFTFrames=642026.04 | 34.8 | |
| Video-R1Frames=162026.04 | 34.6 | |
| GPT-4oFrames=-2026.04 | 34 | |
| VideoChat-R1Base Model=Qwen2.5-VL-7B, Training Strategy=RFT2025.11 | 33.9 | |
| VersaVid-R1Base Model=Qwen2.5-VL-7B, Training Strategy=RFT2025.11 | 33.7 | |
| Video-R1-SFTFrames=322026.04 | 33.3 | |
| LLaVA-OneVision-7BModel Scale=7B, Training Strategy=Open-Source2025.11 | 32.4 | |
| LLaVA-OneVision-7BFrames=-2026.04 | 32.4 | |
| Video-R1-SFTFrames=162026.04 | 31.8 | |
| Baseline (Qwen2.5-VL-7B)Frames=642026.04 | 31.4 | |
| Qwen2.5-VL-7BModel Scale=7B, Training Strategy=Open-Source2025.11 | 30.1 | |
| Baseline (Qwen2.5-VL-7B)Frames=322026.04 | 30.1 | |
| LongVA-7BModel Scale=7B, Training Strategy=Open-Source2025.11 | 29.2 | |
| LongVA-7BFrames=-2026.04 | 29.2 | |
| Time-R1Base Model=Qwen2.5-VL-7B, Training Strategy=RFT2025.11 | 29 | |
| VILA-1.5-8BFrames=-2026.04 | 28.9 | |
| VideoChat-R1Frames=162026.04 | 28.9 | |
| Baseline (Qwen2.5-VL-7B)Frames=162026.04 | 27.7 |