Open-ended Video Reasoning on OpenEQA (test)
58.6Accuracy (< 30s)CLiViS (VideoLLaMA3)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CLiViS (VideoLLaMA3)Paradigm=Video Reasoning, Base VLM=VideoLLaMA32025.06 | 58.6 | 57.2 | 57.3 | |
| VideoLLaMA3Paradigm=End-to-End VLM2025.06 | 57.8 | 57 | 57.1 | |
| Video-R1Paradigm=Video Reasoning2025.06 | 52.6 | 40.6 | 41.9 | |
| CLiViS (Qwen2.5-VL)Paradigm=Video Reasoning, Base VLM=Qwen2.5-VL2025.06 | 52.6 | 46.2 | 46.9 | |
| CLiViS (InternVL3)Paradigm=Video Reasoning, Base VLM=InternVL32025.06 | 51.7 | 55.9 | 55.4 | |
| InternVL3Paradigm=End-to-End VLM2025.06 | 50.9 | 53.8 | 53.6 | |
| Qwen2.5-VLParadigm=End-to-End VLM2025.06 | 49.1 | 39.7 | 40.7 | |
| InternVL2.5Paradigm=End-to-End VLM2025.06 | 46.6 | 35.3 | 36.5 | |
| AlanaVLMParadigm=Video Reasoning2025.06 | 42.2 | 34.9 | 35.7 | |
| Qwen2-VLParadigm=End-to-End VLM2025.06 | 37.1 | 29.9 | 30.7 | |
| Qwen2.5-VL + DeepSeek-V3Paradigm=Socratic-based, Base VLM=Qwen2.5-VL, LLM=DeepSeek-V32025.06 | 32.8 | 19.9 | 23.5 | |
| Qwen2.5-VL + Qwen2.5-MaxParadigm=Socratic-based, Base VLM=Qwen2.5-VL, LLM=Qwen2.5-Max2025.06 | 31.9 | 21.9 | 23 | |
| VideoLLaMA3 + DeepSeek-V3Paradigm=Socratic-based, Base VLM=VideoLLaMA3, LLM=DeepSeek-V32025.06 | 19 | 8.2 | 9.4 | |
| VideoTreeParadigm=Video Reasoning2025.06 | 18.9 | 15.5 | 16.4 | |
| VideoLLaMA3 + Qwen2.5-MaxParadigm=Socratic-based, Base VLM=VideoLLaMA3, LLM=Qwen2.5-Max2025.06 | 16.4 | 8.3 | 9.2 | |
| InternVL3 + DeepSeek-V3Paradigm=Socratic-based, Base VLM=InternVL3, LLM=DeepSeek-V32025.06 | 12.1 | 9.8 | 10 | |
| InternVL3 + Qwen2.5-MaxParadigm=Socratic-based, Base VLM=InternVL3, LLM=Qwen2.5-Max2025.06 | 7.8 | 7.5 | 7.5 | |
| VideoAgentParadigm=Video Reasoning2025.06 | 4.3 | 8.3 | 7.9 |