Video Spatial Understanding on MMSI-Bench 500 QA 1.0
41AccuracyCoCoSI
Evaluation Results
| Method | Links | |
|---|---|---|
| CoCoSIBackbone=Gemini-based, Training-free=true, Agentic Collaboration=true, CogMap=true2026.06 | 41 | |
| CoCoSIBackbone=GPT-based, Training-free=true, Agentic Collaboration=true, CogMap=true2026.06 | 39.1 | |
| Single-Agent BaselineBackbone=Gemini-based, Training-free=true, Agentic Collaboration=false, CogMap=true2026.06 | 36.6 | |
| TraveLERBackbone=Gemini-based, Training-free=true, Agentic Collaboration=true, CogMap=false2026.06 | 36 | |
| VCABackbone=Gemini-based, Training-free=true, Agentic Collaboration=false, CogMap=false2026.06 | 35.8 | |
| VideoAgentBackbone=Gemini-based, Training-free=true, Agentic Collaboration=false, CogMap=false2026.06 | 35.4 | |
| CoCoSIBackbone=Qwen-VL-based, Training-free=true, Agentic Collaboration=true, CogMap=true2026.06 | 35.3 | |
| Single-Agent BaselineBackbone=GPT-based, Training-free=true, Agentic Collaboration=false, CogMap=true2026.06 | 35 | |
| TraveLERBackbone=GPT-based, Training-free=true, Agentic Collaboration=true, CogMap=false2026.06 | 34.7 | |
| VCABackbone=GPT-based, Training-free=true, Agentic Collaboration=false, CogMap=false2026.06 | 34.2 | |
| VideoAgentBackbone=GPT-based, Training-free=true, Agentic Collaboration=false, CogMap=false2026.06 | 33.8 | |
| Single-Agent BaselineBackbone=Qwen-VL-based, Training-free=true, Agentic Collaboration=false, CogMap=true2026.06 | 32.4 | |
| TraveLERBackbone=Qwen-VL-based, Training-free=true, Agentic Collaboration=true, CogMap=false2026.06 | 31.9 | |
| VCABackbone=Qwen-VL-based, Training-free=true, Agentic Collaboration=false, CogMap=false2026.06 | 31.6 | |
| VideoAgentBackbone=Qwen-VL-based, Training-free=true, Agentic Collaboration=false, CogMap=false2026.06 | 31.2 |