Visual Question Answering on ViDoSeek
0.7425Single AccuracyLang2Act
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Lang2ActCategory=Tool-Enhanced VLMs2026.01 | 0.7425 | 0.7535 | 0.7487 | |
| Pixel-ReasonerCategory=Tool-Enhanced VLMs2026.01 | 0.6796 | 0.701 | 0.6889 | |
| EVisRAGCategory=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 0.6651 | 0.7384 | 0.697 | |
| MM-Search-R1Category=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 0.6558 | 0.7223 | 0.6848 | |
| OpenVLThinkerCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 0.6543 | 0.7445 | 0.6935 | |
| VRAG-RLCategory=Tool-Enhanced VLMs2026.01 | 0.6465 | 0.7163 | 0.6769 | |
| DirectCategory=Prompting Methods2026.01 | 0.645 | 0.662 | 0.6524 | |
| ThinkLite-VLCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 0.6388 | 0.7143 | 0.6716 | |
| VisDomCategory=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 0.6295 | 0.7062 | 0.6629 | |
| R1-OnevisionCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 0.6248 | 0.6841 | 0.6506 | |
| VisionMattersCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 0.614 | 0.7223 | 0.6611 | |
| Vision-R1Category=Vision-Language Reasoning Models (VLRMs)2026.01 | 0.6031 | 0.6801 | 0.6366 | |
| TOTCategory=Prompting Methods2026.01 | 0.5736 | 0.6982 | 0.6278 | |
| GOTCategory=Prompting Methods2026.01 | 0.5566 | 0.6237 | 0.5858 |