Visual Question Answering on SlideVQA
78.74Overall AccuracyLang2Act
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Lang2ActCategory=Tool-Enhanced VLMs2026.01 | 78.74 | 83.62 | 64.55 | — | |
| OpenVLThinkerCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 73.45 | 78.94 | 57.5 | — | |
| VRAG-RLCategory=Tool-Enhanced VLMs2026.01 | 73.41 | 81.74 | 49.21 | — | |
| EVisRAGCategory=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 73.18 | 79.55 | 54.67 | — | |
| Pixel-ReasonerCategory=Tool-Enhanced VLMs2026.01 | 72.84 | 78.13 | 57.5 | — | |
| Lang2ActCategory=Tool-Enhanced VLMs2026.01 | 72.78 | 76.78 | 61.73 | — | |
| VISORBackbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=true2026.04 | 72.37 | 78.82 | 53.62 | — | |
| ThinkLite-VLCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 72.05 | 76.88 | 58.02 | — | |
| Vision-R1Category=Vision-Language Reasoning Models (VLRMs)2026.01 | 71.87 | 78.09 | 53.79 | — | |
| MM-Search-R1Category=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 70.7 | 76.64 | 53.44 | — | |
| VisionMattersCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 69.98 | 74.76 | 56.08 | — | |
| EVisRAGBackbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=true2026.04 | 69.09 | 78.21 | 42.32 | — | |
| Vision-R1Category=Vision-Language Reasoning Models (VLRMs)2026.01 | 68.53 | 72.63 | 56.61 | — | |
| VISORBackbone=Qwen2.5-VL-3B, Multi-agent architecture=false, Fine-tuned=true2026.04 | 68.49 | 74.58 | 50.79 | — | |
| EVisRAGBackbone=Qwen2.5-VL-3B, Multi-agent architecture=false, Fine-tuned=true2026.04 | 68.35 | 75.42 | 47.7 | — | |
| EVisRAGCategory=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 68.31 | 73.85 | 52.2 | — | |
| VisDomCategory=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 67.72 | 73.79 | 50.09 | — | |
| DirectCategory=Prompting Methods2026.01 | 66.55 | 71.36 | 52.56 | — | |
| R1-OnevisionCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 66.55 | 71.78 | 51.32 | — | |
| Pixel-ReasonerCategory=Tool-Enhanced VLMs2026.01 | 66.19 | 70.33 | 54.14 | — | |
| M3RAGBackbone=Qwen2.5-VL-7B, Multi-agent architecture=true, Fine-tuned=false, External Source=Published results2026.04 | 65.82 | — | — | — | |
| OpenVLThinkerCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 64.92 | 69.24 | 52.38 | — | |
| ViDoRAGBackbone=Qwen2.5-VL-7B, Multi-agent architecture=true, Fine-tuned=false2026.04 | 63.88 | 72.15 | 39.86 | — | |
| R1-RouterBackbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=true2026.04 | 63.43 | 69.66 | 45.33 | — | |
| GOTCategory=Prompting Methods2026.01 | 63.12 | 69.11 | 45.68 | — | |
| ThinkLite-VLCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 62.8 | 65.84 | 53.97 | — | |
| VRAG-RLBackbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=true, External Source=Published results2026.04 | 62.59 | 69.3 | 43.1 | — | |
| VRAG-RLCategory=Tool-Enhanced VLMs2026.01 | 62.3 | 64.93 | 54.67 | — | |
| MM-Search-R1Category=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 61.53 | 65.47 | 50.09 | — | |
| TOTCategory=Prompting Methods2026.01 | 59.28 | 66.02 | 39.68 | — | |
| R1-RouterBackbone=Qwen2.5-VL-3B, Multi-agent architecture=false, Fine-tuned=true2026.04 | 59.1 | 64.93 | 42.15 | — | |
| VisionMattersCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 59.01 | 61.47 | 51.85 | — | |
| GOTCategory=Prompting Methods2026.01 | 58.47 | 60.92 | 51.32 | — | |
| VRAG-RLBackbone=Qwen2.5-VL-3B, Multi-agent architecture=false, Fine-tuned=true, External Source=Published results2026.04 | 58.45 | 65.3 | 38.6 | — | |
| R1-OnevisionCategory=Vision-Language Reasoning Models (VLRMs)2026.01 | 56.79 | 59.89 | 47.8 | — | |
| DirectCategory=Prompting Methods2026.01 | 56.75 | 59.16 | 49.74 | — | |
| VisDomCategory=Multimodal Retrieval-Augmented Generation Models (MRAGs)2026.01 | 56.48 | 58.92 | 49.38 | — | |
| TOTCategory=Prompting Methods2026.01 | 50.02 | 55.04 | 35.45 | — | |
| MMSearch-R1Backbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=true2026.04 | 49.03 | 52.06 | 40.21 | — | |
| Search-R1-VLBackbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=true, External Source=Published results2026.04 | 46.76 | 48.3 | 42.3 | — | |
| ViDoRAGBackbone=Qwen2.5-VL-3B, Multi-agent architecture=true, Fine-tuned=false2026.04 | 35.94 | 41.44 | 19.93 | — | |
| ReActBackbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=false, External Source=Published results2026.04 | 31.11 | 34.8 | 20.4 | — | |
| Vanilla RAGBackbone=Qwen2.5-VL-7B, Multi-agent architecture=false, Fine-tuned=false, External Source=Published results2026.04 | 26.1 | 29.1 | 17.4 | — | |
| Search-R1-VLBackbone=Qwen2.5-VL-3B, Multi-agent architecture=false, Fine-tuned=true, External Source=Published results2026.04 | 24.71 | 26.3 | 20.1 | — | |
| Vanilla RAGBackbone=Qwen2.5-VL-3B, Multi-agent architecture=false, Fine-tuned=false, External Source=Published results2026.04 | 17.56 | 19.4 | 12.2 | — | |
| ReActBackbone=Qwen2.5-VL-3B, Multi-agent architecture=false, Fine-tuned=false, External Source=Published results2026.04 | 14.47 | 15.7 | 10.9 | — | |
| LongPOCheckpoint=Short Stage2026.02 | — | — | — | 75.5 | |
| LongPOModel family=Qwen3 VL2026.03 | — | — | — | 75.5 | |
| Mistral 3.1 SmallCheckpoint=24B2026.02 | — | — | — | 67.8 | |
| Mistral 3.1 Small 24BModel family=Mistral2026.03 | — | — | — | 67.8 | |
| Mistral Plain Distillation*2026.02 | — | — | — | 71.2 | |
| No-thinkModel family=Qwen3 VL2026.03 | — | — | — | 74.2 | |
| No-thinkModel family=Mistral2026.03 | — | — | — | 68.4 | |
| Plain DistillationModel family=Qwen3 VL2026.03 | — | — | — | 66.8 | |
| Plain DistillationModel family=Mistral2026.03 | — | — | — | 71.2 | |
| Qwen Thinking TracesModel family=Mistral2026.03 | — | — | — | 69.4 | |
| Qwen3 VLCheckpoint=235B A22B2026.02 | — | — | — | 84.5 | |
| Qwen3 VLCheckpoint=32B2026.02 | — | — | — | 77.2 | |
| Qwen3 VL 235B A22B InstructModel family=Qwen3 VL2026.03 | — | — | — | 84.5 | |
| Qwen3 VL 32B InstructModel family=Qwen3 VL2026.03 | — | — | — | 77.2 | |
| Qwen3 VL Plain DistillationCheckpoint=Short Stage2026.02 | — | — | — | 66.8 | |
| Synthetic ReasoningModel family=Qwen3 VL2026.03 | — | — | — | 75.4 | |
| Synthetic ReasoningModel family=Mistral2026.03 | — | — | — | 69.7 |