Visual Search on V* Bench (Accuracy)
90.4AccuracyDeepeyes-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Deepeyes-7BModel Category=Open-Source MLLMs with Tools, reproduced=true2025.12 | 90.4 | |
| Multi-t ViGoRL-7BModel Category=Open-source Models2025.05 | 86.4 | |
| Qwen2.5-VL-32BModel Category=Open-Source MLLMs2025.12 | 85.9 | |
| Qwen2.5-VL-72BModel Category=Open-Source MLLMs2025.12 | 84.8 | |
| CodeDance-7BModel Category=Open-Source MLLMs with Tools2025.12 | 84.8 | |
| Pixel Reasoner-7BModel Category=Open-Source MLLMs with Tools2025.12 | 84.3 | |
| Thyme-VL-7BModel Category=Open-Source MLLMs with Tools2025.12 | 82.2 | |
| IVM-Enhanced GPT-4VModel Category=VLM Tool-Using Pipelines2025.05 | 81.2 | |
| Multi-t ViGoRL-3BModel Category=Open-source Models2025.05 | 81.2 | |
| Multi-t ViGoRL-3BModel Category=Open-source Models, configuration=w/ bounding box outputs2025.05 | 81.2 | |
| Sketchpad-GPT-4oModel Category=VLM Tool-Using Pipelines2025.05 | 80.3 | |
| ViGoRL-3B (Ours)Model Category=Open-source Models2025.05 | 79.1 | |
| Qwen2.5-7B-VLModel Category=Open-source Models2025.05 | 78 | |
| Multi-t ViGoRL-3BModel Category=Open-source Models, ablation=w/o diversity reward2025.05 | 78 | |
| InternVL3-78BModel Category=Open-Source MLLMs2025.12 | 76.4 | |
| Qwen2.5-VL-7BModel Category=Open-Source MLLMs2025.12 | 76.4 | |
| SEALModel Category=Open-source Models2025.05 | 74.8 | |
| Qwen2.5-3B-VLModel Category=Open-source Models2025.05 | 74.2 | |
| Llava-OneVision-72BModel Category=Open-Source MLLMs2025.12 | 73.8 | |
| InternVL2.5-8BModel Category=Open-Source MLLMs2025.12 | 73.7 | |
| Qwen2.5-VL-7B +FINER-TuningBackbone=Qwen2.5-VL-7B, Training Strategy=FINER-Tuning2026.03 | 72.8 | |
| Llava-OneVision-7BModel Category=Open-Source MLLMs2025.12 | 72.7 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.03 | 72.7 | |
| InternVL3.5-8B +FINER-TuningBackbone=InternVL3.5-8B, Training Strategy=FINER-Tuning2026.03 | 71.2 | |
| InternVL3-8BModel Category=Open-Source MLLMs2025.12 | 70.2 | |
| InternVL3.5-14B +FINER-TuningBackbone=InternVL3.5-14B, Training Strategy=FINER-Tuning2026.03 | 70.2 | |
| InternVL3.5-8BBackbone=InternVL3.5-8B2026.03 | 69.1 | |
| InternVL3.5-14BBackbone=InternVL3.5-14B2026.03 | 68 | |
| GPT-4oModel Category=Closed-Source MLLMs2025.12 | 67.5 | |
| GPT-4oModel Category=Proprietary Models2025.05 | 66 | |
| LLaVA-1.6-13BModel Category=Open-source Models2025.05 | 61.8 | |
| LLaVA-1.6-7B +FINER-TuningBackbone=LLaVA-1.6-7B, Training Strategy=FINER-Tuning2026.03 | 55 | |
| GPT-4VModel Category=Proprietary Models2025.05 | 55 | |
| OmniLMM-12B +RLAIF-VBackbone=OmniLMM-12B, Training Strategy=RLAIF-V2026.03 | 54.4 | |
| LLaVA-1.6-7BBackbone=LLaVA-1.6-7B2026.03 | 53.9 | |
| OmniLMM-12BBackbone=OmniLMM-12B2026.03 | 52.9 | |
| LLaVA-1.5-7BModel Category=Open-source Models2025.05 | 48.7 | |
| Gemini-ProModel Category=Proprietary Models2025.05 | 48.2 | |
| VisProgModel Category=VLM Tool-Using Pipelines2025.05 | 41.4 | |
| MM-ReactModel Category=VLM Tool-Using Pipelines2025.05 | 41.4 | |
| VisualChatGPTModel Category=VLM Tool-Using Pipelines2025.05 | 37.6 |