Fine-grained visual understanding on V* Bench
85.5General ScoreSwimBird
Evaluation Results
| Method | Links | |
|---|---|---|
| SwimBirdBackbone=Qwen3-VL 8B2026.02 | 85.5 | |
| SkiLaModel Category=Latent Visual Reasoning Models2026.02 | 84.3 | |
| Pixel ReasonerModel Category=Multimodal Agentic Models2026.02 | 84.3 | |
| Qwen3-VL-8B-InstructModel Category=Textual Reasoning Models, Reproduced=true2026.02 | 83.8 | |
| MonetModel Category=Latent Visual Reasoning Models2026.02 | 83.3 | |
| DeepEyesModel Category=Multimodal Agentic Models2026.02 | 83.3 | |
| ThymeModel Category=Multimodal Agentic Models2026.02 | 82.2 | |
| DeepEyesV2Model Category=Multimodal Agentic Models2026.02 | 81.8 | |
| LVRModel Category=Latent Visual Reasoning Models2026.02 | 81.7 | |
| InternVL3-8BModel Category=Textual Reasoning Models2026.02 | 81.2 | |
| Qwen2.5-VL-32B-InstructModel Category=Textual Reasoning Models2026.02 | 80.6 | |
| Vision-R1Model Category=Textual Reasoning Models2026.02 | 80.1 | |
| Qwen3-VL-8B-ThinkingModel Category=Textual Reasoning Models2026.02 | 77.5 | |
| LLaVA-OneVisonModel Category=Textual Reasoning Models2026.02 | 75.4 | |
| Qwen2.5-VL-7B-InstructModel Category=Textual Reasoning Models2026.02 | 75.3 | |
| SEALModel Category=Multimodal Agentic Models2026.02 | 74.8 | |
| GPT-4oModel Category=Textual Reasoning Models2026.02 | 66 | |
| GPT-5-miniModel Category=Textual Reasoning Models2026.02 | 63.9 |