Fine-grained visual understanding on HR-Bench 4K
79ScoreSwimBird
Evaluation Results
| Method | Links | |
|---|---|---|
| SwimBirdBackbone=Qwen3-VL 8B2026.02 | 79 | |
| DeepEyesV2Model Category=Multimodal Agentic Models2026.02 | 77.9 | |
| ThymeModel Category=Multimodal Agentic Models2026.02 | 77 | |
| Qwen3-VL-8B-InstructModel Category=Textual Reasoning Models, Reproduced=true2026.02 | 76.5 | |
| DeepEyesModel Category=Multimodal Agentic Models2026.02 | 73.2 | |
| MIRROR(ours)Param Size=7B2026.02 | 72.88 | |
| Pixel ReasonerModel Category=Multimodal Agentic Models2026.02 | 72.6 | |
| Qwen3-VL-8B-ThinkingModel Category=Textual Reasoning Models2026.02 | 72.4 | |
| SkiLaModel Category=Latent Visual Reasoning Models2026.02 | 72 | |
| MonetModel Category=Latent Visual Reasoning Models2026.02 | 71 | |
| InternVL3-8BModel Category=Textual Reasoning Models2026.02 | 70 | |
| InternVL3Param Size=8B2026.02 | 70 | |
| LVRModel Category=Latent Visual Reasoning Models2026.02 | 69.6 | |
| Qwen2.5-VL-32B-InstructModel Category=Textual Reasoning Models2026.02 | 69.3 | |
| MIRROR(w/o tool)Param Size=7B2026.02 | 69.13 | |
| Qwen2.5-VL-7BParam Size=7B2026.02 | 68.87 | |
| GPT-5-miniModel Category=Textual Reasoning Models2026.02 | 66.3 | |
| Qwen2.5-VL-7B-InstructModel Category=Textual Reasoning Models2026.02 | 65.5 | |
| Vision-R1Model Category=Textual Reasoning Models2026.02 | 64.8 | |
| LLaVA-OneVisonModel Category=Textual Reasoning Models2026.02 | 63 | |
| LLaVA-OneVisionParam Size=7B2026.02 | 63 | |
| InternVL3Param Size=2B2026.02 | 61.75 | |
| GPT-4oModel Category=Textual Reasoning Models2026.02 | 59 | |
| Qwen2.5-VL-3BParam Size=3B2026.02 | 50.25 |