High-resolution Perception on HR-Bench 8K
82ScoreMetis
Evaluation Results
| Method | Links | |
|---|---|---|
| MetisBackbone=Qwen3-VL-8B-Instruct2026.04 | 82 | |
| Skywork-R1V4-30B-A3BModel Category=Agentic Multimodal Models2026.04 | 79.8 | |
| SenseNova-MARS-8BModel Category=Agentic Multimodal Models2026.04 | 78.4 | |
| Qwen3-VL-8B-InstructModel Category=Open-Source Models2026.04 | 74.6 | |
| DeepLatent-RL-7B*Model Category=Latent Visual reasoning Models, Training Stage=RL, Training Data=visual search data of DeepLatent-180K in SFT Stage 22026.05 | 74.1 | |
| DeepLatent-RL-7BModel Category=Latent Visual reasoning Models, Training Stage=RL, Training Data=full DeepLatent-180K dataset2026.05 | 74 | |
| MiMo-VL-7B +PCBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 73.8 | |
| DeepEyesV2Model Category=Agentic Multimodal Models2026.04 | 73.8 | |
| DeepLatent-SFT-7BModel Category=Latent Visual reasoning Models, Training Stage=SFT2026.05 | 73.6 | |
| Mini o3Model Category=Agentic Multimodal Models2026.04 | 73.3 | |
| DeepEyes-7BModel Category=Tool-based Models2026.05 | 72.6 | |
| ThymeModel Category=Agentic Multimodal Models2026.04 | 72 | |
| Thyme-7BModel Category=Tool-based Models2026.05 | 72 | |
| SegAnswerBase Model=Qwen2.5-VL-7B2026.07 | 71.3 | |
| MiMo-VL-7BBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=No2026.03 | 70.6 | |
| CoVT-7BModel Category=Latent Visual reasoning Models2026.05 | 69.9 | |
| DyFo-7BModel Category=Tool-based Models2026.05 | 69.8 | |
| DeepEyesEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | 69.8 | |
| Qwen2.5-VL-7B +PCBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 69.6 | |
| DeepEyesModel Category=Agentic Multimodal Models2026.04 | 69.5 | |
| InternVL3-8BModel Category=Open-Source Models2026.04 | 69.3 | |
| InternVL3-8BModel Category=Open-source Models2026.05 | 69.3 | |
| Monet-7BModel Category=Latent Visual reasoning Models2026.05 | 68 | |
| Pixel-Reasoner-7BModel Category=Tool-based Models2026.05 | 66.9 | |
| Pixel ReasonerEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | 66.4 | |
| Pixel-ReasonerModel Category=Agentic Multimodal Models2026.04 | 66.1 | |
| Qwen2.5-VL-7BBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=No2026.03 | 65.3 | |
| Qwen2.5-VL-7BModel Category=Open-source Models2026.05 | 65.3 | |
| Qwen2.5-VL-3B +PCBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=Yes2026.03 | 63.7 | |
| Qwen2.5-VL-32B-InstructModel Category=Open-Source Models2026.04 | 63.6 | |
| Qwen2.5-VL-3BBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=No2026.03 | 63.5 | |
| Qwen2.5-VL-7BEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | 63.4 | |
| LVR-7BModel Category=Latent Visual reasoning Models2026.05 | 63 | |
| Qwen2.5-VL-7B-InstructModel Category=Open-Source Models2026.04 | 62.1 | |
| LLaVA-OneVisionModel Category=Open-Source Models2026.04 | 59.8 | |
| LLaVA-OneVision-7BModel Category=Open-source Models2026.05 | 59.8 | |
| LLaVA-OneVision-9BEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | 54.5 |