Visual Reasoning on HR-Bench 8K
72.6Overall ScoreDeepEyes
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DeepEyesSize=7B, Workflow=think-with-images, Evaluation Judge=Qwen2.5-72B-Instruct2025.12 | 72.6 | 86.8 | 58.5 | |
| SubagentVLSize=7B, Workflow=think-through-self-calling, Evaluation Judge=Qwen2.5-7B-Instruct2025.12 | 72.6 | 87 | 58.3 | |
| DeepEyesSize=7B, Workflow=think-with-images, Reproduced=true, Evaluation Judge=Qwen2.5-7B-Instruct2025.12 | 71 | 85 | 57 | |
| Qwen2.5-VLSize=32B, Workflow=baseline2025.12 | 70.4 | 84.5 | 56.3 | |
| ERGOPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods2025.09 | 69.9 | — | — | |
| CoVTReasoning Type=Visual Latent Reasoning2026.05 | 69.7 | 85.4 | 54 | |
| ZoomEyeSize=7B, Workflow=manually-defined-workflow2025.12 | 69.3 | 85.5 | 50 | |
| UniVLRReasoning Type=Visual Latent Reasoning2026.05 | 68.8 | 78.8 | 58.8 | |
| RISBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 68.52 | 79.05 | 57.98 | |
| CoVTBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 68.4 | 79.3 | 57.5 | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=16384x28x282025.09 | 67.1 | — | — | |
| MonetBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 67.01 | 78.59 | 55.43 | |
| PixelReasonerReasoning Type=Tool-based Visual Reasoning2026.05 | 66.9 | 80 | 54.3 | |
| LVRReasoning Type=Visual Latent Reasoning2026.05 | 66.9 | 77.7 | 56.2 | |
| ERGOPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods2025.09 | 66.1 | — | — | |
| Qwen2.5-VL-7BReasoning Type=Textual Reasoning2026.05 | 66 | 80.3 | 51.8 | |
| MiniO3Pixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 65.9 | — | — | |
| VisionThinkPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 65.8 | — | — | |
| Qwen2.5-VLSize=7B, Workflow=baseline2025.12 | 65.3 | 78.8 | 51.8 | |
| DeepEyesReasoning Type=Tool-based Visual Reasoning2026.05 | 65.1 | 77 | 53.3 | |
| Qwen2.5-VL-7B+GLSDBackbone=Qwen2.5-VL-7B, Training=Grounded Latent Supervision Dataset (GLSD)2026.05 | 64.7 | 74.69 | 54.7 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 64.33 | 74.42 | 54.24 | |
| Qwen2.5-VL-7B + vanilla SFTReasoning Type=Textual Reasoning, SFT status=vanilla SFT2026.05 | 63.6 | 73.3 | 54 | |
| LVRBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 63.5 | 75 | 52 | |
| MonetReasoning Type=Visual Latent Reasoning2026.05 | 63.5 | 76.5 | 50.5 | |
| RIS+VLPOBackbone=Qwen2.5-VL-7B, Latent Tokens=5, Optimization=Visual-latent Policy Optimization (VLPO)2026.05 | 63.05 | 72.67 | 53.42 | |
| SkiLaReasoning Type=Visual Latent Reasoning2026.05 | 62.9 | 77.5 | 48.3 | |
| PixelReasonerPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 61.3 | — | — | |
| MGPOPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=reproduction with their code using our data2025.09 | 61.1 | — | — | |
| TreeVGRPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 60.4 | — | — | |
| VisionThinkPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 60.1 | — | — | |
| DeepEyesPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 60 | — | — | |
| PixelReasonerPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 59.9 | — | — | |
| DeepEyesPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 58.3 | — | — | |
| MiniO3Pixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 57.3 | — | — | |
| MGPOPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=reproduction with their code using our data2025.09 | 57.3 | — | — | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=1280x28x282025.09 | 56.5 | — | — | |
| GPT-4oSize=-, Workflow=General2025.12 | 55.5 | 62 | 49 | |
| GPT-4oReasoning Type=Textual Reasoning2026.05 | 55.5 | 62 | 49 | |
| TreeVGRPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 54.4 | — | — | |
| GPT-4oModel Type=Proprietary2026.05 | 51.12 | 57.08 | 45.12 | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=640x28x282025.09 | 49.9 | — | — |