Visual Reasoning on HR-Bench 4K
0.819Overall ScoreP2R-4B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| P2R-4BParameters=4B2026.07 | 0.819 | 0.923 | 0.715 | |
| P2R-8BParameters=8B2026.07 | 0.815 | 0.935 | 0.715 | |
| Qwen3-VL-Instruct-32BParameters=32B2026.07 | 0.786 | 0.86 | 0.713 | |
| SubagentVLSize=7B, Workflow=think-through-self-calling, Evaluation Judge=Qwen2.5-7B-Instruct2025.12 | 0.77 | 0.933 | 0.608 | |
| Thyme-7BParameters=7B2026.07 | 0.77 | 0.91 | 0.63 | |
| DeepEyesSize=7B, Workflow=think-with-images, Reproduced=true, Evaluation Judge=Qwen2.5-7B-Instruct2025.12 | 0.751 | 0.92 | 0.583 | |
| DeepEyesSize=7B, Workflow=think-with-images, Evaluation Judge=Qwen2.5-72B-Instruct2025.12 | 0.751 | 0.913 | 0.59 | |
| DeepEyes-7BParameters=7B2026.07 | 0.751 | 0.913 | 0.59 | |
| P2R-2BParameters=2B2026.07 | 0.751 | 0.898 | 0.601 | |
| Qwen3-VL-Instruct-8BParameters=8B2026.07 | 0.748 | 0.81 | 0.685 | |
| Qwen2.5-VLSize=32B, Workflow=baseline2025.12 | 0.739 | 0.898 | 0.58 | |
| Qwen3-VL-Instruct-4BParameters=4B2026.07 | 0.738 | 0.818 | 0.66 | |
| UniVLRReasoning Type=Visual Latent Reasoning2026.05 | 0.733 | 0.86 | 0.605 | |
| RISBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 0.7323 | 0.8633 | 0.6012 | |
| ERGOPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods2025.09 | 0.73 | — | — | |
| PixelReasonerReasoning Type=Tool-based Visual Reasoning2026.05 | 0.729 | 0.86 | 0.603 | |
| PixelReasoner-7BParameters=7B2026.07 | 0.729 | 0.86 | 0.603 | |
| CoVTBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 0.719 | 0.855 | 0.583 | |
| MonetReasoning Type=Visual Latent Reasoning2026.05 | 0.719 | 0.893 | 0.545 | |
| CoVTReasoning Type=Visual Latent Reasoning2026.05 | 0.719 | 0.842 | 0.595 | |
| RIS+VLPOBackbone=Qwen2.5-VL-7B, Latent Tokens=5, Optimization=Visual-latent Policy Optimization (VLPO)2026.05 | 0.7179 | 0.8283 | 0.6075 | |
| DeepEyesReasoning Type=Tool-based Visual Reasoning2026.05 | 0.713 | 0.838 | 0.588 | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=16384x28x282025.09 | 0.711 | — | — | |
| LVRBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 0.7088 | 0.8325 | 0.575 | |
| Qwen2.5-VL-7B+GLSDBackbone=Qwen2.5-VL-7B, Training=Grounded Latent Supervision Dataset (GLSD)2026.05 | 0.7058 | 0.833 | 0.5786 | |
| Qwen3-VL-Instruct-2BParameters=2B2026.07 | 0.704 | 0.81 | 0.598 | |
| SkiLaReasoning Type=Visual Latent Reasoning2026.05 | 0.703 | 0.847 | 0.557 | |
| MiniO3Pixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 0.7 | — | — | |
| MonetBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 0.6997 | 0.8401 | 0.5593 | |
| LVRReasoning Type=Visual Latent Reasoning2026.05 | 0.699 | 0.847 | 0.557 | |
| MGPOPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=reproduction with their code using our data2025.09 | 0.698 | — | — | |
| ZoomEyeSize=7B, Workflow=manually-defined-workflow2025.12 | 0.696 | 0.843 | 0.55 | |
| ZoomEye-7BParameters=7B2026.07 | 0.696 | 0.843 | 0.55 | |
| Qwen2.5-VL-7B + vanilla SFTReasoning Type=Textual Reasoning, SFT status=vanilla SFT2026.05 | 0.691 | 0.813 | 0.57 | |
| Qwen2.5-VL-7BReasoning Type=Textual Reasoning2026.05 | 0.69 | 0.858 | 0.522 | |
| Qwen2.5-VLSize=7B, Workflow=baseline2025.12 | 0.688 | 0.852 | 0.522 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 0.683 | 0.806 | 0.5603 | |
| ERGOPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods2025.09 | 0.671 | — | — | |
| PixelReasonerPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 0.669 | — | — | |
| VisionThinkPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 0.669 | — | — | |
| PixelReasonerPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 0.665 | — | — | |
| TreeVGRPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 0.664 | — | — | |
| VisionThinkPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 0.661 | — | — | |
| DeepEyesPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 0.66 | — | — | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=1280x28x282025.09 | 0.656 | — | — | |
| DeepEyesPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 0.644 | — | — | |
| MiniO3Pixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 0.628 | — | — | |
| MGPOPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=reproduction with their code using our data2025.09 | 0.628 | — | — | |
| TreeVGRPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 0.624 | — | — | |
| GPT-4oSize=-, Workflow=General2025.12 | 0.59 | 0.7 | 0.48 | |
| GPT-4oReasoning Type=Textual Reasoning2026.05 | 0.59 | 0.7 | 0.48 | |
| GPT-4o2026.07 | 0.59 | 0.7 | 0.48 | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=640x28x282025.09 | 0.57 | — | — | |
| GPT-4oModel Type=Proprietary2026.05 | 0.547 | 0.6493 | 0.4451 |