Visual Perception Reasoning on V* Bench
89.01ScoreZoomEyes
Evaluation Results
| Method | Links | |
|---|---|---|
| ZoomEyesModel=Qwen2.5VL-3B, Training-free=true2026.02 | 89.01 | |
| AdaFocusModel=Qwen2.5VL-3B, Training-free=true2026.02 | 86.38 | |
| Pixel ReasonerModel=Qwen2.5VL-3B, Training-free=false2026.02 | 84.82 | |
| ERGOPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods2025.09 | 83.8 | |
| ZoomEyesModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 83.25 | |
| ERGOPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods2025.09 | 81.7 | |
| MiniO3Pixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 81.2 | |
| DeepEyesPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 78.5 | |
| MGPOPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=reproduction with their code using our data2025.09 | 77.5 | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=16384x28x282025.09 | 77 | |
| TreeVGRPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 76.4 | |
| BaselineModel=Qwen2.5VL-3B, Training-free=true2026.02 | 75.9 | |
| MLLMs-KnowModel=Qwen2.5VL-3B, Training-free=true2026.02 | 75.9 | |
| MiniO3Pixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 74.9 | |
| PixelReasonerPixel Constraint=1280x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 74.5 | |
| VisionThinkPixel Constraint=1280x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 73.8 | |
| MGPOPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=reproduction with their code using our data2025.09 | 67.5 | |
| PixelReasonerPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 67.2 | |
| TreeVGRPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 67 | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=1280x28x282025.09 | 64.9 | |
| DeepEyesPixel Constraint=640x28x28, Post-training Category=Non-efficiency-oriented Post Training Methods2025.09 | 64.9 | |
| AdaFocusModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 62.8 | |
| VisCropModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 62.3 | |
| VisionThinkPixel Constraint=640x28x28, Post-training Category=Efficiency-oriented Post Training Methods, Inference Pipeline=inference with original pipeline2025.09 | 61.8 | |
| DC2Model=LLaVA-v1.5-7B, Training-free=true2026.02 | 57.6 | |
| Qwen2.5-VL-7B-Inst.Pixel Constraint=640x28x282025.09 | 56.5 | |
| MLLMs-KnowModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 56.02 | |
| BaselineModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 48.68 |