High-Resolution Visual Reasoning on HR-Bench 4K
83.4AccuracyGazeVLM (Ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| GazeVLM (Ours)Gaze Bias=Enabled, Base Model=Qwen3-VL-4B2026.05 | 83.4 | |
| Qwen3-VL-4BTraining Paradigm=Vanilla Open-source, Parameters=4B2026.05 | 79.5 | |
| GazeVLM (w/o gaze bias)Gaze Bias=None, Base Model=Qwen3-VL-4B2026.05 | 79 | |
| Qwen3.5-4B-InstTraining Paradigm=Vanilla Open-source, Parameters=4B2026.05 | 78.6 | |
| DeepEyesTraining Paradigm=RL-trained2026.05 | 75.1 | |
| Ground-R1-7BTraining Paradigm=RL-trained, Parameters=7B2026.05 | 75 | |
| MGPOBase Model=Qwen2.5-VL-7B2025.07 | 74.2 | |
| Pixel-Reasoner-7BTraining Paradigm=RL-trained, Parameters=7B2026.05 | 74 | |
| MGPOBase Model=Qwen2.5-VL-3B2025.07 | 70.9 | |
| SPARC-4BTraining Paradigm=SFT-based, Parameters=4B2026.05 | 70.5 | |
| GRPOBase Model=Qwen2.5-VL-7B2025.07 | 69.8 | |
| Qwen2.5-VL-7BTraining Paradigm=Vanilla Open-source, Parameters=7B2026.05 | 68.5 | |
| GRPOBase Model=Qwen2.5-VL-3B2025.07 | 67.9 | |
| Gemini2.5-Flash-LiteTraining Paradigm=Closed-source2026.05 | 67.2 | |
| GPT-4oTraining Paradigm=Closed-source2026.05 | 59 |