High-Resolution Visual Reasoning on HR-Bench 8K
93.5AccuracyS1-VL-32B-RL
Evaluation Results
| Method | Links | |
|---|---|---|
| S1-VL-32B-RLCategory=Thinking-with-Images Specialist Models, Parameters=32B, Training Stage=RL2026.04 | 93.5 | |
| S1-VL-32B-SFTCategory=Thinking-with-Images Specialist Models, Parameters=32B, Training Stage=SFT2026.04 | 85.1 | |
| Gemini 2.5 ProCategory=Proprietary Models2026.04 | 81.5 | |
| Qwen3-VL-235B-A22B-ThinkingCategory=Open-Source Models, Parameters=235B-A22B2026.04 | 80.4 | |
| Skywork-R1V4-30BCategory=Thinking-with-Images Specialist Models, Parameters=30B2026.04 | 79.8 | |
| GazeVLM (Ours)Gaze Bias=Enabled, Base Model=Qwen3-VL-4B2026.05 | 78 | |
| Qwen3-VL-32B-ThinkingCategory=Open-Source Models, Parameters=32B2026.04 | 77 | |
| GazeVLM (w/o gaze bias)Gaze Bias=None, Base Model=Qwen3-VL-4B2026.05 | 74.4 | |
| Qwen3.5-4B-InstTraining Paradigm=Vanilla Open-source, Parameters=4B2026.05 | 74 | |
| GPT-5Category=Proprietary Models2026.04 | 73.75 | |
| Gemini 2.5 FlashCategory=Proprietary Models2026.04 | 73.7 | |
| Qwen3-VL-4BTraining Paradigm=Vanilla Open-source, Parameters=4B2026.05 | 73.6 | |
| DeepEyesTraining Paradigm=RL-trained2026.05 | 72.6 | |
| Thyme-VL (7B)Category=Thinking-with-Images Specialist Models, Parameters=7B2026.04 | 72 | |
| MGPOBase Model=Qwen2.5-VL-7B2025.07 | 71.2 | |
| Ground-R1-7BTraining Paradigm=RL-trained, Parameters=7B2026.05 | 71.1 | |
| Qwen2.5-VL-32BCategory=Open-Source Models, Parameters=32B2026.04 | 70.4 | |
| InternVL3-8BCategory=Open-Source Models, Parameters=8B2026.04 | 69.3 | |
| SPARC-4BTraining Paradigm=SFT-based, Parameters=4B2026.05 | 68.4 | |
| MGPOBase Model=Qwen2.5-VL-3B2025.07 | 68.1 | |
| Pixel-Reasoner-7BTraining Paradigm=RL-trained, Parameters=7B2026.05 | 66.9 | |
| Intern-S1 (235B+6B)Category=Open-Source Models, Parameters=235B+6B2026.04 | 66.38 | |
| GRPOBase Model=Qwen2.5-VL-7B2025.07 | 66.3 | |
| Qwen2.5-VL-7BCategory=Open-Source Models, Parameters=7B2026.04 | 65.3 | |
| GRPOBase Model=Qwen2.5-VL-3B2025.07 | 64.9 | |
| Qwen2.5-VL-7BTraining Paradigm=Vanilla Open-source, Parameters=7B2026.05 | 59.5 | |
| GPT-4oTraining Paradigm=Closed-source2026.05 | 55.5 | |
| Intern-S1-mini (8B)Category=Open-Source Models, Parameters=8B2026.04 | 54.38 |