Visual Reasoning on V* (direct/relative/overall metrics)
95.7Overall Scoreo3
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| o3Size=-, Workflow=General2025.12 | 95.7 | — | — | — | — | |
| SubagentVLSize=7B, Workflow=think-through-self-calling, Evaluation Judge=Qwen2.5-7B-Instruct2025.12 | 91.6 | 93 | 89.5 | — | — | |
| ZoomEyeSize=7B, Workflow=manually-defined-workflow2025.12 | 90.6 | 93.9 | 85.5 | — | — | |
| DeepEyesSize=7B, Workflow=think-with-images, Evaluation Judge=Qwen2.5-72B-Instruct2025.12 | 90.1 | 91.3 | 88.2 | — | — | |
| DeepEyesSize=7B, Workflow=think-with-images, Reproduced=true, Evaluation Judge=Qwen2.5-7B-Instruct2025.12 | 88 | 91.3 | 82.9 | — | — | |
| Chain-of-Focus2026.07 | 88 | — | — | — | — | |
| Qwen2.5-VLSize=32B, Workflow=baseline2025.12 | 87.9 | 87.8 | 88.1 | — | — | |
| SegAnswerbackbone=Qwen2.5-VL-7B2026.07 | 86.4 | — | — | — | — | |
| Pixel Reasonerre-evaluated=using its official model and evaluation code2026.07 | 85.5 | — | — | — | — | |
| DeepEyesre-evaluated=using its official model and evaluation code2026.07 | 84.3 | — | — | — | — | |
| PEARLFine-tuning setting=Single-type, single tool call, Training data=LVR data2026.04 | 81.5 | — | — | 86.1 | 74.5 | |
| DyFoSize=7B, Workflow=manually-defined-workflow2025.12 | 81.2 | 80 | 82.9 | — | — | |
| DyFo2026.07 | 81.2 | — | — | — | — | |
| LVRFine-tuning setting=Single-type, single tool call, Evaluation steps=4 steps2026.04 | 80.1 | — | — | 85.2 | 73.7 | |
| PixelReasonerFine-tuning setting=Single-type, multiple tool calls2026.04 | 80.1 | — | — | 81.7 | 77.6 | |
| SFTFine-tuning setting=Single-type, single tool call, Training data=LVR data2026.04 | 79.1 | — | — | 82.6 | 73.7 | |
| PEARLFine-tuning setting=Single-type, multiple tool calls, Training data=PixelReasoner data2026.04 | 79.1 | — | — | 81.7 | 75 | |
| Qwen2.5-VL-7B-InstructFine-tuning setting=No fine-tuning2026.04 | 78.5 | — | — | 81.7 | 73.7 | |
| CoVTFine-tuning setting=Single-type, single tool call2026.04 | 78 | — | — | — | — | |
| Qwen2.5-VL-7Bre-evaluated=using its official model and evaluation code2026.07 | 77.5 | — | — | — | — | |
| SEALSize=7B, Workflow=manually-defined-workflow2025.12 | 75.4 | 74.8 | 76.3 | — | — | |
| SEAL2026.07 | 75.4 | — | — | — | — | |
| PEARLFine-tuning setting=Multiple-type, single tool call, Training data=ThinkMorph data2026.04 | 73.8 | — | — | 76.5 | 69.7 | |
| LLaVA-OneVision-9Bre-evaluated=using its official model and evaluation code2026.07 | 71.7 | — | — | — | — | |
| Qwen2.5-VLSize=7B, Workflow=baseline2025.12 | 71.2 | 73.9 | 67.1 | — | — | |
| LantErn-RL-82026.03 | 71 | 76 | 67 | — | — | |
| Qwen2.5-VL-3B2026.03 | 70 | 75 | 63 | — | — | |
| GPT-4oSize=-, Workflow=General2025.12 | 66 | — | — | — | — | |
| NTP-RL2026.03 | 66 | 75 | 57 | — | — | |
| SFTFine-tuning setting=Multiple-type, single tool call, Training data=ThinkMorph data2026.04 | 42.4 | — | — | 58.3 | 18.4 |