High-resolution perception on V*
89.53Overall ScoreTikArt
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TikArtParams=8B2026.02 | 89.53 | 91.3 | 86.84 | — | |
| Gemini-3-FlashModel Category=Proprietary Model2026.04 | 86.4 | — | — | — | |
| TikArt w/o ObservationParams=8B2026.02 | 86.39 | 89.57 | 81.58 | — | |
| DeepEyesModel Category=Thinking-with-Images Agent Model2026.04 | 85.6 | — | — | — | |
| DeepEyes-7BModel Category=Tool-based Models2026.05 | 85.6 | — | — | — | |
| MiMo-VL-7B +PCBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 85.3 | — | — | — | |
| DeepLatent-RL-7B*Model Category=Latent Visual reasoning Models, Training Stage=RL, Training Data=visual search data of DeepLatent-180K in SFT Stage 22026.05 | 85.3 | — | — | — | |
| Qwen3-VL-InstructParams=32B2026.02 | 84.82 | 84.35 | 85.53 | — | |
| DeepLatent-RL-7BModel Category=Latent Visual reasoning Models, Training Stage=RL, Training Data=full DeepLatent-180K dataset2026.05 | 84.8 | — | — | — | |
| Qwen3-VL-235B-A22B-InstructParams=235B (22B active)2026.02 | 84.66 | 88.5 | 78.95 | — | |
| Pixel-ReasonerParams=7B2026.02 | 84.3 | — | — | — | |
| SkiLaModel Category=Visual-Latent Model, Evaluation Pipeline=Original2026.04 | 84.3 | — | — | — | |
| Pixel-Reasoner-7BModel Category=Tool-based Models2026.05 | 84.3 | — | — | — | |
| DyFo-7BModel Category=Tool-based Models2026.05 | 84.3 | — | — | — | |
| DeepLatent-SFT-7BModel Category=Latent Visual reasoning Models, Training Stage=SFT2026.05 | 84.3 | — | — | — | |
| LLaVA-OneVision-1.5Params=8B2026.02 | 84.29 | 83.48 | 85.53 | — | |
| HyLaR-7BModel Category=Visual-Latent Model2026.04 | 83.77 | 82.61 | 85.53 | — | |
| Monet-7BModel Category=Latent Visual reasoning Models2026.05 | 83.3 | — | — | — | |
| Monet-7BParams=7B2026.02 | 83.25 | 83.48 | 82.89 | — | |
| MonetModel Category=Visual-Latent Model, Evaluation Pipeline=Original2026.04 | 83.25 | 83.48 | 82.89 | — | |
| TikArt w/o Segment ActionParams=8B2026.02 | 83.24 | 82.6 | 84.21 | — | |
| DeepeyesParams=7B2026.02 | 83.03 | 84.35 | 81.58 | — | |
| ThymeParams=7B2026.02 | 82.2 | — | — | — | |
| ThymeModel Category=Thinking-with-Images Agent Model2026.04 | 82.2 | 83.5 | 80.3 | — | |
| Thyme-7BModel Category=Tool-based Models2026.05 | 82.2 | — | — | — | |
| DeepEyesV2Model Category=Thinking-with-Images Agent Model2026.04 | 81.8 | — | — | — | |
| LVR-7BModel Category=Latent Visual reasoning Models2026.05 | 81.7 | — | — | — | |
| InternVL3-8BModel Category=Open-source Models2026.05 | 81.2 | — | — | — | |
| HyLaR-SFTModel Category=Visual-Latent Model2026.04 | 80.63 | 81.73 | 78.95 | — | |
| LVRParams=7B2026.02 | 80.6 | 81.7 | 79 | — | |
| MiMo-VL-7BBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=No2026.03 | 80.6 | — | — | — | |
| LVRModel Category=Visual-Latent Model2026.04 | 80.6 | 81.7 | 79 | — | |
| TikArt w/o Zoom ActionParams=8B2026.02 | 80.34 | 76.32 | 86.42 | — | |
| Gemini-2.5-flashParams=-2026.02 | 80.1 | 84.35 | 73.68 | — | |
| MonetModel Category=Visual-Latent Model, Evaluation Pipeline=Reproduced2026.04 | 80.1 | 81.73 | 77.63 | — | |
| ZoomEyeModel Category=Thinking-with-Images Agent Model2026.04 | 79.85 | 80.52 | 78.82 | — | |
| Qwen2.5-VL-7B +PCBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 79.7 | — | — | — | |
| Gemini-2.5-proParams=-2026.02 | 79.06 | 84.35 | 71.05 | — | |
| SkiLaModel Category=Visual-Latent Model, Evaluation Pipeline=Reproduced2026.04 | 78.53 | — | — | — | |
| CoVT-7BModel Category=Latent Visual reasoning Models2026.05 | 78.5 | — | — | — | |
| Qwen2.5-VL-7BModel Category=Open-source Models2026.05 | 76.5 | — | — | — | |
| GPT-5Params=-2026.02 | 76.44 | 74.78 | 78.95 | — | |
| Qwen2.5-VL-7BModel Category=Open-Source Model2026.04 | 76.44 | 77.39 | 75 | — | |
| Qwen2.5-VL-7BBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=No2026.03 | 76.4 | — | — | — | |
| Qwen2.5-VL-3BBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=No2026.03 | 75.4 | — | — | — | |
| LLaVA-OneVision-7BModel Category=Open-source Models2026.05 | 75.4 | — | — | — | |
| TikArt w/o GRPOParams=8B2026.02 | 75.39 | 78.26 | 71.05 | — | |
| Qwen3-VL-InstructParams=8B2026.02 | 73.82 | 73.04 | 73.68 | — | |
| Qwen2.5-VL-3B +PCBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=Yes2026.03 | 73.8 | — | — | — | |
| Qwen3-VL-ThinkingParams=8B2026.02 | 72.77 | 71.3 | 75 | — | |
| LLaVA-OneVision-7BModel Category=Open-Source Model2026.04 | 71.25 | 73.48 | 67.89 | — | |
| GPT-4oParams=-2026.02 | 70.68 | 71.3 | 69.74 | — | |
| InternVL3-8BModel Category=Open-Source Model2026.04 | 70.2 | 67.8 | 73.7 | — | |
| GPT-4oModel Category=Proprietary Model2026.04 | 67.5 | 72.2 | 60.5 | — | |
| GPT-5-nanoParams=-2026.02 | 63.87 | 58.26 | 72.37 | — | |
| DeepEyesEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | — | — | — | 84.3 | |
| LLaVA-OneVision-9BEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | — | — | — | 71.7 | |
| Pixel ReasonerEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | — | — | — | 85.5 | |
| Qwen2.5-VL-7BEvaluation protocol=Re-evaluated using its official model and evaluation code2026.07 | — | — | — | 77.5 | |
| SegAnswerBase Model=Qwen2.5-VL-7B2026.07 | — | — | — | 86.4 |