Visual Perception on BLINK
75.9AccuracyTHINKLITE-VL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| THINKLITE-VLMethod=THINKLITE-VL2026.05 | 75.9 | — | |
| QWEN2.5-VL-7B + PGTBackbone=QWEN2.5-VL-7B, Data=PGT-augmented2026.05 | 74.8 | — | |
| QWEN2.5-VL-7B + SPECIALIZED MIXBackbone=QWEN2.5-VL-7B, Data=Specialized Mix2026.05 | 74.7 | — | |
| INTERNVL3-8BBackbone=INTERNVL3-8B2026.05 | 74.6 | — | |
| INTERNVL3-8B + PGTBackbone=INTERNVL3-8B, Data=PGT-augmented2026.05 | 74.5 | — | |
| IMAGE JIGSAWMethod=IMAGE JIGSAW2026.05 | 73.7 | — | |
| QWEN2.5-VL-7BBackbone=QWEN2.5-VL-7B2026.05 | 72.8 | — | |
| L2-VMASBackbone=Qwen3-VL-8B, Mode=Instruct2026.01 | 72.7 | — | |
| Qwen3-VL-8B-MastersSize=8B2025.12 | 72.3 | — | |
| Qwen3.5-27BMode=REASONING, Architecture=Dense, # Total Params=27B, # Activated Params=27B2026.04 | 71.6 | — | |
| QWEN2.5-VL-3B + SPECIALIZED MIXBackbone=QWEN2.5-VL-3B, Data=Specialized Mix2026.05 | 71.5 | — | |
| L2-VMASBackbone=Qwen3-VL-8B, Mode=Thinking2026.01 | 71 | — | |
| VMASBackbone=Qwen3-VL-8B, Mode=Instruct2026.01 | 70.3 | — | |
| Qwen3-VL-8BSize=8B2025.12 | 69.1 | — | |
| Qwen3-VL-4B-MastersSize=4B2025.12 | 69.1 | — | |
| Qwen3-VLModel Size=8B2026.03 | 69.1 | — | |
| EXAONE 4.5 33BMode=REASONING, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 68.8 | — | |
| SingleBackbone=Qwen3-VL-8B, Mode=Instruct2026.01 | 68.7 | — | |
| Qwen3-VL-32BMode=Thinking, Architecture=Dense, # Total Params=33B, # Activated Params=33B2026.04 | 68.5 | — | |
| VMASBackbone=Qwen3-VL-8B, Mode=Thinking2026.01 | 68.4 | — | |
| QWEN2.5-VL-3B + PGTBackbone=QWEN2.5-VL-3B, Data=PGT-augmented2026.05 | 68.4 | — | |
| Qwen3-VL-8BTraining Strategy=Base2026.05 | 68.39 | — | |
| Qwen3-VL-8B-InstructStrategy=Instruct2026.02 | 68.2 | — | |
| InternVL3-8B-MastersSize=8B2025.12 | 68 | — | |
| GPT-4o (0513)2025.06 | 68 | — | |
| GPT-4o-20240513Model Category=Closed-Source MLLMs2026.04 | 68 | — | |
| InternVL3.5-8B-MastersSize=8B2025.12 | 67.8 | — | |
| GPT-5 miniMode=REASONING: HIGH, Architecture=-, # Total Params=-, # Activated Params=-2026.04 | 67.7 | — | |
| QWEN2.5-VL-3BBackbone=QWEN2.5-VL-3B2026.05 | 67.6 | — | |
| EVE (Ours-8B-iter4)Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=42026.04 | 67.3 | — | |
| Qwen2.5-VL-7B-MastersSize=7B2025.12 | 67.2 | — | |
| Qwen3-VL-235B-A22BMode=Thinking, Architecture=MoE, # Total Params=236B, # Activated Params=23B2026.04 | 67.1 | — | |
| GLM-4.1V-9BSize=9B2025.12 | 65.9 | — | |
| MM-Zero-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 65.9 | — | |
| Qwen3-VL-4BSize=4B2025.12 | 65.8 | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2.5-78B2025.06 | 65.3 | — | |
| Qwen3-VL-8B-InstructModel Category=Open-Source MLLMs, Scale=8B2026.04 | 65 | — | |
| VIGORL-3BMethod=VIGORL-3B2026.05 | 65 | — | |
| GPT-4o (0806)Model Version=08062024.12 | 64.7 | — | |
| Jigsaw-R1-8BModel Category=Template-based Self-evolution Methods, Scale=8B2026.04 | 64.7 | — | |
| VisPlay-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 64.5 | — | |
| SingleBackbone=Qwen3-VL-8B, Mode=Thinking2026.01 | 64.3 | — | |
| LLAVA-NEXT-LLAMA3-8B + PGTBackbone=LLAVA-NEXT-LLAMA3-8B, Data=PGT-augmented2026.05 | 64.2 | — | |
| ERNIE 5.0-BaseModel type=pre-trained2026.02 | 64.12 | — | |
| Qwen3-VL-8BTraining Strategy=Staged2026.05 | 64.12 | — | |
| InternVL2.5-78B2025.06 | 63.8 | — | |
| LLAVA-NEXT-LLAMA3-8BBackbone=LLAVA-NEXT-LLAMA3-8B2026.05 | 63.1 | — | |
| EVE-iter2Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=22026.04 | 62.96 | — | |
| EVE-iter1Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=12026.04 | 62.65 | — | |
| SAPStrategy=SAP2026.02 | 62.6 | — | |
| InternVL3.5-4B-MastersSize=4B2025.12 | 62.5 | — | |
| MiMo-VL-8BSize=8B2025.12 | 62.4 | — | |
| EVE-iter3Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=32026.04 | 62.39 | — | |
| MiMo-VL-7B-SFT-2508Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=02026.04 | 62.34 | — | |
| Qwen3-VL-2B-MastersSize=2B2025.12 | 62.3 | — | |
| ReGuLaR-7BModel Category=Latent reasoning LVLMs2026.05 | 61.81 | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=Qwen2-VL-72B2025.06 | 61.4 | — | |
| Qwen3-VL-8BTraining Strategy=Merged2026.05 | 61.34 | — | |
| InternVL3.5-2B-MastersSize=2B2025.12 | 61.1 | — | |
| Qwen2-VL-72BSize=72B2024.12 | 60.5 | — | |
| Qwen2-VL-72B2025.06 | 60.5 | — | |
| Qwen3-VL-8B-ThinkingStrategy=LongCoT2026.02 | 60.3 | — | |
| Claude-3.5-Sonnet2025.06 | 60.1 | — | |
| GPT-4oModel Category=General-purpose LVLMs2026.05 | 60.02 | — | |
| VLSI-7BSize=7B2024.12 | 59.7 | — | |
| InternVL3.5-8BSize=8B2025.12 | 59.5 | — | |
| InternVL-3.5Model Size=8B2026.03 | 59.5 | — | |
| Gemini-1.5-Pro2024.12 | 59.1 | — | |
| Gemini-1.5-Pro2025.06 | 59.1 | — | |
| AutoVModel=Qwen2.5-VL 7B2025.06 | 59 | — | |
| Phantom-7BSize=7B2024.12 | 58.9 | — | |
| LLAVA-NEXT-7B + PGTBackbone=LLAVA-NEXT-7B, Data=PGT-augmented2026.05 | 58.8 | — | |
| InternVL3-9BModel Category=Open-Source MLLMs, Scale=9B2026.04 | 58.6 | — | |
| SPATIAL-LADDER-3BMethod=SPATIAL-LADDER-3B2026.05 | 58.6 | — | |
| LaserModel Category=Latent reasoning LVLMs2026.05 | 58.55 | — | |
| GPT-4V (0409)Model Version=04092024.12 | 58.3 | — | |
| Phi-3.5-Vision-4BSize=4B2025.12 | 58.3 | — | |
| GPT-5 nano (high)Model Category=Closed-Source MLLMs, Scale=high2026.04 | 58.3 | — | |
| Penguin-VLModel Size=8B2026.03 | 58.2 | — | |
| InternVL3.5-4BSize=4B2025.12 | 58.1 | — | |
| InternVL3.5-8BTraining Strategy=Staged2026.05 | 57.71 | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2-76B2025.06 | 57.7 | — | |
| Qwen-ViPER-7BModel Category=Pseudo-label-based Self-evolution Methods, Scale=7B2026.04 | 57.6 | — | |
| InternVL2-76BSize=76B2024.12 | 57.5 | — | |
| APIModel=Qwen2.5-VL 7B2025.06 | 57.3 | — | |
| Qwen2.5-VL-7BModel Category=General-purpose LVLMs2026.05 | 57.02 | — | |
| InternVL2-Llama3-76BModel Size=76B, Backbone=Llama32026.01 | 56.8 | — | |
| InternVL2-76B2025.06 | 56.8 | — | |
| GPT-5 mini (minimal)Model Category=Closed-Source MLLMs, Scale=minimal2026.04 | 56.7 | — | |
| Qwen2.5-VL-7BSize=7B2025.12 | 56.4 | — | |
| BaseModel=Qwen2.5-VL 7B2025.06 | 56.4 | — | |
| Q2.5VL-7BTokens Retained Percentage=100%2025.03 | 56.4 | — | |
| Spatial-SSRL-7BModel Category=Template-based Self-evolution Methods, Scale=7B2026.04 | 56.2 | — | |
| InternVL3.5-8BTraining Strategy=Merged2026.05 | 56.13 | — | |
| MonetModel Category=Latent reasoning LVLMs2026.05 | 56.02 | — | |
| Qwen2.5-VL-7BTraining Strategy=Staged2026.05 | 55.71 | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=NVLM-72B2025.06 | 55.6 | — | |
| InternVL3-8BSize=8B2025.12 | 55.5 | — | |
| LLaVA-OneVision-72B2025.06 | 55.4 | — | |
| LLaVA-OneVision-72BModel Category=Open-Source MLLMs, Scale=72B2026.04 | 55.4 | — |