Multimodal Reasoning on MMStar (Acc., # Tokens)
72.8AccuracyLaRe
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LaReBase VLM=Qwen3-VL-4B-Instruct2025.11 | 72.8 | 82.7 | |
| BaselineRetention level=100%2026.05 | 72.47 | 2,058.5 | |
| FastV-VPRetention level=≤ 10%2026.05 | 72.2 | 1,873.7 | |
| CoF*Base VLM=Qwen3-VL-4B-Instruct2025.11 | 70.6 | 224.4 | |
| FastV-VPRetention level=≤ 5%2026.05 | 70.53 | 1,875.4 | |
| LVR*Base VLM=Qwen3-VL-4B-Instruct2025.11 | 68.5 | 125.3 | |
| MM-CoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 67.8 | 210.9 | |
| Standard CoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 65.3 | 217 | |
| CCoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 64.2 | 230.6 | |
| ICoTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 62.7 | 200.7 | |
| MCOUTBase VLM=Qwen3-VL-4B-Instruct2025.11 | 61.4 | 145.5 | |
| LaReBase VLM=Qwen3-VL-2B-Instruct2025.11 | 61.2 | 70.5 | |
| LOOK-MRetention level=≤ 10%2026.05 | 58.87 | 2,090.6 | |
| MCOUTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 58.2 | 95.4 | |
| CoF*Base VLM=Qwen3-VL-2B-Instruct2025.11 | 58.1 | 208.9 | |
| MM-CoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 57.8 | 120.2 | |
| CCoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 56.7 | 174.4 | |
| Standard CoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 56.6 | 192.1 | |
| LVR*Base VLM=Qwen3-VL-2B-Instruct2025.11 | 56.6 | 108.2 | |
| LOOK-MRetention level=≤ 5%2026.05 | 54.87 | 2,265.4 | |
| ICoTBase VLM=Qwen3-VL-2B-Instruct2025.11 | 53.3 | 164.1 | |
| VisionzipRetention level=≤ 10%2026.05 | 51.53 | 2,496.8 | |
| FastVRetention level=≤ 10%2026.05 | 50.4 | 2,275.3 | |
| FastVRetention level=≤ 5%2026.05 | 50.13 | 2,040.4 | |
| VisionzipRetention level=≤ 5%2026.05 | 44.93 | 2,289.6 |