Multimodal Understanding on MMBench EN v1.1
89.5AccuracyPixelis (Qwen3-VL-8B-Instruct)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Pixelis (Qwen3-VL-8B-Instruct)Size=8B2026.03 | 89.5 | — | |
| Pixelis: SFT + RFTSize=8B2026.03 | 88.6 | — | |
| Pixelis: SFT + TTRLSize=8B2026.03 | 88.3 | — | |
| Seed-1.5-VLSize=20B2026.03 | 88 | — | |
| Pixelis: RFT + TTRLSize=8B2026.03 | 87.9 | — | |
| PRM (process reward; tools; 8B)Size=8B2026.03 | 87.8 | — | |
| Pixelis: SFT onlySize=8B2026.03 | 87.6 | — | |
| InternVL2.5-78BModel Name=InternVL2.5-78B2024.12 | 87.4 | — | |
| Step Self-Consistency (step-level)Size=8B2026.03 | 87.2 | — | |
| Qwen3-VL-30B-A3B-InstructSize=30B2026.03 | 87 | — | |
| Pixel Reasoner (Qwen3-VL-8B-Instruct)Size=8B2026.03 | 86.9 | — | |
| Gemini-2.5-Pro (minimal)Size=-2026.03 | 86.6 | — | |
| RA-TTA (retrieval-augmented)Size=8B2026.03 | 86.4 | — | |
| Pixelis: RFT onlySize=8B2026.03 | 86.2 | — | |
| Realistic TTA of VLMs (StatA)Size=8B2026.03 | 86.1 | — | |
| RV Self-Consistency (answer-only)Size=8B2026.03 | 86 | — | |
| Qwen2-VL-72BModel Name=Qwen2-VL-72B2024.12 | 85.9 | — | |
| Pixelis: TTRL onlySize=8B2026.03 | 85.9 | — | |
| InternVL2.5-38BModel Name=InternVL2.5-38B2024.12 | 85.5 | — | |
| InternVL2-Llama3-76BModel Name=InternVL2-Llama3-76B2024.12 | 85.5 | — | |
| InternVL2-40BModel Name=InternVL2-40B2024.12 | 85.1 | — | |
| LLaVA-OneVision-72BModel Name=LLaVA-OneVision-72B2024.12 | 85 | — | |
| Qwen3-VL-8B-InstructSize=8B2026.03 | 85 | — | |
| InternVL3.5-A3B (with tools)Size=30B2026.03 | 84.8 | — | |
| InternVL2.5-26BModel Name=InternVL2.5-26B2024.12 | 84.2 | — | |
| InternVL2.5-8BModel Name=InternVL2.5-8B2024.12 | 83.2 | — | |
| GPT-4o-20240513Model Name=GPT-4o-202405132024.12 | 83.1 | — | |
| VanillaMode=Upper bound2024.12 | 82.6 | — | |
| VisionZipImage token reduction ratio=66.7%2024.12 | 82 | — | |
| iLLaVAImage token reduction ratio=66.7%2024.12 | 81.7 | — | |
| InternVL2-26BModel Name=InternVL2-26B2024.12 | 81.5 | — | |
| GPT-5 (minimal)Size=-2026.03 | 81.4 | — | |
| SparseVLMImage token reduction ratio=66.7%2024.12 | 81.2 | — | |
| Claude-3.5-SonnetModel Name=Claude-3.5-Sonnet2024.12 | 80.9 | — | |
| FasterVLMImage token reduction ratio=66.7%2024.12 | 80.8 | — | |
| Qwen2-VL-7BModel Name=Qwen2-VL-7B2024.12 | 80.7 | — | |
| VisionZipImage token reduction ratio=77.8%2024.12 | 80.7 | — | |
| PyramidDropImage token reduction ratio=66.7%2024.12 | 80.5 | — | |
| InternVL-Chat-V1.5Model Name=InternVL-Chat-V1.52024.12 | 80.3 | — | |
| GPT-4VModel Name=GPT-4V2024.12 | 80 | — | |
| iLLaVAImage token reduction ratio=77.8%2024.12 | 79.8 | — | |
| iLLaVAImage token reduction ratio=88.9%2024.12 | 79.7 | — | |
| FasterVLMImage token reduction ratio=77.8%2024.12 | 79.6 | — | |
| InternVL2-8BModel Name=InternVL2-8B2024.12 | 79.5 | — | |
| InternVL2.5-4BModel Name=InternVL2.5-4B2024.12 | 79.3 | — | |
| Cambrian-34BModel Name=Cambrian-34B2024.12 | 78.3 | — | |
| SparseVLMImage token reduction ratio=77.8%2024.12 | 78.3 | — | |
| MiniCPM-V2.6Model Name=MiniCPM-V2.62024.12 | 78 | — | |
| PyramidDropImage token reduction ratio=77.8%2024.12 | 77.7 | — | |
| FasterVLMImage token reduction ratio=88.9%2024.12 | 77.1 | — | |
| VisionZipImage token reduction ratio=88.9%2024.12 | 76.1 | — | |
| InternVL2-4BModel Name=InternVL2-4B2024.12 | 75.8 | — | |
| SparseVLMImage token reduction ratio=88.9%2024.12 | 75.8 | — | |
| Gemini-1.5-ProModel Name=Gemini-1.5-Pro2024.12 | 74.6 | — | |
| Qwen2-VL-2BModel Name=Qwen2-VL-2B2024.12 | 72.2 | — | |
| InternVL2.5-2BModel Name=InternVL2.5-2B2024.12 | 72.2 | — | |
| Phi-3.5-Vision-4BModel Name=Phi-3.5-Vision-4B2024.12 | 72.1 | — | |
| PyramidDropImage token reduction ratio=88.9%2024.12 | 71.9 | — | |
| InternVL2-2BModel Name=InternVL2-2B2024.12 | 70.2 | — | |
| InternVL2.5-1BModel Name=InternVL2.5-1B2024.12 | 68.4 | — | |
| InternVL2-1BModel Name=InternVL2-1B2024.12 | 61.6 | — | |
| Claude-3-OpusModel Name=Claude-3-Opus2024.12 | 60.1 | — | |
| LLaVA-OneVision-0.5BModel Name=LLaVA-OneVision-0.5B2024.12 | 59.6 | — | |
| Gemini 2.5 FlashSize=-, mode=instruct2026.04 | — | 86.6 | |
| InternVL3.5Size=8B, mode=instruct2026.04 | — | 79.5 | |
| MiniCPM-o 4.5Size=9B, mode=instruct2026.04 | — | 87.6 | |
| Qwen3-OmniSize=30B-A3B, mode=instruct2026.04 | — | 84.9 | |
| Qwen3-VLSize=8B, mode=instruct2026.04 | — | 84.5 |