Multimodal Reasoning on SEED-Bench Image
78.6ScorePerceptionLM-8B
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PerceptionLM-8BParameter Scale=8B2026.05 | 78.6 | — | — | — | — | — | — | — | — | — | — | |
| PLM-3BModel=PLM-3B2026.05 | 78.3 | — | — | — | — | — | — | — | — | — | — | |
| PerceptionLM-3BParameter Scale=3B2026.05 | 78.3 | — | — | — | — | — | — | — | — | — | — | |
| Molmo2-4BModel=Molmo2-4B2026.05 | 78 | — | — | — | — | — | — | — | — | — | — | |
| Molmo2-4BParameter Scale=4B2026.05 | 78 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8BParameter Scale=8B2026.05 | 77.5 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-4BModel=Qwen3-VL-4B2026.05 | 77.3 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-4BParameter Scale=4B2026.05 | 77.3 | — | — | — | — | — | — | — | — | — | — | |
| Molmo2-8BParameter Scale=8B2026.05 | 77.3 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-8BParameter Scale=8B2026.05 | 77.2 | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek-VL2-16B-A2.4BModel=DeepSeek-VL2-16B-A2.4B2026.05 | 76.8 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-20B-A4BModel=InternVL3.5-20B-A4B2026.05 | 76.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3.5-4BModel=Qwen3.5-4B2026.05 | 76.6 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-4BModel=InternVL3.5-4B2026.05 | 76.3 | — | — | — | — | — | — | — | — | — | — | |
| PerceptionLM-1BParameter Scale=1B2026.05 | 76.3 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-4BParameter Scale=4B2026.05 | 76.3 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3.5-2BModel=Qwen3.5-2B2026.05 | 75.8 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-2BModel=InternVL3.5-2B2026.05 | 75.2 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-2BParameter Scale=2B2026.05 | 75.2 | — | — | — | — | — | — | — | — | — | — | |
| Zamba2-VL-7BParameter Scale=7B2026.05 | 74.9 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-2BModel=Qwen3-VL-2B2026.05 | 74.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-2BParameter Scale=2B2026.05 | 74.8 | — | — | — | — | — | — | — | — | — | — | |
| SphinxPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 74.2 | — | — | — | — | — | — | — | — | — | — | |
| GPT4VPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 74 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL-3B2026.05 | 73.4 | — | — | — | — | — | — | — | — | — | — | |
| Zamba2-VL-2.7BParameter Scale=2.7B2026.05 | 73 | — | — | — | — | — | — | — | — | — | — | |
| Phantom-3.8BModel Size=3.8B2024.09 | 72.8 | — | — | — | — | — | — | — | — | — | — | |
| ZAYA1-VL-8B-A1BModel=ZAYA1-VL-8B-A1B2026.05 | 72.7 | — | — | — | — | — | — | — | — | — | — | |
| GPT4VPrompting Strategy=Zero-Shot CoT2023.11 | 72.5 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-1BParameter Scale=1B2026.05 | 72.5 | — | — | — | — | — | — | — | — | — | — | |
| SphinxPrompting Strategy=Base2023.11 | 71.6 | — | — | — | — | — | — | — | — | — | — | |
| Zamba2-VL-1.2BParameter Scale=1.2B2026.05 | 71.1 | — | — | — | — | — | — | — | — | — | — | |
| TroL-3.8BModel Size=3.8B2024.09 | 70.5 | — | — | — | — | — | — | — | — | — | — | |
| SphinxPrompting Strategy=Zero-Shot CoT2023.11 | 70.3 | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-1.5-13BPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 69.7 | — | — | — | — | — | — | — | — | — | — | |
| CCoTBackbone=LLaVA-1.5-13B, Evaluation Protocol=prompting2023.11 | 69.7 | 59.3 | 76 | 74.4 | 71.8 | 64.3 | 54.5 | 79.2 | 58.8 | 74.2 | — | |
| MM1-MoE-3B×64Model Size=3B x 642024.09 | 69.4 | — | — | — | — | — | — | — | — | — | — | |
| GPT4VPrompting Strategy=Base2023.11 | 69.1 | — | — | — | — | — | — | — | — | — | — | |
| TroL-1.8BModel Size=1.8B2024.09 | 69 | — | — | — | — | — | — | — | — | — | — | |
| VidILBackbone=LLaVA-1.5-13B, Evaluation Protocol=prompting2023.11 | 68.9 | 62.3 | 74.9 | 72.5 | 69.9 | 62.5 | 53.9 | 78 | 49.4 | 71.1 | — | |
| MM1-3BModel Size=3B2024.09 | 68.8 | — | — | — | — | — | — | — | — | — | — | |
| MolmoE-8B-A1BModel=MolmoE-8B-A1B2026.05 | 68.7 | — | — | — | — | — | — | — | — | — | — | |
| Phantom-1.8BModel Size=1.8B2024.09 | 68.6 | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-1.5-13BPrompting Strategy=Base2023.11 | 68.2 | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-1.5-13BPrompting Strategy=Zero-Shot CoT2023.11 | 66.7 | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek-VL-1.3BModel Size=1.3B2024.09 | 66.7 | — | — | — | — | — | — | — | — | — | — | |
| ALLaVA-3B-LongerModel Size=3B, Training Mode=Longer2024.09 | 65.6 | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVision-0.5BModel Size=0.5B2024.09 | 65.5 | — | — | — | — | — | — | — | — | — | — | |
| ALLAVA-3BModel Size=3B2024.09 | 65.2 | — | — | — | — | — | — | — | — | — | — | |
| MiniCPM-2.4BModel Size=2.4B2024.09 | 62.9 | — | — | — | — | — | — | — | — | — | — | |
| Bunny-3BModel Size=3B2024.09 | 62.5 | — | — | — | — | — | — | — | — | — | — | |
| Phantom-0.5BModel Size=0.5B2024.09 | 60.6 | — | — | — | — | — | — | — | — | — | — | |
| QwenVL-ChatPrompting Strategy=Base2023.11 | 58.2 | — | — | — | — | — | — | — | — | — | — | |
| DDCoTBackbone=LLaVA-1.5-13B, Evaluation Protocol=prompting2023.11 | 58 | 47.3 | 63 | 59.8 | 64.1 | 44.6 | 41.4 | 67.1 | 57.7 | 51.6 | — | |
| mPlug-OWL2Prompting Strategy=Base2023.11 | 57.8 | — | — | — | — | — | — | — | — | — | — | |
| InstructBLIP-13BPrompting Strategy=Compositional Chain-of-Thought (CCoT)2023.11 | 56.9 | — | — | — | — | — | — | — | — | — | — | |
| InstructBLIP-13BPrompting Strategy=Base2023.11 | 48.2 | — | — | — | — | — | — | — | — | — | — | |
| BLIP2Prompting Strategy=Base2023.11 | 46.4 | — | — | — | — | — | — | — | — | — | — | |
| InstructBLIP-13BPrompting Strategy=Zero-Shot CoT2023.11 | 37.6 | — | — | — | — | — | — | — | — | — | — | |
| MMCoTEvaluation Protocol=finetuning, Pre-trained on ScienceQA=true2023.11 | 34.4 | 22.1 | 29.5 | 30.2 | 32.8 | 33.6 | 30.3 | 34.1 | 45.9 | 34 | — | |
| Aquila-VLLanguage Backbone=Qwen2.5 1.5B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | — | — | — | — | — | — | — | — | — | — | 73.9 | |
| FineViT-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | — | — | — | — | — | — | — | — | — | — | 77.44 | |
| Intern3.5-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | — | — | — | — | — | — | — | — | — | — | 75.31 | |
| Qwen3-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | — | — | — | — | — | — | — | — | — | — | 74.78 |