Multi-modal Understanding on MMBench EN (Accuracy)
93.53AccuracyMaLoRA
Evaluation Results
| Method | Links | |
|---|---|---|
| MaLoRAModel=Qwen3-VL-8B, Training setting=Ours, Training examples=3.5k2025.10 | 93.53 | |
| MaLoRAModel=Qwen2.5-VL-7B, Training setting=Ours, Training examples=3.5k2025.10 | 92.84 | |
| LoRAModel=Qwen3-VL-8B, Training setting=LoRA, Training examples=3.5k2025.10 | 92.03 | |
| BaseModel=Qwen3-VL-8B, Training setting=Base, Training examples=3.5k2025.10 | 91.34 | |
| BaseModel=Qwen2.5-VL-7B, Training setting=Base, Training examples=3.5k2025.10 | 90.3 | |
| LoRAModel=Qwen2.5-VL-7B, Training setting=LoRA, Training examples=3.5k2025.10 | 90.18 | |
| QLoRAModel=Qwen3-VL-8B, Training setting=QLoRA, Training examples=3.5k2025.10 | 89.03 | |
| QLoRAModel=Qwen2.5-VL-7B, Training setting=QLoRA, Training examples=3.5k2025.10 | 88.91 | |
| Ours:base Qwen2.5Params (B)=14+72, Time (s)=3.22025.08 | 88.54 | |
| InternVL2.5-78BModel Name=InternVL2.5-78B2024.12 | 88.3 | |
| InternVL2-40BModel Name=InternVL2-40B2024.12 | 86.8 | |
| Qwen2.5-VLParams (B)=722025.08 | 86.61 | |
| InternVL2.5-38BModel Name=InternVL2.5-38B2024.12 | 86.5 | |
| Qwen2-VL-72BModel Name=Qwen2-VL-72B2024.12 | 86.5 | |
| InternVL2-Llama3-76BModel Name=InternVL2-Llama3-76B2024.12 | 86.5 | |
| LLaVA-OVParams (B)=722025.08 | 85.9 | |
| LLaVA-OneVision-72BModel Name=LLaVA-OneVision-72B2024.12 | 85.8 | |
| Qwen2.5-VLParams (B)=322025.08 | 85.55 | |
| InternVL3.5-38BRatio=100%, Backbone=InternVL3.5-38B2026.05 | 85.47 | |
| InternVL2.5-26BModel Name=InternVL2.5-26B2024.12 | 85.4 | |
| F3ARatio=60%, Backbone=InternVL3.5-38B2026.05 | 85.21 | |
| VisionZipRatio=60%, Backbone=InternVL3.5-38B2026.05 | 85.01 | |
| SenseNova-SIBase Architecture=InternVL3-8B, Model Category=Ours2025.11 | 84.9 | |
| FastVRatio=60%, Backbone=InternVL3.5-38B2026.05 | 84.79 | |
| InternVL2.5-8B2025.12 | 84.6 | |
| InternVL2.5-8BModel Name=InternVL2.5-8B2024.12 | 84.6 | |
| Qwen3-VL-8BModel Category=Open-source General Models2025.11 | 84.6 | |
| DivPruneRatio=60%, Backbone=InternVL3.5-38B2026.05 | 84.55 | |
| Qwen2.5-VLParams (B)=72025.08 | 84.45 | |
| CDPrunerRatio=60%, Backbone=InternVL3.5-38B2026.05 | 84.36 | |
| VisionZipRatio=40%, Backbone=InternVL3.5-38B2026.05 | 84.01 | |
| FastVRatio=40%, Backbone=InternVL3.5-38B2026.05 | 83.93 | |
| F3ARatio=40%, Backbone=InternVL3.5-38B2026.05 | 83.93 | |
| DivPruneRatio=40%, Backbone=InternVL3.5-38B2026.05 | 83.76 | |
| Qwen2.5-VL-7B2025.12 | 83.5 | |
| SenseNova-SIBase Architecture=Qwen3-VL-8B, Model Category=Ours2025.11 | 83.5 | |
| CDPrunerRatio=40%, Backbone=InternVL3.5-38B2026.05 | 83.42 | |
| InternVL2-26BModel Name=InternVL2-26B2024.12 | 83.4 | |
| GPT-4o-20240513Model Name=GPT-4o-202405132024.12 | 83.4 | |
| SenseNova-SIBase Architecture=Bagel-7B-MoT, Model Category=Ours2025.11 | 83.4 | |
| InternVL-2Params (B)=262025.08 | 83.4 | |
| VST-7B-SFTModel Category=Open-source SI Models2025.11 | 83.3 | |
| GPT-4oTime (s)=1.22025.08 | 83.1 | |
| Qwen2-VL-7B2025.12 | 83 | |
| Qwen2-VL-7BModel Name=Qwen2-VL-7B2024.12 | 83 | |
| Bagel-7B-MoTModel Category=Open-source General Models2025.11 | 82.8 | |
| MaLoRAModel=LLaVA-1.5-7B, Training setting=Ours, Training examples=3.5k2025.10 | 82.8 | |
| Claude-3.5-SonnetModel Name=Claude-3.5-Sonnet2024.12 | 82.6 | |
| InternVL-Chat-V1.5Model Name=InternVL-Chat-V1.52024.12 | 82.2 | |
| F3ARatio=20%, Backbone=InternVL3.5-38B2026.05 | 81.92 | |
| CoT4DET-7B2025.12 | 81.9 | |
| Qwen2.5-OmniParams (B)=7, Time (s)=6.02025.08 | 81.8 | |
| CDPrunerRatio=20%, Backbone=InternVL3.5-38B2026.05 | 81.8 | |
| VisionZipRatio=20%, Backbone=InternVL3.5-38B2026.05 | 81.73 | |
| InternVL2-8B2025.12 | 81.7 | |
| InternVL2-8BModel Name=InternVL2-8B2024.12 | 81.7 | |
| InternVL3-8BModel Category=Open-source General Models2025.11 | 81.7 | |
| InternVL-2Params (B)=82025.08 | 81.7 | |
| DivPruneRatio=20%, Backbone=InternVL3.5-38B2026.05 | 81.62 | |
| MiniCPM-V2.6Model Name=MiniCPM-V2.62024.12 | 81.5 | |
| FastVRatio=20%, Backbone=InternVL3.5-38B2026.05 | 81.37 | |
| InternVL2.5-4BModel Name=InternVL2.5-4B2024.12 | 81.1 | |
| GPT-4VModel Name=GPT-4V2024.12 | 81 | |
| QLoRAModel=LLaVA-1.5-7B, Training setting=QLoRA, Training examples=3.5k2025.10 | 80.95 | |
| LLaVA-OV-7B2025.12 | 80.8 | |
| LLaVA-OVParams (B)=72025.08 | 80.8 | |
| Cambrian-34BModel Name=Cambrian-34B2024.12 | 80.4 | |
| Cambrian-S-7BModel Category=Open-source SI Models2025.11 | 80.4 | |
| LoRAModel=LLaVA-1.5-7B, Training setting=LoRA, Training examples=3.5k2025.10 | 80.25 | |
| InternVL3-2BModel Category=Open-source General Models2025.11 | 79.7 | |
| SenseNova-SIBase Architecture=InternVL3-2B, Model Category=Ours2025.11 | 78.9 | |
| InternVL2-4BModel Name=InternVL2-4B2024.12 | 78.6 | |
| Qwen-VL-Max2025.08 | 77.6 | |
| Phi-3.5-Vision-4BModel Name=Phi-3.5-Vision-4B2024.12 | 76 | |
| GPT-4V2025.08 | 75 | |
| Qwen2-VL-2BModel Name=Qwen2-VL-2B2024.12 | 74.9 | |
| InternVL2.5-2BModel Name=InternVL2.5-2B2024.12 | 74.7 | |
| DPVR-PCTrainable=202M, evaluation=3-seed average2026.06 | 74.2 | |
| DPVR-KVTrainable=202M2026.06 | 74.1 | |
| Gemini-1.5-ProModel Name=Gemini-1.5-Pro2024.12 | 73.9 | |
| DPVR-LFTrainable=202M, evaluation=3-seed average2026.06 | 73.8 | |
| Vanilla LLaVA-1.5-7BTrainable=02026.06 | 73.7 | |
| LoRATrainable=80M, r=642026.06 | 73.4 | |
| FastVTrainable=0, mode=inference-only2026.06 | 73.3 | |
| InternVL2-2BModel Name=InternVL2-2B2024.12 | 73.2 | |
| LoRATrainable=20M, r=162026.06 | 72.7 | |
| BaseModel=LLaVA-1.5-7B, Training setting=Base, Training examples=3.5k2025.10 | 72.52 | |
| VITA1.5Params (B)=7, Time (s)=3.72025.08 | 71.8 | |
| InternVL2.5-1BModel Name=InternVL2.5-1B2024.12 | 70.7 | |
| VanillaBackbone=LLaVA-v1.5-13B2025.11 | 68.3 | |
| FullFTTrainable=7B2026.06 | 66 | |
| InternVL2-1BModel Name=InternVL2-1B2024.12 | 65.4 | |
| VanillaBackbone=LLaVA-v1.5-7B2025.11 | 64.2 | |
| Claude-3-OpusModel Name=Claude-3-Opus2024.12 | 63.3 | |
| LLaVA-OneVision-0.5BModel Name=LLaVA-OneVision-0.5B2024.12 | 61.6 | |
| KL+PDBackbone=LLaVA-v1.5-7B2025.11 | 32.7 | |
| SINEPROJECT(PO+PD)Backbone=LLaVA-v1.5-7B2025.11 | 26.4 | |
| PO+PDBackbone=LLaVA-v1.5-7B2025.11 | 26 | |
| KL+PDBackbone=LLaVA-v1.5-13B2025.11 | 24.7 | |
| SINEPROJECT(PO+PD)Backbone=LLaVA-v1.5-13B2025.11 | 24.1 |