Multimodal Reasoning on MMStar (Score %)
65.2ScoreCAVE-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| CAVE-7BParameters=7B2026.05 | 65.2 | |
| InternVL3.5-8BParameters=8B2026.05 | 64.5 | |
| Qwen3-VL-8B-InstructParameters=8B, Instruct=true2026.05 | 64.3 | |
| DeepEyes-7BParameters=7B2026.05 | 63.6 | |
| Pelican-UnifiedModel Category=Our Unified Model2026.05 | 63.3 | |
| Qwen3-VL-4B-InstructModel Category=Vision-Language Models2026.05 | 62.9 | |
| Qwen2.5-VLParameters=7B2026.05 | 61.5 | |
| IADA +LoRATrainable Params=7.7M2026.03 | 51.2 | |
| AttnRes+LoRATrainable Params=7.7M2026.03 | 48.4 | |
| Pre-trainedTrainable Params=02026.03 | 45.4 | |
| LoRA-onlyTrainable Params=7.6M2026.03 | 44.8 | |
| Gemma3-4B-ITModel Category=Vision-Language Models2026.05 | 37.1 | |
| LLaVA-1.5-7BGating layer configuration=None (Baseline)2026.04 | 33.3 | |
| π0.5Model Category=Vision-Language-Action Models2026.05 | 21.7 | |
| MolmoActModel Category=Vision-Language-Action Models2026.05 | 1.2 | |
| LSGGating layer configuration=L2, Base model=LLaVA-1.5-7B2026.04 | 0.9 | |
| LSGGating layer configuration=L19, Base model=LLaVA-1.5-7B2026.04 | 0.8 | |
| LSGGating layer configuration=L10, Base model=LLaVA-1.5-7B2026.04 | 0.7 |