Interleaved Image Multimodal Understanding on BLINK
66.3ScoreInternVL3-78B
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL3-78Bnumber of parameters=78B2026.02 | 66.3 | |
| InternVL3-38Bnumber of parameters=38B2026.02 | 64.49 | |
| Qwen2.5-VL-32Bnumber of parameters=32B2026.02 | 59.34 | |
| TPRU-32Bnumber of parameters=32B2026.02 | 58.81 | |
| TPRU-7Bnumber of parameters=7B2026.02 | 55.86 | |
| Qwen2.5-VL-7Bnumber of parameters=7B2026.02 | 54.76 | |
| Qwen2.5-VL-3Bnumber of parameters=3B2026.02 | 48.97 | |
| Full ModelSparsity Ratio=0%, Base Model=LLaVA-OneVision2025.04 | 48.4 | |
| TPRU-3Bnumber of parameters=3B2026.02 | 48.13 | |
| SparseGPTSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 46.2 | |
| TAMPSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 45.9 | |
| ECOFLaPSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 45.2 | |
| WandaSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 44 | |
| OWLSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 43.5 | |
| MRoPE-IBackbone Model=Qwen3-VL-4B-Instruct, Positional Encoding Variant=MRoPE-I2025.10 | 37.88 | |
| CircleRoPEBackbone Model=Qwen3-VL-4B-Instruct, Positional Encoding Variant=CircleRoPE2025.10 | 37.52 | |
| MHRoPEBackbone Model=Qwen3-VL-4B-Instruct, Positional Encoding Variant=MHRoPE2025.10 | 37.22 | |
| MRoPEBackbone Model=Qwen3-VL-4B-Instruct, Positional Encoding Variant=MRoPE2025.10 | 36.8 | |
| Vanilla RoPEBackbone Model=Qwen3-VL-4B-Instruct, Positional Encoding Variant=Vanilla RoPE2025.10 | 36.33 | |
| HoPEBackbone Model=Qwen3-VL-4B-Instruct, Positional Encoding Variant=HoPE2025.10 | 34.46 | |
| VideoRoPEBackbone Model=Qwen3-VL-4B-Instruct, Positional Encoding Variant=VideoRoPE2025.10 | 34.44 | |
| MagnitudeSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 0 |