Multi-modal Image Understanding on MMBench
80.78AccuracyMHRoPE
Evaluation Results
| Method | Links | |
|---|---|---|
| MHRoPEBackbone=Qwen3-VL-8B-Instruct2025.10 | 80.78 | |
| HoPEBackbone=Qwen3-VL-8B-Instruct2025.10 | 79.59 | |
| MRoPE-IBackbone=Qwen3-VL-8B-Instruct2025.10 | 79.5 | |
| Vanilla RoPEBackbone=Qwen3-VL-8B-Instruct2025.10 | 79.29 | |
| CircleRoPEBackbone=Qwen3-VL-8B-Instruct2025.10 | 79 | |
| VideoRoPEBackbone=Qwen3-VL-8B-Instruct2025.10 | 78.74 | |
| MRoPEBackbone=Qwen3-VL-8B-Instruct2025.10 | 78.27 | |
| InternVL3.5 + MoE-GRPO (ours)Arch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=MoE-GRPO2026.03 | 77.5 | |
| InternVL3.5 + Stoch-FT-NoiseArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Gaussian Noise2026.03 | 76.3 | |
| InternVL3.5 + Det-FTArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Deterministic Fine-Tuning2026.03 | 75.8 | |
| InternVL3.5 + Stoch-FT-MultiArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Multinomial Sampling2026.03 | 73.9 | |
| Mini-InternVL1.5Arch.=Dense, # activated=2B, # total=2B2026.03 | 70.9 | |
| InternVL2.5Arch.=Dense, # activated=1B, # total=1B2026.03 | 70.7 | |
| InternVL2Arch.=Dense, # activated=1B, # total=1B2026.03 | 65.4 | |
| LLaVA-OVArch.=Dense, # activated=1B, # total=1B2026.03 | 52.1 |