Multimodal Understanding on AI2D (Accuracy)
80.8AccuracyQwen2.5-VL-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-7BToken Retention Ratio=100%, Backbone=Qwen2.5-VL-7B2026.02 | 80.8 | |
| FastVToken Retention Ratio=25%, Backbone=Qwen2.5-VL-7B2026.02 | 74.5 | |
| EntropyPruneToken Retention Ratio=25%, Backbone=Qwen2.5-VL-7B2026.02 | 74.1 | |
| CDPrunerToken Retention Ratio=25%, Backbone=Qwen2.5-VL-7B2026.02 | 71.6 | |
| LongVAArch.=Dense, # activated=7B, # total=7B2026.03 | 70.7 | |
| Mini-InternVL1.5Arch.=Dense, # activated=2B, # total=2B2026.03 | 69.8 | |
| EntropyPruneToken Retention Ratio=12.5%, Backbone=Qwen2.5-VL-7B2026.02 | 69.4 | |
| InternVL2.5Arch.=Dense, # activated=1B, # total=1B2026.03 | 69.3 | |
| CDPrunerToken Retention Ratio=12.5%, Backbone=Qwen2.5-VL-7B2026.02 | 68.3 | |
| InternVL3.5 + MoE-GRPO (ours)Arch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=MoE-GRPO2026.03 | 65.8 | |
| FastVToken Retention Ratio=12.5%, Backbone=Qwen2.5-VL-7B2026.02 | 65.5 | |
| InternVL2Arch.=Dense, # activated=1B, # total=1B2026.03 | 64.1 | |
| InternVL3.5 + Det-FTArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Deterministic Fine-Tuning2026.03 | 62.7 | |
| InternVL3.5 + Stoch-FT-NoiseArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Gaussian Noise2026.03 | 62.4 | |
| InternVL3.5 + Stoch-FT-MultiArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Multinomial Sampling2026.03 | 61.8 | |
| LLaVA-OVArch.=Dense, # activated=1B, # total=1B2026.03 | 57.1 |