Multi-modal Video Understanding on LongVideoBench (accuracy)
47.9AccuracyInternVL2.5
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL2.5Arch.=Dense, # activated=1B, # total=1B2026.03 | 47.9 | |
| InternVL3.5 + Stoch-FT-MultiArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Multinomial Sampling2026.03 | 47 | |
| InternVL3.5 + MoE-GRPO (ours)Arch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=MoE-GRPO2026.03 | 46.5 | |
| LLaVA-OVArch.=Dense, # activated=1B, # total=1B2026.03 | 45.8 | |
| InternVL3.5 + Det-FTArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Deterministic Fine-Tuning2026.03 | 45.3 | |
| InternVL3.5 + Stoch-FT-NoiseArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Gaussian Noise2026.03 | 45.3 | |
| InternVL2Arch.=Dense, # activated=1B, # total=1B2026.03 | 43.3 |