Multimodal Understanding on MMT-Bench
77.2AccuracyGPT-5-high
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5-highModel Category=Proprietary Models2026.06 | 77.2 | |
| Gemini-2.5-ProModel Category=Proprietary Models2026.06 | 75.4 | |
| TVI-CoTModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=8B2026.06 | 65.4 | |
| InternVL3-8BModel Category=Open-Source MLLMs, Model Scale=8B2026.06 | 65 | |
| VAPO-Thinker-7BModel Category=MLLM-based Chain-of-Thought Methods, Model Scale=7B2026.06 | 64.1 | |
| Qwen3-VL-8B (Baseline)Model Category=Baseline, Model Scale=8B2026.06 | 63.3 | |
| Gemma3-27BModel Size=27B2025.12 | 59.2 | |
| Gemma3-4B + AuditDMModel Size=4B, Auditing Component=AuditDM2025.12 | 58.9 | |
| Gemma3-12BModel Size=12B2025.12 | 58.5 | |
| InternVL3.5 + MoE-GRPO (ours)Arch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=MoE-GRPO2026.03 | 54.8 | |
| Gemma3-4BModel Size=4B2025.12 | 53.2 | |
| InternVL3.5 + Stoch-FT-NoiseArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Gaussian Noise2026.03 | 52 | |
| InternVL3.5 + Det-FTArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Deterministic Fine-Tuning2026.03 | 51.8 | |
| InternVL3.5 + Stoch-FT-MultiArch.=MoE, # activated=1.3B, # total=2.9B, Fine-tuning strategy=Stochastic Fine-Tuning with Multinomial Sampling2026.03 | 51.2 | |
| SCLBackbone=MiniCPM-Llama3-V2.5, Strategy=SCL2024.10 | 50.4 | |
| InternVL2.5Arch.=Dense, # activated=1B, # total=1B2026.03 | 50.3 | |
| SFTBackbone=InternLM-XComposer-2-7B, Strategy=SFT2024.10 | 50.2 | |
| SCLBackbone=InternLM-XComposer-2-7B, Strategy=SCL2024.10 | 50.2 | |
| SFTBackbone=MiniCPM-Llama3-V2.5, Strategy=SFT2024.10 | 49.8 | |
| InternVL2Arch.=Dense, # activated=1B, # total=1B2026.03 | 49.5 | |
| InternLM-XComposer-2-7BBackbone=InternLM-XComposer-2-7B, Strategy=Base2024.10 | 49.2 | |
| MiniCPM-Llama3-V2.5Backbone=MiniCPM-Llama3-V2.5, Strategy=Base2024.10 | 49 | |
| SCLBackbone=LLaVA-V1.5-13B, Strategy=SCL2024.10 | 41.16 | |
| SFTBackbone=LLaVA-V1.5-13B, Strategy=SFT2024.10 | 40.48 | |
| LLaVA-V1.5-13BBackbone=LLaVA-V1.5-13B, Strategy=Base2024.10 | 40.08 | |
| SCLBackbone=LLaVA-V1.5-7B, Strategy=SCL2024.10 | 36.96 | |
| SFTBackbone=LLaVA-V1.5-7B, Strategy=SFT2024.10 | 36.36 | |
| LLaVA-V1.5-7BBackbone=LLaVA-V1.5-7B, Strategy=Base2024.10 | 35.4 | |
| POVIDBackbone=LLaVA-V1.5-7B, Strategy=POVID2024.10 | 35.04 | |
| SIMABackbone=LLaVA-V1.5-7B, Strategy=SIMA2024.10 | 34.84 | |
| CSRBackbone=LLaVA-V1.5-7B, Strategy=CSR2024.10 | 33.76 |