Multi-image Reasoning on Mantis
81.71AccuracyQwen3-VL + S2H-DPO
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL + S2H-DPOParameter=2B2026.04 | 81.71 | |
| Qwen3-VLParameter=2B2026.04 | 79.61 | |
| Qwen2.5-VL + S2H-DPOParameter=7B2026.04 | 74.19 | |
| DPS (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DPS2026.01 | 71 | |
| CcDPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 69.1 | |
| VISCModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 69.1 | |
| Qwen2.5-VLParameter=7B2026.04 | 68.66 | |
| Two-stage RL (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=Two-stage RL (DPS and annealing)2026.01 | 68.4 | |
| InternVL2.5Model Category=Open-Source General MLLMs, Parameter Scale=8B2026.01 | 67.7 | |
| DAPO (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DAPO2026.01 | 67.7 | |
| Qwen2.5-VLModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 64.5 | |
| LLaVA-OneVisionModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 64.2 | |
| mPLUG-Owl3Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 63.1 | |
| GPT-4VModel Category=Closed-Source MLLMs2026.01 | 62.7 | |
| LLaVA-NeXT-InterleaveModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 62.7 | |
| GPT-4VParameter=-2026.04 | 62.7 | |
| MIA-DPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 60.4 | |
| Qwen2.5-VL + MIA-DPOParameter=7B2026.04 | 59.45 | |
| Mantis-Idefics2Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 57.1 | |
| VideoRFTModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 56.7 | |
| VILA1.5Model Category=Open-Source General MLLMs, Parameter Scale=8B2026.01 | 51.2 | |
| TW-GRPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 49.8 | |
| Idefics2Parameter=8B2026.04 | 48.9 | |
| LLaVA-v1.5 + S2H-DPOParameter=7B2026.04 | 47.93 | |
| LLaVA 1.6Model Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 45.6 | |
| LLaVA-v1.6Parameter=7B2026.04 | 45.6 | |
| InstructBLIPParameter=13B2026.04 | 45.6 | |
| CogVLMParameter=17B2026.04 | 45.2 | |
| LLaVA-v1.5 + MIA-DPOParameter=7B2026.04 | 44.2 | |
| LLaVA-v1.5Parameter=7B2026.04 | 41.9 | |
| Qwen-VL-ChatParameter=7B2026.04 | 39.2 | |
| Emu2-ChatParameter=37B2026.04 | 37.8 | |
| LLaVA-v1.5 + POVIDParameter=7B2026.04 | 37.8 | |
| VideoLLaVAParameter=7B2026.04 | 35.9 | |
| LLaVA-v1.5 + HA-DPOParameter=7B2026.04 | 34.6 | |
| LLaVA-v1.5 + LLaVA-RLHFParameter=7B2026.04 | 30.4 | |
| FuyuParameter=8B2026.04 | 27.2 | |
| OpenFlamingo-v2Model Category=Open-Source General MLLMs, Parameter Scale=9B2026.01 | 12.4 |