Multimodal Mathematical Reasoning on MathVerse mini
65.9AccuracyMetis
Evaluation Results
| Method | Links | |
|---|---|---|
| MetisModel Category=Agentic Multimodal Models2026.04 | 65.9 | |
| Qwen3-VL-8B-InstructModel Category=Open-source Models2026.04 | 61.3 | |
| VL-Rethinker-7BModel Category=Text-only Reasoning Models2026.04 | 54.2 | |
| D2IjusReasoning Strategy=D2I, Variant=jus, Backbone=Qwen2.5-VL-7B2025.07 | 53.8 | |
| DeepEyesV2Model Category=Agentic Multimodal Models2026.04 | 52.7 | |
| ThinkLite-VL-7BModel Category=Text-only Reasoning Models2026.04 | 52.1 | |
| Qwen2.5-VL-7B w/ GRPO†Reasoning Mode=Intuitive, Backbone=Qwen2.5-VL-7B2025.07 | 51.3 | |
| D2IlocReasoning Strategy=D2I, Variant=loc, Backbone=Qwen2.5-VL-7B2025.07 | 51.1 | |
| Qwen2.5-VL-7B w/ GRPOReasoning Mode=Deliberate, Backbone=Qwen2.5-VL-7B2025.07 | 50.6 | |
| D2IparReasoning Strategy=D2I, Variant=par, Backbone=Qwen2.5-VL-7B2025.07 | 50.6 | |
| MM-Eureka-7BModel Category=Text-only Reasoning Models2026.04 | 50.3 | |
| GPT-4o2025.07 | 50.2 | |
| D2DlocReasoning Strategy=D2D, Variant=loc, Backbone=Qwen2.5-VL-7B2025.07 | 49.1 | |
| D2DjusReasoning Strategy=D2D, Variant=jus, Backbone=Qwen2.5-VL-7B2025.07 | 49.1 | |
| Qwen2.5-VL-7B*Backbone=Qwen2.5-VL-7B2025.07 | 48.2 | |
| OpenVLThinker-7B2025.07 | 47.9 | |
| D2DparReasoning Strategy=D2D, Variant=par, Backbone=Qwen2.5-VL-7B2025.07 | 47.6 | |
| DeepEyesModel Category=Agentic Multimodal Models2026.04 | 47.3 | |
| R1-Onevision-7B2025.07 | 46.4 | |
| Qwen-2.5-VL-7B-InstructModel Category=Open-source Models2026.04 | 45.6 | |
| PEPO_GBackbone=Qwen2.5-VL-3B-Instruct2026.03 | 45.42 | |
| PEPO_DBackbone=InternVL3-2B-Instruct2026.03 | 45.24 | |
| PEPO_GBackbone=InternVL3-2B-Instruct2026.03 | 44.89 | |
| DAPOBackbone=Qwen2.5-VL-3B-Instruct2026.03 | 44.44 | |
| PEPO_DBackbone=Qwen2.5-VL-3B-Instruct2026.03 | 44.23 | |
| GRPOBackbone=InternVL3-2B-Instruct2026.03 | 42.09 | |
| DAPOBackbone=InternVL3-2B-Instruct2026.03 | 41.11 | |
| GRPOBackbone=Qwen2.5-VL-3B-Instruct2026.03 | 40.54 | |
| High-Entropy RLBackbone=Qwen2.5-VL-3B-Instruct2026.03 | 39.84 | |
| InternVL3-8BModel Category=Open-source Models2026.04 | 39.8 | |
| InternVL2.5-8B2025.07 | 39.5 | |
| GPT-4V2025.07 | 39.4 | |
| BaseBackbone=Qwen2.5-VL-3B-Instruct, Mode=zero-shot2026.03 | 37.66 | |
| InternVL2-8B2025.07 | 37 | |
| CPOFine-tuning setting=CPO, Backbone=Qwen3-VL-2B2026.07 | 36.12 | |
| GRPOFine-tuning setting=GRPO, Backbone=Qwen3-VL-2B2026.07 | 32.51 | |
| GSPOFine-tuning setting=GSPO, Backbone=Qwen3-VL-2B2026.07 | 32.16 | |
| Qwen2-VL-7B2025.07 | 31.9 | |
| BaseFine-tuning setting=Base, Backbone=Qwen3-VL-2B2026.07 | 31.12 | |
| LoRAFine-tuning setting=LoRA, Backbone=Qwen3-VL-2B2026.07 | 26.32 | |
| BaseBackbone=InternVL3-2B-Instruct, Mode=zero-shot2026.03 | 23.36 | |
| FFTFine-tuning setting=FFT, Backbone=Qwen3-VL-2B2026.07 | 22.84 | |
| High-Entropy RLBackbone=InternVL3-2B-Instruct2026.03 | 20.7 | |
| LLaVA-CoT-11B2025.07 | 20.3 | |
| LLaVA-OneVisionModel Category=Open-source Models2026.04 | 19.3 |