Multimodal Reasoning on DynaMath
67.2AccuracySwimBird
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SwimBird2026.02 | 67.2 | — | — | |
| Qwen3-VL-8B-Instructreproduced by authors=true2026.02 | 65.3 | — | — | |
| PyVision-ImageTooling Type=Dynamic Tooling2026.02 | 61.6 | — | — | |
| AIRTool=Yes2026.06 | 57.5 | — | — | |
| DeepEyesV22026.02 | 57.2 | — | — | |
| DeepEyes-v2Tooling Type=Dynamic Tooling2026.02 | 57.2 | — | — | |
| DeepEyesV2Tool=General, Param Size=7B2025.11 | 57.2 | — | — | |
| ReLaX-VL-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 55.9 | — | — | |
| VL-Rethinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 55.2 | — | — | |
| Qwen2.5VL-7B w/ GRPOTool=No2026.06 | 55.1 | — | — | |
| DeepEyes2026.02 | 55 | — | — | |
| DeepEyesTooling Type=Static Toolset2026.02 | 55 | — | — | |
| DeepEyesTool=Crop, Param Size=7B2025.11 | 55 | — | — | |
| DeepEyesTool=Yes2026.06 | 55 | — | — | |
| MM-Eureka-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 54.4 | — | — | |
| Qwen2.5VL-7B*Tool=No2026.06 | 53.5 | — | — | |
| Qwen2.5-VL-7B-Instruct2026.02 | 53.3 | — | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.02 | 53.3 | — | — | |
| Qwen-2.5-VLTool=✗, Param Size=7B2025.11 | 53.3 | — | — | |
| Qwen2.5-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 53.2 | — | — | |
| ReLaX-VL-3BModel Category=Reasoning Multimodal LLM, Parameters=3B2025.12 | 52.2 | — | — | |
| Vision-R1-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 52 | — | — | |
| AIR-SFTTool=Yes2026.06 | 48.3 | — | — | |
| R1-VL-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 45.2 | — | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 42.5 | — | — | |
| Qwen2-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 42.1 | — | — | |
| Gemini-2.0-FlashModel Category=Proprietary2025.06 | 42.1 | — | — | |
| Qwen2.5-VL-3BModel Category=General Multimodal LLM, Parameters=3B, reproduced=true2025.12 | 40 | — | — | |
| Intern2-VL-8BModel Category=General Multimodal LLM, Parameters=8B2025.12 | 39.7 | — | — | |
| Claude-3.7-SonnetModel Category=Proprietary2025.06 | 39.7 | — | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 39.6 | — | — | |
| OpenVLThinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B, reproduced=true2025.12 | 38.6 | — | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 38.3 | — | — | |
| Qwen2.5-VL-72B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 37.9 | — | — | |
| Self-consistencyPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 37.6 | — | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 36.5 | — | — | |
| Qwen2.5-VL-72BModel Category=Open-Source2025.06 | 35.9 | — | — | |
| Qwen2.5-VL-72BReasoning Framework=Standard2025.06 | 35.9 | — | — | |
| Qwen2.5-VL-72BPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 35.9 | — | — | |
| InternVL3-38BModel Category=Open-Source2025.06 | 35.3 | — | — | |
| InternVL3-78BModel Category=Open-Source2025.06 | 35.1 | — | — | |
| VL-Rethinker-72BModel Category=Open-Source2025.06 | 34.9 | — | — | |
| GPT-4o-20241120Model Category=Proprietary2025.06 | 34.5 | — | — | |
| MM-Eureka-32BModel Category=Open-Source2025.06 | 33.5 | — | — | |
| Qwen2.5-VL-32BReasoning Framework=Standard2025.06 | 33.3 | — | — | |
| Qwen2.5-VL-7B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 32.7 | — | — | |
| InternVL3-14BModel Category=Open-Source2025.06 | 31.3 | — | — | |
| QVQ-72B-PreviewModel Category=Open-Source2025.06 | 30.7 | — | — | |
| ReVisual-R1-7BModel Category=Open-Source2025.06 | 30.5 | — | — | |
| OmniCaptionerMethod Category=Caption-then-Reason2025.06 | 30.5 | — | — | |
| Qwen2.5-VL-7B (Bo8)Training Strategy=GRPO, Sampling Strategy=Best-of-82025.06 | 29.7 | — | — | |
| Qwen2.5-VL-3B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 29.3 | — | — | |
| Gemma-3-27BModel Category=Open-Source2025.06 | 28.5 | — | — | |
| Ovis2-34BModel Category=Open-Source2025.06 | 27.5 | — | — | |
| Ovis2-16BModel Category=Open-Source2025.06 | 26.3 | — | — | |
| InternVL3-8BModel Category=Open-Source2025.06 | 25.5 | — | — | |
| ECSOMethod Category=Caption-then-Reason2025.06 | 25 | — | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 23.4 | — | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 23.1 | — | — | |
| VisualPRM-8BPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 23 | — | — | |
| Self-consistencyPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 22.9 | — | — | |
| MM-Eureka-7BModel Category=Open-Source2025.06 | 22.6 | — | — | |
| Qwen2.5-VL-7BPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 21.8 | — | — | |
| VL-Rethinker-7BModel Category=Open-Source2025.06 | 21.4 | — | — | |
| Gemma-3-12BModel Category=Open-Source2025.06 | 20.8 | — | — | |
| Ovis2-8BModel Category=Open-Source2025.06 | 20.4 | — | — | |
| Qwen2.5-VL-7BReasoning Framework=Standard2025.06 | 19.4 | — | — | |
| Athena-PRMPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 18.7 | — | — | |
| VisualPRM-8BPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 18 | — | — | |
| Athena-ORMPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 15.2 | — | — | |
| Self-consistencyPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 13.8 | — | — | |
| Qwen2.5-VL-3BReasoning Framework=Standard2025.06 | 13.4 | — | — | |
| InternVL2.5-8BPolicy Model=InternVL2.5-8B, Best-of-N Evaluation=82025.06 | 9.4 | — | — | |
| GRPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=GRPO2026.02 | 0.262 | — | — | |
| GSPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=GSPO2026.02 | 0.259 | — | — | |
| LUSPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=LUSPO2026.02 | 0.246 | — | — | |
| Qwen2.5-VL-7B-Instruct (w/o RLVR)Base model=Qwen2.5-VL-7B-Instruct, Algorithm=w/o RLVR2026.02 | 0.202 | — | — | |
| Gemini-2.0-Flash# Samples=unk2025.08 | — | 55.79 | 43.25 | |
| GPT-4o-mini# Samples=unk2025.08 | — | 40.35 | 37.46 | |
| InternVL2.5-8B# Samples=unk2025.08 | — | 48.88 | 51.24 | |
| MiMo-VL-7B# Samples=unk2025.08 | — | 60.5 | 66.81 | |
| Qwen2.5-VL-72B# Samples=unk2025.08 | — | 51.75 | 53.25 | |
| Qwen2.5-VL-7B# Samples=unk2025.08 | — | 52.81 | 52.89 | |
| Qwen3-VL-4B-Thinking# Samples=unk2025.08 | — | 55.79 | 62.82 | |
| VisualPRM-8B# Samples=400K2025.08 | — | 30 | 62.08 | |
| VRPRM# Samples=53.6K, Explicit Reasoning (CoT)=true, RL Training=true2025.08 | — | 53.51 | 67.95 | |
| VRPRM w/o CoT# Samples=53.6K, Explicit Reasoning (CoT)=false, RL Training=true2025.08 | — | 41.05 | 53.06 | |
| VRPRM w/o RL# Samples=3.6K, Explicit Reasoning (CoT)=true, RL Training=false2025.08 | — | 52.46 | 63.08 | |
| VRPRM w/o RL & w/o CoT# Samples=3.6K, Explicit Reasoning (CoT)=false, RL Training=false2025.08 | — | 50.18 | 55.26 | |
| VRPRM-MiMo# Samples=53.6K, Explicit Reasoning (CoT)=true, RL Training=true2025.08 | — | 59.47 | 74.55 | |
| VRPRM-MiMo w/o CoT# Samples=53.6K, Explicit Reasoning (CoT)=false, RL Training=true2025.08 | — | 59.12 | 73.2 | |
| VRPRM-MiMo w/o RL# Samples=3.6K, Explicit Reasoning (CoT)=true, RL Training=false2025.08 | — | 60.04 | 61.67 | |
| VRPRM-MiMo w/o RL & w/o CoT# Samples=3.6K, Explicit Reasoning (CoT)=false, RL Training=false2025.08 | — | 59.65 | 50.48 | |
| VRPRM-Qwen3# Samples=53.6K, Explicit Reasoning (CoT)=true, RL Training=true2025.08 | — | 57.02 | 72.76 | |
| VRPRM-Qwen3 w/o CoT# Samples=53.6K, Explicit Reasoning (CoT)=false, RL Training=true2025.08 | — | 53.33 | 71.85 | |
| VRPRM-Qwen3 w/o RL# Samples=3.6K, Explicit Reasoning (CoT)=true, RL Training=false2025.08 | — | 59.82 | 57.58 | |
| VRPRM-Qwen3 w/o RL & w/o CoT# Samples=3.6K, Explicit Reasoning (CoT)=false, RL Training=false2025.08 | — | 55.61 | 56.23 |