Multimodal Reasoning on MMStar (Accuracy)
79.2AccuracyTVI-CoT
Evaluation Results
| Method | Links | |
|---|---|---|
| TVI-CoTBackbone=Qwen3-VL-32B2026.06 | 79.2 | |
| Qwen3-VL-32BBackbone=Qwen3-VL-32B2026.06 | 77.7 | |
| LaReBase VLM=Qwen3-VL-8B-Instruct2025.11 | 77.1 | |
| o32025.11 | 75.3 | |
| Qwen3-VL-8B-Thinking2025.11 | 75.2 | |
| Gemini2.5-Pro2025.11 | 74.7 | |
| CapImagineBase VLM=Qwen3-VL-8B-Instruct2025.11 | 74 | |
| PRCRModel=Qwen3-VL, #Params=32B2026.06 | 73.91 | |
| Token-ReplayModel=Qwen3-VL, #Params=32B2026.06 | 73.54 | |
| AutoNPO2026.04 | 72.63 | |
| Vision-R1Base VLM=Qwen3-VL-8B-Instruct2025.11 | 72.6 | |
| BaselineModel=Qwen3-VL, #Params=32B2026.06 | 72.43 | |
| RLEPtype=far future2026.04 | 72.27 | |
| GRPOtype=pure on-policy2026.04 | 72.2 | |
| NPOstage=early + late-stage2026.04 | 72.2 | |
| ExGRPOtype=historical replay2026.04 | 72 | |
| Qwen3-VL-8B-Instruct2026.04 | 71.83 | |
| Qwen3-VL-8B2026.07 | 70.4 | |
| NPOstage=early-stage only2026.04 | 70.3 | |
| OpenMMReasoner-7BEvaluation Source=reproduced2026.05 | 70 | |
| Qwen3-VL-8B + Qwen3-8B / H-OPDStudent=Qwen3-VL-4B-Instruct2026.07 | 70 | |
| AnE-3rdTraining Stage=Round 32026.05 | 69.9 | |
| LUFFYtype=external teacher2026.04 | 69.47 | |
| PRCRModel=InternVL3.5, #Params=8B2026.06 | 69.32 | |
| InternVL3.5-8B2025.11 | 69.3 | |
| Token-ReplayModel=InternVL3.5, #Params=8B2026.06 | 69.05 | |
| PRCRModel=Qwen3-VL, #Params=8B2026.06 | 68.96 | |
| AnE-2ndTraining Stage=Round 22026.05 | 68.9 | |
| Token-ReplayModel=Qwen3-VL, #Params=8B2026.06 | 68.42 | |
| BaselineModel=InternVL3.5, #Params=8B2026.06 | 68.36 | |
| AnE-1stTraining Stage=Round 12026.05 | 68.3 | |
| PStarOpen-Source=Yes, Data=0.5k, Training-Free=Yes2026.05 | 68 | |
| PRCRModel=InternVL3.5, #Params=14B2026.06 | 67.82 | |
| Token-ReplayModel=InternVL3.5, #Params=14B2026.06 | 67.45 | |
| SPARK-VL-7BEvaluation Source=original paper2026.05 | 67.3 | |
| Preliminary RLTraining Phase=Warm-up2026.05 | 67.2 | |
| LVRBase VLM=Qwen3-VL-8B-Instruct2025.11 | 67.1 | |
| BaselineModel=InternVL3.5, #Params=14B2026.06 | 66.76 | |
| Metis-RISEEvaluation Source=reproduced2026.05 | 66.6 | |
| BaselineModel=Qwen3-VL, #Params=8B2026.06 | 65.95 | |
| ThymeModel Type=Thinking-with-Images Agent Model2026.04 | 65.9 | |
| ThymeBase VLM=Qwen3-VL-8B-Instruct2025.11 | 65.9 | |
| Vision-R1Evaluation Source=reproduced2026.05 | 65.8 | |
| MMR1Evaluation Source=reproduced2026.05 | 65.3 | |
| LLaVA-Critic-R1Evaluation Source=original paper2026.05 | 65.1 | |
| GPT-4o2025.09 | 63.9 | |
| VL-RethinkerEvaluation Source=reproduced2026.05 | 63.6 | |
| ZoomEyeModel Type=Thinking-with-Images Agent Model2026.04 | 63.2 | |
| DeFacto (Ours)Backbone=Qwen2.5-VL-7B2025.09 | 63.2 | |
| Pixel Reasoner2025.09 | 62.9 | |
| Visual-SR12025.09 | 62.8 | |
| Unsilencing Latent ReasoningBackbone=VLAA Thinking-7B2026.05 | 62.7 | |
| Qwen2.5-VL-7B-InstructEvaluation Source=original paper2026.05 | 62.5 | |
| HyLaR-7BModel Type=Visual-Latent Model, Number of Parameters=7B2026.04 | 62 | |
| AStarOpen-Source=Yes, Data=0.5k, Training-Free=Yes2026.05 | 61.7 | |
| OpenVLThinkerEvaluation Source=reproduced2026.05 | 61.6 | |
| MulberryOpen-Source=No, Data=260k, Training-Free=No2026.05 | 61.3 | |
| HyLaR-SFTModel Type=Visual-Latent Model, Training Protocol=SFT2026.04 | 60.47 | |
| MonetModel Type=Visual-Latent Model2026.04 | 60.33 | |
| LaserModel Type=Visual-Latent Model2026.04 | 60.27 | |
| Qwen2.5-VL-7BModel Type=Open-Source Model, Number of Parameters=7B2026.04 | 59.7 | |
| LlamaV-o1Open-Source=No, Data=118k, Training-Free=No2026.05 | 59.5 | |
| DMLRBackbone=VLAA Thinking-7B2026.05 | 59.2 | |
| LLaVA-OneVisionModel Type=Open-Source Model2026.04 | 59.13 | |
| CCoTBackbone=VLAA Thinking-7B2026.05 | 59 | |
| VanillaBackbone=VLAA Thinking-7B2026.05 | 58.9 | |
| DeepEyesModel Type=Thinking-with-Images Agent Model2026.04 | 58.73 | |
| ICoTBackbone=VLAA Thinking-7B2026.05 | 58.2 | |
| Chain-of-Focus2025.09 | 58.1 | |
| LVRModel Type=Visual-Latent Model2026.04 | 57.93 | |
| LLaVA-CoTOpen-Source=No, Data=100k, Training-Free=No2026.05 | 57.6 | |
| MCoTBackbone=VLAA Thinking-7B2026.05 | 57.1 | |
| Revisual-R1Evaluation Source=reproduced2026.05 | 57.1 | |
| Unsilencing Latent ReasoningBackbone=R1 OneVision-7B2026.05 | 56.8 | |
| DMLRBackbone=R1 OneVision-7B2026.05 | 56.2 | |
| ESCModel=Qwen2 [92]2026.07 | 56.07 | |
| BaselineModel=Qwen2 [92]2026.07 | 55.93 | |
| ICoTBackbone=R1 OneVision-7B2026.05 | 54 | |
| SID-vModel=GLM4.1V2025.10 | 54 | |
| SID-cModel=GLM4.1V2025.10 | 54 | |
| CCoTBackbone=R1 OneVision-7B2026.05 | 53.5 | |
| InternVL3.5-8BModel Type=Open-Source Model, Number of Parameters=8B2026.04 | 53.33 | |
| H-GRPOBackbone=Qwen2.5-VL-3B2026.06 | 52.4 | |
| VanillaBackbone=R1 OneVision-7B2026.05 | 52.1 | |
| MCoTBackbone=R1 OneVision-7B2026.05 | 51.6 | |
| R1-VLParameters=2B2026.06 | 49.8 | |
| ViGoRLParameters=3B, Setting=UGround2026.06 | 47.5 | |
| MADModel=GLM4.1V2025.10 | 47 | |
| Qwen2.5-VLParameters=3B2026.06 | 46 | |
| DeepEyes2025.09 | 43.6 | |
| Janus-pro-7B2025.11 | 41 | |
| ESCModel=LLaVA [50]2026.07 | 34.13 | |
| BaselineModel=LLaVA [50]2026.07 | 33.53 | |
| IOModel=GLM4.1V2025.10 | 32 | |
| CoTModel=GLM4.1V2025.10 | 29 | |
| Self-ConsisModel=GLM4.1V2025.10 | 29 | |
| SID-vModel=LLaVA1.6-13B2025.10 | 14 | |
| SID-cModel=LLaVA1.6-13B2025.10 | 14 | |
| MADModel=LLaVA1.6-13B2025.10 | 12 | |
| CoTModel=LLaVA1.6-13B2025.10 | 11 |