Mathematical Reasoning on VisuLogic (Accuracy)
30.6Accuracy (VisuLogic Math)V-STAR
Evaluation Results
| Method | Links | |
|---|---|---|
| V-STARData Size=40k2026.04 | 30.6 | |
| OpenVLThinkerData Size=59.2k2026.04 | 29.1 | |
| VL-Cogito-7B + LEADBackbone=VL-Cogito-7B, Decoding Strategy=LEAD2026.03 | 28.9 | |
| ThinkLite-VLData Size=11k2026.04 | 28.9 | |
| VL-Rethinker-7B + LEADBackbone=VL-Rethinker-7B, Decoding Strategy=LEAD2026.03 | 28.5 | |
| VL-Cogito-7BBackbone=VL-Cogito-7B2026.03 | 28.2 | |
| VL-CogitoData Size=80k2026.04 | 28.2 | |
| Vision-R1-7B + LEADBackbone=Vision-R1-7B, Decoding Strategy=LEAD2026.03 | 27.9 | |
| EVE (Ours-8B-iter4)Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=42026.04 | 27.8 | |
| GPT-5 mini (minimal)Model Category=Closed-Source MLLMs, Scale=minimal2026.04 | 27.6 | |
| Jigsaw-R1-8BModel Category=Template-based Self-evolution Methods, Scale=8B2026.04 | 27.4 | |
| VL-Rethinker-7BBackbone=VL-Rethinker-7B2026.03 | 27.3 | |
| VL-RethinkerData Size=39k2026.04 | 27.3 | |
| Qwen2.5VL2026.04 | 26.9 | |
| Vision-R1-7BBackbone=Vision-R1-7B2026.03 | 26.4 | |
| Vision-R1Data Size=210k2026.04 | 26.4 | |
| GPT-4o-20240513Model Category=Closed-Source MLLMs2026.04 | 26.3 | |
| VisPlay-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 26.3 | |
| R1-Onevision-7B + LEADBackbone=R1-Onevision-7B, Decoding Strategy=LEAD2026.03 | 26.1 | |
| MM-Zero-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 25.7 | |
| R1-Onevision-7BBackbone=R1-Onevision-7B2026.03 | 24.9 | |
| R1-OnevisionData Size=155k2026.04 | 24.9 | |
| Qwen3-VL-8B-InstructModel Category=Open-Source MLLMs, Scale=8B2026.04 | 24.6 | |
| GPT-5 nano (high)Model Category=Closed-Source MLLMs, Scale=high2026.04 | 24.5 |