Logical Reasoning on LogicVista (pass@1/avg@8)
64.9Avg Pass@8ADHint
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ADHintBase Model=Qwen3-VL-8B-Instruct2025.12 | 64.9 | 56.3 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2025.12 | 63 | 49.4 | |
| Qwen2.5-VL-32B + VPPOModel Scale=32B, Algorithm=VPPO2025.10 | 59.2 | — | |
| Qwen2.5-VL-32B + DAPOModel Scale=32B, Algorithm=DAPO2025.10 | 58.9 | — | |
| Qwen2.5-VL-32B + GRPOModel Scale=32B, Algorithm=GRPO2025.10 | 58.3 | — | |
| MM-Eureka-32BModel Scale=32B, Prompting Source=Official author-provided2025.10 | 56.8 | — | |
| NoisyRollout-32BModel Scale=32B, Prompting Source=Official author-provided, Note=Trained using the training set of Geo3k2025.10 | 56.2 | — | |
| Qwen2.5-VL-32BModel Scale=32B2025.10 | 52.8 | — | |
| HintGRPOBase Model=Qwen3-VL-8B-Instruct2025.12 | 51.9 | 43.1 | |
| ADHintBase MLLM=Qwen2.5-VL-7B2025.12 | 49.4 | 48.7 | |
| Qwen2.5-VL-7B + VPPOModel Scale=7B, Algorithm=VPPO2025.10 | 47.9 | — | |
| SFT -> RLBase MLLM=Qwen2.5-VL-7B2025.12 | 47.8 | 28.8 | |
| NoisyRollout-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided, Note=Trained using the training set of Geo3k2025.10 | 47.3 | — | |
| Qwen2.5-VL-7B + DAPOModel Scale=7B, Algorithm=DAPO2025.10 | 46.8 | — | |
| PAPO-D-7BModel Scale=7B, Training Strategy=Pure RL2025.10 | 46.7 | — | |
| PAPO_DBackbone=Qwen2.5-VL, Model Size=7B2025.07 | 46.7 | — | |
| GRPOBase MLLM=Qwen2.5-VL-7B2025.12 | 46.4 | 44.9 | |
| MM-Eureka-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 46.3 | — | |
| VL-Rethinker-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 46.3 | — | |
| PAPO_GBackbone=Qwen2.5-VL, Model Size=7B2025.07 | 46.07 | — | |
| SFTBase MLLM=Qwen2.5-VL-7B2025.12 | 46 | 23.7 | |
| GRPOBackbone=Qwen2.5-VL, Model Size=7B2025.07 | 45.62 | — | |
| R1-ShareVL-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 45.6 | — | |
| Qwen2.5-VL-7B + GRPOModel Scale=7B, Algorithm=GRPO2025.10 | 45.6 | — | |
| Qwen2.5-VL-7BModel Scale=7B2025.10 | 42.4 | — | |
| PAPO_DBackbone=Qwen2.5-VL, Model Size=3B2025.07 | 41.67 | — | |
| ADHintBase MLLM=Qwen2.5-VL-3B2025.12 | 41.4 | 42.6 | |
| DAPOBackbone=Qwen2.5-VL, Model Size=3B2025.07 | 40.69 | — | |
| GHPOBase MLLM=Qwen2.5-VL-7B2025.12 | 40.5 | 42.4 | |
| SFT -> RLBase MLLM=Qwen2.5-VL-3B2025.12 | 39.9 | 18.1 | |
| PEPO_GBackbone=Qwen2.5-VL-3B-Instruct, Training Data=ViRL39K2026.03 | 39.85 | — | |
| ThinkLite-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 39.4 | — | |
| LUFFYBase MLLM=Qwen2.5-VL-7B2025.12 | 39.2 | 40.2 | |
| GRPOBase MLLM=Qwen2.5-VL-3B2025.12 | 38.9 | 43.8 | |
| PAPO_GBackbone=Qwen2.5-VL-3B-Instruct, Training Data=ViRL39K2026.03 | 38.67 | — | |
| PAPO_GBackbone=Qwen2.5-VL, Model Size=3B2025.07 | 38.67 | — | |
| LUFFYBase MLLM=Qwen2.5-VL-3B2025.12 | 38.3 | 35.4 | |
| GRPOBackbone=Qwen2.5-VL-3B-Instruct, Training Data=ViRL39K2026.03 | 38.14 | — | |
| GRPOBackbone=Qwen2.5-VL, Model Size=3B2025.07 | 38.14 | — | |
| SFTBase MLLM=Qwen2.5-VL-3B2025.12 | 37.8 | 18.1 | |
| DAPOBackbone=Qwen2.5-VL, Model Size=7B2025.07 | 37.05 | — | |
| HintGRPOBase MLLM=Qwen2.5-VL-7B2025.12 | 34.7 | 36.4 | |
| GHPOBase MLLM=Qwen2.5-VL-3B2025.12 | 33.9 | 35.3 | |
| HintGRPOBase MLLM=Qwen2.5-VL-3B2025.12 | 33.7 | 35.7 | |
| PAPO_GBackbone=Qwen3-VL (thinking), Model Size=2B2025.07 | 32.83 | — | |
| GRPOBackbone=Qwen3-VL (thinking), Model Size=2B2025.07 | 29.84 | — | |
| StepHintBase MLLM=Qwen2.5-VL-3B2025.12 | 27.8 | 12.5 | |
| StepHintBase MLLM=Qwen2.5-VL-7B2025.12 | 22.7 | 10.5 |