Mathematical Reasoning on MathVerse V
67.8AccuracyEASE-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| EASE-7BModel Scale=7B2026.05 | 67.8 | |
| VPPO-RL-7BModel Scale=7B2026.05 | 67.6 | |
| VGPO-7BModel Scale=7B2026.05 | 67.6 | |
| PAPO_D-7BModel Scale=7B2026.05 | 66.6 | |
| PRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 66.3 | |
| VPPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 65.5 | |
| VL-Rethinker-7BModel Scale=7B, Rollout=82026.06 | 65 | |
| PAPO-DModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 64.9 | |
| VL-Rethinker-7BModel Scale=7B2026.05 | 64.8 | |
| R1-ShareVL-7BModel Scale=7B, Rollout=82026.06 | 64.3 | |
| PAPO-GModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 64.1 | |
| NoisyRollout-7BModel Scale=7B, Rollout=82026.06 | 63.8 | |
| MM-Eureka-7BModel Scale=7B2026.05 | 63.6 | |
| MM-Eureka-7BModel Scale=7B, Rollout=82026.06 | 62.4 | |
| GRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 61.7 | |
| PRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 60.7 | |
| PVM-8B (SFT + GRPO)Backbone=8B, Training Strategy=SFT + GRPO, Model Architecture=Persistent Visual Memory2026.05 | 59.8 | |
| PEARL-8BBackbone=8B, Training Strategy=RL-tuned2026.05 | 58.9 | |
| Qwen3-VL-8B (SFT + GRPO)Backbone=8B, Training Strategy=SFT + GRPO2026.05 | 58.5 | |
| CFPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 58.43 | |
| VPPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 58.1 | |
| Euclid-8BBackbone=8B, Training Strategy=RL-tuned2026.05 | 57.9 | |
| PAPO-DModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 57.7 | |
| Qwen3-VL-8B (LoRA-SFT + GRPO)Backbone=8B, Training Strategy=LoRA-SFT + GRPO2026.05 | 57.6 | |
| PVM-8B (SFT)Backbone=8B, Training Strategy=SFT, Model Architecture=Persistent Visual Memory2026.05 | 57.5 | |
| OneThinker-8BBackbone=8B, Training Strategy=RL-tuned2026.05 | 57.4 | |
| Qwen3-VL-8B (SFT)Backbone=8B, Training Strategy=SFT2026.05 | 56.9 | |
| ICoTTraining Strategy=Visual Injection2026.05 | 56.1 | |
| CoMemoTraining Strategy=Visual Injection2026.05 | 55 | |
| Qwen3-VL-8B (LoRA-SFT)Backbone=8B, Training Strategy=LoRA-SFT2026.05 | 55 | |
| PVM-4B (SFT + GRPO)Backbone=4B, Training Strategy=SFT + GRPO, Model Architecture=Persistent Visual Memory2026.05 | 55 | |
| Qwen3-VL-4B (SFT + GRPO)Backbone=4B, Training Strategy=SFT + GRPO2026.05 | 54.6 | |
| PVM-4B (SFT)Backbone=4B, Training Strategy=SFT, Model Architecture=Persistent Visual Memory2026.05 | 54.4 | |
| Qwen3-VL-4B (LoRA-SFT + GRPO)Backbone=4B, Training Strategy=LoRA-SFT + GRPO2026.05 | 54.2 | |
| PAPO-GModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 53.6 | |
| DAPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 53.1 | |
| MemVRTraining Strategy=Visual Injection2026.05 | 52.9 | |
| Qwen3-VL-8B-InstructBackbone=8B2026.05 | 52.9 | |
| Qwen3-VL-4B (SFT)Backbone=4B, Training Strategy=SFT2026.05 | 52.7 | |
| Qwen3-VL-4B-InstructBackbone=4B2026.05 | 52.4 | |
| Qwen3-VL-4B (LoRA-SFT)Backbone=4B, Training Strategy=LoRA-SFT2026.05 | 52.4 | |
| GRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 52.3 | |
| NoisyRollout-7BModel Scale=7B2026.05 | 51.7 | |
| DAPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 51 | |
| PAPOGBackbone=Qwen3-VL-2B-Thinking2026.06 | 50.05 | |
| GRPOBackbone=Qwen3-VL-2B-Thinking2026.06 | 45.41 | |
| Qwen2.5-VL-7BModel Scale=7B, Rollout=82026.06 | 34.3 | |
| ThinkLite-VL-7BModel Scale=7B2026.05 | 30.9 | |
| Qwen2.5-VL-3BModel Scale=3B, Rollout=82026.06 | 30.1 |