Multimodal Understanding on MMMU (dev)
67.3AccuracyPVM-8B (SFT + GRPO)
Evaluation Results
| Method | Links | |
|---|---|---|
| PVM-8B (SFT + GRPO)Backbone=8B, Training Strategy=SFT + GRPO, Model Architecture=Persistent Visual Memory2026.05 | 67.3 | |
| PVM-8B (SFT)Backbone=8B, Training Strategy=SFT, Model Architecture=Persistent Visual Memory2026.05 | 66.7 | |
| PEARL-8BBackbone=8B, Training Strategy=RL-tuned2026.05 | 65.3 | |
| Qwen3-VL-8B (LoRA-SFT + GRPO)Backbone=8B, Training Strategy=LoRA-SFT + GRPO2026.05 | 64.7 | |
| ICoTTraining Strategy=Visual Injection2026.05 | 63.3 | |
| Qwen3-VL-8B (LoRA-SFT)Backbone=8B, Training Strategy=LoRA-SFT2026.05 | 63.3 | |
| PVM-4B (SFT + GRPO)Backbone=4B, Training Strategy=SFT + GRPO, Model Architecture=Persistent Visual Memory2026.05 | 62.7 | |
| CoMemoTraining Strategy=Visual Injection2026.05 | 62 | |
| Euclid-8BBackbone=8B, Training Strategy=RL-tuned2026.05 | 62 | |
| OneThinker-8BBackbone=8B, Training Strategy=RL-tuned2026.05 | 62 | |
| Qwen3-VL-8B (SFT)Backbone=8B, Training Strategy=SFT2026.05 | 60.7 | |
| Qwen3-VL-8B (SFT + GRPO)Backbone=8B, Training Strategy=SFT + GRPO2026.05 | 60.7 | |
| PVM-4B (SFT)Backbone=4B, Training Strategy=SFT, Model Architecture=Persistent Visual Memory2026.05 | 60.7 | |
| MemVRTraining Strategy=Visual Injection2026.05 | 59.3 | |
| Qwen3-VL-4B (SFT)Backbone=4B, Training Strategy=SFT2026.05 | 58 | |
| Qwen3-VL-4B (SFT + GRPO)Backbone=4B, Training Strategy=SFT + GRPO2026.05 | 58 | |
| Qwen3-VL-8B-InstructBackbone=8B2026.05 | 57.3 | |
| Qwen3-VL-4B-InstructBackbone=4B2026.05 | 57.3 | |
| Qwen3-VL-4B (LoRA-SFT)Backbone=4B, Training Strategy=LoRA-SFT2026.05 | 56 | |
| Qwen3-VL-4B (LoRA-SFT + GRPO)Backbone=4B, Training Strategy=LoRA-SFT + GRPO2026.05 | 56 | |
| DefenderIteration=32026.01 | 25.33 | |
| DefenderIteration=22026.01 | 23.33 | |
| Base (M_def^(0)) + Clean DataData=Cleaned2026.01 | 21.33 | |
| Base (M_def^(0))2026.01 | 20.67 | |
| DefenderIteration=12026.01 | 20 |