Vision-Language Understanding on MVista
80.5AccuracyZPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| ZPPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 80.5 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 79.3 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 79 | |
| Qwen3.5Model Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 78.6 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 78.2 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 77.9 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 77.9 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 77.8 | |
| Qwen3-VL-8BTraining Strategy=Staged2026.05 | 75.9 | |
| Qwen3-VL-8BTraining Strategy=Merged2026.05 | 73.8 | |
| ZPPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 73.2 | |
| Qwen3-VL-8BTraining Strategy=Base2026.05 | 72.4 | |
| Qwen2.5-VL-7BTraining Strategy=Staged2026.05 | 71.45 | |
| InternVL3.5-8BTraining Strategy=Staged2026.05 | 70.3 | |
| Qwen2.5-VL-7BTraining Strategy=Merged2026.05 | 69.75 | |
| InternVL3.5-8BTraining Strategy=Merged2026.05 | 69.4 | |
| Qwen2.5-VL-7BTraining Strategy=Base2026.05 | 68.4 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 68.3 | |
| InternVL3-8BTraining Strategy=Staged2026.05 | 65.4 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 65.2 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 63.6 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 62.7 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 62.2 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 62 | |
| InternVL3.5-8BTraining Strategy=Base2026.05 | 60.7 | |
| InternVL3-8BTraining Strategy=Merged2026.05 | 60.7 | |
| Qwen3.5Model Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 60.7 | |
| InternVL3-8BTraining Strategy=Base2026.05 | 18.1 |