Vision-Language Modeling on MMMUPro
53.2AccuracyZPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| ZPPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 53.2 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 49.6 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 49.3 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 49.3 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 48.8 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 47.9 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 47.4 | |
| Qwen3.5Model Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 46.2 | |
| ZPPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 37.6 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 30.5 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 29.9 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 29 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 28.8 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 28.2 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 27.6 | |
| Qwen3.5Model Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 26.8 |