Vision-Language Modeling on MVerse
76AccuracyZPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| ZPPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 76 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 72.8 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 72.3 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 72 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 71.9 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 71.4 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 70.8 | |
| Qwen3.5Model Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 69.7 | |
| ZPPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 59.3 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 51.1 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 47.7 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 47.4 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 45.8 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 45.8 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 45.3 | |
| Qwen3.5Model Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 43.5 |