Vision-Language Modeling on BabyV
18.6AccuracyZPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| ZPPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 18.6 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 14.4 | |
| ZPPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 13.9 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 13.7 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 13.1 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 12.8 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 12.5 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 11.9 | |
| Qwen3.5Model Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 11.6 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 9.8 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 8.6 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 7.8 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 7.5 | |
| Qwen3.5Model Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 6.7 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 6.7 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 6.7 |