Vision-Language Modeling on EmbSp
79.5AccuracyZPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| ZPPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 79.5 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 78.7 | |
| GRPOModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 78.3 | |
| Qwen3.5Model Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 78.2 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 77.6 | |
| On-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 77.5 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=true2026.06 | 77.4 | |
| Off-DistillModel Scale=2B, Prompt Replay Buffer Augmentation=false2026.06 | 77.2 | |
| ZPPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 71.5 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 69.4 | |
| GRPOModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 69.2 | |
| Qwen3.5Model Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 67.9 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 67.1 | |
| On-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 66.7 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=false2026.06 | 65.8 | |
| Off-DistillModel Scale=0.8B, Prompt Replay Buffer Augmentation=true2026.06 | 65.8 |