Text Generation on HH-RLHF and IMDB (test)
3.2Total Training Time (h)PS-PPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PS-PPOBase model=Qwen2.5-3B-Instruct2026.06 | 3.2 | 42.8 | 6.72 | |
| PS-PPOBase model=Llama-3.1-8B-Instruct2026.06 | 4.4 | 45.7 | 7.85 | |
| S-GRPOBase model=Qwen2.5-3B-Instruct2026.06 | 6.7 | 68.9 | 5.9 | |
| DAPOBase model=Qwen2.5-3B-Instruct2026.06 | 7.3 | 73.8 | 7.07 | |
| S-GRPOBase model=Llama-3.1-8B-Instruct2026.06 | 8.1 | 72.6 | 7.09 | |
| DAPOBase model=Llama-3.1-8B-Instruct2026.06 | 11.3 | 77.9 | 8.13 |