Text Generation on HH-RLHF and IMDB
7.12Helpful Assistant ScoreDAPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DAPOBase model=Llama-3.1-8B-Instruct2026.06 | 7.12 | 9.13 | 8.13 | |
| PS-PPOBase model=Llama-3.1-8B-Instruct2026.06 | 7 | 8.69 | 7.85 | |
| DAPO (w/ forking tokens)Base model=Llama-3.1-8B-Instruct2026.06 | 6.9 | 8.65 | 7.86 | |
| Dr.GRPOBase model=Llama-3.1-8B-Instruct2026.06 | 6.82 | 8.75 | 7.79 | |
| GRPOBase model=Llama-3.1-8B-Instruct2026.06 | 6.36 | 8.16 | 7.26 | |
| DAPOBase model=Qwen2.5-3B-Instruct2026.06 | 6.18 | 7.97 | 7.07 | |
| S-GRPOBase model=Llama-3.1-8B-Instruct2026.06 | 6.14 | 7.99 | 7.09 | |
| PS-PPOBase model=Qwen2.5-3B-Instruct2026.06 | 6.02 | 7.42 | 6.72 | |
| DAPO (w/ forking tokens)Base model=Qwen2.5-3B-Instruct2026.06 | 5.95 | 7.46 | 6.71 | |
| RLOOBase model=Llama-3.1-8B-Instruct2026.06 | 5.94 | 7.71 | 6.81 | |
| Dr.GRPOBase model=Qwen2.5-3B-Instruct2026.06 | 5.61 | 7.14 | 6.35 | |
| GRPOBase model=Qwen2.5-3B-Instruct2026.06 | 5.29 | 7.58 | 6.43 | |
| S-GRPOBase model=Qwen2.5-3B-Instruct2026.06 | 5.11 | 6.73 | 5.9 | |
| RLOOBase model=Qwen2.5-3B-Instruct2026.06 | 4.89 | 7 | 5.93 | |
| BaseBase model=Llama-3.1-8B-Instruct2026.06 | 4.71 | 6.44 | 5.59 | |
| BaseBase model=Qwen2.5-3B-Instruct2026.06 | 4.08 | 5.74 | 4.93 |