End-to-end training performance on DAPO-MATH-17k (train)
125.6Step Time (s)Relax
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RelaxBackbone=Qwen3-4B, Training Algorithm=DAPO, Hardware (GPUs)=16×H800, Maximum response length=20,480 tokens, Staleness=12026.04 | 125.6 | 28.7 | 0 | 0 | |
| veRLBackbone=Qwen3-4B, Training Algorithm=DAPO, Hardware (GPUs)=16×H800, Maximum response length=20,480 tokens, Staleness=12026.04 | 150.5 | 23.9 | 38.2 | 27.3 |