Continuous Control on MuJoCo Playground 10 tasks
68610-step RewardReFPO*
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ReFPO*Training stage=0–100M steps, lambda=0.042026.06 | 686 | 690 | 0.0116 | 0.0019 | |
| ReFPOTraining stage=0–100M steps, Trajectory-advantage-weighted=true2026.06 | 671 | 640 | — | — | |
| ReFPOTraining stage=0–100M steps, lambda=0.082026.06 | 646 | 652 | 0.0161 | 0.0057 | |
| ReFPOTraining stage=0–100M steps, lambda=0.022026.06 | 644 | 556 | 0.0187 | 0.0038 | |
| FPOTraining stage=0–100M steps2026.06 | 641 | 565 | 0.0475 | 0.0037 | |
| PPOTraining stage=0–100M steps2026.06 | 622 | — | — | — | |
| ReFPOTraining stage=0–100M steps, lambda=0.122026.06 | 609 | 601 | 0.0193 | 0.0033 | |
| ReinFlowTraining stage=100M–200M steps2026.06 | 523 | 357 | — | — | |
| Flow-GRPOTraining stage=100M–200M steps2026.06 | 507 | 0.0002 | — | — | |
| Flow-GRPOTraining stage=0–100M steps2026.06 | 475 | 0.0001 | — | — | |
| ReinFlowTraining stage=0–100M steps2026.06 | 416 | 339 | — | — |