Online Reinforcement Learning on DMControl FingerSpin (final)
903.92Normalized ReturnGoRL(FM)
Evaluation Results
| Method | Links | |
|---|---|---|
| GoRL(FM)Seeds=5, Decoder=Flow-Matching2025.12 | 903.92 | |
| GoRL(Diff)Seeds=5, Decoder=Diffusion2025.12 | 844.74 | |
| DPPOSeeds=52025.12 | 694.06 | |
| PPOSeeds=52025.12 | 539.03 | |
| FPOSeeds=52025.12 | 56.05 |