Offline Reinforcement Learning on D4RL Adroit v1
87.8Score (pen-human)DOSER
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DOSERMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 87.8 | 79.3 | 83.6 | |
| QGPOMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 73.9 | 54.2 | 64.1 | |
| SVRMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 73.1 | 70.2 | 71.7 | |
| DQLMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 72.8 | 57.3 | 65.1 | |
| IQLMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 71.5 | 37.3 | 54.4 | |
| DTQLMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 64.1 | 81.3 | 72.7 | |
| TD3+BCMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 54.9 | 63.8 | 59.4 | |
| CQLMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 35.2 | 27.2 | 31.2 |