Offline Reinforcement Learning on D4RL Gym-MuJoCo v2 (Medium, Medium-Replay, Medium-Expert)
69.8HalfCheetah-Medium ReturnACL-QL
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ACL-QLMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 69.8 | 97.9 | 79.3 | 55.9 | 99.3 | 96.5 | 87.4 | 107.2 | 113.4 | 89.6 | |
| A2PRMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 68.6 | 100.8 | 89.7 | 56.6 | 101.5 | 94.4 | 98.3 | 112.1 | 114.6 | 93 | |
| DOSERMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 67.5 | 104 | 86.7 | 63 | 104.4 | 94.4 | 96.2 | 111.5 | 110.9 | 93.2 | |
| SVRMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 60.5 | 103.5 | 92.4 | 52.5 | 103.7 | 95.6 | 94.2 | 111.2 | 109.3 | 91.4 | |
| SRPOMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 60.4 | 95.5 | 84.4 | 51.4 | 101.2 | 84.6 | 92.2 | 100.1 | 114 | 87.1 | |
| DTQLMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 57.9 | 99.6 | 89.4 | 50.9 | 100 | 88.5 | 92.7 | 109.3 | 110 | 88.7 | |
| QGPOMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 54.1 | 98 | 86 | 47.6 | 96.9 | 84.4 | 93.5 | 108 | 110.7 | 86.6 | |
| DQLMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 51.5 | 90.5 | 87 | 47.8 | 101.3 | 95.5 | 96.8 | 111.1 | 110.1 | 88 | |
| IDQLMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 51 | 65.4 | 82.5 | 45.9 | 92.1 | 85.1 | 95.9 | 108.6 | 112.7 | 82.1 | |
| TD3+BCMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 48.3 | 59.3 | 83.7 | 44.6 | 60.9 | 81.8 | 90.7 | 98 | 110.1 | 75.3 | |
| IQLMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 47.4 | 66.3 | 78.3 | 44.2 | 94.7 | 73.9 | 86.7 | 91.5 | 109.6 | 83.3 | |
| SfBCMethod category=Diffusion-based methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 45.9 | 57.1 | 77.9 | 37.1 | 86.2 | 65.1 | 92.6 | 108.6 | 109.8 | 75.6 | |
| CQLMethod category=Conventional methods, Evaluation protocol=Average normalized scores at the last training iteration over 4 random seeds2026.05 | 44 | 58.5 | 72.5 | 45.5 | 95 | 77.2 | 91.6 | 105.4 | 108.8 | 77.6 |