Offline-to-Online Reinforcement Learning on D4RL Aggregate
77.6Average Normalized ScoreLoss Smoothing
Evaluation Results
| Method | Links | |
|---|---|---|
| Loss Smoothing2026.07 | 77.6 | |
| ROADMixing Ratio Strategy=ROAD, Algorithm=IQL2026.05 | 71.95 | |
| OPT2026.07 | 65.9 | |
| 0.1Mixing Ratio Strategy=Fixed (0.1), Algorithm=IQL2026.05 | 62.15 | |
| BRMixing Ratio Strategy=BR, Algorithm=IQL2026.05 | 61.8 | |
| DecreasingMixing Ratio Strategy=Decreasing, Algorithm=IQL2026.05 | 59 | |
| 0.3Mixing Ratio Strategy=Fixed (0.3), Algorithm=IQL2026.05 | 58.66 | |
| UniformMixing Ratio Strategy=Uniform, Algorithm=IQL2026.05 | 58.45 | |
| Hard-Switch2026.07 | 58.2 | |
| 0.2Mixing Ratio Strategy=Fixed (0.2), Algorithm=IQL2026.05 | 57.28 | |
| 0.4Mixing Ratio Strategy=Fixed (0.4), Algorithm=IQL2026.05 | 56.49 | |
| 0.5Mixing Ratio Strategy=Fixed (0.5), Algorithm=IQL2026.05 | 56.26 | |
| 0.0Mixing Ratio Strategy=Fixed (0.0), Algorithm=IQL2026.05 | 54.07 | |
| Cal-QL2026.07 | 48.8 | |
| Online2026.07 | 39.4 | |
| Offline2026.07 | 27.3 | |
| AWAC2026.07 | 21.7 |