Offline-to-Online Reinforcement Learning on D4RL Antmaze v2 (All Configurations)
98.9Score (umaze-default)EDIS-Cal-QL
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| EDIS-Cal-QLPolicy Type=Gaussian, Offline training steps=1M, Online fine-tuning steps=200K, Number of seeds=8, Batch size=2562026.05 | 98.9 | 95.9 | 93.9 | 89.3 | 66.1 | 57.1 | 83.5 | |
| Q-FlowPolicy Type=Flow, Offline training steps=1M, Online fine-tuning steps=200K, Number of seeds=8, Batch size=2562026.05 | 96.3 | 96.8 | 82 | 84.3 | 78 | 80.8 | 86.4 | |
| FQLPolicy Type=Flow, Offline training steps=1M, Online fine-tuning steps=200K, Number of seeds=8, Batch size=2562026.05 | 96 | 96.3 | 88.3 | 83.5 | 80.5 | 84 | 88.1 | |
| FBRACPolicy Type=Flow, Offline training steps=1M, Online fine-tuning steps=200K, Number of seeds=8, Batch size=2562026.05 | 95 | 72.1 | 78 | 71 | 36 | 40 | 69.2 | |
| EDIS-IQLPolicy Type=Gaussian, Offline training steps=1M, Online fine-tuning steps=200K, Number of seeds=8, Batch size=2562026.05 | 81.1 | 66.7 | 86.2 | 81.8 | 40 | 52.1 | 68 |