Offline Reinforcement Learning on Walker2D Medium-Replay 1T10S
87.491Average ReturnREAG*MV
Evaluation Results
| Method | Links | |
|---|---|---|
| REAG*MVDistribution Shift=BodyMass (BM), Base Architecture=QT2024.10 | 87.491 | |
| QTDistribution Shift=BodyMass (BM), Base Architecture=QT2024.10 | 87.292 | |
| REAG*MVDistribution Shift=JointNoise (JN), Base Architecture=QT2024.10 | 82.363 | |
| QTDistribution Shift=JointNoise (JN), Base Architecture=QT2024.10 | 82.139 | |
| REAG*DaraDistribution Shift=JointNoise (JN), Base Architecture=QT2024.10 | 79.795 | |
| REAG*DaraDistribution Shift=BodyMass (BM), Base Architecture=QT2024.10 | 76.169 | |
| REAG*MVDistribution Shift=BodyMass (BM), Base Architecture=DT2024.10 | 73.708 | |
| DTDistribution Shift=BodyMass (BM), Base Architecture=DT2024.10 | 73.664 | |
| REAG*DaraDistribution Shift=BodyMass (BM), Base Architecture=DT2024.10 | 67.565 | |
| ReinformerDistribution Shift=BodyMass (BM), Base Architecture=Reinformer2024.10 | 67.032 | |
| REAG*DaraDistribution Shift=BodyMass (BM), Base Architecture=Reinformer2024.10 | 66.658 | |
| REAG*DaraDistribution Shift=JointNoise (JN), Base Architecture=DT2024.10 | 62.226 | |
| DTDistribution Shift=JointNoise (JN), Base Architecture=DT2024.10 | 58.255 | |
| REAG*MVDistribution Shift=JointNoise (JN), Base Architecture=DT2024.10 | 55.722 | |
| REAG*DaraDistribution Shift=JointNoise (JN), Base Architecture=Reinformer2024.10 | 55.438 | |
| ReinformerDistribution Shift=JointNoise (JN), Base Architecture=Reinformer2024.10 | 54.801 | |
| REAG*MVDistribution Shift=BodyMass (BM), Base Architecture=Reinformer2024.10 | 50.296 | |
| REAG*MVDistribution Shift=JointNoise (JN), Base Architecture=Reinformer2024.10 | 47.591 |