Meta-Reinforcement Learning on Meta-World ML1 max 5M steps
6825th Percentile Success RateT-SAC
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| T-SACmax environment steps=5M, success_criterion=final timestep only2025.03 | 68 | 78 | 86 | |
| SACmax environment steps=5M, success_criterion=final timestep only2025.03 | 53 | 60 | 68 | |
| CrossQmax environment steps=5M, success_criterion=final timestep only2025.03 | 50 | 50 | 50 | |
| GTrXL policymax environment steps=5M, success_criterion=final timestep only2025.03 | 18 | 35 | 53 |