Offline Meta Reinforcement Learning on Walker-speed (out-of-distribution)
831.5Average ReturnSPC
Evaluation Results
| Method | Links | |
|---|---|---|
| SPCZero-shot=true2026.03 | 831.5 | |
| CSROZero-shot=true2026.03 | 767.2 | |
| FOCALZero-shot=true2026.03 | 659.6 | |
| UNICORN-SSZero-shot=true2026.03 | 623.7 | |
| UNICORN-SUPZero-shot=true2026.03 | 535.5 | |
| DORAZero-shot=true2026.03 | 425.3 |