Reinforcement Learning Control on DMControl humanoid-walk
868ReturnClean
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Clean2026.06 | 868 | 9.2 | — | — | — | |
| SWAAP (Random)alpha=0.99, rp=0.12026.06 | 826 | — | 10.3 | 9 | — | |
| SWAAPalpha=0.99, rp=0.12026.06 | 775 | — | 11 | 9 | 9 | |
| Direct Model Poisoning2026.06 | 307 | — | 17.4 | — | — |