Reinforcement Learning on Humanoid v5 (Mean Performance)
5,350.9Mean ReturnTFM-S3-TD3
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TFM-S3-TD3Algorithm=TFM-S3-TD3, Variant=proposed2026.04 | 5,350.9 | 7.4 | 38 | |
| TFM-S3-TD3Algorithm=TFM-S3-TD3, Variant=One shot2026.04 | 5,271.2 | 16.6 | 43 | |
| Random Search TD3Algorithm=Random Search, Candidates=322026.04 | 5,245.2 | 12.4 | 48 | |
| TD3Algorithm=TD3, Variant=vanilla2026.04 | 5,127.6 | 13.7 | 48 |