Reinforcement Learning on HalfCheetah Random
103.45Avg Normalized ScoreROAD
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ROADBackbone=Proto2026.05 | 103.45 | — | — | — | — | |
| Fixed-0.1Mixing Ratio=0.1, Backbone=Proto2026.05 | 101.24 | — | — | — | — | |
| Fixed-0.0Mixing Ratio=0.0, Backbone=Proto2026.05 | 100.2 | — | — | — | — | |
| Fixed-0.3Mixing Ratio=0.3, Backbone=Proto2026.05 | 98.69 | — | — | — | — | |
| Fixed-0.5Mixing Ratio=0.5, Backbone=Proto2026.05 | 97.3 | — | — | — | — | |
| Fixed-0.4Mixing Ratio=0.4, Backbone=Proto2026.05 | 95.38 | — | — | — | — | |
| Fixed-0.2Mixing Ratio=0.2, Backbone=Proto2026.05 | 95.35 | — | — | — | — | |
| SLAC-off+S2PBase Algorithm=SLAC-off, S2P Augmentation=true2022.09 | 18.14 | — | — | — | — | |
| SLAC-offBase Algorithm=SLAC-off, S2P Augmentation=false2022.09 | 16.37 | — | — | — | — | |
| IQL+S2PBase Algorithm=IQL, S2P Augmentation=true2022.09 | 12.64 | — | — | — | — | |
| CQL+S2PBase Algorithm=CQL, S2P Augmentation=true2022.09 | 11.77 | — | — | — | — | |
| IQLBase Algorithm=IQL, S2P Augmentation=false2022.09 | 10.28 | — | — | — | — | |
| CQLBase Algorithm=CQL, S2P Augmentation=false2022.09 | 4.89 | — | — | — | — | |
| BCQData source (optimality)=Random, Algorithm variant=Planning2021.02 | — | -1.46 | -1.75 | -1.69 | — | |
| BCQData source (optimality)=Random, Algorithm variant=Acting2021.02 | — | -498.9 | -113.3 | -159.5 | — | |
| CQLData source (optimality)=Random, Algorithm variant=Planning2021.02 | — | -481.5 | -442.2 | -672.4 | — | |
| CQLData source (optimality)=Random, Algorithm variant=Acting2021.02 | — | -0.7 | -2.8 | -0.6 | — | |
| Data Score2021.02 | — | — | — | — | -303 | |
| GELATOvariant=metric2021.02 | — | — | — | — | 2,560 | |
| GELATOvariant=L2021.02 | — | — | — | — | 512 | |
| Imitation2021.02 | — | — | — | — | -41 | |
| MBPO2021.02 | — | — | — | — | 3,527 | |
| MOPO2021.02 | — | — | — | — | 4,114 | |
| MORELData source (optimality)=Random, Algorithm variant=Planning2021.02 | — | -102.3 | -188.7 | -181 | — | |
| MORELData source (optimality)=Random, Algorithm variant=Acting2021.02 | — | -430.8 | -673.7 | -365.5 | — | |
| PE-TS CaDMData source (optimality)=Random2021.02 | — | 754.6 | 744.5 | 767.4 | — | |
| PerSimData source (optimality)=Random2021.02 | — | 2,124 | 2,060 | 472 | — | |
| SAC2021.02 | — | — | — | — | 3,502 | |
| True env+MPCData source (optimality)=Oracle (True Environment)2021.02 | — | 7,459 | 42,893 | 66,675 | — | |
| Vanilla CaDMData source (optimality)=Random2021.02 | — | 288.4 | 362.9 | 351.8 | — |