Locomotion on Hopper
100Convergence (%)DAgger
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DAggerInitial expert dataset size (M)=1,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 3,930,619 | 1,019,895 | |
| DAggerInitial expert dataset size (M)=2,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 3,728,601 | 10,242,247 | |
| DAggerInitial expert dataset size (M)=5,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 3,040,870 | 10,249,170 | |
| DAggerInitial expert dataset size (M)=10,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 3,294,696 | 1,009,159 | |
| DAggerInitial expert dataset size (M)=20,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 4,576,657 | 10,192,135 | |
| EnsembleDAggerInitial expert dataset size (M)=1,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 1,969,178 | 8,879,193 | |
| EnsembleDAggerInitial expert dataset size (M)=2,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 1,839,318 | 8,722,217 | |
| EnsembleDAggerInitial expert dataset size (M)=5,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 2,043,324 | 8,966,155 | |
| EnsembleDAggerInitial expert dataset size (M)=10,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 1,990,446 | 8,806,315 | |
| EnsembleDAggerInitial expert dataset size (M)=20,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 2,804,320 | 8,576,137 | |
| ThriftyDAggerInitial expert dataset size (M)=1,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 1,514,487 | 2,991,118 | |
| ThriftyDAggerInitial expert dataset size (M)=2,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 1,100,403 | 2,873,183 | |
| ThriftyDAggerInitial expert dataset size (M)=10,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 1,566,134 | 2,740,182 | |
| CRSAILInitial expert dataset size (M)=1,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 1,821,103 | 2,928,298 | |
| CRSAILInitial expert dataset size (M)=5,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 2,355,279 | 4,038,978 | |
| CRSAILInitial expert dataset size (M)=10,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 2,694,408 | 57,601,286 | |
| CRSAILInitial expert dataset size (M)=20,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 100 | — | 2,519,628 | 3,388,702 | |
| ThriftyDAggerInitial expert dataset size (M)=5,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 80 | — | 1,537,247 | 2,827,164 | |
| CRSAILInitial expert dataset size (M)=2,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 80 | — | 2,299,300 | 2,578,296 | |
| ThriftyDAggerInitial expert dataset size (M)=20,000, Training steps (T_train)=10,000, K=5, alpha_Hopper=0.952025.11 | 60 | — | 1,806,487 | 265,043 | |
| BC#Preferences=502026.06 | — | 456 | — | — | |
| BC#Preferences=5002026.06 | — | 401 | — | — | |
| CPL#Preferences=502026.06 | — | 391 | — | — | |
| CPL#Preferences=5002026.06 | — | 398 | — | — | |
| CPL+KL#Preferences=502026.06 | — | 451 | — | — | |
| CPL+KL#Preferences=5002026.06 | — | 405 | — | — | |
| IPL#Preferences=502026.06 | — | 436 | — | — | |
| IPL#Preferences=5002026.06 | — | 385 | — | — | |
| P-IQL#Preferences=502026.06 | — | 397 | — | — | |
| P-IQL#Preferences=5002026.06 | — | 571 | — | — | |
| PAWS (MLP)#Preferences=502026.06 | — | 512 | — | — | |
| PAWS (MLP)#Preferences=5002026.06 | — | 637 | — | — | |
| PAWS (Trans.)#Preferences=502026.06 | — | 484 | — | — | |
| PAWS (Trans.)#Preferences=5002026.06 | — | 563 | — | — | |
| Pref Trans.#Preferences=502026.06 | — | 500 | — | — | |
| Pref Trans.#Preferences=5002026.06 | — | 552 | — | — |