Non-Stationary Reinforcement Learning on MuJoCo Non-Stationary
1.24nAUC (Steady)SAC+AES
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| SAC+AESBase RL Algorithm=SAC, Entropy Scheduling (AES)=true2026.01 | 1.24 | 0.87 | 0.94 | 0.94 | 0.94 | 0 | 0.3 | 0.24 | 0.24 | 0.24 | |
| SACBase RL Algorithm=SAC, Entropy Scheduling (AES)=false2026.01 | 1 | 0.67 | 0.76 | 0.68 | 0.65 | 0 | 0.33 | 0.24 | 0.32 | 0.35 | |
| MEow+AESBase RL Algorithm=MEow, Entropy Scheduling (AES)=true2026.01 | 0.98 | 0.95 | 0.88 | 0.87 | 0.89 | 0 | 0.03 | 0.1 | 0.11 | 0.09 | |
| PPO+AESBase RL Algorithm=PPO, Entropy Scheduling (AES)=true2026.01 | 0.94 | 0.69 | 0.79 | 0.82 | 0.77 | 0 | 0.27 | 0.16 | 0.13 | 0.18 | |
| MEowBase RL Algorithm=MEow, Entropy Scheduling (AES)=false2026.01 | 0.92 | 0.67 | 0.71 | 0.79 | 0.63 | 0 | 0.27 | 0.23 | 0.14 | 0.32 | |
| PPOBase RL Algorithm=PPO, Entropy Scheduling (AES)=false2026.01 | 0.88 | 0.66 | 0.64 | 0.67 | 0.57 | 0 | 0.25 | 0.27 | 0.24 | 0.35 | |
| SQL+AESBase RL Algorithm=SQL, Entropy Scheduling (AES)=true2026.01 | 0.81 | 0.65 | 0.79 | 0.71 | 0.64 | 0 | 0.2 | 0.02 | 0.12 | 0.21 | |
| SQLBase RL Algorithm=SQL, Entropy Scheduling (AES)=false2026.01 | 0.8 | 0.52 | 0.53 | 0.5 | 0.44 | 0 | 0.35 | 0.34 | 0.38 | 0.45 |