Non-Stationary Reinforcement Learning on Toy Environments Non-Stationary
1.13nAUC (Steady)SAC+AES
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| SAC+AESBase RL Algorithm=SAC, Entropy Scheduling (AES)=true2026.01 | 1.13 | 0.88 | 0.94 | 0.94 | 0.97 | 0 | 0.22 | 0.17 | 0.17 | 0.14 | |
| SACBase RL Algorithm=SAC, Entropy Scheduling (AES)=false2026.01 | 1 | 0.72 | 0.8 | 0.81 | 0.73 | 0 | 0.28 | 0.2 | 0.19 | 0.27 | |
| SQL+AESBase RL Algorithm=SQL, Entropy Scheduling (AES)=true2026.01 | 0.98 | 0.9 | 0.94 | 0.93 | 0.91 | 0 | 0.08 | 0.04 | 0.05 | 0.07 | |
| MEow+AESBase RL Algorithm=MEow, Entropy Scheduling (AES)=true2026.01 | 0.97 | 1.03 | 1.02 | 1.01 | 0.83 | 0 | -0.06 | -0.05 | -0.04 | 0.14 | |
| SQLBase RL Algorithm=SQL, Entropy Scheduling (AES)=false2026.01 | 0.9 | 0.75 | 0.82 | 0.77 | 0.68 | 0 | 0.17 | 0.09 | 0.14 | 0.24 | |
| MEowBase RL Algorithm=MEow, Entropy Scheduling (AES)=false2026.01 | 0.9 | 0.79 | 0.87 | 0.86 | 0.77 | 0 | 0.12 | 0.03 | 0.04 | 0.14 | |
| PPOBase RL Algorithm=PPO, Entropy Scheduling (AES)=false2026.01 | 0.89 | 0.75 | 0.79 | 0.71 | 0.67 | 0 | 0.16 | 0.11 | 0.2 | 0.25 | |
| PPO+AESBase RL Algorithm=PPO, Entropy Scheduling (AES)=true2026.01 | 0.89 | 0.92 | 0.92 | 0.89 | 0.79 | 0 | -0.03 | -0.03 | 0 | 0.11 |