Reinforcement Learning on CartPole (Average Reward)
1,000Average RewardSALSA-RL
Evaluation Results
| Method | Links | |
|---|---|---|
| SALSA-RLLatent dimension size (hd)=32025.02 | 1,000 | |
| SALSA-RLLatent dimension size (hd)=42025.02 | 1,000 | |
| SALSA-RLLatent dimension size (hd)=62025.02 | 1,000 | |
| SALSA-RLLatent dimension size (hd)=82025.02 | 1,000 | |
| SALSA-RLLatent dimension size (hd)=162025.02 | 1,000 | |
| A2C2025.02 | 1,000 | |
| DDPG2025.02 | 1,000 | |
| DSP2025.02 | 999.6 | |
| TD32025.02 | 998 | |
| PPO2025.02 | 993.9 | |
| AC-Adambatch size=1000, seeds=52026.01 | 989.5 | |
| AC-CGbatch size=1000, seeds=52026.01 | 977.9 | |
| SMACbatch size=1000, seeds=52026.01 | 973.4 | |
| SAC2025.02 | 971.8 | |
| AC-SGDbatch size=1000, seeds=52026.01 | 962.3 | |
| DTSemNetNf (Number of features)=4, Na (Number of actions)=2, Height=42026.05 | 500 | |
| Deep RLNf (Number of features)=4, Na (Number of actions)=2, Height=42026.05 | 500 | |
| DGTNf (Number of features)=4, Na (Number of actions)=2, Height=42026.05 | 500 | |
| VIPERNf (Number of features)=4, Na (Number of actions)=2, Height=42026.05 | 499.95 | |
| ICCTNf (Number of features)=4, Na (Number of actions)=2, Height=42026.05 | 496 | |
| R2PONumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 474.67 | |
| AC-KFACbatch size=1000, seeds=52026.01 | 331.4 | |
| ProPSNumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 258.09 | |
| ProPS+Number of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 253.06 | |
| Average Human2024.05 | 220 | |
| TRPONumber of runs=10, Aggregation protocol=averaged across all training iterations, Source=Best SB32026.05 | 216.92 | |
| DDQN2024.05 | 210 | |
| DQN2024.05 | 200 | |
| FDQN2024.05 | 198 |