Reinforcement Learning on Walker2d v4
39,641,353Avg ReturnMDC-SAN
Evaluation Results
| Method | Links | |
|---|---|---|
| MDC-SAN2026.02 | 39,641,353 | |
| pop-SAN2026.02 | 33,071,514 | |
| Vanilla LIF2026.02 | 18,621,450 | |
| DSN2026.02 | 4,436,196 | |
| ANNBase Algorithm=TD32026.02 | 4,340,383 | |
| PT-LIF2026.02 | 4,314,423 | |
| ANN-SNN2026.02 | 4,235,354 | |
| ILC-SAN2026.02 | 4,200,717 | |
| C-DSACNumber of runs=100, Selection=Best models in training runs2026.04 | 4,808 | |
| S-PLIFFrame=PT2026.01 | 4,497 | |
| PLIFFrame=PT2026.01 | 4,445 | |
| PDAEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 4,367.1 | |
| S-PLIFFrame=Vanilla2026.01 | 4,271 | |
| TRPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 4,128.8 | |
| ReLUFrame=ANN2026.01 | 4,050 | |
| PLIFFrame=Vanilla2026.01 | 4,003 | |
| Fixed βNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 3,995 | |
| PPO-ClipNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 3,717 | |
| per-sample PPO-KLNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 3,717 | |
| Adaptive βNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 3,639 | |
| PPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 3,277 | |
| NPGEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 2,923.2 |