Continuous Control on MuJoCo Walker2d v4
13,060Normalized PerformanceOpti-DICE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Opti-DICE2026.02 | 13,060 | — | |
| Flex-f-DICE2026.02 | 13,030 | — | |
| Flex-f-Q2026.02 | 8,830 | — | |
| IQL2026.02 | 8,560 | — | |
| TD3BC2026.02 | 1,720 | — | |
| CQL2026.02 | 430 | — | |
| VDPODelays=5, Global steps=1M2024.05 | 127 | — | |
| BPQLDelays=5, Global steps=1M2024.05 | 120 | — | |
| AD-SACDelays=5, Global steps=1M2024.05 | 112 | — | |
| TABNAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 90.74 | — | |
| ANNAlgorithm=TD32025.09 | 86.8 | — | |
| CaRe-BNAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 85.92 | — | |
| DC/ACDelays=5, Global steps=1M2024.05 | 85 | — | |
| ANN-SNNAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 84.7 | — | |
| TEBNAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 84.7 | — | |
| ILC-SANAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 84 | — | |
| MDC-SANAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 79.28 | — | |
| A-SACDelays=5, Global steps=1M2024.05 | 76 | — | |
| BNTTAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 73.78 | — | |
| AD-SACDelays=25, Global steps=1M2024.05 | 72 | — | |
| tdBNAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 69.28 | — | |
| pop-SANAlgorithm=TD3, Simulation time steps=5, Neuron type=CLIF spiking neurons2025.09 | 66.14 | — | |
| DIDADelays=5, Global steps=1M2024.05 | 61 | — | |
| BPQLDelays=25, Global steps=1M2024.05 | 59 | — | |
| FQL2025.09 | 56.5986 | — | |
| TD32025.09 | 51.0441 | — | |
| SAC2025.09 | 48.155 | — | |
| VDPODelays=25, Global steps=1M2024.05 | 27 | — | |
| DC/ACDelays=25, Global steps=1M2024.05 | 26 | — | |
| BPQLDelays=50, Global steps=1M2024.05 | 23 | — | |
| AD-SACDelays=50, Global steps=1M2024.05 | 23 | — | |
| MEOW2025.09 | 19.1923 | — | |
| DDPG2025.09 | 17.4466 | — | |
| A-SACDelays=25, Global steps=1M2024.05 | 12 | — | |
| A-SACDelays=50, Global steps=1M2024.05 | 11 | — | |
| DC/ACDelays=50, Global steps=1M2024.05 | 11 | — | |
| VDPODelays=50, Global steps=1M2024.05 | 11 | — | |
| DIDADelays=25, Global steps=1M2024.05 | 10 | — | |
| DIDADelays=50, Global steps=1M2024.05 | 8 | — | |
| DUPOMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 2.59 | — | |
| DUPOMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1.62 | — | |
| DUPOMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1.42 | — | |
| State AugmentationMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1.19 | — | |
| State AugmentationMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1.16 | — | |
| DC/ACMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 1 | — | |
| DC/ACMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 1 | — | |
| DC/ACMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 1 | — | |
| State AugmentationMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 0.82 | — | |
| State PredictionMaximum observation delay (ΔTmax)=25, Training steps=1M2026.07 | 0.66 | — | |
| State PredictionMaximum observation delay (ΔTmax)=10, Training steps=1M2026.07 | 0.45 | — | |
| State PredictionMaximum observation delay (ΔTmax)=5, Training steps=1M2026.07 | 0.28 | — | |
| BROEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 3,432 | |
| CrossQEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 6,257 | |
| DIMENFE=162026.05 | — | 6,447.8 | |
| DIPONFE=202026.05 | — | 2,181.7 | |
| DPMDNFE=202026.05 | — | 5,023.2 | |
| DR.QEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 6,422 | |
| DreamerV3Environment Steps=1M, Number of Random Seeds=102026.05 | — | 4,519 | |
| DroQEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 4,781 | |
| FlowRLNFE=12026.05 | — | 4,564.7 | |
| FoGEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 5,124 | |
| FPMD-MNFE=12026.05 | — | 4,723.6 | |
| FPMD-RNFE=12026.05 | — | 4,134.5 | |
| MaxEntDPNFE=202026.05 | — | 5,142.2 | |
| MR.QEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 6,039 | |
| PPOEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 2,487 | |
| PPONFE=12026.05 | — | 3,751.5 | |
| QSMNFE=202026.05 | — | 3,613.4 | |
| QVPONFE=202026.05 | — | 3,826.3 | |
| REDQEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 5,228 | |
| SACNFE=12026.05 | — | 4,988.1 | |
| SAC Flow-GNFE=42026.05 | — | 5,094.1 | |
| SAC Flow-TNFE=42026.05 | — | 5,676.5 | |
| SimBaEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 4,290 | |
| SimbaV2Environment Steps=1M, Number of Random Seeds=102026.05 | — | 6,938 | |
| SMFPNFE=12026.05 | — | 6,776.5 | |
| SPONFE=12026.05 | — | 3,321.8 | |
| TD3NFE=12026.05 | — | 3,513.9 | |
| TD3+OFEEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 5,195 | |
| TD7Environment Steps=1M, Number of Random Seeds=102026.05 | — | 6,096 | |
| TDMPC2Environment Steps=1M, Number of Random Seeds=102026.05 | — | 3,008 | |
| TQCEnvironment Steps=1M, Number of Random Seeds=102026.05 | — | 5,321 |