ResearchBenchmarksReinforcement Learning on DoublePendulumFollow2,896.43IQM ReturnA2ER-52.8852712.80241,478.492,244.1776Apr 29, 2025Evaluation ResultsMethodMethodLinksIQM ReturnA2ERblock strategy=trueblock strategy=true2025.042,896.43A2ER-Bblock strategy=falseblock strategy=false2025.041,905.71FIFO2025.04873.66DERalpha=1alpha=12025.0460.55