ResearchBenchmarksReinforcement Learning on 25 environment-dataset combinationsFollow77.94Normalized ScoreSOPE52.56459.15265.7472.328May 7, 2026Evaluation ResultsMethodMethodLinksNormalized ScoreSOPE2026.0577.94SPEQ O2O2026.0572.29Cal-QL2026.0568.49SACfD2026.0567.62RLPD2026.0553.54