ResearchBenchmarksOffline Reinforcement Learning on D4RL Maze2D Umaze v1Follow67.8Avg Normalized ReturnGAC-0.039217.572935.18552.7971Dec 25, 2025Evaluation ResultsMethodMethodLinksAvg Normalized ReturnGACStrategy=Exploitation...Strategy=Exploitation query (GAC-E[y])2025.1267.8LPT2025.1265.4GACStrategy=Exploration q...Strategy=Exploration query (GAC-p(y+))2025.1264.2GACStrategy=Fixed target...Strategy=Fixed target steering (GAC-y*)2025.1259.2GACStrategy=Sampling from...Strategy=Sampling from prior (p(y|z)p(z))2025.1235.7DT2025.1228.4IQL2025.124.5CQL2025.123.9QDT2025.122.57