ResearchTasksIn-Context Reinforcement LearningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedMiniworld environment ε = 0.4 (test)PPO123.5Average Episode Reward12Jun 9, 2026MW 70-1 DR9IC-IQL33NAUC4May 27, 2026MW 40-1 DR9IC-CQL35NAUC4May 27, 2026MW 20-1 DR9AD32NAUC4May 27, 2026