ResearchTasksAdversarial Reinforcement LearningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedConnect Four 100% optimal adversary (test-time)ESPER-0.98Avg Return3Feb 26, 2026Connect Four 70% optimal adversary (test)ARDT0.02Average Return3Feb 26, 2026Connect Four 50% optimal adversary (test-time)ARDT0.11Average Return3Feb 26, 2026Connect Four 30% optimal adversary (test-time)ARDT0.55Average Return3Feb 26, 2026