ResearchBenchmarksReinforcement Learning on Office World Map 2Follow16,240Training StepsQR-MAXRM-22,212.48237,341.76496,896756,450.24Dec 16, 2025Evaluation ResultsMethodMethodLinksTraining StepsQR-MAXRM2025.1216,240R-MAXRM2025.1260,058QR-MAX2025.1269,494QRM2025.12282,943R-MAX2025.12307,805Q-Learning2025.12977,552