ResearchBenchmarksContinual Reinforcement Learning on Atari Two-cycle (train)Follow0.722C1 Forward ScoreDV3-0.066320.138340.3430.54766Mar 12, 2026Evaluation ResultsMethodMethodLinksC1 Forward ScoreC2 Forward ScoreMax Forward ScoreC1 Forward TransferC2 Forward TransferRecovery ScoreAccuracyMinimum AccuracyWorst-Case AccuracyDV3Training Mode=Two-cycleTraining Mode=Two-cycle2026.030.7220.3780.735-0.514-0.750.610.9-0.393-0.299TES-SACTraining Mode=Two-cycleTraining Mode=Two-cycle2026.030.1940.1120.089-0.898-0.8820.7674.4-0.203-0.168ARROWTraining Mode=Two-cycleTraining Mode=Two-cycle2026.03-0.0360.030.012-0.5540.3091.41879.60.4420.388