ResearchBenchmarksReinforcement Learning on MLGym BreakoutFollow100RewardHuman Best46.772860.591474.4188.2286Jun 20, 2026Evaluation ResultsMethodMethodLinksRewardHuman BestMethod Type=Human Base...Method Type=Human Baseline2026.06100MLEvolveMethod Type=Prior WorksMethod Type=Prior Works2026.0683.28ARTSBase Model=o3Base Model=o32026.0678ARTS*Base Model=Qwen 4B, Te...Base Model=Qwen 4B, Test-time trained=true2026.0675.95LinearMethod Type=Prior WorksMethod Type=Prior Works2026.0664.03AIRAMethod Type=Prior WorksMethod Type=Prior Works2026.0657.94ARTSBase Model=Qwen 4BBase Model=Qwen 4B2026.0648.82