ResearchBenchmarksLanguage Reasoning on LangR unseen tasks (test)Follow60.8Pass@1SGE41.87246.78651.756.614Mar 2, 2026Evaluation ResultsMethodMethodLinksPass@1SGERL Training=Strategy-G...RL Training=Strategy-Guided Exploration (SGE), Aggregation=Average ± standard error across 3 random seeds2026.0360.8GRPORL Training=GRPO, Aggr...RL Training=GRPO, Aggregation=Average ± standard error across 3 random seeds2026.0346Zero-ShotEvaluation Protocol=Di...Evaluation Protocol=Direct base-LLM evaluation2026.0342.6