ResearchBenchmarksReinforcement Learning on TomatoFollow6.28True ScoreORPO3.90884.52445.145.7556Apr 13, 2026Evaluation ResultsMethodMethodLinksTrue ScoreProxy ScoreWorst ScoreOccurrenceWorst* ScoreORPOTraining Reward=Proxy...Training Reward=Proxy only2026.046.286.83-1.510.0003-1.51Max-MinTraining Reward=Proxy...Training Reward=Proxy only2026.044.564.68-1.370-1.37ORPO*Training Reward=Proxy...Training Reward=Proxy only2026.0443.98-1.090-1.09