Loading the SOTA2 catalog…
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1 · SOTA2 Research