Loading the SOTA2 catalog…
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning · SOTA2 Research