Loading the SOTA2 catalog…
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning · SOTA2 Research