Loading the SOTA2 catalog…
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning · SOTA2 Research