Loading the SOTA2 catalog…
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning · SOTA2 Research