Loading the SOTA2 catalog…
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization · SOTA2 Research