Loading the SOTA2 catalog…
H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning · SOTA2 Research