Loading the SOTA2 catalog…
Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models · SOTA2 Research