Loading the SOTA2 catalog…
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models · SOTA2 Research