Loading the SOTA2 catalog…
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models · SOTA2 Research