Loading the SOTA2 catalog…
Generative RLHF-V: Learning Principles from Multi-modal Human Preference · SOTA2 Research