Loading the SOTA2 catalog…
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning · SOTA2 Research