Loading the SOTA2 catalog…
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned · SOTA2 Research