Loading the SOTA2 catalog…
VLRM: Vision-Language Models act as Reward Models for Image Captioning · SOTA2 Research