Loading the SOTA2 catalog…
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation · SOTA2 Research