Loading the SOTA2 catalog…
Spatio-Temporal Attention Models for Grounded Video Captioning · SOTA2 Research