Loading the SOTA2 catalog…
Self-Chained Image-Language Model for Video Localization and Question Answering · SOTA2 Research