Loading the SOTA2 catalog…
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding · SOTA2 Research