Loading the SOTA2 catalog…
Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models · SOTA2 Research