Loading the SOTA2 catalog…
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens · SOTA2 Research