Loading the SOTA2 catalog…
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues · SOTA2 Research