Loading the SOTA2 catalog…
Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding · SOTA2 Research