Loading the SOTA2 catalog…
Distilled Dual-Encoder Model for Vision-Language Understanding · SOTA2 Research