Loading the SOTA2 catalog…
A3VLM: Actionable Articulation-Aware Vision Language Model · SOTA2 Research