Loading the SOTA2 catalog…
Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding · SOTA2 Research