Loading the SOTA2 catalog…
4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration · SOTA2 Research