Loading the SOTA2 catalog…
Vision-language models lag human performance on physical dynamics and intent reasoning · SOTA2 Research