Loading the SOTA2 catalog…
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning · SOTA2 Research