Loading the SOTA2 catalog…
AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering · SOTA2 Research