Loading the SOTA2 catalog…
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning · SOTA2 Research