Loading the SOTA2 catalog…
PrLM: Learning Explicit Reasoning for Personalized RAG via Contrastive Reward Optimization · SOTA2 Research