Loading the SOTA2 catalog…
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization? · SOTA2 Research