Loading the SOTA2 catalog…
Toward Trustworthy Difficulty Assessments: Large Language Models as Judges in Programming and Synthetic Tasks · SOTA2 Research