Loading the SOTA2 catalog…
Quantifying construct validity in large language model evaluations · SOTA2 Research