Loading the SOTA2 catalog…
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks · SOTA2 Research