Loading the SOTA2 catalog…
Large Language Model Evaluation on 10 tasks average benchmark leaderboard · SOTA2 Research