Loading the SOTA2 catalog…
Benchmarking Knowledge Boundary for Large Language Models: A Different Perspective on Model Evaluation · SOTA2 Research