Language Understanding and Generation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
92.7SST-2 Accuracy
6
May 4, 2026
LLaMA3-3B Language Tasks Suite (SST-2, RTE, CB, BoolQ, WSC, WIC, MultiRC, COPA, ReCoRD, SQuAD, DROP)
92.6SST-2 Accuracy
6
May 4, 2026