General Benchmarks
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
General Benchmarks MMLU, HellaSwag, OBQA, WinoGrande, ARC-C, PiQA, SciQ, LogiQA
35.68MMLU Accuracy
70
General Benchmarks MMLU, CEval, CMMLU, HumanEval, MBPP, MATH, GSM8K, MultiArith, BBH
59.3MMLU
18
General Benchmarks
74Average Score
12
General Benchmarks Llama 3.1 8B
66.5Generation Quality Score
11
General Benchmarks (C-Eval, IFEval, MATH-500, LCB, H-Swag, SST-5, CrossNER) (test)
80.2C-Eval Score
6
General Benchmarks
57.8Top-1 Accuracy
6
General Benchmarks Italian
37.47ARC-C-it
6
General Benchmarks Aggregated (test)
76.54General Capability Score
5
General Benchmarks MMLU-R, MMLU-P, GPQA
92.3MMLU-R (multi@5 Acc)
4
General Benchmarks (MMLU, AlpacaEval, Arena-Hard)
73.41MMLU Accuracy
4
12 general benchmarks Avg
68.24General Average Score
3