Massive Multitask Language Understanding on MMLU (Sub-category Performance)
82.7STEM AccuracyLeanQuant
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LeanQuantModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 82.7 | 83.2 | 90.6 | 87.7 | 86.1 | |
| OSAQ+GPTQModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 82.6 | 83.2 | 90.8 | 87.7 | 86.1 | |
| GPTQModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 82.3 | 82.6 | 90.5 | 87.5 | 85.7 | |
| OSAQ+GPTQModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 76.7 | 77.4 | 89.3 | 85.7 | 82.3 | |
| LeanQuantModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 76.6 | 77.3 | 89.2 | 85.9 | 82.3 | |
| GPTQModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 76.3 | 77.2 | 89.3 | 85.2 | 82 |