World Knowledge on MMLU
82.4AccuracyLlama-3.3-70b-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-3.3-70b-InstructSize=70B, Few-shot settings=5-shot2026.03 | 82.4 | |
| KarnakSize=40B, Few-shot settings=5-shot2026.03 | 82.37 | |
| Qwen3-32BSize=32B, Few-shot settings=5-shot2026.03 | 82.25 | |
| Fanar-27BSize=27B, Few-shot settings=5-shot2026.03 | 78.89 | |
| AceGPT-v2-70B-ChatSize=70B, Few-shot settings=5-shot2026.03 | 77.98 | |
| Gemma-3-27B-itSize=27B, Few-shot settings=5-shot2026.03 | 77.38 | |
| AceGPT-v2-32B-ChatSize=32B, Few-shot settings=5-shot2026.03 | 75.72 | |
| Jais-2-70B-ChatSize=70B, Few-shot settings=5-shot2026.03 | 73.86 | |
| Fanar-1-9b-instructSize=9B, Few-shot settings=5-shot2026.03 | 71.32 | |
| Allam-7B-Instruct-preview-v2Size=7B, Few-shot settings=5-shot2026.03 | 63.8 | |
| ID-LoRABase Model=Mistral-7B, Trainable Parameters (%)=0.62%2026.02 | 60.6 | |
| HydraLoRABase Model=Mistral-7B, Trainable Parameters (%)=1.22%2026.02 | 60.4 | |
| MoELoRABase Model=Mistral-7B, Trainable Parameters (%)=1.21%2026.02 | 60.4 | |
| DoRABase Model=Mistral-7B, Trainable Parameters (%)=1.16%2026.02 | 59.4 | |
| LoRABase Model=Mistral-7B, Trainable Parameters (%)=1.15%2026.02 | 59.3 | |
| ID-LoRABase Model=LLaMA-3-8B, Trainable Parameters (%)=0.56%2026.02 | 51 | |
| LoRABase Model=LLaMA-3-8B, Trainable Parameters (%)=1.03%2026.02 | 48.1 | |
| HydraLoRABase Model=LLaMA-3-8B, Trainable Parameters (%)=1.10%2026.02 | 47.5 | |
| FFTBase Model=Mistral-7B, Trainable Parameters (%)=100%2026.02 | 46.7 | |
| LLaMA2-7B# tokens for training=2T, Shot number=5, Training data source=Different from RedPajama2023.10 | 46.6 | |
| NITPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 46.14 | |
| NTPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 43.71 | |
| DoRABase Model=LLaMA-3-8B, Trainable Parameters (%)=1.05%2026.02 | 41.5 | |
| MoELoRABase Model=LLaMA-3-8B, Trainable Parameters (%)=1.09%2026.02 | 41.5 | |
| FFTBase Model=LLaMA-3-8B, Trainable Parameters (%)=100%2026.02 | 37.5 | |
| NITPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 37.37 | |
| NTPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 34.6 | |
| NITPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 31.68 | |
| NTPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 31.05 | |
| INCITE-Base-3B# tokens for training=800B, Shot number=5, Training data source=RedPajama2023.10 | 27 | |
| Open-LLaMA-3B-v1# tokens for training=1T, Shot number=5, Training data source=RedPajama2023.10 | 27 | |
| Pythia-2.8B# tokens for training=300B, Shot number=5, Training data source=RedPajama2023.10 | 26.9 | |
| Open-LLaMA-3B-v2# tokens for training=1T, Shot number=5, Training data source=RedPajama2023.10 | 26.9 | |
| Sheared-LLaMA-2.7B# tokens for training=50B, Shot number=5, Training data source=RedPajama2023.10 | 26.4 | |
| OPT-2.7B# tokens for training=300B, Shot number=5, Training data source=Different from RedPajama2023.10 | 25.9 | |
| Pythia-1.4B# tokens for training=300B, Shot number=5, Training data source=Different from RedPajama2023.10 | 25.7 | |
| Sheared-LLaMA-1.3B# tokens for training=50B, Shot number=5, Training data source=RedPajama2023.10 | 25.7 | |
| TinyLlama-1.1B# tokens for training=3T, Shot number=5, Training data source=RedPajama2023.10 | 25.5 | |
| OPT-1.3B# tokens for training=300B, Shot number=5, Training data source=Different from RedPajama2023.10 | 24.7 |