Multitask Language Understanding on ArabicMMLU
72.5AccuracyGPT-4
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4Setting=Few-shot2024.12 | 72.5 | |
| LLaMA3-Tamed-70BSetting=Few-shot2024.12 | 66.56 | |
| Llama3-70BSetting=Few-shot2024.12 | 65.51 | |
| Qwen1.5-72BSetting=Few-shot2024.12 | 61.23 | |
| ChatGPT 3.5 TurboSetting=Few-shot2024.12 | 57.7 | |
| Qwen1.5-32BSetting=Few-shot2024.12 | 55.94 | |
| LLaMA3-Tamed-8BSetting=Few-shot2024.12 | 50.17 | |
| Qwen2.5scenario=zero-shot2025.12 | 47.2 | |
| Gamayunscenario=zero-shot2025.12 | 47 | |
| Qwen1.5-7BSetting=Few-shot2024.12 | 46.41 | |
| Qwen3scenario=zero-shot, thinking_mode=no-thinking2025.12 | 46.3 | |
| Llama3-8BSetting=Few-shot2024.12 | 45.78 | |
| Jais-30B-v3Setting=Few-shot2024.12 | 44.47 | |
| Gemma3scenario=zero-shot2025.12 | 39.8 | |
| Llama3.2scenario=zero-shot2025.12 | 37.2 | |
| EuroLMscenario=zero-shot2025.12 | 26.8 |