Multitask Language Understanding on CMMLU (test)
78.3AccuracyGPT 4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT 4oModel Type=Proprietary2024.10 | 78.3 | |
| Baichuan-omniParameters=7B, Modality=Omni-modal2024.10 | 72.2 | |
| Qwen-14BShots=5-shot, Model Size Group=13 ~ 20B Models2024.03 | 70.1 | |
| InternLM2-20BShots=5-shot, Model Size Group=13 ~ 20B Models2024.03 | 68.7 | |
| Qwen1.5-ChatParameters=7B, Modality=Pure text2024.10 | 68 | |
| ChatGLM3-6B-BaseShots=5-shot, Model Size Group=<= 7B Models2024.03 | 66.5 | |
| InternLM2-7BShots=5-shot, Model Size Group=<= 7B Models2024.03 | 66.3 | |
| InternLM2-7B-BaseShots=5-shot, Model Size Group=<= 7B Models2024.03 | 63 | |
| InternLM2-20B-BaseShots=5-shot, Model Size Group=13 ~ 20B Models2024.03 | 63 | |
| Qwen-7BShots=5-shot, Model Size Group=<= 7B Models2024.03 | 62.5 | |
| Baichuan2-13B-BaseShots=5-shot, Model Size Group=13 ~ 20B Models2024.03 | 61.3 | |
| Baichuan2-7B-BaseShots=5-shot, Model Size Group=<= 7B Models2024.03 | 57 | |
| MAP-NeoParameters=7B, Modality=Pure text2024.10 | 55.1 | |
| Mixtral-8x7B-v0.1Shots=5-shot, Model Size Group=13 ~ 20B Models2024.03 | 53.3 | |
| CLOBase Model=Qwen2.5-3B, Training Data=English + Chinese2025.05 | 52.1 | |
| Llama3-InstructParameters=8B, Modality=Pure text2024.10 | 51.7 | |
| VITAParameters=8x7B, Modality=Omni-modal2024.10 | 46.6 | |
| SFTBase Model=Qwen2.5-3B, Training Data=English + Chinese2025.05 | 46.34 | |
| Mistral-7B-v0.1Shots=5-shot, Model Size Group=<= 7B Models2024.03 | 44.6 | |
| CLOBase Model=Llama-3-8B, Training Data=English + Chinese2025.05 | 41.99 | |
| SFT+DPOBase Model=Llama-3-8B, Training Data=English + Chinese2025.05 | 40.91 | |
| SFTBase Model=Llama-3-8B, Training Data=English + Chinese2025.05 | 39.36 | |
| Llama2-13BShots=5-shot, Model Size Group=13 ~ 20B Models2024.03 | 38.8 | |
| SFT-tgtBase Model=Llama-3-8B, Training Data=Target language only2025.05 | 38.55 | |
| SFT+DPOBase Model=Mistral-7B, Training Data=English + Chinese2025.05 | 37.02 | |
| CLOBase Model=Llama-2-13B, Training Data=English + Chinese2025.05 | 34.17 | |
| CLOBase Model=Mistral-7B, Training Data=English + Chinese2025.05 | 34.05 | |
| SFTBase Model=Mistral-7B, Training Data=English + Chinese2025.05 | 33.74 | |
| SFT+DPOBase Model=Llama-2-13B, Training Data=English + Chinese2025.05 | 33.26 | |
| SFT-tgtBase Model=Mistral-7B, Training Data=Target language only2025.05 | 33.16 | |
| Llama2-7BShots=5-shot, Model Size Group=<= 7B Models2024.03 | 31.9 | |
| SFT-tgtBase Model=Llama-2-13B, Training Data=Target language only2025.05 | 31.51 | |
| SFTBase Model=Llama-2-13B, Training Data=English + Chinese2025.05 | 31.17 | |
| CLOBase Model=Llama-2-7B, Training Data=English + Chinese2025.05 | 28.11 | |
| SFT-tgtBase Model=Llama-2-7B, Training Data=Target language only2025.05 | 27.19 | |
| SFTBase Model=Llama-2-7B, Training Data=English + Chinese2025.05 | 27.12 | |
| SFT+DPOBase Model=Llama-2-7B, Training Data=English + Chinese2025.05 | 26.59 | |
| OLMoParameters=7B, Modality=Pure text2024.10 | 25.6 |