Comprehensive Examination on CMMLU (test)
68.1AccuracyQwen-14B-Chat
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-14B-ChatParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 68.1 | |
| InternLM2-Chat-20B-SFTParameter size group=13~20B Models, Evaluation protocol=5-shot, Alignment=SFT2024.03 | 65.3 | |
| InternLM2-Chat-20BParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 65.1 | |
| InternLM2-Chat-7B-SFTParameter size group=< 7B Models, Evaluation protocol=5-shot, Alignment=SFT2024.03 | 63.2 | |
| InternLM2-Chat-7BParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 63 | |
| Qwen-7B-ChatParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 57.9 | |
| ChatGLM3-6BParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 57.8 | |
| Baichuan2-13B-ChatParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 54.8 | |
| GPT-3.5Model type=API Models, Evaluation protocol=5-shot2024.03 | 53.9 | |
| Baichuan2-7B-ChatParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 53.4 | |
| Mixtral-8x7B-Instruct-v0.1Parameter size group=13~20B Models, Evaluation protocol=5-shot, Alignment=Instruct2024.03 | 50.6 | |
| Mistral-7B-Instruct-v0.2Parameter size group=< 7B Models, Evaluation protocol=5-shot, Alignment=Instruct2024.03 | 42 | |
| Llama2-13B-ChatParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 33.8 | |
| Llama2-7B-ChatParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 30.7 |