Comprehensive Examination on C-Eval (test)
71.5AccuracyQwen-14B-Chat
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-14B-ChatParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 71.5 | |
| InternLM2-Chat-20B-SFTParameter size group=13~20B Models, Evaluation protocol=5-shot, Alignment=SFT2024.03 | 63.7 | |
| InternLM2-Chat-20BParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 63 | |
| InternLM2-Chat-7B-SFTParameter size group=< 7B Models, Evaluation protocol=5-shot, Alignment=SFT2024.03 | 60.9 | |
| InternLM2-Chat-7BParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 60.8 | |
| Qwen-7B-ChatParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 59.8 | |
| ChatGLM3-6BParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 59.1 | |
| Baichuan2-13B-ChatParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 56.3 | |
| Mixtral-8x7B-Instruct-v0.1Parameter size group=13~20B Models, Evaluation protocol=5-shot, Alignment=Instruct2024.03 | 54 | |
| Baichuan2-7B-ChatParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 53.9 | |
| GPT-3.5Model type=API Models, Evaluation protocol=5-shot2024.03 | 52.5 | |
| Mistral-7B-Instruct-v0.2Parameter size group=< 7B Models, Evaluation protocol=5-shot, Alignment=Instruct2024.03 | 42.4 | |
| Llama2-13B-ChatParameter size group=13~20B Models, Evaluation protocol=5-shot2024.03 | 35 | |
| Llama2-7B-ChatParameter size group=< 7B Models, Evaluation protocol=5-shot2024.03 | 34.9 |