Python Coding on HumanEval (test)
67.7AccuracyInternLM2-Chat-20B
Evaluation Results
| Method | Links | |
|---|---|---|
| InternLM2-Chat-20BShots=4-shot, Model Size Group=13~20B Models2024.03 | 67.7 | |
| InternLM2-Chat-20B-SFTShots=4-shot, Model Size Group=13~20B Models2024.03 | 67.1 | |
| InternLM2-Chat-7B-SFTShots=4-shot, Model Size Group=< 7B Models2024.03 | 61.6 | |
| InternLM2-Chat-7BShots=4-shot, Model Size Group=< 7B Models2024.03 | 59.2 | |
| ChatGLM3-6BShots=4-shot, Model Size Group=< 7B Models2024.03 | 53.1 | |
| Qwen-14B-ChatShots=4-shot, Model Size Group=13~20B Models2024.03 | 41.5 | |
| Qwen-7B-ChatShots=4-shot, Model Size Group=< 7B Models2024.03 | 36 | |
| Mistral-7B-Instruct-v0.2Shots=4-shot, Model Size Group=< 7B Models2024.03 | 35.4 | |
| Mixtral-8x7B-Instruct-v0.1Shots=4-shot, Model Size Group=13~20B Models2024.03 | 32.3 | |
| Baichuan2-13B-ChatShots=4-shot, Model Size Group=13~20B Models2024.03 | 19.5 | |
| Baichuan2-7B-ChatShots=4-shot, Model Size Group=< 7B Models2024.03 | 17.7 | |
| Llama2-7B-ChatShots=4-shot, Model Size Group=< 7B Models2024.03 | 15.2 | |
| Llama2-13B-ChatShots=4-shot, Model Size Group=13~20B Models2024.03 | 8.5 |