Python Coding on HumanEval-X (test)
43.9AccuracyInternLM2-Chat-7B-SFT
Evaluation Results
| Method | Links | |
|---|---|---|
| InternLM2-Chat-7B-SFTShots=5-shot, Model Size Group=< 7B Models2024.03 | 43.9 | |
| InternLM2-Chat-20B-SFTShots=5-shot, Model Size Group=13~20B Models2024.03 | 42 | |
| InternLM2-Chat-7BShots=5-shot, Model Size Group=< 7B Models2024.03 | 41.7 | |
| InternLM2-Chat-20BShots=5-shot, Model Size Group=13~20B Models2024.03 | 39.8 | |
| Mixtral-8x7B-Instruct-v0.1Shots=5-shot, Model Size Group=13~20B Models2024.03 | 38.3 | |
| Qwen-14B-ChatShots=5-shot, Model Size Group=13~20B Models2024.03 | 29.9 | |
| Mistral-7B-Instruct-v0.2Shots=5-shot, Model Size Group=< 7B Models2024.03 | 27.1 | |
| Qwen-7B-ChatShots=5-shot, Model Size Group=< 7B Models2024.03 | 24.4 | |
| Baichuan2-13B-ChatShots=5-shot, Model Size Group=13~20B Models2024.03 | 18.3 | |
| ChatGLM3-6BShots=5-shot, Model Size Group=< 7B Models2024.03 | 17.6 | |
| Baichuan2-7B-ChatShots=5-shot, Model Size Group=< 7B Models2024.03 | 15.4 | |
| Llama2-13B-ChatShots=5-shot, Model Size Group=13~20B Models2024.03 | 12.9 | |
| Llama2-7B-ChatShots=5-shot, Model Size Group=< 7B Models2024.03 | 10.6 |