Code Generation on MBPP v1 (test)
68.9Pass@1ChatGLM3-6B-Base
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ChatGLM3-6B-Baseshot count=5-shot, Parameter size group=< 7B Models2024.03 | 68.9 | — | |
| InternLM2-20Bshot count=5-shot, Parameter size group=13~20B Models2024.03 | 63 | — | |
| Unnatural-CodeLlamaParams=34B2023.06 | 61.2 | — | |
| WizardCoder (34B)Params=34B, temperature=0.2, top_p=0.95, n (number of samples)=2002023.06 | 61.2 | — | |
| InternLM2-20B-Baseshot count=5-shot, Parameter size group=13~20B Models2024.03 | 59.9 | — | |
| Mixtral-8x7B-v0.1shot count=5-shot, Parameter size group=13~20B Models2024.03 | 59.5 | — | |
| Code-Davinci-002Params=Unknown2023.06 | 58.1 | — | |
| CodeLlama-InstructParams=34B2023.06 | 57 | — | |
| CodeLlama-PythonParams=34B2023.06 | 56.2 | — | |
| CodeLlamaParams=34B2023.06 | 55 | — | |
| InternLM2-7B-Baseshot count=5-shot, Parameter size group=< 7B Models2024.03 | 54.1 | — | |
| Qwen-14Bshot count=5-shot, Parameter size group=13~20B Models2024.03 | 52.9 | — | |
| GPT-3.5 (ChatGPT)Params=Unknown2023.06 | 52.2 | — | |
| InternLM2-7Bshot count=5-shot, Parameter size group=< 7B Models2024.03 | 51.8 | — | |
| WizardCoder (15B)Params=15B, temperature=0.2, top_p=0.95, n (number of samples)=2002023.06 | 51.8 | — | |
| PaLM 2-SParams=Unknown2023.06 | 50 | — | |
| Mistral-7B-v0.1shot count=5-shot, Parameter size group=< 7B Models2024.03 | 47.5 | — | |
| PaLM-CoderSize (B)=540, Python Code (GB)=~20, Other Code (GB)=~200, Other (GB)=~4000, Code License=Permissive, Infill?=false2022.04 | 47 | — | |
| PaLM-CoderParams=540B2023.06 | 47 | — | |
| Code-Cushman-001Params=Unknown2023.06 | 45.9 | — | |
| Llama2Params=70B2023.06 | 45 | — | |
| Baichuan2-13B-Baseshot count=5-shot, Parameter size group=13~20B Models2024.03 | 44 | — | |
| StarCoderParams=15B2023.06 | 43.6 | — | |
| Llama2-13Bshot count=5-shot, Parameter size group=13~20B Models2024.03 | 38.9 | — | |
| LlamaParams=65B2023.06 | 37.7 | — | |
| PaLMParams=540B2023.06 | 36.8 | — | |
| CodeGen-MonoParams=16B2023.06 | 35.3 | — | |
| Baichuan2-7B-Baseshot count=5-shot, Parameter size group=< 7B Models2024.03 | 35 | — | |
| Qwen-7B-Chatshot count=5-shot, Parameter size group=< 7B Models2024.03 | 33.9 | — | |
| Llama2-7Bshot count=5-shot, Parameter size group=< 7B Models2024.03 | 28.8 | — | |
| CodeGeeXParams=13B2023.06 | 24.4 | — | |
| INCODER-6.7BSize (B)=6.7, Python Code (GB)=52, Other Code (GB)=107, Other (GB)=57, Code License=Permissive, Infill?=true2022.04 | 19.4 | — | |
| LaMDASize (B)=137, Infill?=false2022.04 | 14.8 | — | |
| Llama-3-8BInstruction-tuned=true2024.07 | — | 67.9 | |
| No-SFTSFT Training Domain=None (Base Model)2026.05 | — | 58 | |
| Qwen1.5-7BInstruction-tuned=true2024.07 | — | 48.9 | |
| Qwen2-7BInstruction-tuned=true2024.07 | — | 67.2 | |
| SFT-GTSFT Training Domain=Code SFT (Ling-Coder-SFT)2026.05 | — | 58 | |
| SFT-GTSFT Training Domain=Math SFT (MixChain-Z-PRM12K)2026.05 | — | 58 | |
| SFT-SDSFT Training Domain=Code SFT (Ling-Coder-SFT)2026.05 | — | 59.2 | |
| SFT-SDSFT Training Domain=Math SFT (MixChain-Z-PRM12K)2026.05 | — | 58.6 | |
| TABOMSFT Training Domain=Code SFT (Ling-Coder-SFT)2026.05 | — | 60.6 | |
| TABOMSFT Training Domain=Math SFT (MixChain-Z-PRM12K)2026.05 | — | 59.2 |