Code Generation on MBPP (Pass@1)
83.1Pass@1MapCoder
Evaluation Results
| Method | Links | |
|---|---|---|
| MapCoderLLM=GPT-42024.05 | 83.1 | |
| CoTLLM=GPT-42024.05 | 82.4 | |
| DirectLLM=GPT-42024.05 | 81.1 | |
| LATS (ReAct)Model=GPT-3.5, Solutions sampled during expansion=5, Iterations=82023.10 | 81.1 | |
| MapCoderLLM=ChatGPT2024.05 | 78.3 | |
| ReflexionLLM=GPT-42024.05 | 78.3 | |
| DeepSeek-R12025.12 | 77.43 | |
| Self-PlanningLLM=GPT-42024.05 | 75.8 | |
| CompassMax-V3-Thinking2025.12 | 73.54 | |
| DeepSeekCoder-33Bdecoding=greedy2024.02 | 73.2 | |
| ReflexionLLM=ChatGPT2024.05 | 73 | |
| RAPModel=GPT-3.52023.10 | 71.4 | |
| AnalogicalLLM=ChatGPT2024.05 | 70.5 | |
| DeepSeekCoder-6.7Bdecoding=greedy2024.02 | 70.2 | |
| ReflexionModel=GPT-3.52023.10 | 70 | |
| Self-collaborationLLM=ChatGPT2024.05 | 68.2 | |
| CompassMax-V32025.12 | 67.7 | |
| ReActModel=GPT-3.52023.10 | 67 | |
| StarCoder2-15Bdecoding=greedy2024.02 | 66.2 | |
| ToTModel=GPT-3.52023.10 | 65.8 | |
| CodeLlama-34Bdecoding=greedy2024.02 | 65.4 | |
| CodeLlama-13Bdecoding=greedy2024.02 | 62.4 | |
| Mixtral 8x7BActive Params=13B2024.01 | 60.7 | |
| AnalogicalLLM=GPT-42024.05 | 58.4 | |
| StarCoder2-3Bdecoding=greedy2024.02 | 57.4 | |
| Self-PlanningLLM=ChatGPT2024.05 | 55.7 | |
| DeepSeekCoder-1.3Bdecoding=greedy2024.02 | 55.4 | |
| Phi-2# Non-Emb Params=2.5B2024.07 | 55 | |
| CoTModel=GPT-3.52023.10 | 54.9 | |
| CoTLLM=ChatGPT2024.05 | 54.5 | |
| StarCoder2-7Bdecoding=greedy2024.02 | 54.4 | |
| BreadcrumbsLM=true, Math=false, Code=true2025.02 | 53.4 | |
| StableCode-3Bdecoding=greedy2024.02 | 53.1 | |
| CodeLlama-7Bdecoding=greedy2024.02 | 52.1 | |
| StarCoderBase-15Bdecoding=greedy2024.02 | 50.6 | |
| Ties-MergingLM=true, Math=true, Code=true2025.02 | 50.2 | |
| Mistral 7BActive Params=7B2024.01 | 50.2 | |
| LLaMA 2 70BActive Params=70B2024.01 | 49.8 | |
| DirectLLM=ChatGPT2024.05 | 49.8 | |
| BreadcrumbsLM=true, Math=true, Code=true2025.02 | 49.4 | |
| Model StockLM=true, Math=true, Code=true2025.02 | 47.8 | |
| StarCoderBase-7Bdecoding=greedy2024.02 | 47.4 | |
| LED-MergingLM=true, Math=false, Code=true2025.02 | 47.2 | |
| Model StockLM=true, Math=false, Code=true2025.02 | 47 | |
| Llama-3-8Bevaluation_mode=Few-shot2024.07 | 47 | |
| DeepSeekMoE Chat 16B# Shot=3-shot, Total Params=16.4B, Activated Params=2.8B, FLOPs per 4K Tokens=74.4T2024.01 | 46.2 | |
| Llama-3-SynEevaluation_mode=Few-shot2024.07 | 45.6 | |
| LED-MergingLM=true, Math=true, Code=true2025.02 | 44.6 | |
| StarCoderBase-3Bdecoding=greedy2024.02 | 42.6 | |
| Ties-MergingLM=true, Math=false, Code=true2025.02 | 41.6 | |
| LLaMA 1 33BActive Params=33B2024.01 | 40.9 | |
| DeepSeekMoE 16B# Shot=3-shot, # Total Params=16.4B, # Activated Params=2.8B, FLOPs per 4K Tokens=74.4T, # Training Tokens=2T2024.01 | 39.2 | |
| DeepSeekMoE 16B# Shot=3-shot, # Total Params=16.4B, # Activated Params=2.8B, FLOPs per 4K Tokens=74.4T, # Training Tokens=2T2024.01 | 39.2 | |
| DeepSeek Chat 7B# Shot=3-shot, Total Params=6.9B, Activated Params=6.9B, FLOPs per 4K Tokens=183.5T2024.01 | 39 | |
| DeepSeek 7B (Dense)# Shot=3-shot, # Total Params=6.9B, # Activated Params=6.9B, FLOPs per 4K Tokens=183.5T, # Training Tokens=2T2024.01 | 39 | |
| MAmmoTH2-8Bevaluation_mode=Few-shot2024.07 | 38.8 | |
| Task ArithmeticLM=true, Math=false, Code=true2025.02 | 37.8 | |
| Qwen2-1.5B# Non-Emb Params=1.2B2024.07 | 37.4 | |
| Mistral-7B-v0.3evaluation_mode=Few-shot2024.07 | 36 | |
| LLaMA 2 13BActive Params=13B2024.01 | 35.4 | |
| LED-MergingWizardLM-13B=true, WizardMath-13B=false, LLama-2-13B-Code-Alpaca=true2025.02 | 33.8 | |
| w/o MergingLM=false, Math=false, Code=true2025.02 | 33.6 | |
| DeepSeek 67B (Dense)# Shot=3-shot2024.01 | 33.6 | |
| Task ArithmeticWizardLM-13B=true, WizardMath-13B=false, LLama-2-13B-Code-Alpaca=true2025.02 | 33.2 | |
| DeepSeekMoE 145B# Shot=3-shot2024.01 | 33.2 | |
| Ties-MergingWizardLM-13B=true, WizardMath-13B=false, LLama-2-13B-Code-Alpaca=true2025.02 | 33 | |
| DCLM-7Bevaluation_mode=Few-shot2024.07 | 32.6 | |
| BreadcrumbsWizardLM-13B=true, WizardMath-13B=false, LLama-2-13B-Code-Alpaca=true2025.02 | 32 | |
| DeepSeekMoE 142B (Half Activated)# Shot=3-shot2024.01 | 32 | |
| w/o MergingWizardLM-13B=true, WizardMath-13B=false, LLama-2-13B-Code-Alpaca=false2025.02 | 31.4 | |
| Ties-MergingWizardLM-13B=true, WizardMath-13B=true, LLama-2-13B-Code-Alpaca=true2025.02 | 31.4 | |
| Gemma-2B# Non-Emb Params=2.0B2024.07 | 29.2 | |
| BreadcrumbsWizardLM-13B=true, WizardMath-13B=true, LLama-2-13B-Code-Alpaca=true2025.02 | 28.4 | |
| LLaMA2 SFT 7B# Shot=3-shot, Total Params=6.7B, Activated Params=6.7B, FLOPs per 4K Tokens=187.9T2024.01 | 27.8 | |
| GShard 137B# Shot=3-shot2024.01 | 27.6 | |
| LLaMA 2 7BActive Params=7B2024.01 | 26.1 | |
| Task ArithmeticWizardLM-13B=true, WizardMath-13B=true, LLama-2-13B-Code-Alpaca=true2025.02 | 25.8 | |
| LED-MergingWizardLM-13B=true, WizardMath-13B=true, LLama-2-13B-Code-Alpaca=true2025.02 | 23.8 | |
| w/o MergingWizardLM-13B=false, WizardMath-13B=false, LLama-2-13B-Code-Alpaca=true2025.02 | 22.82 | |
| Qwen2-0.5B# Non-Emb Params=0.3B2024.07 | 22 | |
| Task ArithmeticLM=true, Math=true, Code=true2025.02 | 21.8 | |
| LLaMA2 7B# Shot=3-shot, # Total Params=6.7B, # Activated Params=6.7B, FLOPs per 4K Tokens=187.9T, # Training Tokens=2T2024.01 | 21.8 | |
| Qwen1.5-1.8B# Non-Emb Params=1.2B2024.07 | 18 | |
| Llama-3-Chinese-8Bevaluation_mode=Few-shot2024.07 | 14.8 | |
| Model StockWizardLM-13B=true, WizardMath-13B=false, LLama-2-13B-Code-Alpaca=true2025.02 | 14 | |
| Model StockWizardLM-13B=true, WizardMath-13B=true, LLama-2-13B-Code-Alpaca=true2025.02 | 6.2 | |
| Galactica-6.7Bevaluation_mode=Few-shot2024.07 | 2 | |
| w/o MergingLM=true, Math=false, Code=false2025.02 | 1 |