Code Generation on MBPP
79.8AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4oSource Model=GPT-4o, Evaluation Context=Source Performance Reference2025.12 | 79.8 | |
| Direct TransferSource Model=GPT-4o, Target Model=Gemma3-27B-it2025.12 | 79.01 | |
| Qwen2.5-14B (Teacher)Role=Teacher, Parameters=14B2026.02 | 78.9 | |
| PromptBridgeSource Model=GPT-4o, Target Model=Gemma3-27B-it2025.12 | 78.59 | |
| GPT-5 OptimizerSource Model=GPT-4o, Target Model=Gemma3-27B-it2025.12 | 78.34 | |
| Qwen2-72B2024.07 | 76.9 | |
| AgentVerseImplementation Framework=MASFactory2026.03 | 75.15 | |
| AgentVerseImplementation Framework=original2026.03 | 74.54 | |
| ChatDevImplementation Framework=MASFactory2026.03 | 74.2 | |
| Vibe Graphing-ChatDevBase Multi-Agent System=ChatDev, Implementation Framework=Vibe Graphing2026.03 | 74.2 | |
| Llama 3 405BParameters=405B2024.07 | 73.4 | |
| Vibe Graphing-Task SpecificImplementation Framework=Vibe Graphing2026.03 | 72.37 | |
| Qwen2-57B-A14BArchitecture=MoE, # Act Params=14B, # Params=57B2024.07 | 71.9 | |
| Mixtral-8x22B2024.07 | 71.7 | |
| ChatDevImplementation Framework=original2026.03 | 71.4 | |
| Mixtral 8x22BParameters=8x22B2024.07 | 71.2 | |
| Qwen1.5-110B2024.07 | 70.9 | |
| Llama-3-70B2024.07 | 70.4 | |
| Direct TransferSource Model=GPT-4o, Target Model=Qwen3-32B2025.12 | 69.1 | |
| GPT-5 OptimizerSource Model=GPT-4o, Target Model=Qwen3-32B2025.12 | 68.85 | |
| HuggingGPTImplementation Framework=original2026.03 | 68.6 | |
| PromptBridgeSource Model=GPT-4o, Target Model=Qwen3-32B2025.12 | 67.25 | |
| SMModel=Dream-Coder-7B (instruct), NFE budget=1/1, FT steps=33.5k2025.10 | 67 | |
| Qwen1.5-72B2024.07 | 66.9 | |
| PromptBridgeSource Model=GPT-4o, Target Model=Llama3.1-8B-Instruct2025.12 | 66.25 | |
| Llama 3 70BParameters=70B2024.07 | 66.2 | |
| BinaryModel=Dream-Coder-7B (instruct), NFE budget=1/1, FT steps=-2025.10 | 65.8 | |
| BinaryModel=Dream-Coder-7B (instruct), NFE budget=1/1, FT steps=33.5k2025.10 | 65.6 | |
| Yi-1.5-34BArchitecture=Dense, # Act Params=32B, # Params=32B2024.07 | 65.5 | |
| GPT-5 OptimizerSource Model=GPT-4o, Target Model=Llama3.1-8B-Instruct2025.12 | 65.41 | |
| HuggingGPTImplementation Framework=MASFactory2026.03 | 64.4 | |
| Qwen1.5-32BArchitecture=Dense, # Act Params=34B, # Params=34B2024.07 | 64.2 | |
| Mixtral-8x7BArchitecture=MoE, # Act Params=12B, # Params=47B2024.07 | 63.9 | |
| CAMELImplementation Framework=original2026.03 | 60.6 | |
| MetaGPTImplementation Framework=MASFactory2026.03 | 59.14 | |
| BinaryModel=Dream-7B (instruct), NFE budget=1/1, FT steps=33.5k2025.10 | 58.3 | |
| BinaryModel=Dream-7B (instruct), NFE budget=1/1, FT steps=-2025.10 | 57.8 | |
| CAMELImplementation Framework=MASFactory2026.03 | 57.8 | |
| SMModel=Dream-7B (instruct), NFE budget=1/1, FT steps=33.5k2025.10 | 56.4 | |
| SMModel=Dream-Coder-7B (instruct), NFE budget=1/2, FT steps=33.5k2025.10 | 56.2 | |
| DAREUse Sens=true, Backbone=Mistral-7B2025.02 | 55.1 | |
| Task ArithmeticUse Sens=true, Backbone=Mistral-7B2025.02 | 54.4 | |
| Code-Llama 7BModality=Finetuned2023.10 | 52.5 | |
| BinaryModel=Dream-Coder-7B (instruct), NFE budget=1/2, FT steps=-2025.10 | 51.6 | |
| CodeUse Sens=false, Backbone=Mistral-7B2025.02 | 50.9 | |
| Direct TransferSource Model=GPT-4o, Target Model=Llama3.1-8B-Instruct2025.12 | 50.88 | |
| Ties-MergingUse Sens=true, Backbone=Mistral-7B2025.02 | 49.9 | |
| BinaryModel=Dream-Coder-7B (instruct), NFE budget=1/2, FT steps=33.5k2025.10 | 49.8 | |
| ChatUse Sens=false, Backbone=Mistral-7B2025.02 | 49.6 | |
| DAREUse Sens=false, Backbone=Mistral-7B2025.02 | 49.4 | |
| Ties-MergingUse Sens=false, Backbone=Mistral-7B2025.02 | 48.4 | |
| Info-Gain SamplerK=1, Backbone=Dream-7B, Block Size=16, Attention Type=full-attention2026.02 | 48.4 | |
| SMModel=Dream-7B (instruct), NFE budget=1/2, FT steps=33.5k2025.10 | 48.4 | |
| Llama 3 8BParameters=8B2024.07 | 47.6 | |
| Mistral 7BModality=Pretrained2023.10 | 47.5 | |
| Mistral 7BParameters=7B2024.07 | 47.5 | |
| Model-GLUE2024.10 | 47.2 | |
| PC-SamplerK=1, Backbone=Dream-7B, Block Size=16, Attention Type=full-attention2026.02 | 46.2 | |
| EntropyK=1, Backbone=Dream-7B, Block Size=16, Attention Type=full-attention2026.02 | 45 | |
| Model-GLUEModel Zoo=Mistral2024.10 | 44.6 | |
| Gemma 7BParameters=7B2024.07 | 44.4 | |
| Info-GainModel=SDAR-8B-Chat, K (Decoding Rate)=12026.02 | 44.4 | |
| GIFTmax generation length=256 tokens, fine-tuning dataset=KodCode, training epochs=52026.05 | 44.4 | |
| LIFT3max generation length=256 tokens, fine-tuning dataset=KodCode, training epochs=52026.05 | 44 | |
| LIFT2max generation length=256 tokens, fine-tuning dataset=KodCode, training epochs=52026.05 | 43.6 | |
| Vanillamax generation length=256 tokens, fine-tuning dataset=KodCode, training epochs=52026.05 | 43.2 | |
| BinaryModel=Dream-7B (instruct), NFE budget=1/2, FT steps=33.5k2025.10 | 43.1 | |
| BinaryModel=Dream-7B (instruct), NFE budget=1/2, FT steps=-2025.10 | 42.8 | |
| Info-GainModel=TraDo-8B-Instruct, K (Decoding Rate)=12026.02 | 42.4 | |
| MiniLLM2026.02 | 42.2 | |
| DistiLLM2026.02 | 42.1 | |
| TSD-KD2026.02 | 42.1 | |
| F-L-SModel Zoo=Mistral2024.10 | 42 | |
| Best Single Model2024.10 | 42 | |
| CARTmax generation length=256 tokens, fine-tuning dataset=KodCode, training epochs=52026.05 | 41.9 | |
| GKDbeta=0.92026.02 | 41.8 | |
| LLaMA-3-70B-Instruct2025.06 | 41.6 | |
| Speculative KD2026.02 | 41.6 | |
| Instruct (Base)max generation length=256 tokens, fine-tuning dataset=KodCode, training epochs=52026.05 | 41.1 | |
| RCOBase Model=LLaMA-3-8B-Instruct2025.06 | 41.08 | |
| Supervised-KD2026.02 | 41 | |
| KLASSK=1, Backbone=Dream-7B, Block Size=16, Attention Type=full-attention2026.02 | 40.6 | |
| Sequence-Level KD2026.02 | 40.6 | |
| MarginK=1, Backbone=Dream-7B, Block Size=16, Attention Type=full-attention2026.02 | 40.4 | |
| RCOBase Model=LLaMA-2-13B-Chat2025.06 | 39.92 | |
| ConfidenceK=1, Backbone=Dream-7B, Block Size=16, Attention Type=full-attention2026.02 | 39.8 | |
| RCOBase Model=LLaMA-2-7B-Chat2025.06 | 39.72 | |
| Info-GainModel=SDAR-8B-Chat, K (Decoding Rate)=22026.02 | 39.4 | |
| Info-Gain SamplerK=2, Backbone=Dream-7B, Block Size=16, Attention Type=full-attention2026.02 | 39.4 | |
| LLaMA-3-8B-Instruct2025.06 | 39.32 | |
| DPCOBase Model=LLaMA-3-8B-Instruct2025.06 | 39.04 | |
| DAREModel Zoo=Mistral2024.10 | 39.02 | |
| LLAMA-2-70B-Chat2025.06 | 39 | |
| RCOBase Model=UltraCM-13B2025.06 | 38.92 | |
| RCOBase Model=Auto-J-13B2025.06 | 38.8 | |
| DPCOBase Model=LLaMA-2-13B-Chat2025.06 | 38.56 | |
| Qwen2.5-1.5B (Student)Role=Student, Parameters=1.5B2026.02 | 38.4 | |
| LLaMA-2-13B-Chat2025.06 | 38.24 | |
| MathUse Sens=false, Backbone=Mistral-7B2025.02 | 38.1 | |
| Info-GainModel=TraDo-8B-Instruct, K (Decoding Rate)=22026.02 | 37.4 |