Code Generation on MBPP (pass@1, pass@80)
89.1Pass@1Vanilla (Ref)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Vanilla (Ref)LLM=Claude-3.5-Sonnet, Multi-agent=No, Dynamic Routing=No2026.03 | 89.1 | — | |
| NBDiff-7B-INSTRUCTParameters=7B, Training Protocol=Instruct, Sampling Strategy=p=0.9, T=12025.12 | 87.6 | — | |
| Vanilla (Ref)LLM=GPT-4o, Multi-agent=No, Dynamic Routing=No2026.03 | 87.2 | — | |
| AMRO-SLLM=LLM Pool*, Multi-agent=Yes, Dynamic Routing=Yes2026.03 | 86.3 | — | |
| Low-Rankα=1/16, Backbone Model=Qwen2.5-Coder-Instruct2025.06 | 86.2 | — | |
| MasRouterLLM=LLM Pool*, Multi-agent=Yes, Dynamic Routing=Yes2026.03 | 84 | — | |
| BitDeltaα=1/16, Backbone Model=Qwen2.5-Coder-Instruct2025.06 | 83.9 | — | |
| PRINMIXα=1/16, Backbone Model=Qwen2.5-Coder-Instruct2025.06 | 83.1 | — | |
| GPT-4Base Model=-, Params=-, Instruction Data=-, Model Weight=-2024.06 | 83 | — | |
| Alignedα=1, Backbone Model=Qwen2.5-Coder-Instruct2025.06 | 82.8 | — | |
| Delta-CoMeα=1/16, Backbone Model=Qwen2.5-Coder-Instruct2025.06 | 82.7 | — | |
| AFlowLLM=GPT-4o-mini, Multi-agent=Yes, Dynamic Routing=No2026.03 | 82.2 | — | |
| GPT-3.5Base Model=-, Params=-, Instruction Data=-, Model Weight=-2024.06 | 81.6 | — | |
| OneFlowMethod category=Single-LLM implementation, Executor LLM=GPT-4o-mini, Execution protocol=single-agent execution2026.01 | 81.4 | — | |
| OneFlowMethod category=Automatically designed multi-agent frameworks, Executor LLM=GPT-4o-mini, Execution protocol=multi-agent2026.01 | 81.1 | — | |
| Backboneα=1, Backbone Model=Qwen2.5-Coder-Instruct2025.06 | 80.7 | — | |
| PromptBridgeSource Model=GPT-4o, Target Model=o4-mini2025.12 | 80.6 | — | |
| PromptBridgeSource Model=GPT-4o, Target Model=o32025.12 | 80.44 | — | |
| DCDModel=NBDiff, Cache=None2026.01 | 80.2 | — | |
| GEPASource Model=GPT-4o, Target Model=o32025.12 | 80.01 | — | |
| GPT-5 OptimizerSource Model=GPT-4o, Target Model=o32025.12 | 79.93 | — | |
| GPT-4oModel Role=Source Model2025.12 | 79.8 | — | |
| Direct TransferSource Model=GPT-4o, Target Model=o4-mini2025.12 | 79.09 | — | |
| AFlowMethod category=Automatically designed multi-agent frameworks, Executor LLM=GPT-4o-mini, Execution protocol=multi-agent2026.01 | 78.8 | — | |
| AFlowMethod category=Single-LLM implementation, Executor LLM=GPT-4o-mini, Execution protocol=single-agent execution2026.01 | 78.8 | — | |
| Sub-block-basedModel=NBDiff, Cache=None2026.01 | 78.2 | — | |
| Direct TransferSource Model=GPT-4o, Target Model=o32025.12 | 77.92 | — | |
| MIPROv2Source Model=GPT-4o, Target Model=o4-mini2025.12 | 77.92 | — | |
| LLaDA2.0-mini preview 16BA1BParameters=16BA1B, Sampling Strategy=p=0.9, T=12025.12 | 77.8 | — | |
| GEPASource Model=GPT-4o, Target Model=o4-mini2025.12 | 76.49 | — | |
| AFlowLLM=Gemini-1.5-flash, Multi-agent=Yes, Dynamic Routing=No2026.03 | 76 | — | |
| GPTSwarmLLM=GPT-4o-mini, Multi-agent=Yes, Dynamic Routing=No2026.03 | 75.4 | — | |
| RouterDCLLM=LLM Pool*, Multi-agent=No, Dynamic Routing=Yes2026.03 | 75.2 | — | |
| VanillaLLM=Claude-3.5-haiku, Multi-agent=No, Dynamic Routing=No2026.03 | 74.5 | — | |
| AoT (Atom)LLM=GPT-4o-mini, Multi-agent=No, Dynamic Routing=No2026.03 | 74 | — | |
| PromptBridgeSource Model=GPT-4o, Target Model=Llama3.1-70B-Instruct2025.12 | 73.64 | — | |
| LLM-DebateLLM=GPT-4o-mini, Multi-agent=Yes, Dynamic Routing=No2026.03 | 73.6 | — | |
| GPT-5 OptimizerSource Model=GPT-4o, Target Model=Llama3.1-70B-Instruct2025.12 | 73.38 | — | |
| CoTMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard2026.01 | 73.3 | — | |
| MultiPersonaMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard2026.01 | 73.3 | — | |
| GoT (Graph)LLM=GPT-4o-mini, Multi-agent=No, Dynamic Routing=No2026.03 | 73.2 | — | |
| VanillaLLM=Gemini-1.5-flash, Multi-agent=No, Dynamic Routing=No2026.03 | 73 | — | |
| ToT (Tree)LLM=GPT-4o-mini, Multi-agent=No, Dynamic Routing=No2026.03 | 72.8 | — | |
| IOMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard2026.01 | 72.6 | — | |
| RouteLLMLLM=LLM Pool*, Multi-agent=No, Dynamic Routing=Yes2026.03 | 72.6 | — | |
| VanillaLLM=GPT-4o-mini, Multi-agent=No, Dynamic Routing=No2026.03 | 72.2 | — | |
| SDAR 8BParameters=8B, Sampling Strategy=p=0.9, T=12025.12 | 72 | — | |
| CoT SCMethod category=Manual baselines, Executor LLM=GPT-4o-mini, Execution protocol=standard, shots=5-shot2026.01 | 71.9 | — | |
| VanillaLLM=Llama-3.1-70b, Multi-agent=No, Dynamic Routing=No2026.03 | 71.8 | — | |
| GPT-5 OptimizerSource Model=GPT-4o, Target Model=o4-mini2025.12 | 70.03 | — | |
| LLaDA-MoE 7B-A1BParameters=7B-A1B, Sampling Strategy=p=0.9, T=12025.12 | 70 | — | |
| CoTLLM=GPT-4o-mini, Multi-agent=No, Dynamic Routing=No2026.03 | 69.6 | — | |
| LowRankCR=32, Backbone=Codellama-13B, Fine-tuned Model=WizardCoder-13B2025.04 | 68.8 | — | |
| MIPROv2Source Model=GPT-4o, Target Model=Llama3.1-70B-Instruct2025.12 | 68.26 | — | |
| IMPARTCR=32, Backbone=Codellama-13B, Fine-tuned Model=WizardCoder-13B2025.04 | 68 | — | |
| Fine-tunedCR=1, Backbone=Codellama-13B, Fine-tuned Model=WizardCoder-13B2025.04 | 67.7 | — | |
| Direct TransferSource Model=GPT-4o, Target Model=Llama3.1-70B-Instruct2025.12 | 65.57 | — | |
| UNICODERBase Model=Code Llama, Params=7B, Instruction Data=true, Model Weight=true2024.06 | 65.2 | — | |
| DARECR=32, Backbone=Codellama-13B, Fine-tuned Model=WizardCoder-13B2025.04 | 64.6 | — | |
| UNICODERBase Model=Deepseek-Coder, Params=6.7B, Instruction Data=true, Model Weight=true2024.06 | 64.3 | — | |
| Magicoder-CLBase Model=Code Llama, Params=7B, Instruction Data=true, Model Weight=true2024.06 | 64.2 | — | |
| MIPROv2Source Model=GPT-4o, Target Model=o32025.12 | 63.81 | — | |
| WaveCoder-DSBase Model=Deepseek-Coder, Params=6.7B, Instruction Data=true, Model Weight=true2024.06 | 62.8 | — | |
| BackboneCR=1, Backbone=Codellama-13B, Fine-tuned Model=WizardCoder-13B2025.04 | 62.7 | — | |
| DeepseekCoderBase Model=-, Params=6.7B, Instruction Data=false, Model Weight=true2024.06 | 60.6 | — | |
| DCDModel=Dream-v0-Instruct-7B, Cache=Dual2026.01 | 58.8 | — | |
| Dream-v0 Instruct-7BParameters=7B, Training Protocol=Instruct, Sampling Strategy=p=0.9, T=12025.12 | 58.8 | — | |
| CCDModel=Dream-v0-Instruct-7B, Cache=Custom2026.01 | 58 | — | |
| DCDModel=Dream-v0-Instruct-7B, Cache=Prefix2026.01 | 57.4 | — | |
| DCDModel=Dream-v0-Instruct-7B, Cache=None2026.01 | 56.8 | — | |
| Block-basedModel=Dream-v0-Instruct-7B, Cache=None2026.01 | 55 | — | |
| IOAStudent Model=Qwen2.5-14B, Teacher LLM=DeepSeek-R12026.02 | 54.94 | — | |
| IOA (Ours)Model=Qwen2.5-7B2026.02 | 54.94 | — | |
| ProphetModel=Dream-v0-Instruct-7B, Cache=None2026.01 | 54.6 | — | |
| Block-basedModel=Dream-v0-Instruct-7B, Cache=Prefix2026.01 | 53.6 | — | |
| Block-basedModel=Dream-v0-Instruct-7B, Cache=Dual2026.01 | 52.8 | — | |
| WizardCoderBase Model=StarCoder, Params=15B, Instruction Data=true, Model Weight=true2024.06 | 51.8 | — | |
| WaveCoder-SCBase Model=StarCoder, Params=15B, Instruction Data=true, Model Weight=true2024.06 | 51 | — | |
| Sub-block-basedModel=Fast-dLLM-v2-7B, Cache=None2026.01 | 50.2 | — | |
| PaLM 2-S*Backbone=PaLM 2-S, Training=Trained with additional code-related tokens2023.05 | 50 | 86.6 | |
| IOAStudent Model=Qwen2.5-3B, Teacher Model=DeepSeek-R12026.02 | 49.32 | — | |
| DCDModel=Fast-dLLM-v2-7B, Cache=Dual2026.01 | 49 | — | |
| DCDModel=Fast-dLLM-v2-7B, Cache=None2026.01 | 48.6 | — | |
| Block-basedModel=Fast-dLLM-v2-7B, Cache=None2026.01 | 48.4 | — | |
| GKDModel=Qwen2.5-7B2026.02 | 48.16 | — | |
| IOAModel=Qwen2.5-3B, Teacher LLM=OpenAI o12026.02 | 47.86 | — | |
| WaveCoder-CLBase Model=Code Llama, Params=7B, Instruction Data=true, Model Weight=true2024.06 | 47.2 | — | |
| PaLM-Coder-540BParameters=540B2023.05 | 47 | 80.8 | |
| Sigma-MoE-Tiny Base# Shots=3-shot, Architecture=MoE, # Activated Params=0.5B, # Total Params=20B2025.12 | 47 | — | |
| POCLModel=Qwen2.5-7B2026.02 | 46.57 | — | |
| Gemma-3 4B Base# Shots=3-shot, Architecture=Dense, # Activated Params=4B, # Total Params=4B2025.12 | 46.4 | — | |
| Sub-block-basedModel=Fast-dLLM-v2-7B, Cache=Dual2026.01 | 46 | — | |
| DistiLLM-2Model=Qwen2.5-7B2026.02 | 45.63 | — | |
| SuperCorrectModel=Qwen2.5-7B2026.02 | 45.05 | — | |
| IOAStudent Model=LLaMA3.2-3B, Teacher Model=DeepSeek-R12026.02 | 44.97 | — | |
| ABKDModel=Qwen2.5-7B2026.02 | 44.87 | — | |
| Code-Llama-InstructBase Model=Code Llama, Params=7B, Instruction Data=true, Model Weight=true2024.06 | 44.4 | — | |
| IOAModel=LLaMA3.2-3B, Teacher LLM=OpenAI o12026.02 | 44.25 | — | |
| Qwen-1.5 14BRole=Teacher2024.07 | 44 | — | |
| Qwen-1.5 14BModel Type=Teacher2024.07 | 44 | — |