Coding on MBPP (Accuracy)
98.4AccuracyGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Evaluation Protocol=Closed-Source2026.01 | 98.4 | |
| Qwen2.5-CoderModel Size=14B, GPU Hour=∼1.8M2025.05 | 85.4 | |
| InfiGFusionModel Size=14B, GPU Hour=1952025.05 | 85.2 | |
| Fine-tunedBase LLM=Qwen-2.5-7B-Instruct, Parameters=x4, Evaluation Protocol=0 shot2026.05 | 84.8 | |
| Gemini2.5-ProEvaluation Protocol=Closed-Source2026.01 | 83.7 | |
| ATLAS (cluster)Evaluation Protocol=In-Distribution2026.01 | 83.6 | |
| ATLAS (cluster)Evaluation Protocol=Out-of-Distribution2026.01 | 83.6 | |
| DiDi-Merg.-LBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.0, Evaluation Protocol=0 shot2026.05 | 83.3 | |
| GPT-4oEvaluation Protocol=Closed-Source2026.01 | 82.6 | |
| FREE-MergingBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.08, Evaluation Protocol=0 shot2026.05 | 82.5 | |
| ATLAS (RL)Evaluation Protocol=Out-of-Distribution2026.01 | 81.8 | |
| FuseChatModel Size=14B, GPU Hour=6502025.05 | 81.8 | |
| Qwen2.5-InstructModel Size=14B, GPU Hour=∼1.8M2025.05 | 81.7 | |
| Twin-MergingBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.25, Evaluation Protocol=0 shot2026.05 | 81.6 | |
| Zero-shotBase LLM=Qwen-2.5-7B-Instruct, Parameters=x1, Evaluation Protocol=0 shot2026.05 | 80.2 | |
| InfiFusionModel Size=14B, GPU Hour=1602025.05 | 79.63 | |
| FuseLLMModel Size=14B, GPU Hour=2252025.05 | 79.28 | |
| BertRouterEvaluation Protocol=Out-of-Distribution2026.01 | 79 | |
| RouterDCEvaluation Protocol=Out-of-Distribution2026.01 | 78.7 | |
| MiniLogitModel Size=14B, GPU Hour=2202025.05 | 78.49 | |
| MAS-GPTBackbone=Qwen-2.5-72B-Instruct2025.05 | 78 | |
| Pivot-SFTModel Size=14B, GPU Hour=1202025.05 | 77.86 | |
| RouterDCEvaluation Protocol=In-Distribution2026.01 | 77.7 | |
| DebateBackbone=Qwen-2.5-72B-Instruct2025.05 | 77.6 | |
| AgentVerseBackbone=Qwen-2.5-72B-Instruct2025.05 | 77.6 | |
| ESModel=Qwen2.5-3B-Inst2026.03 | 77.2 | |
| GRPOModel=Qwen2.5-3B-Inst2026.03 | 77 | |
| DyLANBackbone=Qwen-2.5-72B-Instruct2025.05 | 76.8 | |
| SingleBackbone=Qwen-2.5-72B-Instruct2025.05 | 76.5 | |
| PPOModel=Qwen2.5-3B-Inst2026.03 | 76.3 | |
| ES + TT-MVModel=Qwen2.5-3B-Inst2026.03 | 76.3 | |
| Yuan3.0-1T Base#Shots=3-shot, Architecture=MoE, # activated params=68.5B, # total params=1010B2026.01 | 75.9 | |
| RandOptModel=Qwen2.5-3B-Inst2026.03 | 75.9 | |
| CoTBackbone=Qwen-2.5-72B-Instruct2025.05 | 75.8 | |
| SCBackbone=Qwen-2.5-72B-Instruct2025.05 | 75.8 | |
| MacNetBackbone=Qwen-2.5-72B-Instruct2025.05 | 75.5 | |
| AFlow-MathBackbone=Qwen-2.5-72B-Instruct2025.05 | 75.5 | |
| DeepSeek-V3-Base#Shots=3-shot, Architecture=MoE, # activated params=37B, # total params=671B2026.01 | 75.4 | |
| RandOptModel=OLMo3-7B-Inst2026.03 | 75.1 | |
| TT-MV†Model=Qwen2.5-3B-Inst2026.03 | 74.5 | |
| MAVBackbone=Qwen-2.5-72B-Instruct2025.05 | 74 | |
| Best-of-N‡Model=Qwen2.5-3B-Inst2026.03 | 73 | |
| Fine-tunedBase LLM=Llama-3.1-8B-Instruct, Parameters=x4, Evaluation Protocol=0 shot2026.05 | 73 | |
| ES + TT-MVModel=OLMo3-7B-Inst2026.03 | 72.5 | |
| AdaRASCategory=Steering2026.01 | 72.22 | |
| BertRouterEvaluation Protocol=In-Distribution2026.01 | 72.1 | |
| ESModel=OLMo3-7B-Inst2026.03 | 72 | |
| DiDi-Merg.-LBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.0, Evaluation Protocol=0 shot2026.05 | 71.8 | |
| AgentVerseBackbone=Llama-3.3-70B-Instruct2025.05 | 71.3 | |
| FREE-MergingBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.08, Evaluation Protocol=0 shot2026.05 | 71.2 | |
| GRPOModel=OLMo3-7B-Inst2026.03 | 70.8 | |
| Phi-4Model Size=14B, GPU Hour=∼1.0M2025.05 | 70.8 | |
| ProbingCategory=Steering2026.01 | 70.37 | |
| Twin-MergingBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.25, Evaluation Protocol=0 shot2026.05 | 70.3 | |
| MAS-GPTBackbone=Llama-3.3-70B-Instruct2025.05 | 70.3 | |
| GRPOModel=Qwen2.5-1.5B-Inst2026.03 | 70.2 | |
| ES + TT-MVModel=Qwen2.5-1.5B-Inst2026.03 | 70.2 | |
| DyLANBackbone=Llama-3.3-70B-Instruct2025.05 | 70.1 | |
| ESModel=Qwen2.5-1.5B-Inst2026.03 | 69.9 | |
| CoTBackbone=Llama-3.3-70B-Instruct2025.05 | 69.7 | |
| SCBackbone=Llama-3.3-70B-Instruct2025.05 | 69.7 | |
| DebateBackbone=Llama-3.3-70B-Instruct2025.05 | 69.7 | |
| AFlow-MathBackbone=Llama-3.3-70B-Instruct2025.05 | 69.7 | |
| RandOptModel=Qwen2.5-1.5B-Inst2026.03 | 69.6 | |
| STMBackbone=Llama-3.2-3B-Instruct2026.05 | 69.6 | |
| Best-of-N‡Model=Qwen2.5-1.5B-Inst2026.03 | 69.5 | |
| BaseModel=Qwen2.5-3B-Inst2026.03 | 69.5 | |
| MAVBackbone=Llama-3.3-70B-Instruct2025.05 | 69.1 | |
| Mistral-SmallModel Size=24B, GPU Hour=∼1.6M2025.05 | 68.8 | |
| CoTPrompting=Vanilla CoT, Base Model=Qwen3-1.7B2026.01 | 68.78 | |
| MLPRouterEvaluation Protocol=In-Distribution2026.01 | 68.7 | |
| PPOModel=Qwen2.5-1.5B-Inst2026.03 | 68.5 | |
| LLaMA-3.1-405B Base#Shots=3-shot, Architecture=Dense, # activated params=405B, # total params=405B2026.01 | 68.4 | |
| TT-MV†Model=Qwen2.5-1.5B-Inst2026.03 | 68 | |
| SingleBackbone=Llama-3.3-70B-Instruct2025.05 | 67.9 | |
| MLPRouterEvaluation Protocol=Out-of-Distribution2026.01 | 67.7 | |
| PPOModel=OLMo3-7B-Inst2026.03 | 67.7 | |
| MacNetBackbone=Llama-3.3-70B-Instruct2025.05 | 67.1 | |
| Zero-shotBase LLM=Llama-3.1-8B-Instruct, Parameters=x1, Evaluation Protocol=0 shot2026.05 | 66.9 | |
| MADBackbone=Qwen-2.5-72B-Instruct2025.05 | 66.3 | |
| BaseModel=OLMo3-7B-Inst2026.03 | 65.9 | |
| RandOptModel=Llama3.1-8B-Inst2026.03 | 65.2 | |
| ESModel=Llama3.1-8B-Inst2026.03 | 64.8 | |
| FS RouterEvaluation Protocol=Training-free, Prompting Strategy=Few-shot2026.01 | 64.7 | |
| ZS RouterEvaluation Protocol=Training-free, Prompting Strategy=Zero-shot2026.01 | 64.2 | |
| Best-of-N‡Model=OLMo3-7B-Inst2026.03 | 64 | |
| ES + TT-MVModel=Llama3.1-8B-Inst2026.03 | 63.1 | |
| AutoGenBackbone=Llama-3.3-70B-Instruct2025.05 | 62.9 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.05 | 62.7 | |
| BaseModel=Qwen2.5-1.5B-Inst2026.03 | 62.3 | |
| DFTBackbone=Llama-3.2-3B-Instruct2026.05 | 62.2 | |
| ESModel=OLMo3-7B2026.03 | 61.8 | |
| GRPOModel=Llama3.1-8B-Inst2026.03 | 61 | |
| TT-MV†Model=OLMo3-7B-Inst2026.03 | 60.5 | |
| Best-of-N‡Model=Llama3.1-8B-Inst2026.03 | 60 | |
| TT-MV†Model=Llama3.1-8B-Inst2026.03 | 59.5 | |
| HySparse# Shots=3-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=HySparse2026.02 | 59.3 | |
| Anchored LearningBackbone=Llama-3.2-3B-Instruct2026.05 | 58.7 | |
| GRPOModel=OLMo3-7B2026.03 | 58.5 | |
| PPOModel=OLMo3-7B2026.03 | 57.9 |