Code Generation on MBPP (accuracy (%))
92.2Accuracy (%)MegaAgent
Evaluation Results
| Method | Links | |
|---|---|---|
| MegaAgentBackbone=GPT-4o API2024.08 | 92.2 | |
| Low-Rankα=1/16, Model Architecture=Qwen2.5-Coder-Instruct-14B2025.06 | 88.6 | |
| PRINMIXα=1/16, Model Architecture=Qwen2.5-Coder-Instruct-14B2025.06 | 86.9 | |
| BitDeltaα=1/16, Model Architecture=Qwen2.5-Coder-Instruct-14B2025.06 | 86.5 | |
| Delta-CoMeα=1/16, Model Architecture=Qwen2.5-Coder-Instruct-14B2025.06 | 86.5 | |
| Alignedα=1, Model Architecture=Qwen2.5-Coder-Instruct-14B2025.06 | 85.4 | |
| AutoGenBackbone=GPT-4o API2024.08 | 85.3 | |
| Backboneα=1, Model Architecture=Qwen2.5-Coder-Instruct-14B2025.06 | 84.7 | |
| openPangu-EmbeddedModel=7B, CoT Mode=auto_think, Precision=INT82025.12 | 83.27 | |
| AgentVerseBackbone=GPT-4o API2024.08 | 82.4 | |
| MetaGPTBackbone=GPT-4o API2024.08 | 81.7 | |
| openPangu-EmbeddedModel=7B, CoT Mode=auto_think, Precision=FP162025.12 | 80.16 | |
| openPangu-EmbeddedModel=7B, CoT Mode=slow_think, Precision=INT82025.12 | 79.77 | |
| Seed DiffusionSystem Optimization=true, Throughput (TPS)=1600, # Shots=32025.12 | 79.4 | |
| openPangu-EmbeddedModel=7B, CoT Mode=no_think, Precision=INT82025.12 | 78.21 | |
| CamelBackbone=GPT-4o API2024.08 | 78.1 | |
| PASC-GRPObase_model=qwen3-8B-Base2026.01 | 77.49 | |
| openPangu-EmbeddedModel=7B, CoT Mode=slow_think, Precision=FP162025.12 | 77.43 | |
| Mercury Coder MiniSystem Optimization=true, Throughput (TPS)=1109, # Shots=32025.12 | 77.1 | |
| openPangu-EmbeddedModel=7B, CoT Mode=no_think, Precision=FP162025.12 | 77.04 | |
| qwen3-8B2026.01 | 76.7 | |
| Mercury Coder SmallSystem Optimization=true, Throughput (TPS)=737, # Shots=32025.12 | 76.6 | |
| Gemini DiffusionSystem Optimization=true, Throughput (TPS)=1479, # Shots=32025.12 | 76 | |
| GRP-SFTbase_model=qwen3-8B-Base2026.01 | 75.45 | |
| Gemini 2.0 Flash LiteSystem Optimization=true, Throughput (TPS)=201, # Shots=32025.12 | 75 | |
| GPT 4o MiniSystem Optimization=true, Throughput (TPS)=59, # Shots=32025.12 | 74.6 | |
| Qwen3-Coder-30B-A3B-Instruct (Base)Model=Qwen3-Coder-30B-A3B-Instruct, Strategy=Base2026.02 | 73.8 | |
| qwen3-8B-Base2026.01 | 73.27 | |
| REFUSION# Total Params=8B, # Activated Params=8B, System Optimization=false, Throughput (TPS)=98, # Shots=32025.12 | 68.2 | |
| LLaDA-MoE# Total Params=7B, # Activated Params=1B, System Optimization=true, Throughput (TPS)=884, # Shots=32025.12 | 67.45 | |
| Qwen2.5-7B-InstructType=Autoregressive2026.01 | 66.93 | |
| Nova MicroSystem Optimization=true, Throughput (TPS)=148, # Shots=32025.12 | 65.4 | |
| Llama3.1-8B-InstructType=Autoregressive2026.01 | 65.37 | |
| BF16Bit=BF16, Model=Qwen3-8B, Group size (g)=N/A2026.02 | 65 | |
| Ouro2.6B-Thinking + RLTT2026.02 | 64.6 | |
| OJBKQBit=4-bit, Model=Qwen3-8B, Group size (g)=1282026.02 | 63.2 | |
| AWQBit=4-bit, Model=Qwen3-8B, Group size (g)=1282026.02 | 62.8 | |
| BF16Bit=BF16, Model=Qwen3-4B, Group size (g)=N/A2026.02 | 62.4 | |
| openPangu-EmbeddedModel=1B, CoT Mode=no_think, Precision=INT82025.12 | 62.26 | |
| openPangu-EmbeddedModel=1B, CoT Mode=no_think, Precision=FP162025.12 | 61.87 | |
| Ouro2.6B-Thinking2026.02 | 61.3 | |
| Ouro2.6B-Thinking + GRPO2026.02 | 61.3 | |
| GPTQBit=4-bit, Model=Qwen3-8B, Group size (g)=1282026.02 | 60.2 | |
| Ouro2.6B-Thinking + SFT2026.02 | 59.9 | |
| AWQBit=4-bit, Model=Qwen3-4B, Group size (g)=1282026.02 | 57.8 | |
| OJBKQBit=4-bit, Model=Qwen3-4B, Group size (g)=1282026.02 | 56.8 | |
| Qwen3 – 4B2026.02 | 56.5 | |
| Dream-v0Configuration=Base-7B2025.12 | 56.2 | |
| Fast-dLLMModel=Dream-v0-base-7B, Partition=Fixed, Cache=Prefix2026.02 | 55.8 | |
| SwordsmanModel=Dream-v0-base-7B, Partition=Adaptive, Cache=Prefix2026.02 | 55.8 | |
| NBDiff-7B-BASEConfiguration=Base2025.12 | 55.8 | |
| AMA ResamplingStrategy=Resampling2025.05 | 55.76 | |
| SwordsmanModel=Dream-v0-base-7B, Partition=Adaptive, Cache=None2026.02 | 55.6 | |
| Fast-dLLMModel=Dream-v0-base-7B, Partition=Fixed, Cache=None2026.02 | 55.4 | |
| Qwen3-4B-Instruct-2507 (C/E Weighted)Model=Qwen3-4B-Instruct-2507, Strategy=C/E Weighted2026.02 | 55.2 | |
| SwordsmanModel=Dream-v0-base-7B, Partition=Adaptive, Cache=Dual2026.02 | 54.8 | |
| Fast-dLLMModel=Dream-v0-base-7B, Partition=Fixed, Cache=Dual2026.02 | 54.4 | |
| AMA ReweightingStrategy=Reweighting2025.05 | 54.32 | |
| D2FModel=Dream-v0-base-7B, Partition=Fixed, Cache=Prefix2026.02 | 53.48 | |
| Qwen3-4B-Instruct-2507 (Base)Model=Qwen3-4B-Instruct-2507, Strategy=Base2026.02 | 52.6 | |
| LLaDA-MoE-7BConfiguration=A1B-Base2025.12 | 52.4 | |
| CodeUltraFeedback Specialist2025.05 | 52.16 | |
| Model Averaging2025.05 | 52.16 | |
| LLaDA1.5-8BDecoding Strategy=PC-Sampler2026.01 | 51.36 | |
| LLaDA1.5-8BDecoding Strategy=FourierSampler2026.01 | 50.58 | |
| Standard2025.05 | 50 | |
| Zephyr-7b-sft-full2025.05 | 49.9 | |
| LLaDA-8B-InstructDecoding Strategy=PC-Sampler2026.01 | 49.81 | |
| DeepSeekR1 – 7B2026.02 | 49.7 | |
| SafeRLHF Specialist2025.05 | 48.56 | |
| BF16Bit=BF16, Model=LLaMA3-8B, Group size (g)=N/A2026.02 | 48.4 | |
| LLaDA-8B-InstructDecoding Strategy=FourierSampler2026.01 | 47.86 | |
| Qwen3-4B-Instruct-2507 (Edge Only)Model=Qwen3-4B-Instruct-2507, Strategy=Edge Only2026.02 | 47.2 | |
| AWQBit=4-bit, Model=LLaMA3-8B, Group size (g)=1282026.02 | 46.4 | |
| P-TrojanModel=Qwen2.5-1.5B2025.12 | 46.2 | |
| cleanModel=Qwen2.5-1.5B2025.12 | 46 | |
| GPTQBit=4-bit, Model=LLaMA3-8B, Group size (g)=1282026.02 | 45 | |
| OJBKQBit=4-bit, Model=LLaMA3-8B, Group size (g)=1282026.02 | 44.8 | |
| Chatbot Arena 2024 Specialist2025.05 | 44.6 | |
| Qwen3 – 1.7B2026.02 | 43.7 | |
| LLaDA1.5-8BDecoding Strategy=RWS2026.01 | 43.19 | |
| BadNet-CEModel=Qwen2.5-1.5B2025.12 | 43 | |
| LLaDA1.5-8BDecoding Strategy=Vanilla2026.01 | 42.02 | |
| LLaDA-8B-InstructDecoding Strategy=RWS2026.01 | 42.02 | |
| LLaDA-8B-InstructDecoding Strategy=Vanilla2026.01 | 41.25 | |
| SwordsmanModel=LLaDA-1.5, Partition=Adaptive, Cache=None2026.02 | 41 | |
| QUIPBit=4-bit, Model=LLaMA3-8B, Group size (g)=1282026.02 | 40.4 | |
| Fast-dLLMModel=LLaDA-1.5, Partition=Fixed, Cache=Prefix2026.02 | 40 | |
| AdaBlockModel=LLaDA-8B-Instruct, Partition=Adaptive, Cache=None2026.02 | 39.8 | |
| SwordsmanModel=LLaDA-1.5, Partition=Adaptive, Cache=Prefix2026.02 | 39.8 | |
| Qwen3-4B-Instruct-2507 (Basic Only)Model=Qwen3-4B-Instruct-2507, Strategy=Basic Only2026.02 | 39.8 | |
| Fast-dLLMModel=LLaDA-1.5, Partition=Fixed, Cache=None2026.02 | 39.6 | |
| LLaDA-8B-InstructSettings=Official scoreboard evaluation settings2026.01 | 39.6 | |
| SwordsmanModel=LLaDA-1.5, Partition=Adaptive, Cache=Dual2026.02 | 39.4 | |
| BadNetModel=Qwen2.5-1.5B2025.12 | 39.2 | |
| CTC+CETraining=CTC-trained2026.01 | 38.4 | |
| LLaDA-8BConfiguration=Base2025.12 | 38.2 | |
| D2FModel=LLaDA-8B-Instruct, Partition=Fixed, Cache=Prefix2026.02 | 38 | |
| AdaBlockModel=LLaDA-8B-Instruct, Partition=Adaptive, Cache=Dual2026.02 | 38 | |
| AdaBlockModel=LLaDA-1.5, Partition=Adaptive, Cache=None2026.02 | 37.6 |