Coding on HumanEval (accuracy)
95.62AccuracyDMoA
Evaluation Results
| Method | Links | |
|---|---|---|
| DMoABackbone=gpt-oss-120b, Evaluation Paradigm=few-shot2026.05 | 95.62 | |
| CoTBackbone=Qwen-2.5-72B-Instruct2025.05 | 93.9 | |
| SpecReasonBackbone=gpt-oss-120b2026.05 | 91.22 | |
| SCBackbone=Qwen-2.5-72B-Instruct2025.05 | 90.2 | |
| DebateBackbone=Qwen-2.5-72B-Instruct2025.05 | 90.2 | |
| SafeSieveBackbone=gpt-oss-120b2026.05 | 90.15 | |
| G-DesignerBackbone=gpt-oss-120b, Evaluation Paradigm=few-shot2026.05 | 89.4 | |
| AFlowBackbone=gpt-oss-120b2026.05 | 89.25 | |
| STEERBackbone=gpt-oss-120b2026.05 | 89.02 | |
| MAS-GPTBackbone=Qwen-2.5-72B-Instruct2025.05 | 89 | |
| SingleBackbone=Qwen-2.5-72B-Instruct2025.05 | 88.4 | |
| ARG-DesignerBackbone=gpt-oss-120b, Evaluation Paradigm=few-shot2026.05 | 88.25 | |
| AgentVerseBackbone=Llama-3.3-70B-Instruct2025.05 | 87.8 | |
| AFlow-MathBackbone=Qwen-2.5-72B-Instruct2025.05 | 87.8 | |
| MacNetBackbone=Qwen-2.5-72B-Instruct2025.05 | 87.2 | |
| GPTSwarmBackbone=gpt-oss-120b2026.05 | 87.14 | |
| ICaRusModel=Qwen3-8B-Base, KV Sharing=O2026.02 | 86.6 | |
| MacNetBackbone=Llama-3.3-70B-Instruct2025.05 | 86.6 | |
| MAS-GPTBackbone=Llama-3.3-70B-Instruct2025.05 | 86.6 | |
| AgentVerseBackbone=Qwen-2.5-72B-Instruct2025.05 | 86.6 | |
| Fine-tunedBase LLM=Qwen-2.5-7B-Instruct, Parameters=x4, Evaluation Protocol=0 shot2026.05 | 85.7 | |
| Zero-shotBase LLM=Qwen-2.5-7B-Instruct, Parameters=x1, Evaluation Protocol=0 shot2026.05 | 85.4 | |
| SingleBackbone=Llama-3.3-70B-Instruct2025.05 | 85.4 | |
| CoTBackbone=Llama-3.3-70B-Instruct2025.05 | 85.4 | |
| LLM-DebateBackbone=gpt-oss-120b2026.05 | 85.24 | |
| DebateBackbone=Llama-3.3-70B-Instruct2025.05 | 84.8 | |
| DiDi-Merg.-LBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.0, Evaluation Protocol=0 shot2026.05 | 84.4 | |
| Complete GraphBackbone=gpt-oss-120b2026.05 | 84.25 | |
| AFlow-MathBackbone=Llama-3.3-70B-Instruct2025.05 | 84.2 | |
| Random GraphBackbone=gpt-oss-120b2026.05 | 84.14 | |
| FREE-MergingBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.08, Evaluation Protocol=0 shot2026.05 | 84.1 | |
| AutoGenBackbone=gpt-oss-120b2026.05 | 83.5 | |
| DyLANBackbone=Llama-3.3-70B-Instruct2025.05 | 82.9 | |
| MoABackbone=gpt-oss-120b2026.05 | 82.48 | |
| SCBackbone=Llama-3.3-70B-Instruct2025.05 | 82.3 | |
| Twin-MergingBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.25, Evaluation Protocol=0 shot2026.05 | 81.9 | |
| Multi ModelModel=Qwen3-8B-Base, KV Sharing=X2026.02 | 81.7 | |
| DyLANBackbone=Qwen-2.5-72B-Instruct2025.05 | 79.9 | |
| MAVBackbone=Llama-3.3-70B-Instruct2025.05 | 78 | |
| StarBackbone=gpt-oss-120b2026.05 | 76.28 | |
| MAVBackbone=Qwen-2.5-72B-Instruct2025.05 | 76.2 | |
| SCBackbone=gpt-oss-120b2026.05 | 75.82 | |
| AutoGenBackbone=Qwen-2.5-72B-Instruct2025.05 | 75.6 | |
| TreeBackbone=gpt-oss-120b2026.05 | 75.33 | |
| MADBackbone=Llama-3.3-70B-Instruct2025.05 | 75 | |
| CoTBackbone=gpt-oss-120b2026.05 | 74.05 | |
| ChainBackbone=gpt-oss-120b2026.05 | 73.4 | |
| VanillaBackbone=gpt-oss-120b2026.05 | 73.28 | |
| Fine-tunedBase LLM=Llama-3.1-8B-Instruct, Parameters=x4, Evaluation Protocol=0 shot2026.05 | 72 | |
| MADBackbone=Qwen-2.5-72B-Instruct2025.05 | 72 | |
| DiDi-Merg.-LBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.0, Evaluation Protocol=0 shot2026.05 | 70.4 | |
| FREE-MergingBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.08, Evaluation Protocol=0 shot2026.05 | 69.1 | |
| Twin-MergingBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.25, Evaluation Protocol=0 shot2026.05 | 68.8 | |
| Base ModelModel=Qwen3-8B-Base, KV Sharing=.2026.02 | 68.3 | |
| Zero-shotBase LLM=Llama-3.1-8B-Instruct, Parameters=x1, Evaluation Protocol=0 shot2026.05 | 68.3 | |
| STMBackbone=Llama-3.2-3B-Instruct2026.05 | 57.3 | |
| Anchored LearningBackbone=Llama-3.2-3B-Instruct2026.05 | 57.3 | |
| Iter-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 57.1 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.05 | 54.9 | |
| DFTBackbone=Llama-3.2-3B-Instruct2026.05 | 54.3 | |
| Self-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 54.3 | |
| Low-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 52.4 | |
| AutoGenBackbone=Llama-3.3-70B-Instruct2025.05 | 51.2 | |
| SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 50.7 | |
| KL-SFTBackbone=Llama-3.2-3B-Instruct2026.05 | 49.4 | |
| Multi ModelModel=LLaMA-3.1-8B, KV Sharing=X2026.02 | 48.2 | |
| ICaRusModel=LLaMA-3.1-8B, KV Sharing=O2026.02 | 48.2 | |
| EGSPO-SASeq. Len.=5122026.03 | 44.5 | |
| EGSPO-SASeq. Len.=Best2026.03 | 44.5 | |
| EGSPO-SASeq. Len.=2562026.03 | 41.5 | |
| EGSPOSeq. Len.=2562026.03 | 40.2 | |
| EGSPOSeq. Len.=Best2026.03 | 40.2 | |
| EGSPOSeq. Len.=5122026.03 | 39.6 | |
| LLaDA-8B-InstructSeq. Len.=5122026.03 | 37.8 | |
| LLaDA-8B-InstructSeq. Len.=Best2026.03 | 37.8 | |
| d1Seq. Len.=5122026.03 | 37.8 | |
| d1Seq. Len.=Best2026.03 | 37.8 | |
| Base ModelModel=LLaMA-3.1-8B, KV Sharing=.2026.02 | 36.6 | |
| LLaDA-8B-InstructSeq. Len.=2562026.03 | 35.3 | |
| d1Seq. Len.=2562026.03 | 32.9 | |
| EGSPOSeq. Len.=1282026.03 | 32.3 | |
| EGSPO-SASeq. Len.=1282026.03 | 32.3 | |
| d1Seq. Len.=1282026.03 | 31.1 | |
| LLaDA-8B-InstructSeq. Len.=1282026.03 | 27.4 |