Code Generation on HumanEval (Acc, TPF, AUP)
97.56AccuracyAgent Q-Mix
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Agent Q-MixBase Model=GPT-OSS:120B2026.04 | 97.56 | — | — | |
| LobsterBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 97.56 | — | — | |
| AutoGenBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 96.95 | — | — | |
| LobsterBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 96.34 | — | — | |
| TopoDIMBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 95.73 | — | — | |
| Agent FrameworkBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 95.73 | — | — | |
| Agent Q-MixBase Model=Gemini-3.1-Flash-Lite2026.04 | 95.73 | — | — | |
| AutoGenBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 95.12 | — | — | |
| LangGraphBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 95.12 | — | — | |
| SMCS2025.07 | 95.12 | — | — | |
| AFlowMethod Category=Workflow induction methods, Optimizer LLM=Claude-3.5 Sonnet, Executor LLM=GPT-4o mini2026.04 | 94.7 | — | — | |
| AFlow (our reimplementation)Optimizer LLM=GPT-4.1 mini, Executor LLM=GPT-4.1 mini2026.04 | 94.65 | — | — | |
| Agent FrameworkBase Model=Gemini-3.1-Flash-Lite, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 94.51 | — | — | |
| GPT-4.12025.07 | 94.51 | — | — | |
| LangGraphBase Model=GPT-OSS:120B, Framework Category=Commercial Framework (multi-agent mode)2026.04 | 93.9 | — | — | |
| G-DesignerBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 93.9 | — | — | |
| TopoDIMBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 93.9 | — | — | |
| FLOWBOTOptimizer LLM=GPT-4.1 mini, Executor LLM=GPT-4.1 mini2026.04 | 93.74 | — | — | |
| GPTSwarmBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 93.7 | — | — | |
| LLM-DebateBase Model=Gemini-3.1-Flash-Lite, Framework Category=Static multi-agent2026.04 | 92.68 | — | — | |
| MaASBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 92.68 | — | — | |
| GTDBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 92.68 | — | — | |
| QwQ-32B2025.07 | 92.68 | — | — | |
| G-DesignerBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 92.08 | — | — | |
| MaASBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 92.07 | — | — | |
| Base (direct)Base Model=Gemini-3.1-Flash-Lite, Framework Category=Single-agent baseline2026.04 | 92.07 | — | — | |
| Self-MoA2025.07 | 92.07 | — | — | |
| CoT + Self-ConsistencyMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 91.6 | — | — | |
| MedPromptMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 91.6 | — | — | |
| GTDBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 91.46 | — | — | |
| AgentDropoutBase Model=Gemini-3.1-Flash-Lite, Framework Category=Adaptive topology methods2026.04 | 91.46 | — | — | |
| GPTSwarmBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 90.24 | — | — | |
| Base (direct)Base Model=GPT-OSS:120B, Framework Category=Single-agent baseline2026.04 | 90.04 | — | — | |
| LLM-DebateBase Model=GPT-OSS:120B, Framework Category=Static multi-agent2026.04 | 89.63 | — | — | |
| MultiPersonaMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 89.3 | — | — | |
| Disagreement-Guided Strategy Routing (Qwen2.5-Coder-7B-Instruct)Backbone=Qwen2.5-Coder-7B-Instruct, Sampling strategy=Ours2026.04 | 89 | — | — | |
| Chain-of-Thought (CoT)Method Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 88.6 | — | — | |
| AgentDropoutBase Model=GPT-OSS:120B, Framework Category=Adaptive topology methods2026.04 | 88.41 | — | — | |
| Qwen2.5-Coder-7B-Instruct (Majority)Backbone=Qwen2.5-Coder-7B-Instruct, Sampling strategy=Majority2026.04 | 88.1 | — | — | |
| Qwen2.5-Coder-7B-Instruct (DV)Backbone=Qwen2.5-Coder-7B-Instruct, Sampling strategy=DV2026.04 | 88.1 | — | — | |
| Qwen2.5-Coder-7B-Instruct (BoN)Backbone=Qwen2.5-Coder-7B-Instruct, Sampling strategy=BoN2026.04 | 87.9 | — | — | |
| Self RefineMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 87.8 | — | — | |
| Disagreement-Guided Strategy Routing (Qwen3-8B)Backbone=Qwen3-8B, Sampling strategy=Ours2026.04 | 87.8 | — | — | |
| Qwen2.5-Coder-7B-Instruct (SCoP)Backbone=Qwen2.5-Coder-7B-Instruct, Sampling strategy=SCoP2026.04 | 87.1 | — | — | |
| Direct IO promptingMethod Category=Prompt-based methods, Executor LLM=GPT-4o mini2026.04 | 87 | — | — | |
| Qwen2.5-Coder-7B-InstructBackbone=Qwen2.5-Coder-7B-Instruct, Sampling strategy=Base2026.04 | 86.6 | — | — | |
| Qwen3-8B (Majority)Backbone=Qwen3-8B, Sampling strategy=Majority2026.04 | 85.5 | — | — | |
| Qwen3-8B (BoN)Backbone=Qwen3-8B, Sampling strategy=BoN2026.04 | 85.4 | — | — | |
| Qwen3-8B (DV)Backbone=Qwen3-8B, Sampling strategy=DV2026.04 | 85.3 | — | — | |
| Qwen3-8BBackbone=Qwen3-8B, Sampling strategy=Base2026.04 | 84.8 | — | — | |
| MCNIGModel Size=8B2026.03 | 84.6 | — | — | |
| IGModel Size=8B2026.03 | 84.5 | — | — | |
| ORMModel Size=8B2026.03 | 84 | — | — | |
| Qwen3-8B (SCoP)Backbone=Qwen3-8B, Sampling strategy=SCoP2026.04 | 83.7 | — | — | |
| ADASMethod Category=Workflow induction methods, Optimizer LLM=Claude-3.5 Sonnet, Executor LLM=GPT-4o mini2026.04 | 82.4 | — | — | |
| BoostLoRA# Additional Params=122026.04 | 80.4 | — | — | |
| STRATAGEMModel=STRATAGEM (Ours)2026.04 | 77.93 | — | — | |
| ImplicitPRM2026.03 | 77.8 | — | — | |
| SPIRALModel=SPIRAL2026.04 | 77.44 | — | — | |
| OVM2026.03 | 76.3 | — | — | |
| SDAR-8B-b32Model Category=Block-wise dLLMs, Evaluation Protocol=Zero-shot2026.03 | 73.5 | 2.39 | 123.8 | |
| LightningRL-8B-b32Model Category=Block-wise dLLMs, Evaluation Protocol=Zero-shot2026.03 | 72.6 | 6.3 | 450.1 | |
| Base modelmode=zero-shot, # Additional Params=02026.04 | 72.6 | — | — | |
| QwenPRMModel Size=7B2026.03 | 69.7 | — | — | |
| Qwen3-4B-BaseModel=Qwen3-4B-Base2026.04 | 67.93 | — | — | |
| Qwen-2.5-7B-itModel Category=AR Models, Evaluation Protocol=Zero-shot2026.03 | 67.7 | 1 | 67.7 | |
| TinyLoRA# Additional Params=129,0242026.04 | 67.7 | — | — | |
| EAGLE-3 (LLaMA-3.1)Model Category=AR Models, Evaluation Protocol=Zero-shot2026.03 | 67.6 | 5.98 | 344.8 | |
| Majority voting2026.03 | 66.1 | — | — | |
| BaselineModel=Qwen3-8B, Precision=FP162026.05 | 64.63 | — | — | |
| TinyLoRA# Additional Params=8,0642026.04 | 64.6 | — | — | |
| TinyLoRA# Additional Params=2522026.04 | 64 | — | — | |
| TinyLoRA# Additional Params=162026.04 | 63.4 | — | — | |
| Fast-dLLM-v2Model Category=Block-wise dLLMs, Evaluation Protocol=Zero-shot2026.03 | 61.7 | 2.58 | 128.9 | |
| Granite PRM v22026.03 | 61.6 | — | — | |
| TORQModel=Qwen3-8B2026.05 | 61.15 | — | — | |
| Attn-Sampler (Parallel)Model=Fast-dLLM v2 7B2026.03 | 58.54 | — | — | |
| Attn-Sampler (Sequential)Model=Fast-dLLM v2 7B2026.03 | 57.93 | — | — | |
| Full FT# Additional Params=3.09B2026.04 | 57.9 | — | — | |
| d3LLM-DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 57.1 | 3.2 | 129.5 | |
| Single sampling2026.03 | 55.5 | — | — | |
| Entropy SamplerModel=Fast-dLLM v2 7B2026.03 | 55.49 | — | — | |
| DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 55.2 | 1 | 55.2 | |
| KLASSModel=Fast-dLLM v2 7B2026.03 | 54.88 | — | — | |
| Fast-dLLM-DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 54.3 | 1.33 | 63.5 | |
| dParallel-DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 54.3 | 2.57 | 98.8 | |
| Disagreement-Guided Strategy Routing (DS-Llama-8B)Backbone=DS-Llama-8B, Sampling strategy=Ours2026.04 | 54.3 | — | — | |
| Fast-dLLMModel=Fast-dLLM v2 7B2026.03 | 54.27 | — | — | |
| Confidence SamplerModel=Fast-dLLM v2 7B2026.03 | 54.27 | — | — | |
| QuaRotModel=Qwen3-8B2026.05 | 53.37 | — | — | |
| UnsafeChainBase Model=R1-8B, Data Subset=random2025.07 | 53.05 | — | — | |
| UnsafeChainBase Model=R1-8B, Data Subset=full2025.07 | 53.05 | — | — | |
| DS-Llama-8B (BoN)Backbone=DS-Llama-8B, Sampling strategy=BoN2026.04 | 52.8 | — | — | |
| MathShepherd2026.03 | 52.5 | — | — | |
| SelfOrg*Backbone=Qwen2.5-1.5B-Instruct, Number of agents=10, Communication=Single sequential round2026.05 | 52.03 | — | — | |
| OriginalBackbone=Qwen3-8b2025.10 | 51.83 | — | — | |
| NEXABackbone=Qwen2.5-1.5B-Instruct, Number of agents=102026.05 | 51.42 | — | — | |
| R1Model Size=8B2025.07 | 51.22 | — | — | |
| Margin SamplerModel=Fast-dLLM v2 7B2026.03 | 51.22 | — | — | |
| SingleBackbone=Qwen2.5-1.5B-Instruct, Number of agents=12026.05 | 50.41 | — | — |