Code Generation on HumanEval (HumanEval and Average metrics)
95.22HumanEval ScoreSIGMA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SIGMAMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=full, Backbone=DeepSeek-V3.22026.05 | 95.22 | 93.91 | |
| G-DesignerMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 94.82 | 93.22 | |
| FinTool-Qwen3-14BModel Variant=domain-tuned2026.03 | 94.51 | — | |
| GoAMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=partial, Backbone=DeepSeek-V3.22026.05 | 94.25 | 92.86 | |
| GPTSwarmMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=full, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 94.16 | 92.39 | |
| MoAMulti-agent collaboration (Mul.)=full, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 94.03 | 89.83 | |
| CoTMulti-agent collaboration (Mul.)=no, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 93.41 | 90.68 | |
| ComplexCoTMulti-agent collaboration (Mul.)=no, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 93.14 | 91.01 | |
| VanillaMulti-agent collaboration (Mul.)=no, Inter-agent relations (Rel.)=no, Conflicting signals (Conf.)=no, Backbone=DeepSeek-V3.22026.05 | 93.13 | 87.69 | |
| Qwen3-14BModel Variant=base2026.03 | 92.68 | — | |
| Qwen3-8BModel Variant=base2026.03 | 89.63 | — | |
| FinTool-Qwen3-8BModel Variant=domain-tuned2026.03 | 89.02 | — | |
| AlphaTokenModel=Qwen-3.5-9B2026.06 | 78.88 | 78.88 | |
| Standard FTModel=Qwen-3.5-9B2026.06 | 75.94 | 75.94 | |
| ssTOKEN ICLR 2026Model=Qwen-3.5-9B2026.06 | 75.21 | 75.21 | |
| XTF ICLR 2026Model=Qwen-3.5-9B2026.06 | 73.61 | 73.61 | |
| Token Cleaning ICML 2025Model=Qwen-3.5-9B2026.06 | 73.58 | 73.58 | |
| LESS ICML 2024Model=Qwen-3.5-9B2026.06 | 72.82 | 72.82 | |
| STM NeurIPS 2025Model=Qwen-3.5-9B2026.06 | 71.93 | 71.93 | |
| LoRA ICLR 2022Model=Qwen-3.5-9B2026.06 | 70.2 | 70.2 | |
| AlphaTokenModel=Gemma-3-4B2026.06 | 62.15 | 62.15 | |
| Pre-trainedModel=Qwen-3.5-9B2026.06 | 60.96 | 60.96 | |
| Standard FTModel=Gemma-3-4B2026.06 | 58.79 | 58.79 | |
| ssTOKEN ICLR 2026Model=Gemma-3-4B2026.06 | 57.45 | 57.45 | |
| LESS ICML 2024Model=Gemma-3-4B2026.06 | 57.17 | 57.17 | |
| Token Cleaning ICML 2025Model=Gemma-3-4B2026.06 | 57.12 | 57.12 | |
| XTF ICLR 2026Model=Gemma-3-4B2026.06 | 56.64 | 56.64 | |
| LightMoERatio=30%, # Params=0.45B2026.03 | 55.7 | 55.3 | |
| OriginalRatio=0%, # Params=-2026.03 | 55 | 37.7 | |
| LoRA ICLR 2022Model=Gemma-3-4B2026.06 | 54.62 | 54.62 | |
| LoRARatio=0%, # Params=0.45B2026.03 | 54.2 | 55.5 | |
| STM NeurIPS 2025Model=Gemma-3-4B2026.06 | 54.01 | 54.01 | |
| Replace (w/o shared)Ratio=30%, # Params=0.45B2026.03 | 53.8 | 52.6 | |
| Replace (w. shared)Ratio=30%, # Params=0.45B2026.03 | 52 | 52.3 | |
| MC-SMoERatio=30%, # Params=0.45B2026.03 | 51.7 | 53.4 | |
| Replace (w/o shared)Ratio=40%, # Params=0.45B2026.03 | 51.3 | 50.1 | |
| LightMoERatio=40%, # Params=0.45B2026.03 | 51.3 | 53 | |
| MC-SMoE*Ratio=40%, # Params=1.65B2026.03 | 50.6 | 50.2 | |
| Replace (w. shared)Ratio=40%, # Params=0.45B2026.03 | 50.1 | 49.3 | |
| Full FTRatio=0%, # Params=6.92B2026.03 | 49.7 | 56.3 | |
| MC-SMoERatio=40%, # Params=0.45B2026.03 | 48.4 | 47.3 | |
| LightMoERatio=50%, # Params=0.45B2026.03 | 48.4 | 48.1 | |
| Replace (w/o shared)Ratio=50%, # Params=0.45B2026.03 | 46.3 | 44.3 | |
| Standard FTModel=Llama-3.2-3B2026.06 | 44.6 | 44.6 | |
| Replace (w. shared)Ratio=50%, # Params=0.45B2026.03 | 44.5 | 41.6 | |
| AlphaTokenModel=Llama-3.2-3B2026.06 | 43.98 | 43.98 | |
| MC-SMoE*Ratio=50%, # Params=1.65B2026.03 | 43.8 | 45.3 | |
| MC-SMoERatio=50%, # Params=0.45B2026.03 | 43.1 | 42.5 | |
| ssTOKEN ICLR 2026Model=Llama-3.2-3B2026.06 | 42.93 | 42.93 | |
| LESS ICML 2024Model=Llama-3.2-3B2026.06 | 42.24 | 42.24 | |
| XTF ICLR 2026Model=Llama-3.2-3B2026.06 | 42.02 | 42.02 | |
| HC-SMoERatio=30%, # Params=0.45B2026.03 | 41.7 | 46.8 | |
| STM NeurIPS 2025Model=Llama-3.2-3B2026.06 | 41.54 | 41.54 | |
| LoRA ICLR 2022Model=Llama-3.2-3B2026.06 | 41.46 | 41.46 | |
| MoBERatio=30%, # Params=0.45B2026.03 | 41.3 | 50.4 | |
| Token Cleaning ICML 2025Model=Llama-3.2-3B2026.06 | 40.93 | 40.93 | |
| Pre-trainedModel=Gemma-3-4B2026.06 | 35.36 | 35.36 | |
| BaselineBase Model=Qwen1.5-MoE, Weight Bits=162026.04 | 34.76 | 46.1 | |
| HC-SMoERatio=40%, # Params=0.45B2026.03 | 33.5 | 38.4 | |
| Instruction TunedModel Category=Base Models2026.04 | 30.49 | — | |
| DARE TiesExpert Combination=Code + Instruction2026.04 | 30.49 | — | |
| Precise FusionExpert Combination=Code + Instruction2026.04 | 30.49 | — | |
| Task ArithmeticExpert Combination=Code + Instruction2026.04 | 29.27 | — | |
| Pre-trainedModel=Llama-3.2-3B2026.06 | 28.66 | 28.66 | |
| MoBERatio=40%, # Params=0.45B2026.03 | 28.6 | 43 | |
| Avg BaselineExpert Combination=Code + Instruction2026.04 | 28.05 | — | |
| Ties MergingExpert Combination=Instruction + Math2026.04 | 28.05 | — | |
| MoBiEBase Model=Qwen1.5-MoE, Weight Bits=1.352026.04 | 27.12 | 37.54 | |
| GPTQBase Model=Qwen1.5-MoE, Weight Bits=32026.04 | 27.05 | 37.48 | |
| DARE Task ArithmeticExpert Combination=Code + Instruction2026.04 | 26.83 | — | |
| HC-SMoERatio=50%, # Params=0.45B2026.03 | 26.5 | 34 | |
| WIDENExpert Combination=Code + Instruction2026.04 | 26.22 | — | |
| MoBiEModel=DeepSeekMoE-16B-Chat, #Bits(W)=1.462026.04 | 25.66 | — | |
| WIDENExpert Combination=Code + Instruction + Math2026.04 | 25 | — | |
| BaselineModel=DeepSeekMoE-16B-Chat, #Bits(W)=162026.04 | 24.39 | — | |
| WIDENExpert Combination=Instruction + Math2026.04 | 24.39 | — | |
| ARB-LLMBase Model=Qwen1.5-MoE, Weight Bits=1.112026.04 | 23.86 | 30.6 | |
| Avg BaselineExpert Combination=Code + Instruction + Math2026.04 | 23.17 | — | |
| Precise FusionExpert Combination=Instruction + Math2026.04 | 22.56 | — | |
| Precise FusionExpert Combination=Code + Instruction + Math2026.04 | 22.56 | — | |
| BiLLMBase Model=Qwen1.5-MoE, Weight Bits=1.112026.04 | 22.12 | 28.8 | |
| BaselineModel=Qwen-MoE-14B-Chat, #Bits(W)=162026.04 | 21.34 | — | |
| Ties MergingExpert Combination=Code + Instruction + Math2026.04 | 21.34 | — | |
| Avg BaselineExpert Combination=Instruction + Math2026.04 | 20.73 | — | |
| NoWagBase Model=Qwen1.5-MoE, Weight Bits=2.042026.04 | 20.23 | 26.94 | |
| CodeModel Category=Base Models2026.04 | 19.51 | — | |
| ARB-LLMModel=DeepSeekMoE-16B-Chat, #Bits(W)=1.112026.04 | 19.31 | — | |
| BaseModel Category=Base Models2026.04 | 18.9 | — | |
| Task ArithmeticExpert Combination=Code + Math2026.04 | 18.29 | — | |
| MoEQuantBase Model=Qwen1.5-MoE, Weight Bits=22026.04 | 18.12 | 22.56 | |
| Precise FusionExpert Combination=Code + Math2026.04 | 17.68 | — | |
| BiLLMModel=DeepSeekMoE-16B-Chat, #Bits(W)=1.112026.04 | 17.2 | — | |
| DARE TiesExpert Combination=Code + Math2026.04 | 17.07 | — | |
| WIDENExpert Combination=Code + Math2026.04 | 17.07 | — | |
| Ties MergingExpert Combination=Code + Instruction2026.04 | 16.46 | — | |
| DARE Task ArithmeticExpert Combination=Code + Math2026.04 | 16.46 | — | |
| Avg BaselineExpert Combination=Code + Math2026.04 | 15.85 | — | |
| Ties MergingExpert Combination=Code + Math2026.04 | 15.85 | — | |
| MoBiEModel=Qwen-MoE-14B-Chat, #Bits(W)=1.352026.04 | 15.75 | — | |
| MathModel Category=Base Models2026.04 | 14.02 | — |