Code Generation on HumanEval (Score)
93HumanEval ScoreMaAS
Evaluation Results
| Method | Links | |
|---|---|---|
| MaASCost=0.08, CP len=1810.82026.01 | 93 | |
| LAMaSCost=0.1, CP len=1042.72026.01 | 92.11 | |
| CoT*5+SCCost=0.37, CP len=952.52026.01 | 90.84 | |
| Gen-CoTCost=0.07, CP len=734.52026.01 | 90.08 | |
| GenerateCost=0.07, CP len=797.92026.01 | 88.55 | |
| ExpertWeaverBase Model=Qwen2.5-7B, Sparsity=25%2026.02 | 64.6 | |
| Qwen3-4BParams=4B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 62.2 | |
| YuLan-Mini-2.4BParams=2.4B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 61.6 | |
| FLAPBase Model=Qwen2.5-7B, Sparsity=25%2026.02 | 57.6 | |
| Qwen3-1.7BParams=1.7B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 52.44 | |
| PCMind-2.1-Kaiyuan-2BParams=2B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 42.68 | |
| Qwen2.5-3BParams=3B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 42.1 | |
| SmolLM3-3BParams=3B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 39.63 | |
| Qwen2.5-1.5BParams=1.5B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 37.2 | |
| LLaDASampling steps=256, Tokens=2562026.02 | 33.54 | |
| ExpertWeaverBase Model=Qwen2.5-7B, Sparsity=12.5%2026.02 | 32.9 | |
| LLaDA-MDLMSampling steps=256, Tokens=2562026.02 | 32.32 | |
| LLaDA-XDLMSampling steps=256, Tokens=2562026.02 | 31.71 | |
| Qwen2-1.5BParams=1.5B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 31.1 | |
| Qwen3-0.6BParams=0.6B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 29.88 | |
| Llama-3.2-3BParams=3B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 29.88 | |
| LLaDA-MDLMSampling steps=128, Tokens=2562026.02 | 25 | |
| LLaDA-XDLMSampling steps=128, Tokens=2562026.02 | 25 | |
| LLaDASampling steps=128, Tokens=2562026.02 | 24.39 | |
| SmolLM2-1.7BParams=1.7B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 22.6 | |
| Llama-3.2-1BParams=1B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 18.9 | |
| Gemma2-2BParams=2B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 17.7 | |
| LLaDA-XDLMSampling steps=64, Tokens=2562026.02 | 15.85 | |
| LLaDA-MDLMSampling steps=64, Tokens=2562026.02 | 14.63 | |
| LLM-PrunerBase Model=Qwen2.5-7B, Sparsity=25%2026.02 | 14.6 | |
| LLaDASampling steps=64, Tokens=2562026.02 | 12.8 | |
| LLaDA-XDLMSampling steps=32, Tokens=2562026.02 | 10.98 | |
| LLM-PrunerBase Model=Qwen2.5-7B, Sparsity=12.5%2026.02 | 10.4 | |
| FLAPBase Model=Qwen2.5-7B, Sparsity=12.5%2026.02 | 9.8 | |
| LLaDA-XDLM-inferSampling steps=128, Tokens=2562026.02 | 8.54 | |
| LLaDA-MDLMSampling steps=32, Tokens=2562026.02 | 6.71 | |
| LLaDA-XDLM-inferSampling steps=256, Tokens=2562026.02 | 6.71 | |
| OLMo-2-0425-1BParams=1B, Shot setting=3 shot, Evaluation mode=generation2025.12 | 6.71 | |
| LLaDASampling steps=32, Tokens=2562026.02 | 5.49 | |
| LLaDA-XDLM-inferSampling steps=64, Tokens=2562026.02 | 4.88 | |
| LLaDA-XDLM-inferSampling steps=32, Tokens=2562026.02 | 1.22 | |
| Qwen3-Coder-30B-A3B-Instruct (Base)Model=Qwen3-Coder-30B-A3B-Instruct, Strategy=Base2026.02 | 0.9451 | |
| Qwen3-4B-Instruct-2507 (C/E Weighted)Model=Qwen3-4B-Instruct-2507, Strategy=C/E Weighted2026.02 | 0.9146 | |
| Qwen3-4B-Instruct-2507 (Complex Only)Model=Qwen3-4B-Instruct-2507, Strategy=Complex Only2026.02 | 0.9085 | |
| Qwen3-4B-Instruct-2507 (B/I Weighted)Model=Qwen3-4B-Instruct-2507, Strategy=B/I Weighted2026.02 | 0.8963 | |
| Qwen3-4B-Instruct-2507 (C/E Weighted (Rev))Model=Qwen3-4B-Instruct-2507, Strategy=C/E Weighted (Rev)2026.02 | 0.8963 | |
| Qwen3-4B-Instruct-2507 (Basic Only)Model=Qwen3-4B-Instruct-2507, Strategy=Basic Only2026.02 | 0.8963 | |
| Qwen3-4B-Instruct-2507 (Edge Only)Model=Qwen3-4B-Instruct-2507, Strategy=Edge Only2026.02 | 0.8963 | |
| Qwen3-4B-Instruct-2507 (Base)Model=Qwen3-4B-Instruct-2507, Strategy=Base2026.02 | 0.8902 | |
| Qwen3-4B-Instruct-2507 (Uniform)Model=Qwen3-4B-Instruct-2507, Strategy=Uniform2026.02 | 0.8841 |