Code Generation on HumanEval+ (Pass@1)
86Pass@1Qwen2.5-Coder-7B-Inst
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-Coder-7B-InstRole=Teacher Model2026.04 | 86 | |
| BASEBase model=Qwen2.5-Coder-7B-Instruct2026.06 | 83.1 | |
| OriginalBackbone=Qwen3-4B-Instruct2026.06 | 82.93 | |
| FULL TOKENSBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=100.02026.06 | 82.9 | |
| CODEBLOCKBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.92026.06 | 82.9 | |
| CLAMBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=3.52026.06 | 81.7 | |
| Weight Averaging w/ HARCMerging Strategy=Weight Averaging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 81.55 | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.92026.06 | 81.3 | |
| IndividualModel Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 81.29 | |
| TIES-MergingMerging Strategy=TIES-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 81.1 | |
| DS2Base model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=4.62026.06 | 81.1 | |
| TIES-Merging w/ HARCMerging Strategy=TIES-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 80.76 | |
| WUDI-Merging w/ HARCMerging Strategy=WUDI-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 80.61 | |
| CODEBLOCKBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.92026.06 | 80.5 | |
| Weight AveragingMerging Strategy=Weight Averaging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 80.3 | |
| WUDI-MergingMerging Strategy=WUDI-Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 80.22 | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-7B-Instruct, Eff. Tokens (%)=1.72026.06 | 79.9 | |
| DARE w/ HARCMerging Strategy=DARE, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 79.85 | |
| BASEBase model=Qwen2.5-Coder-3B-Instruct2026.06 | 79.3 | |
| RANDOM SELECTIONBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.92026.06 | 79.3 | |
| DAREMerging Strategy=DARE, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 79.15 | |
| PriFT-massBackbone=Qwen3-4B-Instruct2026.06 | 78.66 | |
| OriginalBackbone=Qwen2.5-Coder-3B2026.06 | 78.66 | |
| Fisher Merging w/ HARCMerging Strategy=Fisher Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=true2026.06 | 78.58 | |
| Fisher MergingMerging Strategy=Fisher Merging, Model Architecture=Qwen3-30B-A3B, HARC Calibration=false2026.06 | 78.35 | |
| PriFT-massBackbone=Qwen2.5-Coder-3B2026.06 | 77.44 | |
| PriFT-probBackbone=Qwen3-4B-Instruct2026.06 | 76.83 | |
| DS2Base model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=4.62026.06 | 76.8 | |
| CLAMBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=3.52026.06 | 76.8 | |
| TOKEN CLEANINGBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=1.72026.06 | 75.6 | |
| Qwen3Architecture=Dense, # Total Params=8.0B, # Trained Tokens=36T2025.10 | 75.3 | |
| PriFT-probBackbone=Qwen2.5-Coder-3B2026.06 | 75 | |
| FULL TOKENSBase model=Qwen2.5-Coder-3B-Instruct, Eff. Tokens (%)=100.02026.06 | 75 | |
| ASFTBackbone=Qwen2.5-Coder-3B2026.06 | 73.78 | |
| EAFTBackbone=Qwen2.5-Coder-3B2026.06 | 73.17 | |
| IDFTBackbone=Qwen2.5-Coder-3B2026.06 | 71.34 | |
| ASFTBackbone=Qwen3-4B-Instruct2026.06 | 70.73 | |
| Qwen3Architecture=Dense, # Total Params=4.0B, # Trained Tokens=36T2025.10 | 70.7 | |
| Ouro 2.6B R4Architecture=LoopLM, # Total Params=2.6B, # Trained Tokens=7.7T2025.10 | 70.7 | |
| Qwen3Model Version=4B, Architecture=Dense, # Params=4.0B, # Tokens=36T2025.10 | 70.7 | |
| Qwen2.5Architecture=Dense, # Total Params=7.0B, # Trained Tokens=18T2025.10 | 70.6 | |
| SFTBackbone=Qwen2.5-Coder-3B2026.06 | 70.12 | |
| DFTBackbone=Qwen2.5-Coder-3B2026.06 | 70.12 | |
| TALRBackbone=Qwen2.5-Coder-3B2026.06 | 70.12 | |
| SFTBackbone=Qwen3-4B-Instruct2026.06 | 68.29 | |
| EAFTBackbone=Qwen3-4B-Instruct2026.06 | 68.29 | |
| OuroModel Version=1.4B R4, Architecture=LoopLM, # Params=1.4B, # Tokens=7.7T, recurrent steps=42025.10 | 67.4 | |
| ExpGraphBackbone Model=Llama-3.1-8B-Instruct (Large LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 66 | |
| IDFTBackbone=Qwen3-4B-Instruct2026.06 | 65.85 | |
| Qwen2.5-Coder-1.5BRole=Student Model2026.04 | 64 | |
| ABKDTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 64 | |
| VCRDTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 64 | |
| GKDTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 62.8 | |
| DistilLLMTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 62.8 | |
| Qwen2.5Architecture=Dense, # Total Params=3.0B, # Trained Tokens=18T2025.10 | 62.2 | |
| Qwen2.5Model Version=3B, Architecture=Dense, # Params=3.0B, # Tokens=18T2025.10 | 62.2 | |
| DistillLM-2Teacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 61.6 | |
| KDTeacher Model=Qwen2.5-Coder-7B-Inst, Student Model=Qwen2.5-Coder-1.5B2026.04 | 61 | |
| ISTBase Model=Qwen3-8B, Params=334M2026.05 | 61 | |
| TALRBackbone=Qwen3-4B-Instruct2026.06 | 60.98 | |
| DFTBackbone=Qwen3-4B-Instruct2026.06 | 60.37 | |
| DomLoRABase Model=Qwen3-8B, Params=2.1M2026.05 | 59.8 | |
| Qwen3Model Version=1.7B, Architecture=Dense, # Params=1.7B, # Tokens=36T2025.10 | 59.8 | |
| PLoPBase Model=Qwen3-8B, Params=131M2026.05 | 58.5 | |
| LoRABase Model=Qwen3-8B, Params=334M2026.05 | 56.1 | |
| ExpGraphBackbone Model=Llama-3.2-3B-Instruct (Small LLM), Method Category=LLM-Centric Experience Learning Baselines2026.05 | 56 | |
| ME-DLM Stage-3Budget=1/12026.05 | 53 | |
| DomLoRABase Model=LLaMA-3.1-8B, Params=2.3M2026.05 | 50 | |
| Soft MaskBudget=1/12026.05 | 50 | |
| ME-DLM Stage-2Budget=1/12026.05 | 50 | |
| ME-DLM Stage-3Budget=1/22026.05 | 46.3 | |
| Qwen2.5Model Version=1.5B, Architecture=Dense, # Params=1.5B, # Tokens=18T2025.10 | 46.3 | |
| ME-DLM Stage-2Budget=1/22026.05 | 42.7 | |
| ESPOSequence Length=5122026.05 | 42.7 | |
| LoRABase Model=LLaMA-3.1-8B, Params=321M2026.05 | 42.1 | |
| ISTBase Model=LLaMA-3.1-8B, Params=321M2026.05 | 42.1 | |
| LLaDASequence Length=5122026.05 | 41.5 | |
| GDSDSequence Length=5122026.05 | 41.5 | |
| PLoPBase Model=LLaMA-3.1-8B, Params=125M2026.05 | 40.2 | |
| GDSD w/ TLCSequence Length=2562026.05 | 39.6 | |
| GDSD w/ TLCSequence Length=5122026.05 | 39.6 | |
| GDSD w/ TLCSequence Length=Avg.2026.05 | 39.2 | |
| ME-DLM Stage-3Budget=1/42026.05 | 39 | |
| GDSDSequence Length=Avg.2026.05 | 38.6 | |
| LLaDA-InstructBudget=1/22026.05 | 38.4 | |
| GDSDSequence Length=2562026.05 | 38.4 | |
| GDSD w/ TLCSequence Length=1282026.05 | 38.4 | |
| LLaDA-InstructBudget=1/12026.05 | 37.8 | |
| SPGSequence Length=5122026.05 | 37.8 | |
| diffu-GRPO (d1)Sequence Length=5122026.05 | 37.2 | |
| Gemma3Architecture=Dense, # Total Params=12.0B, # Trained Tokens=12T2025.10 | 37.2 | |
| ESPOSequence Length=2562026.05 | 36.6 | |
| GDSDSequence Length=1282026.05 | 36 | |
| SPGSequence Length=Avg.2026.05 | 35 | |
| ESPOSequence Length=Avg.2026.05 | 34.6 | |
| SPGSequence Length=2562026.05 | 34.2 | |
| Soft MaskBudget=1/22026.05 | 33.8 | |
| ME-DLM Stage-2Budget=1/42026.05 | 32.9 | |
| wd1Sequence Length=5122026.05 | 32.9 | |
| SPGSequence Length=1282026.05 | 32.9 |