Code on HumanEval (Accuracy)
96.34HumanEval AccuracyQwen3.5-9B + AR-SFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3.5-9B + AR-SFTModel Size=9B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 96.34 | — | |
| Qwen3.5-9B (released)Model Size=9B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 95.12 | — | |
| Qwen3-Next-80B-A3B#Token=6002026.04 | 95.1 | — | |
| JoyAI-LLM Flash#Token=9002026.04 | 94.5 | — | |
| Qwen3.5-35B-A3B#Token=3002026.04 | 93.9 | — | |
| GLM-4.7-Flash-T#Token=72002026.04 | 93.9 | — | |
| GPT-5Evaluation Protocol=Closed-Source2026.01 | 93.4 | — | |
| FLARE-4BModel Size=4B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 93.29 | — | |
| Qwen3.5-4B + AR-SFTModel Size=4B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 92.68 | — | |
| GPT-4.1Evaluation Protocol=Closed-Source2026.01 | 92.1 | — | |
| Qwen3-30B-A3B#Token=4002026.04 | 92.1 | — | |
| FLARE-9BModel Size=9B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 92.07 | — | |
| ATLAS (cluster)Evaluation Protocol=In-Distribution2026.01 | 91.5 | — | |
| ATLAS (cluster)Evaluation Protocol=Out-of-Distribution2026.01 | 91.5 | — | |
| Qwen3-14BModel Scale=14B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 88.41 | — | |
| Qwen3-14B + NGMModel Scale=14B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 88.41 | — | |
| SUNBackbone=Qwen3-14B-Base, Decode Execution=Shared2026.03 | 88.4 | — | |
| Qwen3-30B-A3B-Baseshots=5-shot, Decoding=Greedy2026.04 | 87.8 | — | |
| Qwen3.5-4B (released)Model Size=4B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 87.8 | — | |
| Qwen3-8B + NGMModel Scale=8B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 86.59 | — | |
| GPT-4oEvaluation Protocol=Closed-Source2026.01 | 85.4 | — | |
| ATLAS (RL)Evaluation Protocol=Out-of-Distribution2026.01 | 85.4 | — | |
| Qwen3-8BModel Scale=8B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 85.37 | — | |
| Full-FTBackbone=Qwen3-14B-Base, Decode Execution=Independent2026.03 | 85.3 | — | |
| JoyAI-LLM Flash-Baseshots=5-shot, Decoding=Greedy2026.04 | 85.3 | — | |
| SDAR-30B-A3Bgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 84.15 | — | |
| SUNBackbone=Qwen3-8B-Base, Decode Execution=Shared2026.03 | 84.1 | — | |
| Full-FTBackbone=Qwen3-8B-Base, Decode Execution=Independent2026.03 | 83.5 | — | |
| Gemini2.5-ProEvaluation Protocol=Closed-Source2026.01 | 81.5 | — | |
| LLaDA2.0-minigeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 81.1 | — | |
| LLaDA2.1-minigeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 81.1 | — | |
| Qwen3-4B + NGMModel Scale=4B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 81.1 | — | |
| RouterDCEvaluation Protocol=In-Distribution2026.01 | 80.5 | — | |
| Qwen3-4BModel Scale=4B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 80.49 | — | |
| SDAR-8B-Chatgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 79.88 | — | |
| Qwen3.5-35B-A3B-Baseshots=5-shot, Decoding=Greedy2026.04 | 79.8 | — | |
| N-3-Super 120B-A12B-BaseShots=0, Sample size (n)=322026.04 | 79.4 | — | |
| RouterDCEvaluation Protocol=Out-of-Distribution2026.01 | 79.2 | — | |
| AdaRASCategory=Steering2026.01 | 79.19 | — | |
| BertRouterEvaluation Protocol=Out-of-Distribution2026.01 | 78.7 | — | |
| Dream-7B-Instructgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 78.05 | — | |
| ProbingCategory=Steering2026.01 | 77.85 | — | |
| BaselineModel=LLaDA2.0-mini2026.05 | 77.44 | — | |
| ElasticModel=LLaDA2.0-mini2026.05 | 77.44 | — | |
| CoTPrompting=Vanilla CoT, Base Model=Qwen3-1.7B2026.01 | 77.18 | — | |
| GLM-4.5 Air-BaseShots=0, Sample size (n)=322026.04 | 76.3 | — | |
| MLPRouterEvaluation Protocol=In-Distribution2026.01 | 76.2 | — | |
| BertRouterEvaluation Protocol=In-Distribution2026.01 | 75.4 | — | |
| MLPRouterEvaluation Protocol=Out-of-Distribution2026.01 | 75 | — | |
| GRPO (RLVRR)#Data=10K2026.01 | 73 | — | |
| GRPO (BLEU)#Data=10K2026.01 | 72.8 | — | |
| Instruct#Data=n/a2026.01 | 72.6 | — | |
| GRPO (RLPR)#Data=10K2026.01 | 72.3 | — | |
| DPO#Data=10K2026.01 | 72.2 | — | |
| GRPO (RM)#Data=10K2026.01 | 72.1 | — | |
| GRPO (GRM)#Data=10K2026.01 | 70.9 | — | |
| SFT#Data=100K2026.01 | 70.8 | — | |
| Yuan3.0-1T Base#Shots=0-shot, Architecture=MoE, # activated params=68.5B, # total params=1010B2026.01 | 70.7 | — | |
| Ling-flash base-2.0Shots=0, Sample size (n)=322026.04 | 70.1 | — | |
| SFT#Data=10K2026.01 | 69.9 | — | |
| FS RouterEvaluation Protocol=Training-free, Prompting Strategy=Few-shot2026.01 | 68.9 | — | |
| Qwen3-14B-BaseDecode Execution=N/A2026.03 | 68.9 | — | |
| Qwen3-8B-BaseDecode Execution=N/A2026.03 | 68.3 | — | |
| Full-FTBackbone=Qwen3-1.7B-Base, Decode Execution=Independent2026.03 | 67.1 | — | |
| SUNBackbone=Qwen3-1.7B-Base, Decode Execution=Shared2026.03 | 67.1 | — | |
| Qwen3.5-2B + AR-SFTModel Size=2B, Training Protocol=AR-SFT, Decoding Strategy=Native AR decoding2026.06 | 67.07 | — | |
| GRPO (Random)#Data=10K2026.01 | 66.8 | — | |
| DeepSeek-V3-Base#Shots=0-shot, Architecture=MoE, # activated params=37B, # total params=671B2026.01 | 65.2 | — | |
| FLARE-2BModel Size=2B, Training Protocol=FLARE, Decoding Strategy=AR-Trust sampling2026.06 | 64.02 | — | |
| Qwen3-1.7B + NGMModel Scale=1.7B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 62.2 | — | |
| Qwen3-1.7BModel Scale=1.7B, NGM Configuration=False, Decoding Settings=identical decoding settings2026.05 | 60.37 | — | |
| OpenThinker-3-1.5BCategory=Post-training2026.01 | 59.06 | — | |
| OpenReasoning-Nemotron-1.5BCategory=Post-training2026.01 | 57.72 | — | |
| LLaMA-3.1-405B Base#Shots=0-shot, Architecture=Dense, # activated params=405B, # total params=405B2026.01 | 54.9 | — | |
| AtteNTModel=Gemma-7B2026.02 | 54.26 | 2,012 | |
| Standard Fine-tuningModel=Gemma-7B2026.02 | 53.83 | 2,282 | |
| ZS RouterEvaluation Protocol=Training-free, Prompting Strategy=Zero-shot2026.01 | 53 | — | |
| WINASparsity=40%, Backbone=Phi-4-14B2025.05 | 53 | — | |
| WINASparsity=50%, Backbone=Phi-4-14B2025.05 | 51.83 | — | |
| Baseline (full model)Sparsity=0%, Backbone=Phi-4-14B2025.05 | 50.61 | — | |
| WINASparsity=25%, Backbone=Phi-4-14B2025.05 | 50 | — | |
| SUNBackbone=LLaMA3.1-8B, Decode Execution=Shared2026.03 | 49.4 | — | |
| Full-FTBackbone=LLaMA3.1-8B, Decode Execution=Independent2026.03 | 48.2 | — | |
| Qwen3-1.7B-BaseDecode Execution=N/A2026.03 | 48.2 | — | |
| Qwen3.5-2B (released)Model Size=2B, Training Protocol=Released, Decoding Strategy=Native AR decoding2026.06 | 48.17 | — | |
| TEALSparsity=25%, Backbone=Phi-4-14B2025.05 | 46.95 | — | |
| AtteNTModel=Mistral-7B2026.02 | 46.55 | 1,802 | |
| LLaDA-8B-Instructgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 46.34 | — | |
| TEALSparsity=40%, Backbone=Phi-4-14B2025.05 | 45.73 | — | |
| DeepSeek-R1-Distill-Qwen-1.5BCategory=Post-training2026.01 | 45.64 | — | |
| Standard Fine-tuningModel=Mistral-7B2026.02 | 43.42 | 2,042 | |
| TEALSparsity=50%, Backbone=Phi-4-14B2025.05 | 41.46 | — | |
| WINASparsity=65%, Backbone=Phi-4-14B2025.05 | 41.46 | — | |
| DARTBase Model=Qwen3-0.6B2026.05 | 38.11 | — | |
| Random RouterEvaluation Protocol=Training-free2026.01 | 37.8 | — | |
| BaseBase Model=Qwen3-0.6B2026.05 | 37.65 | — | |
| GRPOBase Model=Qwen3-0.6B2026.05 | 37.5 | — | |
| Qwen3-0.6B + NGMModel Scale=0.6B, NGM Configuration=True, Decoding Settings=identical decoding settings2026.05 | 37.2 | — | |
| STaRBase Model=Qwen3-0.6B2026.05 | 37.2 | — | |
| LLaMA3.1-8BDecode Execution=N/A2026.03 | 36.6 | — |