Code on MBPP (Pass@1)
91.05Pass@1FLARE-9B
Evaluation Results
| Method | Links | |
|---|---|---|
| FLARE-9BParams=9B, Sampling mode=AR-Trust2026.06 | 91.05 | |
| Qwen3.5-9BParams=9B, Sampling mode=AR2026.06 | 89.11 | |
| FLAREParams=4B, Sampling mode=AR-Trust2026.06 | 89.11 | |
| LLaDA-2.0-flashParams=100B-A5B, Sampling mode=Diffusion2026.06 | 88.29 | |
| Qwen3.5Params=4B, Sampling mode=AR2026.06 | 82.49 | |
| FLARE-9BParams=9B, Sampling mode=Diffusion-Trust2026.06 | 82.1 | |
| LLaDA-2.0-miniParams=16B-A1B, Sampling mode=Diffusion2026.06 | 81.5 | |
| PrivCodeBackbone=Qwen2.5-Coder-7B, Privacy Epsilon=42025.12 | 77.9 | |
| FLAREParams=4B, Sampling mode=Diffusion-Trust2026.06 | 77.82 | |
| NonDPFTBackbone=Qwen2.5-Coder-7B, Privacy Epsilon=None2025.12 | 77.2 | |
| NonDPFTBackbone=CodeQwen1.5-7B, Privacy Epsilon=None2025.12 | 74.3 | |
| NonDPFTBackbone=DS-Coder-6.7B, Privacy Epsilon=None2025.12 | 73.8 | |
| PrivCodeBackbone=CodeQwen1.5-7B, Privacy Epsilon=42025.12 | 72.5 | |
| SDARParams=8B, Sampling mode=Diffusion2026.06 | 72 | |
| SDAR-8B-Chatgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 71.6 | |
| SDAR 30B-A3BParams=30B-A3B, Sampling mode=Diffusion2026.06 | 71.6 | |
| DP-AdapterBackbone=CodeQwen1.5-7B, Privacy Epsilon=42025.12 | 69.6 | |
| PrivCodeBackbone=DS-Coder-6.7B, Privacy Epsilon=42025.12 | 69 | |
| JFTBackbone=CodeQwen1.5-7B, Privacy Epsilon=42025.12 | 68.1 | |
| FLAREParams=2B, Sampling mode=AR-Trust2026.06 | 68.09 | |
| JFTBackbone=Qwen2.5-Coder-7B, Privacy Epsilon=42025.12 | 67.6 | |
| Qwen3Size=4B, Type=Base, Prompt Setting=3-shot2025.12 | 67.5 | |
| Youtu-LLMSize=2B, Type=Base, Prompt Setting=3-shot2025.12 | 66.6 | |
| JFTBackbone=DS-Coder-6.7B, Privacy Epsilon=42025.12 | 66.4 | |
| DPFTBackbone=CodeQwen1.5-7B, Privacy Epsilon=42025.12 | 66.4 | |
| PrivCodeBackbone=CodeGemma-7B, Privacy Epsilon=42025.12 | 66.1 | |
| DPFTBackbone=DS-Coder-6.7B, Privacy Epsilon=42025.12 | 65.8 | |
| SDARParams=4B, Sampling mode=Diffusion2026.06 | 65.4 | |
| NonDPFTBackbone=CodeGemma-7B, Privacy Epsilon=None2025.12 | 64.8 | |
| LLaDA2.0-minigeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 64.8 | |
| DP-AdapterBackbone=DS-Coder-6.7B, Privacy Epsilon=42025.12 | 62.9 | |
| LLaDA2.1-minigeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 62.6 | |
| DP-AdapterBackbone=Qwen2.5-Coder-7B, Privacy Epsilon=42025.12 | 61.4 | |
| SDARParams=1.7B, Sampling mode=Diffusion2026.06 | 61.1 | |
| Llama3.1-InstructCorrector Sampling=false, Number of shots=32026.02 | 57.8 | |
| Dream-7B-Instructgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 56.4 | |
| DPFTBackbone=Qwen2.5-Coder-7B, Privacy Epsilon=42025.12 | 56.3 | |
| Qwen3Size=1.7B, Type=Base, Prompt Setting=3-shot2025.12 | 55.6 | |
| FLAREParams=2B, Sampling mode=Diffusion-Trust2026.06 | 55.25 | |
| Qwen3.5Params=2B, Sampling mode=AR2026.06 | 53.31 | |
| SDAR-30B-A3Bgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 52 | |
| JFTBackbone=CodeGemma-7B, Privacy Epsilon=42025.12 | 51.7 | |
| SmolLM3Size=3B, Type=Base, Prompt Setting=3-shot2025.12 | 51 | |
| ProSeCo SamplingCorrector Sampling=true, Number of shots=32026.02 | 50.2 | |
| Llama3.1Size=8B, Type=Base, Prompt Setting=3-shot2025.12 | 49.4 | |
| DP-AdapterBackbone=CodeGemma-7B, Privacy Epsilon=42025.12 | 48.8 | |
| Engram-27BShots=3-shot2026.01 | 48.2 | |
| MoE-27BShots=3-shot2026.01 | 46.6 | |
| Engram-40BShots=3-shot2026.01 | 46.2 | |
| Gemma3Size=4B, Type=Base, Prompt Setting=3-shot2025.12 | 45.8 | |
| DPFTBackbone=CodeGemma-7B, Privacy Epsilon=42025.12 | 45 | |
| ProSeCo SFTCorrector Sampling=false, Number of shots=32026.02 | 44 | |
| Vanilla SFTCorrector Sampling=false, Training strategy=SFT, Number of shots=32026.02 | 43.2 | |
| Vanilla SFT + ReMDMCorrector Sampling=true, Training strategy=SFT, Number of shots=32026.02 | 42.4 | |
| LLaDA-BaseCorrector Sampling=false, Number of shots=32026.02 | 40.4 | |
| dLLM-VarToken budget=256, Number of shots=32026.04 | 40.2 | |
| VSBToken budget=256, Number of shots=3, Backbone=LLaDA-1.52026.04 | 39.8 | |
| LLaDA-8B-Instructgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 38.8 | |
| VSBToken budget=256, Number of shots=3, Backbone=LLaDA-8B2026.04 | 38.6 | |
| AdaBlockToken budget=256, Number of shots=32026.04 | 37.6 | |
| DAEDALToken budget=256, Number of shots=32026.04 | 36 | |
| Blockwise-SFTToken budget=256, Number of shots=32026.04 | 35.6 | |
| LLaDA-1.5Token budget=256, Number of shots=32026.04 | 35.6 | |
| Dense-4BShots=3-shot2026.01 | 35.4 | |
| LLaDA-Instruct + ReMDMCorrector Sampling=true, Number of shots=32026.02 | 35.2 | |
| LLaDA-8BToken budget=256, Number of shots=32026.04 | 35.2 | |
| D2FToken budget=256, Number of shots=32026.04 | 34.6 | |
| LLaDA-Instruct + PRISMCorrector Sampling=true, Number of shots=32026.02 | 32.3 | |
| LLaDA-InstructCorrector Sampling=false, Number of shots=32026.02 | 29.4 | |
| LLaDA1.5Corrector Sampling=false, Number of shots=32026.02 | 27.2 | |
| CDLMRefinement steps=4, Confidence threshold=0.92025.12 | 17.5 | |
| MDLMRefinement steps=4, Confidence threshold=0.92025.12 | 11.4 | |
| BaseRefinement steps=4, Confidence threshold=0.92025.12 | 11.4 |