Code Generation on MBPP (tau, Speedup)
7.68SpeedupDFlash+DDTree
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DFlash+DDTreeModel=Qwen3-Coder-30B-A3B-Instruct, Temperature=0.0, Node Budget=Best from {16, 32, 64, 128, 256, 512, 1024}2026.04 | 7.68 | 9.94 | — | — | |
| DFlash+DDTreeModel=Qwen3-Coder-30B-A3B-Instruct, Temperature=1.0, Node Budget=Best from {16, 32, 64, 128, 256, 512, 1024}2026.04 | 7.55 | 9.76 | — | — | |
| DFlash+DDTreeModel=Qwen3-4B, Temperature=0.0, Node Budget=Best from {16, 32, 64, 128, 256, 512, 1024}2026.04 | 6.42 | 9.16 | — | — | |
| DFlash+DDTreeModel=Qwen3-8B, Temperature=0.0, Node Budget=Best from {16, 32, 64, 128, 256, 512, 1024}2026.04 | 6.39 | 9.07 | — | — | |
| DFlash+DDTreeModel=Qwen3-4B, Temperature=1.0, Node Budget=Best from {16, 32, 64, 128, 256, 512, 1024}2026.04 | 6.09 | 8.72 | — | — | |
| DFlash+DDTreeModel=Qwen3-8B, Temperature=1.0, Node Budget=Best from {16, 32, 64, 128, 256, 512, 1024}2026.04 | 5.87 | 8.34 | — | — | |
| Draft-OPDModel=Q3-8B, Thinking Mode=Disabled, Temperature=02026.05 | 5.64 | — | — | 6.57 | |
| DFlashModel=Qwen3-Coder-30B-A3B-Instruct, Temperature=0.02026.04 | 5.61 | 7.19 | — | — | |
| DFlashModel=Qwen3-Coder-30B-A3B-Instruct, Temperature=1.02026.04 | 5.47 | 7.03 | — | — | |
| Draft-OPDModel=Q3-4B, Thinking Mode=Disabled, Temperature=02026.05 | 5.4 | — | — | 6.6 | |
| DFlashModel=Q3-8B, Thinking Mode=Disabled, Temperature=02026.05 | 5.17 | — | — | 6.04 | |
| Draft-OPDModel=Q3-8B, Thinking Mode=Disabled, Temperature=0.62026.05 | 5.17 | — | — | 6.02 | |
| DFlashModel=Q3-4B, Thinking Mode=Disabled, Temperature=02026.05 | 5 | — | — | 6.04 | |
| Draft-OPDModel=Q3-4B, Thinking Mode=Disabled, Temperature=0.62026.05 | 4.98 | — | — | 6.13 | |
| Draft-OPDModel=Q3-8B, Thinking Mode=Enabled, Temperature=02026.05 | 4.86 | — | — | 5.73 | |
| Draft-OPDModel=Q3-4B, Thinking Mode=Enabled, Temperature=02026.05 | 4.85 | — | — | 5.96 | |
| EAGLE-3Model=Q3-8B, Thinking Mode=Disabled, Temperature=02026.05 | 4.83 | — | — | 5.99 | |
| DFlashModel=Q3-8B, Thinking Mode=Disabled, Temperature=0.62026.05 | 4.74 | — | — | 5.62 | |
| DFlashModel=Q3-4B, Thinking Mode=Disabled, Temperature=0.62026.05 | 4.72 | — | — | 5.73 | |
| EAGLE-3Model=Q3-8B, Thinking Mode=Enabled, Temperature=02026.05 | 4.55 | — | — | 5.64 | |
| EAGLE-3Model=Q3-4B, Thinking Mode=Disabled, Temperature=02026.05 | 4.55 | — | — | 5.84 | |
| DFlashModel=Q3-8B, Thinking Mode=Enabled, Temperature=02026.05 | 4.5 | — | — | 5.19 | |
| EAGLE-3Model=Q3-8B, Thinking Mode=Disabled, Temperature=0.62026.05 | 4.4 | — | — | 5.69 | |
| DFlashModel=Q3-4B, Thinking Mode=Enabled, Temperature=02026.05 | 4.39 | — | — | 5.51 | |
| DFlashModel=Qwen3-4B, Temperature=0.02026.04 | 4.38 | 6.1 | — | — | |
| DFlashModel=Qwen3-8B, Temperature=0.02026.04 | 4.33 | 5.99 | — | — | |
| EAGLE-3Model=Q3-4B, Thinking Mode=Enabled, Temperature=02026.05 | 4.29 | — | — | 5.33 | |
| Draft-OPDModel=Q3-4B, Thinking Mode=Enabled, Temperature=0.62026.05 | 4.21 | — | — | 5.13 | |
| Draft-OPDModel=Q3-8B, Thinking Mode=Enabled, Temperature=0.62026.05 | 4.17 | — | — | 5.05 | |
| DFlashModel=Qwen3-4B, Temperature=1.02026.04 | 4.03 | 5.56 | — | — | |
| EAGLE-3Model=Q3-4B, Thinking Mode=Disabled, Temperature=0.62026.05 | 3.98 | — | — | 5.66 | |
| EAGLE-3Model=Q3-8B, Thinking Mode=Enabled, Temperature=0.62026.05 | 3.93 | — | — | 5.33 | |
| DFlashModel=Q3-4B, Thinking Mode=Enabled, Temperature=0.62026.05 | 3.92 | — | — | 4.77 | |
| DFlashModel=Q3-8B, Thinking Mode=Enabled, Temperature=0.62026.05 | 3.85 | — | — | 4.66 | |
| DFlashModel=Qwen3-8B, Temperature=1.02026.04 | 3.83 | 5.3 | — | — | |
| EAGLE-3Model=Q3-4B, Thinking Mode=Enabled, Temperature=0.62026.05 | 3.75 | — | — | 5.03 | |
| EAGLE-2Model=Vicuna-13B, GPU=NVIDIA A402026.04 | 3.41 | 5.3 | — | — | |
| EDA (Ours)Temperature=0, Target Model=Qwen2.5-Coder-7B2026.03 | 3.18 | 5.43 | — | — | |
| GOOSEModel=Llama-3-8B, GPU=NVIDIA A402026.04 | 3 | 3.61 | — | — | |
| EAGLE-2Model=Vicuna-7B, GPU=NVIDIA A402026.04 | 2.92 | 4.83 | — | — | |
| GOOSEModel=Vicuna-7B, GPU=NVIDIA A402026.04 | 2.84 | 3.4 | — | — | |
| Full-FTTemperature=0, Target Model=Qwen2.5-Coder-7B2026.03 | 2.82 | 4.93 | — | — | |
| EDA (Base)Temperature=0, Target Model=Qwen2.5-Coder-7B2026.03 | 2.8 | 4.9 | — | — | |
| LoRATemperature=0, Target Model=Qwen2.5-Coder-7B2026.03 | 2.76 | 4.78 | — | — | |
| GOOSEModel=Qwen2-8B, GPU=NVIDIA A402026.04 | 2.7 | 3.2 | — | — | |
| GOOSEModel=Vicuna-13B, GPU=NVIDIA A402026.04 | 2.66 | 3.11 | — | — | |
| TRModel=Llama-3-8B, GPU=NVIDIA A402026.04 | 2.6 | 3.11 | — | — | |
| HyperDFlashTemperature=0, Drafted steps per verification round=6, Decoding mode=Think-high mode2026.06 | 2.56 | 3.35 | — | — | |
| EDA (Ours)Temperature=1, Target Model=Qwen2.5-Coder-7B2026.03 | 2.48 | 4.92 | — | — | |
| TRModel=Qwen2-8B, GPU=NVIDIA A402026.04 | 2.43 | 2.81 | — | — | |
| TRModel=Vicuna-7B, GPU=NVIDIA A402026.04 | 2.41 | 2.86 | — | — | |
| EAGLE-2Model=Llama-3-8B, GPU=NVIDIA A402026.04 | 2.41 | 4.83 | — | — | |
| TRModel=Vicuna-13B, GPU=NVIDIA A402026.04 | 2.4 | 2.78 | — | — | |
| GOOSEModel=Vicuna-33B, GPU=NVIDIA A1002026.04 | 2.35 | 2.74 | — | — | |
| Full-FTTemperature=1, Target Model=Qwen2.5-Coder-7B2026.03 | 2.23 | 4.32 | — | — | |
| EDA (Base)Temperature=1, Target Model=Qwen2.5-Coder-7B2026.03 | 2.22 | 4.3 | — | — | |
| HyperDFlashTemperature=1, Drafted steps per verification round=6, Decoding mode=Think-high mode2026.06 | 2.19 | 2.87 | — | — | |
| LoRATemperature=1, Target Model=Qwen2.5-Coder-7B2026.03 | 2.18 | 4.25 | — | — | |
| MTPTemperature=0, Drafted steps per verification round=3, Decoding mode=Think-high mode2026.06 | 2.09 | 2.71 | — | — | |
| TRModel=Vicuna-33B, GPU=NVIDIA A1002026.04 | 1.95 | 2.29 | — | — | |
| MTPTemperature=1, Drafted steps per verification round=3, Decoding mode=Think-high mode2026.06 | 1.91 | 2.48 | — | — | |
| LookaheadModel=Llama-3-8B, GPU=NVIDIA A402026.04 | 1.79 | 1.86 | — | — | |
| LookaheadModel=Vicuna-7B, GPU=NVIDIA A402026.04 | 1.75 | 1.78 | — | — | |
| LookaheadModel=Qwen2-8B, GPU=NVIDIA A402026.04 | 1.75 | 1.68 | — | — | |
| PLDModel=Llama-3-8B, GPU=NVIDIA A402026.04 | 1.67 | 1.72 | — | — | |
| LookaheadModel=Vicuna-13B, GPU=NVIDIA A402026.04 | 1.64 | 1.67 | — | — | |
| PLDModel=Vicuna-7B, GPU=NVIDIA A402026.04 | 1.59 | 1.62 | — | — | |
| Vanilla DFlashTemperature=0, Drafted steps per verification round=6, Decoding mode=Think-high mode2026.06 | 1.56 | 1.98 | — | — | |
| MTPTemperature=0, Drafted steps per verification round=6, Decoding mode=Think-high mode2026.06 | 1.53 | 2.82 | — | — | |
| LookaheadModel=Vicuna-33B, GPU=NVIDIA A1002026.04 | 1.51 | 1.49 | — | — | |
| PLDModel=Vicuna-13B, GPU=NVIDIA A402026.04 | 1.49 | 1.51 | — | — | |
| Vanilla DFlashTemperature=1, Drafted steps per verification round=6, Decoding mode=Think-high mode2026.06 | 1.46 | 1.87 | — | — | |
| PLDModel=Qwen2-8B, GPU=NVIDIA A402026.04 | 1.45 | 1.5 | — | — | |
| MTPTemperature=1, Drafted steps per verification round=6, Decoding mode=Think-high mode2026.06 | 1.42 | 2.54 | — | — | |
| PLDModel=Vicuna-33B, GPU=NVIDIA A1002026.04 | 1.27 | 1.3 | — | — | |
| Training-FreeTemperature=0, Target Model=Qwen2.5-Coder-7B2026.03 | 1.21 | 1.85 | — | — | |
| RESTModel=Qwen2-8B, GPU=NVIDIA A402026.04 | 1.11 | 1.04 | — | — | |
| RESTModel=Vicuna-33B, GPU=NVIDIA A1002026.04 | 1.08 | 1.05 | — | — | |
| RESTModel=Vicuna-7B, GPU=NVIDIA A402026.04 | 1.05 | 1.05 | — | — | |
| RESTModel=Vicuna-13B, GPU=NVIDIA A402026.04 | 1.04 | 1.04 | — | — | |
| RESTModel=Llama-3-8B, GPU=NVIDIA A402026.04 | 1.01 | 1.01 | — | — | |
| ARModel=Vicuna-7B, GPU=NVIDIA A402026.04 | 1 | 1 | — | — | |
| ARModel=Llama-3-8B, GPU=NVIDIA A402026.04 | 1 | 1 | — | — | |
| ARModel=Qwen2-8B, GPU=NVIDIA A402026.04 | 1 | 1 | — | — | |
| ARModel=Vicuna-13B, GPU=NVIDIA A402026.04 | 1 | 1 | — | — | |
| ARModel=Vicuna-33B, GPU=NVIDIA A1002026.04 | 1 | 1 | — | — | |
| Training-FreeTemperature=1, Target Model=Qwen2.5-Coder-7B2026.03 | 0.79 | 1.3 | — | — | |
| D-PACETarget=Llama-3.1-8B-Instruct, Temperature (T)=0, Drafter=3L, Epoch=62026.05 | — | 3.68 | 2.84 | — | |
| D-PACETarget=Qwen3-8B, Temperature (T)=0, Drafter=3L, Epoch=62026.05 | — | 3.56 | 2.65 | — | |
| DFlashTarget=Llama-3.1-8B-Instruct, Temperature (T)=0, Drafter=3L, Epoch=62026.05 | — | 3.28 | 2.53 | — | |
| DFlashTarget=Qwen3-8B, Temperature (T)=0, Drafter=3L, Epoch=62026.05 | — | 3.22 | 2.4 | — | |
| FP16Model=Llama-3.2-1B-Instruct2025.11 | — | — | 33 | — | |
| IntAttentionModel=Llama-3.2-1B-Instruct2025.11 | — | — | 34.2 | — | |
| MDMSampling steps=256, Zero-shot=true, Backbone=LLaDA-8B-Instruct2025.10 | — | — | 0.182 | — | |
| MDMSampling steps=512, Zero-shot=true, Backbone=LLaDA-8B-Instruct2025.10 | — | — | 0.278 | — | |
| MDMSampling steps=1024, Zero-shot=true, Backbone=LLaDA-8B-Instruct2025.10 | — | — | 0.319 | — | |
| PRISMSampling steps=256, Zero-shot=true, Backbone=LLaDA-8B-Instruct, Trainable parameters=≈ 250M2025.10 | — | — | 0.218 | — | |
| PRISMSampling steps=512, Zero-shot=true, Backbone=LLaDA-8B-Instruct, Trainable parameters=≈ 250M2025.10 | — | — | 0.291 | — | |
| PRISMSampling steps=1024, Zero-shot=true, Backbone=LLaDA-8B-Instruct, Trainable parameters=≈ 250M2025.10 | — | — | 0.323 | — | |
| Quant-OnlyModel=Llama-3.2-1B-Instruct2025.11 | — | — | 31.4 | — |