Speculative Decoding on Qwen2.5-Instruct Evaluation Set
60MTP AccuracyCLP
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CLPModel Size=0.5B, k=3, Draft Head Architecture=MLP head, Quantization (int8)=int82026.06 | 60 | 1.546 | 0 | — | |
| CLPModel Size=1.5B, k=2, Draft Head Architecture=MLP head, Quantization (int8)=int82026.06 | 18.1 | 1.244 | 0.9 | — | |
| CLPModel Size=7B, k=2, Draft Head Architecture=MLP head, Quantization (int8)=int82026.06 | 18 | 1.143 | 0 | — | |
| CLPModel Size=1.5B, k=3, Draft Head Architecture=MLP head, Quantization (int8)=int82026.06 | 14.5 | 1.294 | 1.8 | — | |
| CLPModel Size=7B, k=3, Draft Head Architecture=MLP head, Quantization (int8)=int82026.06 | 14.4 | 1.199 | 0 | — | |
| CLPModel Size=0.5B, k=3, Draft Head Architecture=EAGLE-style feature-level draft head, Quantization (int8)=int82026.06 | 4.6 | 1.088 | 1.2 | — |