Language Modeling on MMLU (Accuracy, TPS, Speedup)
216.6Throughput (tok/s)PARSE + EAGLE3
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PARSE + EAGLE3Model=Qwen3-235B-A22B, Draft Model=Qwen3-8B, Decoding Strategy=PARSE with EAGLE3 on both draft and target, Precision=FP8, Inference Framework=SGLang2026.05 | 216.6 | — | 2.42 | |
| Eagle3Model=Qwen3-235B-A22B, Decoding Strategy=Speculative decoding, Precision=FP8, Inference Framework=SGLang2026.05 | 145.2 | — | 1.62 | |
| PARSEModel=Qwen3-235B-A22B, Draft Model=Qwen3-8B, Decoding Strategy=Full and partial verify, Precision=FP8, Inference Framework=SGLang2026.05 | 130.4 | 82.8 | 1.46 | |
| Qwen3-235B-A22BModel=Qwen3-235B-A22B, Decoding Strategy=Plain autoregressive, Precision=FP8, Inference Framework=SGLang2026.05 | 89.5 | 86.8 | — | |
| Qwen3-8BModel=Qwen3-8B, Decoding Strategy=Standalone (draft model)2026.05 | — | 77.6 | — |