Language Modeling on MMLU-Pro
191.3Tokens Per Second (TPS)PARSE + EAGLE3
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PARSE + EAGLE3Model=Qwen3-235B-A22B, Draft Model=Qwen3-8B, Decoding Strategy=PARSE with EAGLE3 on both draft and target, Precision=FP8, Inference Framework=SGLang2026.05 | 191.3 | — | 2.07 | |
| PARSEModel=Qwen3-235B-A22B, Draft Model=Qwen3-8B, Decoding Strategy=Full and partial verify, Precision=FP8, Inference Framework=SGLang2026.05 | 149.5 | 63.6 | 1.62 | |
| Eagle3Model=Qwen3-235B-A22B, Decoding Strategy=Speculative decoding, Precision=FP8, Inference Framework=SGLang2026.05 | 137 | — | 1.48 | |
| Qwen3-235B-A22BModel=Qwen3-235B-A22B, Decoding Strategy=Plain autoregressive, Precision=FP8, Inference Framework=SGLang2026.05 | 92.3 | 66.8 | — | |
| Qwen3-8BModel=Qwen3-8B, Decoding Strategy=Standalone2026.05 | — | 56.8 | — |