Inference Performance Evaluation on 1,424 simulated conversations on AWS EC2 P5 (test)
0.92Latency (s)π_θ^EAGLE+FP8
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| π_θ^EAGLE+FP8Decoding=Greedy, Precision=FP8, Quantization=FP8, Speculative Decoding=EAGLE2026.06 | 0.92 | 3.8 | 6.54 | 4.48 | |
| Baseline (π_θ^FP8)Decoding=N/A, Precision=FP82026.06 | 1.6 | — | 3.66 | 2.5 | |
| Baseline (π_θ^BF16)Decoding=N/A, Precision=BF162026.06 | 1.69 | — | 3.4 | 2.33 | |
| Llama 3 70B (π_T)Decoding=N/A, Precision=BF162026.06 | 3.92 | — | 1.46 | 1 |