LLM Inference on ShareGPT Vicuna unfiltered
117.5Throughput (tok/s)ENTMTPτ
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ENTMTPτTree Configuration=per-step switching between conservative and aggressive trees2026.06 | 117.5 | 3.47 | 2.99 | |
| ENTMTP*Tree Configuration=per-task throughput-optimal tree2026.06 | 116.9 | 3.42 | 2.97 | |
| Hydra (default)Tree Configuration=authors’ published default trees2026.06 | 109 | 2.89 | 3.06 | |
| Medusa (default)Tree Configuration=authors’ published default trees2026.06 | 94.6 | 2.72 | 2.86 | |
| Vanilla VicunaBase Model=Vicuna-7B v1.32026.06 | 35.4 | 1 | 1 |