End-to-end inference tuning on LLaMA
29.5Tuning Time (s)STOF
Evaluation Results
| Method | Links | |
|---|---|---|
| STOFInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 29.5 | |
| STOFInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 43.6 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 48.8 | |
| BoltInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 52.1 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 110.8 | |
| BoltInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 124.6 | |
| STOFInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 264.6 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 820.6 | |
| BoltInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 837 |