End-to-end inference tuning on GPT
23.8Tuning Time (s)STOF
Evaluation Results
| Method | Links | |
|---|---|---|
| STOFInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 23.8 | |
| STOFInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 40.9 | |
| BoltInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 48.8 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 49.5 | |
| BoltInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 99.8 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 100.8 | |
| STOFInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 122.2 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 664.4 | |
| BoltInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 738.6 |