End-to-end inference tuning on T5
43.1Tuning Time (s)STOF
Evaluation Results
| Method | Links | |
|---|---|---|
| STOFInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 43.1 | |
| BoltInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 70.7 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 71.9 | |
| STOFInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 80.3 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 239 | |
| BoltInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 244.7 | |
| STOFInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 388.3 | |
| BoltInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 1,860.8 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 1,987.6 |