End-to-end inference tuning on BERT-Large
22.6Tuning Time (s)STOF
Evaluation Results
| Method | Links | |
|---|---|---|
| STOFInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 22.6 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 52.4 | |
| STOFInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 55 | |
| BoltInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 57.3 | |
| BoltInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 126.1 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 132.3 | |
| STOFInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 225.3 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 1,049.7 | |
| BoltInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 1,067.7 |