End-to-end inference tuning on ViT
93.9Tuning Time (s)STOF
Evaluation Results
| Method | Links | |
|---|---|---|
| STOFInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 93.9 | |
| STOFInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 99.3 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 100.2 | |
| BoltInput Size (Batch Size, Sequence Length)=(1, 128), GPU=NVIDIA A1002025.06 | 120.7 | |
| STOFInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 412.8 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 437.8 | |
| BoltInput Size (Batch Size, Sequence Length)=(8, 512), GPU=NVIDIA A1002025.06 | 468.8 | |
| BoltInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 3,848.6 | |
| MCFuserInput Size (Batch Size, Sequence Length)=(16, 2048), GPU=NVIDIA A1002025.06 | 4,264.3 |