P99 Per-Token Latency on LLaMA-2 7B (inference)
8.23P99 Per-Token Latency (ms)JIT+CUDA
Evaluation Results
| Method | Links | |
|---|---|---|
| JIT+CUDAGeneration Length=10, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 8.23 | |
| JIT+CUDAGeneration Length=50, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 9.44 | |
| JIT+CUDAGeneration Length=100, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 11.36 | |
| JIT+CUDAGeneration Length=150, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 12.54 | |
| JIT+CUDAGeneration Length=200, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 14.11 | |
| JIT+CUDAGeneration Length=250, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 15.65 | |
| TensorRT–LLMGeneration Length=350, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 16.53 | |
| TensorRT–LLMGeneration Length=300, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 16.54 | |
| TensorRT–LLMGeneration Length=400, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 16.54 | |
| TensorRT–LLMGeneration Length=250, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 16.56 | |
| TensorRT–LLMGeneration Length=450, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 16.61 | |
| TensorRT–LLMGeneration Length=500, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 16.65 | |
| TensorRT–LLMGeneration Length=200, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 16.73 | |
| JIT+CUDAGeneration Length=300, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 17.31 | |
| HuggingFaceGeneration Length=50, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.3 | |
| HuggingFaceGeneration Length=500, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.38 | |
| HuggingFaceGeneration Length=10, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.39 | |
| HuggingFaceGeneration Length=450, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.39 | |
| HuggingFaceGeneration Length=400, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.41 | |
| HuggingFaceGeneration Length=350, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.42 | |
| HuggingFaceGeneration Length=300, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.44 | |
| HuggingFaceGeneration Length=250, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.46 | |
| HuggingFaceGeneration Length=100, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.49 | |
| HuggingFaceGeneration Length=200, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.49 | |
| JIT+CUDAGeneration Length=350, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.9 | |
| HuggingFaceGeneration Length=150, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 18.94 | |
| JIT+CUDAGeneration Length=400, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 20.5 | |
| JIT+CUDAGeneration Length=450, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 22.09 | |
| JIT+CUDAGeneration Length=500, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 23.68 | |
| TensorRT–LLMGeneration Length=150, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 24.03 | |
| TensorRT–LLMGeneration Length=100, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 31.59 | |
| TensorRT–LLMGeneration Length=50, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 52.28 | |
| TensorRT–LLMGeneration Length=10, Hardware=NVIDIA H100, Precision=FP16, Batch size=1, Number of trials=10002026.04 | 68.84 |