Multi-modal to Text Generation Latency on ServeGen 16-GPU setup (real-world workload)
7.1Latency P50Cornserve
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CornserveInput Modality=T+I+V+A, Hardware=16-GPU2025.12 | 7.1 | 18.4 | 19.8 | |
| CornserveInput Modality=T+I+V, Hardware=16-GPU2025.12 | 7.3 | 18.7 | 22.6 | |
| CornserveInput Modality=T+I+A, Hardware=16-GPU2025.12 | 7.9 | 19.1 | 24.8 | |
| CornserveInput Modality=T+I, Hardware=16-GPU2025.12 | 8.5 | 16.3 | 21 | |
| MonolithInput Modality=T+I+A, Hardware=16-GPU2025.12 | 20.1 | 114.7 | 128.6 | |
| MonolithInput Modality=T+I+V, Hardware=16-GPU2025.12 | 20.8 | 88.8 | 150.4 | |
| MonolithInput Modality=T+I, Hardware=16-GPU2025.12 | 28.9 | 95.8 | 125.2 | |
| MonolithInput Modality=T+I+V+A, Hardware=16-GPU2025.12 | 33.6 | 128.9 | 181.4 |