Efficiency Evaluation on Efficiency Profiling Workload G.7 Detailed Efficiency Results
13.75End-to-End Latency (s)LLMBoost
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LLMBoostModel Composition=7B+7B, Execution Strategy=near-parallel, Number of GPUs=22025.12 | 13.75 | 0.075 | 33.18 | 20.72 | |
| LLMBoostModel Composition=7B+7B, Execution Strategy=near-parallel, Number of GPUs=12025.12 | 14.87 | 0.081 | 54.11 | 39.75 | |
| T-CopilotModel Composition=7B+7B2025.12 | 17.21 | 0.097 | 64.41 | 42.05 | |
| Qwen14BModel Composition=14B2025.12 | 17.4 | 0.098 | 56.74 | 36.97 | |
| LLMBoostModel Composition=7B+7B+7B, Execution Strategy=near-parallel, Number of GPUs=32025.12 | 18.88 | 0.103 | 33.18 | 20.13 | |
| LLMBoostModel Composition=7B+7B+7B, Execution Strategy=near-parallel, Number of GPUs=12025.12 | 21.67 | 0.118 | 73.42 | 58.48 | |
| LLMBoostModel Composition=7B+7B, Execution Strategy=sequential2025.12 | 22.53 | 0.125 | 60.29 | 38.72 | |
| VOTEModel Composition=7B+7B2025.12 | 22.61 | 0.127 | 53.39 | 39.14 | |
| UNITEModel Composition=7B+7B2025.12 | 22.66 | 0.128 | 53.39 | 39.12 |