LLM Generation Efficiency on Qwen 2.5 14B (2048 input + 256 generation tokens)
8.7End-to-end Latency (s)GRIFFIN
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRIFFINSparsity=50%, r=256, Hardware=NVIDIA L40 GPU, Precision=BF16, Model Portion=First half of the model2025.05 | 8.7 | 34 | |
| CapreseSparsity=50%, r=256, Hardware=NVIDIA L40 GPU, Precision=BF162025.05 | 8.7 | 34 | |
| LoRASparsity=50%, r=256, Hardware=NVIDIA L40 GPU, Precision=BF16, Configuration=GRIFFIN paired with LoRA weights2025.05 | 9.5 | 37 | |
| FullHardware=NVIDIA L40 GPU, Precision=BF162025.05 | 10.5 | 41 |