LLM Generation Efficiency on Qwen 14B (2048 input + 2048 generation tokens) 2.5
70.3End-to-end Latency (s)GRIFFIN
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRIFFINSparsity=50%, r=256, Hardware=NVIDIA L40 GPU, Precision=BF16, Model Portion=First half of the model2025.05 | 70.3 | 34 | |
| CapreseSparsity=50%, r=256, Hardware=NVIDIA L40 GPU, Precision=BF162025.05 | 70.4 | 34 | |
| LoRASparsity=50%, r=256, Hardware=NVIDIA L40 GPU, Precision=BF16, Configuration=GRIFFIN paired with LoRA weights2025.05 | 76.5 | 37 | |
| FullHardware=NVIDIA L40 GPU, Precision=BF162025.05 | 84.4 | 41 |