Text Generation on OPT 512 prompt + 32 generation tokens 6.7B
50.5Throughput (token/s)H2O-c (20%)
Evaluation Results
| Method | Links | |
|---|---|---|
| H2O-c (20%)Seq. length=512+32, Model size=6.7B, Batch size=70, Offloading status=without offloading, Hardware=single T4 GPU, KV cache budget=20%, Weights compression=4-bit2023.06 | 50.5 | |
| H2O (20%)Seq. length=512+32, Model size=6.7B, Batch size=4, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 35.1 | |
| AccelerateSeq. length=512+32, Model size=6.7B, Batch size=2, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 20.4 | |
| FlexGenSeq. length=512+32, Model size=6.7B, Batch size=2, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 20.2 |