Text Generation on OPT 512 prompt + 1024 generation tokens 6.7B
62.3Throughput (token/s)H2O-c (20%)
Evaluation Results
| Method | Links | |
|---|---|---|
| H2O-c (20%)Seq. length=512+1024, Model size=6.7B, Batch size=44, Offloading status=without offloading, Hardware=single T4 GPU, KV cache budget=20%, Weights compression=4-bit2023.06 | 62.3 | |
| H2O (20%)Seq. length=512+1024, Model size=6.7B, Batch size=4, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 52.1 | |
| FlexGenSeq. length=512+1024, Model size=6.7B, Batch size=1, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 16.9 |