Text Generation on OPT 512 prompt + 512 generation tokens 6.7B
72.5Throughput (tokens/sec)H2O-c (20%)
Evaluation Results
| Method | Links | |
|---|---|---|
| H2O-c (20%)Seq. length=512+512, Model size=6.7B, Batch size=52, Offloading status=without offloading, Hardware=single T4 GPU, KV cache budget=20%, Weights compression=4-bit2023.06 | 72.5 | |
| H2O (20%)Seq. length=512+512, Model size=6.7B, Batch size=4, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 51.7 | |
| FlexGenSeq. length=512+512, Model size=6.7B, Batch size=1, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 16.8 | |
| AccelerateSeq. length=512+512, Model size=6.7B, Batch size=1, Offloading status=without offloading, Hardware=single T4 GPU2023.06 | 15.5 |