Loading the SOTA2 catalog…
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · SOTA2 Research