Loading the SOTA2 catalog…
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget · SOTA2 Research