Loading the SOTA2 catalog…
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding · SOTA2 Research