Loading the SOTA2 catalog…
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills · SOTA2 Research