Loading the SOTA2 catalog…
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration · SOTA2 Research