Loading the SOTA2 catalog…
SDFP: Speculative Decoding with FIT-Pruned Models for Training-Free and Plug-and-Play LLM Acceleration · SOTA2 Research