1 post tagged llm-serving.
DeepSeek's open-source speculative decoding framework speeds up per-user generation by up to 85% without any quality loss.