Skip to yearly menu bar Skip to main content


Memory-Efficient Speculative Decoding with Quantized Draft KV Cache

Saksham Gupta ⋅ Michael Li

Abstract

Chat is not available.