LiveLoRA: Exact KV Cache Migration Across LoRA Updates for LLM Agents
Tianyi Huang ⋅ Samuel Xu ⋅ Samuel Xu ⋅ Ivy Gu ⋅ Nathan Huang ⋅ Siqi Zhang
Abstract
Updating an LLM adapter should not require rebuilding the inference state of every active session from its entire interaction history. Because affected key-value (KV) cache entries were computed under the previous model version, reusing them mixes versions, whereas restoring coherence through full-history re-prefill incurs work that grows with context length and fleet size. We introduce LiveLora, a model-runtime co-design that makes final-block K/V migration exact for a specific LoRA adapter family under explicit model and arithmetic assumptions. LiveLora freezes the lower decoder blocks and fixes the down-projection matrices of the final-block K/V adapters, leaving each token's base projections and low-rank coordinates invariant across adapter versions. To publish an update, the runtime reconstructs final-block K/V tensors for the new version, recomputes only the latest token through the final block, and commits the adapter, cache, and next-token logits as one coherent state. On Qwen2.5-3B-Instruct, the within-run median paired speedup over full-history re-prefill averages $68.98\times$ across five runs for 32 active sessions with 8K tokens each. Across all seven evaluated settings, migrated final-block K/V tensors match independently captured references with an observed maximum absolute error of zero; we also quantify the resulting storage and cached-decoding overheads. LiveLora turns online adapter publication into a version-coherent model-runtime state transition, allowing active sessions to adopt new adapter versions without replaying their interaction histories.
Chat is not available.
Successful Page Load