ClinKV: Clinical-Risk-Aware KV Cache Compression for Long-Context Medical Language Models
Abstract
Long-context medical language models increasingly rely on key-value (KV) cache compression to reduce memory use and decoding cost. Existing compressors typically rank cached states by attention, recency, norms, reconstruction fidelity, or generic task utility. Recent work shows that these objectives can preserve average benchmark performance while selectively degrading instructions, safety behavior, or factual grounding. We argue that this mismatch is especially consequential in medicine: a clinically decisive fact can be historically low-attention, temporally distant, and linguistically brief, yet become critical for a later recommendation. Examples include a drug allergy, anticoagulant use, pregnancy status, renal impairment, a prior adverse reaction, or a suicidal statement. Our position is that KV compression for medical LLMs should optimize expected clinical harm from forgetting, not only average task loss. We formalize this objective and propose ClinKV, a modular risk-aware wrapper that combines a base compressor with a clinical-risk reserve and a mixed-precision fallback for uncertain high-consequence spans. We then specify an evaluation protocol spanning LongHealth, MedOdyssey, HealthBench, MedSafetyBench, and controlled delayed-critical-fact stress tests, with metrics that measure harm-weighted critical information retention and counterfactual action preservation alongside memory and latency. The central claim is methodological: clinical KV compression should be judged on a risk-efficiency frontier rather than an accuracy-memory frontier alone.