Cluster the Cache, Not the Data: Preserving Informative Missingness and Utility Under Microaggregation for In-Context Tabular Learning
Abstract
Machine learning is widely applied to clinical tabular data to uncover new diagnostic and prognostic capabilities. However, many models are vulnerable to attacks that put a patient's private data at risk. Microaggregation is a straightforward approach to reduce this risk, but it can significantly degrade model performance. Missingness patterns are often informative as they may represent clinical approaches to distinct patient types, which can similarly be lost through clustering. To address this, we evaluate microaggregation in the Key-Value representation space (kv-cache), using the tabular foundation model TabPFN-3. We perform a controlled comparison of kv-cache microaggregation with two raw data microaggregation approaches: one clustering the raw data, and another imposing majority-vote missingness on raw clustered values. Across three healthcare datasets (one private and two public) with naturally occurring informative missingness, we observe that cache-space clustering retains more predictive utility at every compression level, and the advantage widens under heavier compression. Interestingly, we observe that the benefit is largely diminished in datasets without informative missingness. Membership-inference risk, assessed through a calibrated likelihood-ratio attack, diminishes to close to random chance with moderate compression in a similar manner for all 3 methods. Cache-space clustering is also markedly more reproducible across seeds.