Editing Large Language Models with Geometry-Aware Regularization
Abstract
Knowledge editing has become a key technique for updating knowledge in large language models. However, existing methods neglect the geometric structure of the knowledge representation space, leading to two key limitations: they cannot distinguish which pre-trained knowledge should be preferentially preserved (core vs. non-core) nor which new edits are inherently difficult (hard vs. easy). To address this, we propose \textsc{GEdit}, a \textit{geometry-aware} framework built on a Low-Rank Plus Diagonal (LR+D) factor model that captures the intrinsic low-rank structure of Transformer FFN hidden representations. \textsc{GEdit} features two complementary mechanisms. First, the \textit{low-rank regularization} applies targeted protection on a low-rank core knowledge subspace. Second, the \textit{adaptive regularization} dynamically adjusts the regularization strength based on each edit request's editability: relaxing constraints for hard edits to ensure success, and tightening them for easy edits to better preserve pretrained knowledge. Extensive experiments on \texttt{GPT2-XL}, \texttt{GPT-J}, and \texttt{LLaMA-3-8B} under large-scale editing scenarios—including both mini-batch and fully sequential editing settings with up to 10,000 continuous updates—show that \textsc{GEdit} consistently and significantly outperforms state-of-the-art baselines in both editing success and knowledge preservation.