KNOT: A Knowledge Entanglement Benchmark for Robust Unlearning Evaluation
Abstract
Large language model (LLM) unlearning aims to selectively remove specific knowledge while preserving model utility. Current benchmarks such as TOFU, WMDP, and MUSE evaluate unlearning on semantically disjoint forget and retain sets, making the task artificially easy. We formalize this gap through a taxonomy of knowledge entanglement at three levels: entity, concept, and skill. Based on this taxonomy, we construct KNOT, a benchmark with approximately 5,300 forget, 9,800 retain, and 5,000 boundary QA pairs, each annotated with a quantitative entanglement score. Evaluation of seven methods on KNOT reveals that (a) all methods degrade severely under high entanglement, (b) the coreset effect reported on standard benchmarks vanishes, and (c) existing metrics miss critical boundary behavior. We propose Boundary Precision and Boundary Recall to fill this gap. Cross-benchmark comparison confirms that KNOT's entanglement, rather than surface difficulty, drives the observed degradation.