Measuring Stability via Minimal Context Interventions
Abstract
Large language models (LLMs) increasingly rely on external context supplied at inference time to ground their outputs in current and verifiable information, not solely on their stale parametric knowledge. Yet, models trust this context without any guarantee that their answers depend on it in a stable or meaningful way. A small, seemingly innocuous change to the given context can silently flip a highly confident answer. In this paper, we study this sensitivity to both quantify LLM robustness and identify the context fragments driving model answers. We frame this as an optimization problem where the goal is to find a minimal-size intervention applied to the context to flip the model answer. To address this problem, we introduce \sysName, a lightweight, model-agnostic framework for context sensitivity analysis that integrates seamlessly with existing open-source LLMs without requiring retraining. To bypass the computational bottlenecks of brute-force context perturbation, \sysName\ leverages attribution functions to target context interventions across varying context lengths. Extensive experiments across three LLMs on question-answering benchmarks demonstrate the effectiveness of \sysName. Our findings yield novel insights into LLM robustness behavior and systematically benchmark which attribution methods yield reliable sensitivity profiles across diverse settings. For example, we discover that while Mistral-7B is robust to small changes in a given context, Qwen3-Instruct is highly sensitive.