Runtime Mitigation of Weight-Induced Misalignment through Corrective Retrieval for Clinical LLM Agents
Darrell Cheng ⋅ Vinh Pham ⋅ Tenzing J Lunn ⋅ Jonathan Lugo ⋅ Ryan Bui ⋅ Kevin Zhu
Abstract
Clinical institutions may operate language models within retrieval and memory scaffolds without access to the model weights or fine-tuning history; retraining may be infeasible when weights are available. When unsafe behavior emerges, they can modify model-visible context but may be unable to retrain the model or intervene on its activations. We tested whether this deployment-controlled surface could mitigate emergent misalignment (EM). Using a Qwen2.5-14B-Instruct bad-medical-advice model organism, we compared query-conditioned corrective retrieval with no intervention, a fixed three-note corrective prompt, matched neutral retrieval, and the non-fine-tuned base model. On 180 harmful clinical requests, corrective retrieval reduced the rate of responses classified as harmful by an automated judge from 63.4\% to 15.1\% (paired difference $-48.3$ percentage points; bias-corrected and accelerated (BCa) 95\% interval $[-53.3,-42.7]$), closing 76.3\% of the judged-harm gap relative to the base model; however, C3 classified 21.1\% of Tier-D responses as refusals versus 2.5\% for C1, so the aggregate reduction includes an 18.6-point increase in refusal. It nevertheless remained 15.0 points above that reference. On an amended 180-question benign clinical instrument used to test a preregistered over-refusal hypothesis, none of the 10{,}800 responses across six conditions was classified as a refusal; a secondary harmful-response analysis decreased from 9.2\% to 2.2\%. Retrieval did not significantly outperform the fixed corrective prompt under the primary protocol; episodic retrieval also allowed self-authored session records to compete with corrective notes. A non-clinical estimate was sensitive to the endpoint definition. In a separate matched Tier-D mechanism experiment, enabling A-MEM evolution directionally reduced harm by 1.76 points, but the primary interval included zero. All primary outcomes came from one automated judge; a post-hoc second-model audit of 60 stratified Tier-D responses found 45.0\% exact four-way verdict agreement and 82.5\% agreement on binary harm status among 40 non-derailed responses; it did not validate clinical correctness or recompute treatment effects. In this single-model, single-seed study, corrective context reduced the rate of responses judged harmful under the study rubric, but the findings do not establish a retrieval-specific advantage or true model repair.
Chat is not available.
Successful Page Load