SIREN: Evidence-Coupled Rule Editing for Selective Intervention
Abstract
Personalized assistants must not only retain a user's rules, but also determine when those rules should influence an interaction. We formulate selective intervention as a rule-editing problem: condition-action rules are stored persistently, evidence refines their applicability, and the system can activate one rule or defer to its unedited behavior. We propose SIREN, an evidence-coupled framework with separate interfaces for rule installation, local refinement, and selective use. Each rule occupies a key-value codebook entry attached to a frozen language model; a condition recognizer, per-rule threshold, and activation policy determine whether its value influences generation. We construct a synthetic intervention-rule testbed and instantiate the framework on Llama-3.1-8B-Instruct. In a controlled replay of 82 rules with a pre-existing roster-wide contrastive bank, mean condition recognition AUC rises from 0.743 at initialization to 0.922 after 16 positive and 16 hard-negative examples per rule. Full-roster routing selects the correct slot on 67/90 seed-rule and 267/432 original expanded-rule positives, with 0 and 3 wrong-slot selections; coverage falls on turn-transformed and freshly generated contexts. Withholding under-evidenced entries reduces peak wrong-slot decisions from 18 to 3 on one replay sequence while delaying their use; both arms end level once every entry is eligible. These results establish a proof of concept for separating persistent rule storage from selective activation. Reliable response execution remains a distinct, unresolved component.