DIAL: A Bounded-Monotone Adapter with Closed-Form Residual $L_\infty$ Bounds for Scalar-Controlled Refusal Modulation
Nguyen Vu Nguyen
Abstract
We introduce DIAL (Directional Intervention, Asymptotically Limited), a parameter-efficient adapter that adds a bounded, monotone, scalar-controlled residual $\delta(s,c) = \sigma(g) \odot \tanh(c/T) \odot \tanh(Dz)$ with $z = D^\top \phi(s)$ to a base language model's hidden state. The architectural form yields a closed-form upper bound $\lVert \delta \rVert_\infty \le \sigma(g_{\max})$ on the residual $L_\infty$ norm, computable from model weights alone. Across 5 base models (3.8B to 70B parameters, 4 architecture families; one checkpoint per family, with 5-seed csweep variance at Llama-3.1-8B), the residual saturates at the predicted bound with mean ratio $1.000 \pm 0.001$ (worst-case family deviation 0.2 percentage points). On Llama-3.1-8B under GPT-4o judge ($n=100$ per bench, zero content-filter blocks), DIAL exceeds a FiLM baseline trained with the same multi-$c$ objective on every benchmark and judge. Semantic vs strict judge: HarmBench DIAL +0.52 / +0.55 vs FiLM +0.44 / +0.39; AdvBench DIAL +0.63 / +0.66 vs FiLM +0.46 / +0.50; XSTest DIAL +0.33 / +0.25 vs FiLM +0.26 / +0.22; both judges agree on the ranking. Outside training range, DIAL's residual saturates at $\sigma(g_{\max})$ while FiLM's grows unboundedly: at $c = 10^4$, FiLM's residual $L_\infty$ norm is 534 times larger than DIAL's. The bound holds across all control values, not only trained ones; concrete deployment scenarios (configuration error, multi-tenant policy ceilings, adversarial composition) are where this distinction matters in practice. The bounded-monotone form produces analogous saturation behavior on two continuous-control RL environments (HC-Vel $n=35$, Ant-Goal $n=30$) under matched oracle-$c$; the prior fails when $c$ must be inferred online (meta-RL). Ablations confirm both the bound and the saturating temperature are necessary; a frozen-gate variant retains the bound at lower controllability magnitude. We release adapter train reports and canonical bench generations, algebraic derivations (with a SymPy reproducibility script), the full GPT-4o judge cache (semantic and strict), and a human-labeled validation set ($n=50$, Cohen's kappa = 0.895). Full adapter checkpoints are deferred to camera-ready release.
Chat is not available.
Successful Page Load