Locality-Controlled OOD Guidance for Robust Regulatory DNA Sequence Design
Abstract
Designing regulatory DNA sequences for targeted gene-expression control is an important challenge. Recent approaches frame this task as property-conditioned generation, where diffusion models trained on natural sequences are guided toward high predicted activity through reward-based fine-tuning or sampling-time guidance. We show that single-reward evaluation can obscure reward-hacking behavior. By decomposing evaluation into property, novelty, and plausibility, we identify an empirically underexplored region of the evaluation space in which all three objectives are satisfied, but no existing baseline succeeds. {We postulate that this gap arises because prior methods treat guidance primarily as a matter of magnitude.} In regulatory DNA, however, sequence grammar is spatially nonuniform: motif positions are sensitive to perturbation, while background regions allow more exploration. We therefore propose \textsc{LocoGen} by reformulate guidance as a problem of \textbf{locality}. \textsc{LocoGen} learns an invariant motif mask through sequence-level invariance learning to protect motif positions, while applying OOD damping only to background positions. It operates at sampling time on a frozen diffusion backbone, with no retraining. Across HepG2, K562, and SK-N-SH, \textsc{LocoGen} reaches this previously underexplored region of the property--novelty--plausibility space. Code is available at \url{https://huggingface.co/AnonymousAuthor42/LocoGen}.