Personalized Safety in Federated Fine-Tuning of Large Language Models
Tianzhe Xiao ⋅ Gaozhuo Liu ⋅ Yichen Li ⋅ Haozhao Wang ⋅ Hanlin Cai ⋅ Ruixuan Li ⋅ Ozgur Akan
Abstract
Large language models (LLMs) are increasingly adapted inside privacy- and regulation-constrained institutions where local post-training data cannot be centrally pooled. Federated learning therefore becomes a natural mechanism for collaborative LLM adaptation, but it also raises a safety question that is not captured by either classical federated learning or standard centralized alignment: clients may share a model while still requiring different safe behaviors. We formalize this setting as \emph{personalized federated safety}, where safety correctness is conditioned on the target client's local policy rather than treated as a single global refusal rule. To make this problem transferable rather than benchmark-specific, we introduce a compact, interpretable client policy space and a policy-conditioned benchmark construction framework, then instantiate it with five representative client archetypes spanning regulated healthcare, legal compliance, financial safety, youth safety, and enterprise general safety. The resulting benchmark contains 1{,}500 client-conditioned training examples and 750 held-out evaluation examples. We further propose \emph{Safety-Aware Weighted LoRA} (SAW-LoRA), a lightweight add-on for federated LoRA fine-tuning that lets each client selectively absorb incoming global task updates according to estimated task benefit and local-safety risk. Across two task datasets and two backbones, this add-on substantially lowers client-specific ASR for the strongest local safety mechanisms, reaching ASR near $1$--$2\%$ while preserving high benign acceptance. These findings position personalized federated safety as a concrete research problem with a practical update mechanism for federated LLM adaptation.
Chat is not available.
Successful Page Load