EfficientXpert: Efficient Domain Adaptation for Large Language Models via Propagation-Aware Pruning
Abstract
Deploying domain-specialized large language models on resource-constrained hardware motivates reducing both adaptation cost and the number of retained weights. LoRA lowers adaptation cost but leaves a dense backbone, while separate pruning can discard weights that become important during adaptation. We propose \textbf{EfficientXpert}, a framework that co-adapts low-rank updates and sparse support. Its ForeSight Mask scores weights through a downstream reconstruction surrogate using the evolving LoRA-augmented weights; Partial Brain Surgeon (PBS) realigns the adapter through a closed-form correction. At 40\% sparsity, the strongest variant retains 98.55--108.99\% of dense-LoRA aggregate performance across the evaluated LLaMA settings. On Qwen3-8B, adaptation adds 17.9--27.0\% training time and at most 4.8\% peak GPU memory. Our analysis connects the domain- and rank-dependent effects of PBS to its constrained reconstruction objective, providing guidance for selecting adapter recovery capacity. Together, these results demonstrate efficient production of sparse, domain-specialized experts.