Felid: A Flexible and Efficient Design for Transformer Fine-Tuning over Encrypted Data
Linru Zhang ⋅ Jun J Sim ⋅ Xiangning Wang ⋅ Jiahao Zhong ⋅ Kaiyu Zhou ⋅ Huanyi Ye ⋅ Yongsen Zheng ⋅ Lushan Song ⋅ Xiaojian Liang ⋅ Yingting Liu ⋅ Yujing Sun ⋅ Huaxiong Wang ⋅ Pu Duan ⋅ Kwok-Yan Lam
Abstract
Fully Homomorphic Encryption (FHE) enables non-interactive, privacy-preserving transformer fine-tuning, protecting highly sensitive data as inputs to large language models. Prior work has predominantly pursued efficiency gains through architectural refinements, yet overall performance remains bounded by the substantial cost inherent in encrypted computation. To address these issues, we target the bottleneck at the ciphertext operation level. Concretely, we propose a novel packing strategy that supports batching over a flexible number of inputs, making it well-suited for fine-tuning workloads. In addition, we design operation-specific matrix evaluation algorithms that significantly reduce (i) the number of ciphertext rotations and relinearizations, and (ii) the number of bootstrapping invocations. Additionally, we exploit functional bootstrapping to fold the evaluation of element-wise activation functions directly into the bootstrapping step, thereby reducing overall computational complexity. As a result, our framework achieves a $30\times$ reduction in ciphertext rotations and relinearizations across all matrix operations, as well as a $2.5\times$ reduction in FHE bootstrapping invocations over the entire computation flow, compared to the latest state-of-the-art work. It is worth noting that our system completes fine-tuning of a 2-layer BERT-style model with a batch size of 16 in 650 seconds, achieving a $8\times$ speedup over the state-of-the-art work.
Chat is not available.
Successful Page Load