Capricorn: Highly Efficient and Secure Mixture of Experts Inference Framework
Lushan Song ⋅ Xiaojian Liang ⋅ Shishuai Du ⋅ Jun J Sim ⋅ Yingting Liu ⋅ Xin Zhang ⋅ Jiang-Ming Yang ⋅ Pu Duan
Abstract
The rapid advancement of Large Language Models (LLMs) based on Mixture-of-Experts (MoE) architecture has enhanced a growing demand for secure inference frameworks that protect both client inputs and server model weights. However, existing secure MoE inference frameworks suffer from two limitations: (1) a costly two-stage secure routing pipeline that uses expensive secure Top-$K$ and secure equality test protocols, resulting in significant computational and communication overhead, and (2) secure evaluation of nonlinear activations like SiLU requires several multiplication and comparison operations, resulting in significant communication overhead. In this paper, we present Capricorn, a highly efficient and secure MoE inference framework that overcomes the two limitations above. Firstly, Capricorn utilizes a novel one-stage routing pipeline that reduces both the total number of Top-$K$ protocol invocations and their input data size. At the same time, we propose a lightweight method to generate one-hot token selection matrices locally, thus avoiding expensive secure equality tests. Secondly, we design a secure and precise SiLU protocol, which uses secure lookup tables (LUTs) with dynamic bit-width and precision to reduce communication overhead. Extensive experiments demonstrate that Capricorn achieves up to $50.93 \times$ and $2.10 \times$ speedup in secure routing and secure SiLU protocols, respectively, compared to the state-of-the-art (SOTA) framework CryptoMoE (NeurIPS'25), while maintaining the original model inference accuracy.
Chat is not available.
Successful Page Load