Jaguar: Fast Private CNN Inference with Power-of-Two Homomorphic Arithmetic
Yewon Jeong ⋅ Nayoung Jung ⋅ Hyeri Roh ⋅ Woo-Seok Choi
Abstract
Hybrid HE/2PC private CNN inference remains bottlenecked by prime-modulus homomorphic arithmetic in convolution and by a precision flow that runs ReLU at doubled bitwidth before invoking a separate truncation protocol. We present Jaguar, a system built on a single design choice---a power-of-two ciphertext ring---that addresses both. The choice enables SPA-Conv, a coefficient-domain convolution kernel that replaces NTT-centric polynomial multiplication with scalar--polynomial accumulation, and an exact ciphertext-side truncation by local right shifts that lets ReLU run directly at the target fixed-point precision and eliminates the post-ReLU truncation protocol. Where NTT remains genuinely useful---at the client, for the single polynomial multiplication during decryption---we recover it through an auxiliary NTT prime, preserving the power-of-two protocol substrate while keeping decryption $O(N\log N)$. On ImageNet-scale ResNet-18, ResNet-50, and MobileNetV2 with AVX disabled, Jaguar achieves 2.07--3.72$\times$ lower end-to-end latency than Cheetah and 2.16--3.36$\times$ lower than Rhombus, with 1.16--1.76$\times$ lower communication than Cheetah.
Chat is not available.
Successful Page Load