LITHE: Lattice-Indexed Twin Hadamard Encoding for Diffusion Personalization
Abstract
Diffusion personalization deployments serve growing catalogues of LoRA adapters, but resident-LoRA serving is capped near the GPU memory budget (about 150 rank-16 adapters on an A100 80 GB GPU running FLUX.1-dev). We introduce LITHE, a LoRA-compatible deployment representation that stores each adapter as an integer-index stream over a deterministic Hadamard codebook, and that pushes this resident-adapter ceiling to a verified 10,000 adapters on the same single GPU at sub-80 GB peak. The same on-disk indices drive two serving modes: an index-resident mode that routes a heterogeneous batch of adapters in one forward, and a decoded compiled mode within 1% of an optimized LoRA+compile baseline at 28-step inference at 1024² resolution on FLUX.1-dev (and 1.07–1.43× faster at lower resolutions). On disk, each rank-16 FLUX.1-dev adapter occupies 1.81 MB on average across the benchmark pool, a 50×/200× reduction over the matched r=16 / r=64 LoRA disk payload after the same coder.