BQ-LoRA: Binary-Quantized Low-Rank Adapters as Implicit Regularizers for Parameter-Efficient Fine-Tuning
Juyoung Park ⋅ Heejun Lee
Abstract
Parameter-efficient fine-tuning of quantized large language models is dominated by QLoRA, which combines 4-bit weight quantization with FP16 low-rank adapters. We show the FP16 adapter precision is unnecessary: a single binary bit per adapter parameter, combined with a learned per-layer scale $s$ and a Scaled Straight-Through Estimator, matches or exceeds QLoRA across a broad range of benchmarks while reducing adapter storage by $14$--$16\times$. We introduce two variants--NormScale, which clips STE gradients to the unit ball, and LearnThresh, which makes the binarization threshold learnable---unified under a framework we call BQ-LoRA. Our main theoretical contribution is a tight characterization of the BQ-LoRA hypothesis class: it has bounded entrywise norm, a high-probability spectral norm bound that is $\sqrt{r}$ times sharper than the worst case, a generically full-rank update, and consequently a Rademacher complexity that yields an explicit $\tilde{O}(s_{\max}\sqrt{d_{\text{in}} d_{\text{out}}/n})$ generalization bound. None of these properties hold a priori for FP16 LoRA. On Qwen2.5 (1.5B/7B/72B) and LLaMA-2 (7B/70B), BQ-LoRA matches or beats QLoRA and LoftQ at $\sim 1/16$ the adapter footprint; the predicted regularization signature appears empirically in dataset ablations and training-loss curves. We extend Punica's SGMV kernel with a binary inner-product primitive (BQ-SGMV) and prove a cache phase-transition result: shrinking adapters $16\times$ enlarges the HBM cache from $2{,}500$ to $39{,}000$ adapters, eliminating misses under skewed workloads.
Chat is not available.
Successful Page Load