Angular Networks: Low-Bit Learning from Randomized Similarity Estimators
Abstract
We propose a new framework for low-precision neural networks based on angular geometry and randomized low-bit estimators. Rather than approximating Euclidean linear operations under limited precision, we reinterpret neural computation in terms of directional similarity and construct quantized random-feature estimators of cosine interactions. Our approach introduces \textit{Angular Layers}, which replace standard linear transformations with low-bit projections that estimate angular similarity in the forward pass, while gradients are computed with respect to the underlying continuous cosine geometry. This decoupling enables stable optimization without straight-through estimators or heuristic surrogate gradients, even under aggressive quantization of both weights and activations. On the theory side, we introduce \emph{G-Gradient} as a gradient proxy for angular layers and show that, in a one-layer planted model, a sufficiently small \emph{G}-gradient can guarantee exact recovery of a target binary solution. We further prove uniform approximation and convergence guarantees that provide a rigorous path from tractable continuous optimization to exact recovery in the original discrete model. Empirically, we instantiate these ideas in \textit{Angular Nets}, including low-precision variants of ResNet, Vision Transformers, and BERT-style encoders, and obtain strong performance across vision and language tasks at substantially reduced precision. Overall, our results suggest that preserving angular structure provides a principled foundation for low-bit deep learning. Our full implementation is available at: \url{https://github.com/GGM2026/GGM}.