$\mathcal{P}$Torch: Narrowing the Gap Between Projection and Gradient-Based Learning
Jeff Cyuzuzo Jambé ⋅ Jan Quan ⋅ Panagiotis Patrinos
Abstract
Projection-based optimization offers a gradient-free alternative to conventional gradient-based methods, but existing frameworks suffer from high computational costs and poor scaling with depth. We introduce $\mathcal{P}$Torch, a projection-based framework that builds directly on top of PyTorch's autograd engine, eliminating the overhead of prior approaches. We further improve performance through nonlinear relaxation and memory-efficient projections. Together, these contributions make the algorithm competitive with gradient-based methods for shallow MLPs on MNIST and CIFAR-10. Finally, we prove a vanishing target theorem showing that the layer-wise learning signal decays with increased depth, serving as an analogue of the vanishing gradients problem for projection-based methods.
Chat is not available.
Successful Page Load