TreeProp: Training $N$-Layer Deep Networks with $\mathcal{O}(\log N)$ Parallel Time Complexity
Neeraj Mohan Sushma ⋅ Aditya Nagarsekar ⋅ Cabrel Teguemne Fokam ⋅ Robin Schiewer ⋅ Amit K Pal ⋅ Christian Mayr ⋅ Anand Subramoney ⋅ David Kappel
Abstract
Modern deep neural networks are trained with error backpropagation, whose sequential forward and backward dependencies limit parallel execution across layers. We propose *TreeProp*, an architecture-agnostic variational learning framework that reorganizes network layers into a tree-structured hierarchy. TreeProp replaces sequential training paths with hierarchical computation, reducing the critical path for both the forward pass and the backward gradient propagation from ${\mathcal{O}(N)\}$ to ${\mathcal{O}(\log N)\}$ for a network of $N$ layers. To the best of our knowledge, TreeProp is the first general DNN learning framework with logarithmic parallel time complexity for both forward and backward training computations. We evaluate TreeProp on vision classification and autoregressive language modeling, obtaining competitive performance with end-to-end backpropagation and outperforming contrastive training baselines.
Chat is not available.
Successful Page Load