Emergence, Retention and Mitigation of Ill-conditioning due to Basis Lifting in KANs
Abstract
Kolmogorov–Arnold Networks (KANs) expand each input coordinate through a fixed basis family, yet the optimization role of this basis lifting remains poorly understood. We identify basis lifting as a structural source of ill-conditioning in KANs and provide a mechanism-level analysis linking lifted feature geometry to convergence, accuracy, and seed sensitivity. Our analysis studies the evolution of lifted empirical Gram matrices under stochastic gradient descent (SGD) and shows that fixed basis lifting induces and preserves near-rank deficiency through eigenvalue and determinant drift. It further concentrates feature energy into a smaller number of directions, amplifying dominant eigenvalues and worsening condition numbers. These results show that basis selection is not merely an expressivity choice, but also an optimization and conditioning choice. Motivated by this mechanism, we introduce T-KAN, a learnable basis-coordinate transformation that aligns basis directions before the learned linear map. It reshapes lifted Gram geometries while preserving the expressivity and introducing negligible parameter overhead. In controlled small-to-moderate KAN regimes, T-KAN often improves conditioning and optimization with B-spline, Chebyshev, and RBF KANs.