AuGhostmentation: The Eyes Never Stand Still—$\textit{Why Should CNNs?}$
Emirhan Inan ⋅ Suayb Arslan
Abstract
Convolutional Neural Networks (CNNs) draw broad inspiration from the primate visual system yet prove notoriously brittle under common image corruptions, whereas the latter suffers little from such degradations. Extending beyond this initial inspiration, architectures incorporating quantifiable biological priors that enhance network resilience are increasingly gaining traction. Nevertheless, most of these approaches have so far been confined to emulating the higher visual areas, seldom utilizing the upstream retinal stages that constitute the very basis of $\textit{seeing}$. At its core, primate vision relies on photons captured by retinal photoreceptors, the signals of whose conversion are then relayed through the Lateral Geniculate Nucleus (LGN) to primary visual cortex (V1) and further downstream for additional processing. Because cone photoreceptors are densest at the center of the retina (fovea), the eyes must be repositioned continuously (saccade) to direct the high-acuity macula toward objects of interest (fixation). Furthermore, the resultant rapid eye movements induce intra-saccadic motion streaks that putatively aid in the spatiotemporal processing of inputs and provision of translation-invariant object correspondence across frames. Motivated by this mechanism, we propose $\texttt{AuGhostmentation}$, a portmanteau of $\textit{augmentation}$ and $\textit{ghosting}$, which delivers the intra-saccadic retinal smear as a frontend transform. Superimposing the $\textit{ghosted}$ images—fainter counterparts—atop the original, based on human oculomotor statistics, the resultant datum elicits $\textit{retinal persistence}$ by integrating varied recompositions of the scene in a single representation, closely mimicking the foveal exploration scheme intrinsic to dynamic primate vision. When prepended to ResNet-18 trained on Tiny ImageNet, $\texttt{AuGhostmentation}$ retains 98% validation accuracy while delivering a striking 60% relative gain in robustness on Tiny ImageNet-C. Applied to VOneNet, it likewise preserves 98% classification accuracy and yields a substantial 25% relative increase in compounded robustness. For ImageNet-100, the introduced $\textit{ghosts}$ not only improve the corruption accuracy relative to the ResNet-50 baseline by 23%, but the clean accuracy by 4%, also. Building on these resilience gains, $\texttt{AuGhostmentation}$ outpaces competing augmentations by up to 18% and baseline ResNet-50 by up to 37% on stylized tests, suggesting that retinal priors naturally foster shape-biased processing and thereby provide greater alignment with human object recognition.
Chat is not available.
Successful Page Load