DINOv2 Feature Distillation for Recurrent Local-Update Image Classifiers
Abdullah Al Maruf
Abstract
Recurrent local-update vision models can remain compact in parameter count because the same update rule is reused across multiple computation steps. It is less clear, however, whether representations from large pretrained vision models can be transferred effectively to this type of student. We study feature distillation from a frozen DINOv2 ViT-S/14 teacher to NCA-Lite, a recurrent local-update classifier, on CIFAR-100. A training-only linear projector aligns the student’s latent representation with the teacher’s normalized CLS feature using a cosine loss alongside cross-entropy. Under a controlled three-seed protocol, Feature-KD improves NCA-Lite test accuracy from $51.71 \pm 0.19\%$ to $54.77 \pm 0.62\%$. A closely parameter-count-matched feed-forward CNN also improves, from $53.99 \pm 1.06\%$ to $56.13 \pm 0.75\%$, showing that the benefit is not unique to the recurrent architecture. A validation-only truncation diagnostic indicates that the prediction advantage for NCA-Lite appears mainly at later recurrent states. Despite its compact parameter count, NCA-Lite requires substantially more computation than the feed-forward comparator, highlighting a trade-off between recurrent parameter sharing and computational cost.
Chat is not available.
Successful Page Load