BasicLT: Basic-Level Abstraction and Selective Differentiation for Long-Tailed Recognition
Abstract
Long-tailed recognition is commonly formulated as a problem of imbalanced supervision, where rare classes suffer from insufficient training examples. In this work, we study a complementary source of difficulty: under severe scarcity, tail classes may be forced into fine-grained discrimination before the model has acquired sufficiently reliable coarse semantic structure. We refer to this phenomenon as granularity mismatch under scarcity. To address it, we propose BasicLT, a two-stage framework that augments an underlying fine-grained recognition model with basic-level abstraction and commitment-guided selective differentiation. In Stage 1, BasicLT learns an explicit basic-level branch and aligns it with the family-level distribution induced by the fine-grained predictor, encouraging stable coarse semantic organization before additional fine-grained refinement is introduced. In Stage 2, a frozen Stage 1 teacher provides stable family-level commitment signals, and a family-conditioned residual refinement module is trained only on committed medium-shot and few-shot samples. At inference time, the residual correction is activated conservatively according to student-side commitment, so that fine-grained refinement acts as a conditional within-family correction rather than an unconditional auxiliary classifier. We evaluate BasicLT on CIFAR-100-LT, CIFAR-10-LT, ImageNet-LT, iNaturalist 2018, and Places-LT. Across these benchmarks, BasicLT achieves competitive performance against strong long-tailed baselines, with the most consistent gains appearing in medium-shot and few-shot regimes. Ablations and analyses further show that both basic-level abstraction and commitment-guided selective refinement are important, supporting the view that controlling the semantic granularity of discrimination can complement conventional rebalancing strategies for long-tailed learning. Code is available at Supplement.