Seen-Constrained Model-Order Selection for Unknown-$K$ Generalized Category Discovery
Mingfu Yan ⋅ Jiancheng Huang ⋅ HAIPENG LUO ⋅ Yi Huang ⋅ Yuqi Peng ⋅ Qiang Zhou ⋅ Shifeng Chen
Abstract
Generalized category discovery (GCD) becomes substantially more challenging when the number of novel classes is unknown. We argue that, after the backbone is fixed, this stage should be treated as seen-constrained model-order selection rather than as an auxiliary clustering detail. The selector must estimate the class count without novel labels while remaining consistent with the labeled seen categories available in the standard GCD setting. We instantiate this view by constructing a nested merge path from a seen-constrained overclustering and selecting the coarsest near-optimal cut under observable surrogate decision risks. These risks combine seen-class consistency, cross-seed stability, unlabeled-side compactness, and a tiny-cluster degeneracy penalty. The resulting protocol separates class-count error, downstream clustering quality, and robustness. Where an aligned public unknown-$K$ baseline is available, namely on SelEx, the selector reduces class-count error on all four fine-grained datasets, including a reduction from 342 to 15 classes on Herbarium-19. Broader-domain diagnostics further distinguish ImageNet100, where the default selector can underestimate the class count toward the seen-class regime, from CIFAR100, where the selector landscape exposes an accuracy-count operating-point trade-off. The mixed accuracy outcomes reveal distinct count-, path-, and representation-level bottlenecks, supporting unknown-$K$ selection as a decision layer that should be evaluated separately from representation learning.
Chat is not available.
Successful Page Load