Beyond Generation: Unlocking Discriminative Representations from Diffusion Models
Haowen Cui ⋅ Ge Wu ⋅ Shuo Chen ⋅ Yikai Ge ⋅ Ge Gao ⋅ Xiang Li ⋅ Jun Li ⋅ Jian Yang
Abstract
Recent advancements in generative models have increasingly leveraged self-supervised visual representations to guide the synthesis process, significantly improving generation quality and semantic consistency. However, the reverse paradigm of harnessing the generative process to enhance visual representations remains largely underexplored. In this paper, we investigate the semantic properties of class tokens synthesized by generative models. We observe that generative models equipped with representation entanglement can generate class tokens that exhibit stronger discriminative capabilities than those extracted from the original pretrained visual models. Motivated by this observation, we propose a completely new Generation-to-Perception Knowledge Distillation (GPKD) framework, where our method generates class tokens as teacher signals to instill global semantics into the student model. To prevent the degradation of local details, we further incorporate a new masked patch-level distillation objective. This dual distillation strategy enhances global representations while mitigating the forgetting of local details, thereby producing more robust representations. Extensive experiments demonstrate that GPKD obtains consistent improvements across various downstream tasks compared with the base visual encoder, achieving a gain of more than 2\% in $k$-NN accuracy on ImageNet-1K.
Chat is not available.
Successful Page Load