Point Clustering Encoders
Abstract
Learning expressive and transferable representations from unstructured 3D data remains a fundamental challenge. Existing backbones mirror the design of image-based networks by relying on symmetric U-Net architectures with large parametric decoders and feature skip connections. While successful in 2D, these designs are ill-suited for 3D domain where point coordinates, that define the underlying geometric manifold on which feature representations are learned, are sensitive to sensor-specific sampling patterns. In such architectures, skip connections allow low-level coordinate cues to bypass the semantic bottleneck, leading the model to overfit to local spatial patterns rather than learning robust, transferable semantic abstractions. We re-think this design and introduce Point Clustering Encoders (PCE), a minimal, decoder-free network that treats point cloud processing as a hierarchy of end-to-end learned spatial supertokens. At the core of PCE is our Time-Reversible Attention Pooling (TRAP) layer, which formalizes hierarchical downsampling as a time-reversible Markov chain. By coupling upsampling to the reverse chain, we eliminate the need for traditional decoders; the task of representation learning is fully delegated to the encoder. PCE not only surpasses state-of-the-art w.r.t. segmentation accuracy in indoor/outdoor datasets across a variety of tasks, but also reduces the number of trainable parameters, and can process scenes consisting of up to 6.5M points.