Interpretable Geometric Tokens: Training-Free LiDAR Scene Representations via Conformal Geometric Algebra
Ilya Afanasyev
Abstract
Transformers and graph neural networks for 3D perception require a tokenization step: dense unstructured LiDAR points must be reduced to a bounded set of discrete units before attention or message passing. Learned tokenizers produce latent vectors whose coordinates carry no physical meaning. We propose a training-free alternative: a real-time tokenizer that abstracts a LiDAR scan into a sparse set of geometric tokens, 8.7 per frame on average. Each token is a Conformal Geometric Algebra (CGA) primitive, a plane or a sphere, stored as a 4- or 5-coefficient vector of the conformal space $\mathbb{R}^{4,1}$ with physically interpretable entries. Point-to-token assignment reduces to a single algebraic inner product, and CGA motors carry tokens across frames (identity motors in our benchmark). Carried-over tokens explain most points of the next frame without re-clustering, which contributes a 10× speedup and a 1.3–1.5× compactness gain in ablations. On 28 KITTI sequences (8,307 scans), the token stream is over 2500× smaller than the raw point stream. A hybrid mode (tokens plus raw-stored residuals) reaches 17.7× mean compression at 53.8 dB D1-PSNR and about 44 FPS on a laptop CPU. Treated as sparse measures over a primitive manifold, token sets support optimal transport directly: distances computed on about 140 bytes per scene recover drive identity (retrieval mAP 0.27, chance 0.03) and structural scene regime (mAP 0.42, chance 0.32). We position the token stream as an interpretable, transport-ready substrate for geometric transformers, GNNs, and distributional objectives.
Chat is not available.
Successful Page Load