fpsketch: low-dimensional embeddings of molecular fingerprints without inflated similarities
Austin Tripp
Abstract
Molecular fingerprints are sparse, high-dimensional vectors that are typically ``folded'' into a lower dimension $d$ before use. An undesired side effect of this folding is that Tanimoto similarities between fingerprints are systematically inflated. % We show that a well-known algorithm called CountSketch can compress fingerprints without this bias. We furthermore extend this to count fingerprints using a unary encoding of counts, allowing Tanimoto similarities of count fingerprints to be approximated using dot products instead of min/max operations. The resulting method, \textsc{fpsketch}, requires only hash functions to implement, needs no fitting, and is unbiased regardless of $d$ or the number of substructures represented. On drug-like molecules from ZINC, \textsc{fpsketch} matches true Tanimoto similarity far more closely than traditional folding at every dimension, removing the traditional trade-off between fingerprint richness and output dimension. We release \textsc{fpsketch} as an open-source python package.
Chat is not available.
Successful Page Load