WaveSem: Frequency-Adaptive Tokenization for Disentangling Semantics and Noise in Genomics
Shou Z Chen ⋅ Bing He ⋅ Zhenchao Tang ⋅ Jun Zhu ⋅ Minghao Yang ⋅ Tianxu Lv ⋅ Yu Wang ⋅ Jiayang Wu ⋅ Fang Wang ⋅ Yu Zhao ⋅ Chenchen Qin ⋅ Jianhua Yao
Abstract
High-throughput scientific sensing, exemplified by genomics, faces a critical representation bottleneck: discrete biological events are often obscured by continuous, high-entropy stochastic noise. Standard neural codecs fail to distinguish deterministic semantics from stochastic sensor noise, inefficiently allocating their information budget to high-entropy fluctuations rather than to low-entropy structural motifs. We introduce \textsc{WaveSem}, a frequency-adaptive framework that imposes a physically motivated inductive bias on discrete representation learning. Through integrated wavelet decomposition, \textsc{WaveSem} explicitly disentangles each signal into a structured semantic core and a residual texture component. We validate this approach on large-scale nanopore sequencing data, a domain characterized by extreme sequence lengths and low signal-to-noise ratios (SNRs). \textsc{WaveSem} achieves a $14\times$ storage reduction while maintaining competitive downstream basecalling accuracy. Moreover, linear probing of the frozen encoder reaches $\mathrm{AUROC}=0.928$ for 5mC methylation detection, demonstrating that the learned tokens preserve semantic information beyond nucleotide identity. These results establish disentangled vector-quantized representations as a scalable foundation for efficient genomic analysis and long-term archival.
Chat is not available.
Successful Page Load