Can Bits Seal Language?
Abstract
We present an encoding method that maps each input token to a randomized sequence of binary values (0s and 1s). By operating at the finest representational granularity and applying compression, this approach reduces the effective vocabulary to just two symbols, yielding a compact, task-agnostic representation. Across eleven main tasks spanning six modalities (human text, genetic/clinical data, programming code, system logs, biological signals, and tabular medical data), it achieves up to two orders of magnitude stronger privacy guarantees under six gradient-based attack settings compared to seven encoder baselines and four defense baselines, including up to 549× improvement on FILM and 62× on TAG. At the same time, it reduces embedding model size by up to 98%, while maintaining competitive accuracy and computational efficiency. It also improves multilingual generalization, achieving up to a 4× increase in accuracy in low-resource settings.