Bottleneck Benchmark: Sequence Compression Under Controlled Difficulty and Fixed Latent Size
Abstract
One key application of autoencoders is to learn how to compress token sequences. Since the difficulty of compression depends on factors such as sequence length, vocabulary size, structural regularity, and input noise, it would be valuable to understand how these factors affect the performance of different architectures. Surprisingly, however, there is no benchmark for comparing autoencoder architectures that allows control over these aspects of difficulty. We introduce the Bottleneck Benchmark (BB), an open experiment framework in which encoder and decoder communicate only through a fixed-size latent and that allows to configure the previously mentioned difficulty aspects for any customizable data generator. As a case study, we evaluate BiLSTM, Transformer, and Bidirectional SSM encoders on copy, sort, and set tasks, finding that performance differences arise from inductive biases under noise and length rather than raw capacity.