What matters in a bridge? Analysing Bridge Networks for Molecular Property Prediction
Abstract
Large language models (LLMs) often struggle to predict molecular properties on unseen datasets. Bridge networks address this limitation by mapping representations from a frozen encoder into soft tokens for a frozen LLM, but the contribution of each component remains unclear. We present a controlled factorial study of bridge networks for molecular property prediction. Across 72 training runs, we vary four encoders, three bridge architectures, and two LLMs over three seeds, and evaluate them on eight recent tasks selected to reduce data-contamination risk. The best bridged configurations outperform zero-shot LLM baselines on all eight tasks and encoder-only probes on six. The simple MLP bridge performs best overall, whereas the vocabulary-anchored bridge is usually the weakest. Decoder choice explains only about 1\% of the error variance, while seed variation is the largest source of variance on most regression tasks. We also find that a general-purpose BERT encoder remains surprisingly competitive with chemistry-specific encoders, raising questions about the information transferred through these bridges.