Learning Shared Crystal-Language Representations through Progressive Alignment
Abstract
Text encoders are the workhorse of materials information retrieval, yet they have no notion of a periodic crystal structure. A chemical formula, a space-group symbol, and a set of fractional coordinates are, to a sentence encoder, ordinary words disconnected from the physical properties they determine. In this work, we explore which forms of supervision are required to turn a text encoder into a shared crystal–language representation. Rather than presenting a single monolithic method, we decompose training into three sequential stages: masked language modeling, contrastive alignment to geometric representations from a frozen interatomic potential, and contrastive alignment between structural serializations and natural-language descriptions. All three stages update a shared ModernBERT backbone, yielding a single text encoder enriched for materials structures and text. We freeze the checkpoint after each stage and evaluate its representations using linear probes on Matbench and bidirectional retrieval across six tasks drawn from three separate data releases. Within this training suite, masked language modeling has little effect in isolation but provides a beneficial initialization for the complete alignment pipeline, while potential and textual alignment primarily enhance property decodability and cross-modal retrieval, respectively. With these three stages, we introduce MATCHAi, a text-based encoder model, that achieves the best aggregate Matbench rank among the evaluated frozen text encoders after atomistic alignment, while the final encoder ranks first among the 19 evaluated encoders across all six retrieval tasks and remains competitive for materials-property prediction.