Aligning Sparse 5G Telemetry with Diagnostic Text: A Leakage-Aware Study of Contrastive Cross-Modal Retrieval
Abstract
Operating a 5G Radio Access Network produces two records of every fault that are never joined: machine-generated Performance Management (PM) telemetry, a sparse high-dimensional multivariate time series that captures the fault's counter signature, and a human-written Trouble Report (TR) that captures the diagnosis. We study whether contrastive learning can align these two modalities in a shared space for cross-modal retrieval, under conditions that depart from the image-caption setting where such alignment was established: the reports form many-to-one fault families, and the paired corpus is modest (~4,300 examples). A dual-encoder pipeline pairs a time-series encoder with a frozen domain-pretrained text encoder, evaluated on a leakage-aware split in which whole fault families are held out atomically. Two controlled ablations test whether self-supervised encoder pretraining helps and which contrastive objective suits the domain. The selected model reaches Hit@10 of 15.4% (95% CI 12.4-18.4), a ten-fold gain over an unaligned baseline, yet most queries still retrieve no relevant report in the top ten. Pretraining accelerates convergence without a measurable change in final retrieval quality, and no fault-family-aware objective improves on plain CLIP. Embedding diagnostics locate the binding constraint on the telemetry side: learned telemetry embeddings separate fault families far more weakly (permutation z ~ 8-10) than the frozen text embeddings (z = 37.8). The binding constraint is the temporal representation this pipeline learns, not the contrastive objective layered on top.