VVTRec: Radio Interferometric Reconstruction through Visual and Textual Modality Enrichment
Abstract
Radio astronomy is an indispensable discipline for studying distant celestial objects. Measurements of wave signals from radio telescopes, called visibility, need to be transformed into images for astronomical observations. These dirty images blend information from real sources and artifacts. Therefore, astronomers usually perform reconstruction before imaging to obtain cleaner images. Existing methods consider only a single modality of sparse visibility data, resulting in images with remaining artifacts and insufficient modeling of correlation. We propose VVTRec, a multimodal radio interferometric data reconstruction method, to enhance visibility information extraction and improve image-domain output quality. Since the information in sparse visibility is inherently limited, we transform it into image-form and text-form features. Therefore, we can leverage vision-language models to process the two derived modalities, providing knowledge banks as a supplement for reconstruction. The sparse visibility is accordingly utilized as the query to perform the integration and selection of multimodal external knowledge. Consequently, VVTRec improves the structural integrity and accuracy of reconstructed images by enriching spatial and semantic information. Our experiments demonstrate that VVTRec effectively enhances imaging results by exploiting multimodal information without introducing excessive computational overhead.