Rad-VL-JEPA: A Unified Visual Tool for Agentic 3D Chest CT Reporting
Abstract
Chest CT reporting agents increasingly use tool calling to combine global context with patient-specific visual evidence. However, they often rely on separate models for different tool functions, many of which are evaluated only through end-to-end agent performance rather than directly on their intended tool tasks. As a result, errors from intermediate tools may introduce unsupported evidence into the final report. We introduce Rad-VL-JEPA, a unified queryable visual tool for report retrieval, finding recognition, and fine-grained attribute characterization. Rad-VL-JEPA maps CT volumes and text queries into a joint embedding space. The learned representation supports different downstream tasks through a unified visual tool interface. We show that Rad-VL-JEPA performs strongly across all three tasks compared with specialized baselines. We further show how these capabilities can be composed into a CT reporting agent, outperforming existing RRG baselines.