Locating and Repairing Domain Shift in VLM Trajectory Planning
Zhihong Cui ⋅ Hengyu Liu ⋅ Michael A. Riegler ⋅ Guandong Xu ⋅ Amir Taherkordi ⋅ Tor Skeie
Abstract
Vision-Language Models (VLMs) are increasingly adopted as end-to-end driving planners, yet they suffer from cross-city domain shift. Existing methods treat this shift as a single quantity addressed by uniform alignment or single-site editing. We refute this view on two counts: ***(i) domain shift in VLM trajectory planners is inherently structured across modality and layer, and (ii) a locate-then-repair recipe matched to this structure outperforms every uniform-objective baseline***. We analyze domain shift via activation patching and uncover two regularities. **(F1) Modality axis.** Image and text tokens exhibit distinct shift patterns in each VLM. **(F2) Layer axis.** The layer-wise shift partition is architecture-determined and dataset-stable. These findings recast domain shift as a **two-dimensional tensor** indexed by *modality* (image vs. text tokens) and *layer*, motivating a simple principle: ***locate where shift occurs, and repair only there***. We instantiate this principle in **MoLaRx** (**Mo**dality–**La**yer **R**epair), a locate-then-repair framework. The Causal Importance Score (CIS) estimates the domain shift as a modality–layer tensor and partitions the layers into functional zones. CIS-Guided Selective Repair (CISR) then applies a matched operator (Maximum Mean Discrepancy (MMD) / low-rank adaptation (LoRA) / freezing) to repair each zone, with the composition adapting automatically to each architecture. We evaluate **MoLaRx** on three architecturally distinct VLMs (LFM2-VL-1.6B, Qwen2-VL-2B, and PaliGemma2-3B) across two cross-city benchmarks (nuScenes and Argoverse 2). MoLaRx attains the lowest cross-city trajectory-error gap on every (VLM, dataset) cell, while using $50$–$60\%$ fewer trainable parameters and $8.4$–$10.9$ ms lower inference latency. In contrast, the strongest scalar baseline (MMD-only) collapses on every cell—direct evidence that no scalar objective can repair structured shift. Code: [anonymous.4open.science/r/MoLaRx-E3B4](https://anonymous.4open.science/r/MoLaRx-E3B4/).
Chat is not available.
Successful Page Load