Intervene3D: Intervention-Based Controlled Inference for Multimodal Perception under Partial Observability
Constantino Msigwa ⋅ Denis Bernard ⋅ Jaeseok Yun
Abstract
Multimodal 3D detection becomes unreliable under partial observability caused by occlusion, LiDAR sparsity, and camera--LiDAR miscalibration. Most fusion detectors are fully feed-forward: once trained, they provide no explicit mechanism to incorporate structured test-time constraints or to bound how such constraints can alter predictions. We propose Intervene3D, a controlled-inference wrapper for BEV-based detectors that performs metric-bounded latent refinement at inference time. Given sensor inputs, the detector produces an initial BEV latent $\mathbf z_0$ and an evidence-conditioned trust metric $\mathbf H$ (a diagonal precision derived from predicted uncertainty). A separate control stream is mapped to a fully specified differentiable constraint energy (count intervals and spatial admissibility). Inference then updates $\mathbf z$ using $K\in\{1,2,3\}$ diagonal-metric preconditioned projected steps that reduce constraint energy while remaining within an uncertainty-gated trust region defined by $\mathbf H$. To improve robustness to extrinsic drift, we additionally introduce \textbf{phase-aware spectral alignment} that normalizes cross-modal Fourier-phase discrepancies prior to attention. We report clean accuracy and a severity-swept robustness protocol (robustness-AUC, feasibility diagnostics, and conflict-induced FP inflation).
Chat is not available.
Successful Page Load