CARVE: Counterfactual Video Editing for Auditing and Hardening Video Detectors
Wei Zhou ⋅ YIMING CHEN ⋅ Xu Jinwei ⋅ Yang zhou ⋅ Quan Gan ⋅ Li Yang
Abstract
Video detectors are usually evaluated on observational splits, where target events co-occur with environmental and capture factors. A high AUC under this setup can reflect either the event itself or its surrounding conditions, and standard reporting cannot tell them apart. This work addresses the entanglement with CARVE, a counterfactual video editing protocol. For each source clip, CARVE generates a matched quartet $(V^0, V^E, V^A, V^{AE})$ that holds camera pose, layout, and surrounding traffic fixed while independently intervening on environment and event. A two-phase generator first produces a layout-conditioned reference image, then performs reference-guided video editing; a three-layer reference / objective / VLM-panel protocol filters the output. Each quartet supports thresholded and continuous diagnostics for false-positive purity, event faithfulness, environment-induced drift, and held-out factor compositions. Thresholded scores are paired with continuous ones so that a conservative detector cannot suppress scores everywhere and look robust. The same audit drives training: CGAA-IC samples weak factors weighted by measured brittleness and adds quartet-level purity, consistency, and faithfulness losses; the inference graph is unchanged. On CCTV accident detection, the audit shows that clean AUC does not predict robustness rankings: a VideoMAE detector with $91.4\%$ AUC on TAD scores only $72.3\%$ CPS on the night slice of CARVE-Q. CGAA-IC raises purity from $84.2\%$ to $89.9\%$ and cross-dataset average AUC from $80.3\%$ to $84.5\%$.
Chat is not available.
Successful Page Load