Causal Gradient Steering: Exposing Shortcut Gradients by Destroying Causal Signal
Ahmed Radwan ⋅ Ahmad Abdel-Qader ⋅ Mahmoud Soliman ⋅ Omar Abdelaziz ⋅ Ahmed Elgazwy ⋅ Shan Du ⋅ Mohamed Shehata
Abstract
It is often easier to destroy the causal content of a training sample than to isolate it. We use this observation to expose shortcut gradients during training. If an intervention removes the semantic evidence that justifies a sample's label while preserving nuisance structure, then any predictive gradient the model still produces on the destroyed view points toward features that can support the label without the causal signal. The destroyed view is constructed from the same sample as the clean view, and the implemented optimizer estimates this shortcut direction on the current minibatch and applies the steering step periodically. Thus CGS injects a recurring within-sample shortcut-correction signal while remaining domain-label-free and close to ERM in cost. We formalize this principle through a nuisance-only Distributionally Robust Optimization geometry, in which an adversary may perturb nuisance factors while holding the semantic component fixed. The resulting local Danskin majorizer induces a first-order robust-feasible cone in parameter space, and Causal Gradient Steering (CGS) is the closed-form projection of the ERM step onto a tractable single-normal approximation of this cone. We further prove target-risk and probe-admissibility guarantees, showing that exact counterfactual generation in pixel space is unnecessary. CGS is compatible with standard first-order training and can act either as a standalone optimizer or as a plug-in correction for existing DG algorithms. It achieves state-of-the-art performance across five standard domain generalization benchmarks. Moreover, using a single locked CGS hyperparameter configuration, it improves seventeen prior DG methods by up to +15.8 percentage points without method-specific retuning, improves OSTrack-256 across GOT-10k, UAV123, and OTB2015, and establishes new matched-protocol state of the art for domain-generalized semantic segmentation by improving SCSD from 51.42 to 52.35 mIoU on GTAV+SYNTHIA$\rightarrow$\{Cityscapes, BDD, Mapillary\} and from 35.66 to 36.99 mIoU on GTAV$\rightarrow$ACDC. Beyond this algorithm, causal signal destruction defines a research direction parallel to data augmentation. The destruction operator is a design axis, and task-specific destroyers can turn shortcut suppression into a reusable tool for broader vision problems.
Chat is not available.
Successful Page Load