CAST: Certifiable Aggregation of Smoothed Teachers for Robust Policy Adaptation
Abstract
Transfer learning (TL)-based policy adaptation in deep reinforcement learning (DRL) usually relies on target-side data information for retraining when facing new tasks. However, two important issues remain underexplored in practical scenarios: existing DRL transfer learning methods usually lack theoretical guarantees against adversarial attacks, and target-side data information may be unavailable for retraining. In this paper, we propose Certifiable Aggregation of Smoothed Teachers (CAST), which adapts source policies by aggregating multiple certified smoothed teachers without any form of retraining, and further certifies the robustness of the adapted policy at both the action level and the cumulative reward level. CAST certification faces three key challenges: (i) value-function transfer in DRL cannot preserve certified robustness without robust retraining; (ii) reward changes break the direct reusability of source teachers' action-level certificates; and (iii) existing certified robustness transfer results provide limited guidance for certifying the student's cumulative reward. These challenges prevent us from directly using existing TL techniques in DRL. To address them, CAST constructs robustness signals from source certificates, incorporates them into policy aggregation to obtain action-level certificates, and bridges the adapted policy's action-level certificates to its target-task cumulative-reward certificate. Experiments on five multi-objective DRL benchmarks show that CAST exceeds the best source teacher in 17 out of 40 attack configurations and largely remains within the performance range of the source teachers. When testing whether smoothing benefits can be transferred, CAST improves in 39 out of 40 configurations, with a maximum relative gain of 183.9\%.