Timezone: »
We develop algorithms for imitation learning from policy data that was corrupted by unobserved confounders. Sources of such confounding include (a) persistent perturbations to actions or (b) the expert responding to a part of the state that the learner does not have access to. When a confounder affects multiple timesteps of recorded data, it can manifest as spurious correlations between states and actions that a learner might latch on to, leading to poor policy performance. To break up these spurious correlations, we apply modern variants of the classical instrumental variable regression (IVR) technique, enabling us to recover the causally correct underlying policy without requiring access to an interactive expert. In particular, we present two techniques, one of a generative-modeling flavor (DoubIL) that can utilize access to a simulator and one of a game-theoretic flavor (ResiduIL) that can be run entirely offline. We discuss, from the perspective of performance, the types of confounding under which it is better to use an IVR-based technique instead of behavioral cloning and vice versa. We find both of our algorithms compare favorably to behavioral cloning on a simulated rocket landing task.
Author Information
Gokul Swamy (Carnegie Mellon University)
Sanjiban Choudhury (Aurora Innovation)
James Bagnell (Aurora Innovation)
Steven Wu (Carnegie Mellon University)
More from the Same Authors
-
2021 : What Would the Expert do()?: Causal Imitation Learning »
Gokul Swamy · Sanjiban Choudhury · James Bagnell · Steven Wu -
2021 : Iterative Methods for Private Synthetic Data: Unifying Framework and New Methods »
Terrance Liu · Giuseppe Vietri · Steven Wu -
2021 : What Would the Expert do()?: Causal Imitation Learning »
Gokul Swamy · Sanjiban Choudhury · James Bagnell · Steven Wu -
2021 : What Would the Expert $do(\cdot)$?: Causal Imitation Learning »
Gokul Swamy · Sanjiban Choudhury · James Bagnell · Steven Wu -
2021 : Bayesian Persuasion for Algorithmic Recourse »
Keegan Harris · Valerie Chen · Joon Sik Kim · Ameet Talwalkar · Hoda Heidari · Steven Wu -
2021 : What Would the Expert $do(\cdot)$?: Causal Imitation Learning »
Gokul Swamy · Sanjiban Choudhury · James Bagnell · Steven Wu -
2021 : Information Discrepancy in Strategic Learning »
Yahav Bechavod · Chara Podimata · Steven Wu · Juba Ziani -
2021 : Gaming Helps! Learning from Strategic Interactions in Natural Dynamics »
Yahav Bechavod · Katrina Ligett · Steven Wu · Juba Ziani -
2021 : Bayesian Persuasion for Algorithmic Recourse »
Keegan Harris · Valerie Chen · Joon Kim · Ameet S Talwalkar · Hoda Heidari · Steven Wu -
2021 : What Would the Expert $do(\cdot)$?: Causal Imitation Learning »
Gokul Swamy · Sanjiban Choudhury · James Bagnell · Steven Wu -
2021 : What Would the Expert $do(\cdot)$?: Causal Imitation Learning »
Gokul Swamy · Sanjiban Choudhury · James Bagnell · Steven Wu -
2021 : Information Discrepancy in Strategic Learning »
Yahav Bechavod · Chara Podimata · Steven Wu · Juba Ziani -
2021 : Gaming Helps! Learning from Strategic Interactions in Natural Dynamics »
Yahav Bechavod · Katrina Ligett · Steven Wu · Juba Ziani -
2021 : Bayesian Persuasion for Algorithmic Recourse »
Keegan Harris · Valerie Chen · Joon Kim · Ameet S Talwalkar · Hoda Heidari · Steven Wu -
2022 : Strategy-Aware Contextual Bandits »
Keegan Harris · Chara Podimata · Steven Wu -
2022 : Choosing Public Datasets for Private Machine Learning via Gradient Subspace Distance »
Xin Gu · Gautam Kamath · Steven Wu -
2022 : Strategy-Aware Contextual Bandits »
Keegan Harris · Chara Podimata · Steven Wu -
2022 : Strategy-Aware Contextual Bandits »
Keegan Harris · Chara Podimata · Steven Wu -
2022 : Differentially Private Gradient Boosting on Linear Learners for Tabular Data »
Saeyoung Rho · Shuai Tang · Sergul Aydore · Michael Kearns · Aaron Roth · Yu-Xiang Wang · Steven Wu · Cedric Archambeau -
2022 : Counterfactual Decision Support Under Treatment-Conditional Outcome Measurement Error »
Luke Guerdan · Amanda Coston · Kenneth Holstein · Steven Wu -
2022 Poster: On Privacy and Personalization in Cross-Silo Federated Learning »
Ken Liu · Shengyuan Hu · Steven Wu · Virginia Smith -
2022 Poster: Brownian Noise Reduction: Maximizing Privacy Subject to Accuracy Constraints »
Justin Whitehouse · Aaditya Ramdas · Steven Wu · Ryan Rogers -
2022 Poster: Incentivizing Combinatorial Bandit Exploration »
Xinyan Hu · Dung Ngo · Aleksandrs Slivkins · Steven Wu -
2022 Poster: Sequence Model Imitation Learning with Unobserved Contexts »
Gokul Swamy · Sanjiban Choudhury · J. Bagnell · Steven Wu -
2022 Poster: Private Synthetic Data for Multitask Learning and Marginal Queries »
Giuseppe Vietri · Cedric Archambeau · Sergul Aydore · William Brown · Michael Kearns · Aaron Roth · Ankit Siva · Shuai Tang · Steven Wu -
2022 Poster: Minimax Optimal Online Imitation Learning via Replay Estimation »
Gokul Swamy · Nived Rajaraman · Matt Peng · Sanjiban Choudhury · J. Bagnell · Steven Wu · Jiantao Jiao · Kannan Ramchandran -
2022 Poster: Bayesian Persuasion for Algorithmic Recourse »
Keegan Harris · Valerie Chen · Joon Kim · Ameet Talwalkar · Hoda Heidari · Steven Wu -
2021 : What Would the Expert $do(\cdot)$?: Causal Imitation Learning (Gokul Swamy) »
Gokul Swamy -
2021 : Leveraging strategic interactions for causal discovery »
Steven Wu -
2021 : Bayesian Persuasion for Algorithmic Recourse »
Keegan Harris · Valerie Chen · Joon Sik Kim · Ameet Talwalkar · Hoda Heidari · Steven Wu -
2021 : Contributed Talk 2: What Would the Expert do?: Causal Imitation Learning »
Gokul Swamy -
2021 : What Would the Expert do()?: Causal Imitation Learning »
Gokul Swamy · Sanjiban Choudhury · James Bagnell · Steven Wu -
2021 Poster: Iterative Methods for Private Synthetic Data: Unifying Framework and New Methods »
Terrance Liu · Giuseppe Vietri · Steven Wu -
2021 Poster: Stateful Strategic Regression »
Keegan Harris · Hoda Heidari · Steven Wu -
2020 Poster: Metric-Free Individual Fairness in Online Learning »
Yahav Bechavod · Christopher Jung · Steven Wu -
2020 Poster: Understanding Gradient Clipping in Private SGD: A Geometric Perspective »
Xiangyi Chen · Steven Wu · Mingyi Hong -
2020 Poster: Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based Algorithms »
Xiangyi Chen · Tiancong Chen · Haoran Sun · Steven Wu · Mingyi Hong -
2020 Spotlight: Understanding Gradient Clipping in Private SGD: A Geometric Perspective »
Xiangyi Chen · Steven Wu · Mingyi Hong -
2020 Oral: Metric-Free Individual Fairness in Online Learning »
Yahav Bechavod · Christopher Jung · Steven Wu -
2020 Session: Orals & Spotlights Track 20: Social/Adversarial Learning »
Steven Wu · Miro Dudik -
2019 Poster: Equal Opportunity in Online Classification with Partial Feedback »
Yahav Bechavod · Katrina Ligett · Aaron Roth · Bo Waggoner · Steven Wu -
2019 Poster: Random Quadratic Forms with Dependence: Applications to Restricted Isometry and Beyond »
Arindam Banerjee · Qilong Gu · Vidyashankar Sivakumar · Steven Wu -
2019 Poster: Private Hypothesis Selection »
Mark Bun · Gautam Kamath · Thomas Steinke · Steven Wu -
2019 Poster: Locally Private Gaussian Estimation »
Matthew Joseph · Janardhan Kulkarni · Jieming Mao · Steven Wu -
2018 : Invited Talk: Drew Bagnell, CMU and Aurora »
James Bagnell -
2018 : Drew Bagnell / Wen Sun »
James Bagnell · Wen Sun -
2017 : Spotlights »
Antti Kangasrääsiö · Richard Everett · Yitao Liang · Yang Cai · Steven Wu · Vidya Muthukumar · Sven Schmit -
2017 Poster: Accuracy First: Selecting a Differential Privacy Level for Accuracy Constrained ERM »
Katrina Ligett · Seth Neel · Aaron Roth · Bo Waggoner · Steven Wu -
2016 Poster: Learning from Rational Behavior: Predicting Solutions to Unknown Linear Programs »
Shahin Jabbari · Ryan Rogers · Aaron Roth · Steven Wu