Timezone: »
Medical studies frequently require to extract the relationship between each covariate and the outcome with statistical confidence measures. To do this, simple parametric models are frequently used (e.g. coefficients of linear regression) but always fitted on the whole dataset. However, it is common that the covariates may not have a uniform effect over the whole population and thus a unified simple model can miss the heterogeneous signal. For example, a linear model may be able to explain a subset of the data but fail on the rest due to the nonlinearity and heterogeneity in the data. In this paper, we propose DDGroup (data-driven group discovery), a data-driven method to effectively identify subgroups in the data with a uniform linear relationship between the features and the label. DDGroup outputs an interpretable region in which the linear model is expected to hold. It is simple to implement and computationally tractable for use. We show theoretically that, given a large enough sample, DDGroup recovers a region where a single linear model with low variance is well-specified (if one exists), and experiments on real-world medical datasets confirm that it can discover regions where a local linear model has improved performance. Our experiments also show that DDGroup can uncover subgroups with qualitatively different relationships which are missed by simply applying parametric approaches to the whole dataset.
Author Information
Zachary Izzo (Stanford University)
Ruishan Liu (Stanford University)
James Zou (Stanford)
More from the Same Authors
-
2022 : Predicting Immune Escape with Pretrained Protein Language Model Embeddings »
Kyle Swanson · Howard Chang · James Zou -
2022 : Is Unsupervised Performance Estimation Impossible When Both Covariates and Labels shift? »
Lingjiao Chen · Matei Zaharia · James Zou -
2022 : DrML: Diagnosing and Rectifying Vision Models using Language »
Yuhui Zhang · Jeff Z. HaoChen · Shih-Cheng Huang · Kuan-Chieh Wang · James Zou · Serena Yeung -
2022 : Provable Re-Identification Privacy »
Zachary Izzo · Jinsung Yoon · Sercan Arik · James Zou -
2022 : Recommendation for New Drugs with Limited Prescription Data »
Zhenbang Wu · Huaxiu Yao · Zhe Su · David Liebovitz · Lucas Glass · James Zou · Chelsea Finn · Jimeng Sun -
2022 : An Electrocardiogram-Based Risk Score for Cardiovascular Mortality »
John Hughes · David Ouyang · Pierre Elias · James Zou · Euan Ashley · Marco Perez -
2022 : An Electrocardiogram-Based Risk Score for Cardiovascular Mortality »
John Hughes · David Ouyang · Pierre Elias · James Zou · Euan Ashley · Marco Perez -
2022 Poster: Estimating and Explaining Model Performance When Both Covariates and Labels Shift »
Lingjiao Chen · Matei Zaharia · James Zou -
2022 Poster: SkinCon: A skin disease dataset densely annotated by domain experts for fine-grained debugging and analysis »
Roxana Daneshjou · Mert Yuksekgonul · Zhuo Ran Cai · Roberto Novoa · James Zou -
2022 Poster: HAPI: A Large-scale Longitudinal Dataset of Commercial ML API Predictions »
Lingjiao Chen · Zhihua Jin · Evan Sabri Eyuboglu · Christopher RĂ© · Matei Zaharia · James Zou -
2022 Poster: Uncalibrated Models Can Improve Human-AI Collaboration »
Kailas Vodrahalli · Tobias Gerstenberg · James Zou -
2022 Poster: C-Mixup: Improving Generalization in Regression »
Huaxiu Yao · Yiping Wang · Linjun Zhang · James Zou · Chelsea Finn -
2022 Poster: Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning »
Victor Weixin Liang · Yuhui Zhang · Yongchan Kwon · Serena Yeung · James Zou -
2022 Poster: WeightedSHAP: analyzing and improving Shapley based feature attributions »
Yongchan Kwon · James Zou -
2021 Poster: Adversarial Training Helps Transfer Learning via Better Representations »
Zhun Deng · Linjun Zhang · Kailas Vodrahalli · Kenji Kawaguchi · James Zou -
2021 Poster: Dimensionality Reduction for Wasserstein Barycenter »
Zachary Izzo · Sandeep Silwal · Samson Zhou -
2020 Session: Orals & Spotlights Track 02: COVID/Health/Bio Applications »
Tristan Naumann · James Zou -
2019 Poster: Making AI Forget You: Data Deletion in Machine Learning »
Antonio Ginart · Melody Guan · Gregory Valiant · James Zou -
2019 Spotlight: Making AI Forget You: Data Deletion in Machine Learning »
Antonio Ginart · Melody Guan · Gregory Valiant · James Zou -
2017 Workshop: Machine Learning in Computational Biology »
James Zou · Anshul Kundaje · Gerald Quon · Nicolo Fusi · Sara Mostafavi -
2017 Poster: NeuralFDR: Learning Discovery Thresholds from Hypothesis Features »
Fei Xia · Martin J Zhang · James Zou · David Tse