Revisiting RegressionMSR: A Reproducibility Study
Abstract
Regression-adjusted Monte Carlo estimators for Shapley values and probabilistic values combine surrogate modeling with maximum-sample-reuse estimation to reduce the cost of feature-attribution computation. We present a reproducibility study of RegressionMSR in the interventional feature-attribution setting, focusing on artifact fidelity, practical algorithmic choices, and empirical robustness. We first identify implementation-level differences between the estimator displayed in the paper and the released artifact: the code computes separated conditional means over coalitions containing and excluding each feature, rather than the displayed all-row signed residual average, and it is also not guaranteed to use a probability mass in its residual denominator. We characterize these conventions algebraically and verify their predicted targets with synthetic diagnostics. On real data, the stored denominator strongly attenuates the residual correction: correcting both conventions changes LinearMSR little but reduces TreeMSR error by more than half in the tested Shapley settings. We then test the authors' practical single-surrogate simplification, finding that it substantially reduces runtime while maintaining and often slightly improving accuracy. Next, we qualitatively reproduce the headline Shapley benchmark among the original comparison methods, where TreeMSR remains the lowest-error estimator on average; however, OddSHAP performs best in our modern benchmark extension, and rank-stability metrics provide a complementary ordering. Finally, we extend the original Gaussian-noise experiment to Laplacian and coalition-size-dependent heteroskedastic noise, finding that TreeMSR's low-noise advantage persists across other low-noise regimes. Overall, our results support RegressionMSR's empirical value while clarifying implementation details needed for reproducible use.