Attribution-Guided Exit Policy for Reliable Early-Exit Inference
Abstract
Early-exit neural networks improve inference efficiency by enabling predictions at intermediate layers, but their effectiveness critically depends on the exit policy that decides whether to stop or continue computation. Existing policies, often based on confidence thresholds, enable early termination but offer limited insight into the reliability of such decisions. To address this, we introduce an explainability-aware framework that defines an attribution-grounded reliability signal for early-exit decisions. We propose Progressive Feature Attribution Maps (PFAM), which capture how feature importance evolves across exits, and the Interpretability-Based Early-Exit Score (IEES), a unified decision metric that integrates prediction confidence with attribution strength and cross-exit stability. To enable efficient deployment, we further introduce a lightweight proxy predictor that estimates IEES using inexpensive, attribution-free features, allowing attribution-informed decisions without incurring inference-time overhead. Experiments on convolutional architectures (ResNet, MobileNet, MSDNet) show that our approach improves accuracy–efficiency trade-offs, achieving up to 2.5\% accuracy gains while reducing inference time by up to 24\% compared to full-depth inference. Similar trends are observed on a vision transformer and a BERT-based language model, suggesting applicability beyond convolutional architectures.