Beyond FLOPs: Train-Full, Deploy-Partial Multi-Exit Inference via Selective Lightweight IC Ensemble
Abstract
Early-exit networks promise inference savings by terminating computation at intermediate classifiers, but FLOPs-based gains rarely translate into on-device latency reduction under static-graph compilation. On NVIDIA Jetson Orin Nano, the per-sample exit policy of Shallow-Deep Networks (SDN) runs approximately 2× slower than vanilla ResNet-56 and misses the 60 fps deadline despite using only 40% of the FLOPs, because dynamic branching precludes layer fusion. We propose the Selective Lightweight IC Ensemble Network (SLIENet), a train-full, deploy-partial framework compiled as a single static FP16 engine. SLIENet trains the full backbone with light self-distillation, using the final classifier as a stop-gradient teacher to preserve error diversity across internal classifiers (ICs). At deployment, a sub-second calibration search selects an IC subset for softmax averaging, while overthinking late blocks are truncated at depth k. On CIFAR-100 across ResNet-56, VGG-16, and MobileNetV1, SLIENet outperforms SDN-based variants and remains competitive with ZTW. On Jetson Orin Nano, the recommended SLIE5 configuration improves accuracy by +2.46 pp while reducing p50 latency by 14% and energy by 15% over vanilla, with zero deadline misses across 10,000 test samples. Thus, multi-exit deployment should be a pre-compiled deterministic inference plan, not a sample-wise routing policy.