Small Model Portfolios for Many Deployment Profiles: Submodular Coverage under Bundled Constraints
Ji Cheng
Abstract
Deploying deep learning models across heterogeneous hardware requires each model to jointly satisfy a bundle of coupled constraints on accuracy, latency, memory, and energy. Maintaining a specialist for every deployment scenario is operationally prohibitive, so practitioners need a small portfolio of $K$ models that collectively covers many deployment profiles. Existing hardware-aware NAS and Pareto-based approaches target marginal objectives and do not directly address this downstream decision. We formalize it as Deployment Portfolio Selection over a fixed pre-measured archive, where a profile is covered only when a single retained model simultaneously satisfies its entire constraint bundle. The induced objective is weighted maximum coverage, which is NP-hard but monotone submodular, so greedy selection inherits a $(1-1/e)$ guarantee, and sample-then-optimize admits a finite-sample bound. Across an 18-family sweep on HW-NAS-Bench, direct bundled-coverage optimization is most valuable in the small-portfolio, low-coverability regime: gains concentrate on the hardest profiles, are most robust at $K = 3$, and persist under three non-uniform demand models. Complementary analyses of remeasurement noise, runtime, and compact feasibility storage clarify when bundled coverage is preferable to scalarized or Pareto-based alternatives.
Chat is not available.
Successful Page Load