vLLM Hook: Programming Model Internals is the New Harness
Abstract
Inference technology is essential for deploying and scaling generative artificial intelligence (AI) models. However, current AI inference engines operate like non-configurable circuits. While they offer fast inference, they have limited programmability of the internal states of deployed models. This limitation restricts support for advanced monitoring and control methods needed for safety, interpretability, and governance. To address this issue, this paper proposes programmable AI inference engines that incorporate a modular configuration layer to reuse or adapt inference traces on existing engines. Our design, implemented as an extension to vLLM, improves latency by up to 8.2x and reduces memory usage by three orders of magnitude when extracting the embedding of the last token. Furthermore, we present three native programming applications: hallucination detection, prompt injection detection, and activation steering.