Position: The MLOps-to-LLMOps Transition Demands a New Operational Paradigm
Abstract
Enterprises with mature MLOps practices for predictive ML face a concrete question: which practices still work for generative AI, and where do they fall short? We argue that LLMOps cannot be treated as an incremental extension of MLOps. Drawing on the MLOps and LLMOps literature and on design patterns visible in nine widely adopted LLM application platforms (LangSmith, Weights & Biases Weave, Arize Phoenix, Braintrust, PromptLayer, Humanloop, Helicone, LangFuse, and Portkey), we classify 23 core operational capabilities into three transfer classes, driven by generative AI's architectural demands: natural language as a control surface, non-deterministic outputs, unbounded context, and autonomous multi-step execution. Eleven capabilities, including CI/CD, version control, and incident response, transfer directly. Seven require substantial re-engineering: testing shifts from single-metric accuracy to multi-dimensional rubrics, feature stores become prompt registries, and threshold-based alerting gives way to LLM-judged monitoring. Five have no MLOps precedent: prompt lifecycle management, context optimization, guardrail engineering, model routing, and agent orchestration. These five constitute the operational substrate that enterprise agent platforms need in place before agent orchestration and workflow evolution can be governed safely at scale. We present this 11-7-5 classification as a practitioner checklist and testable hypothesis, not a validated result, and outline the fielded study needed to confirm it.