A Latent-Load Framework for Reliability Analysis and Intervention Design in LLM Pipelines
Abstract
Multi-step LLM systems increasingly combine generation, retrieval, tool use, code execution, and verification into pipelines whose components consume and transform an evolving semantic state; end-to-end reliability is therefore governed by pipeline dynamics rather than by component-level accuracy alone. In such pipelines, an early ambiguity, hallucination, or reasoning error may be amplified, masked, partially corrected, or reintroduced by later components. Existing evaluations and failure taxonomies identify where LLM pipelines fail, but they do not provide a predictive theory of how semantic corruption propagates across dependent steps or how limited interventions should be allocated. To fill this gap, we introduce a latent-load framework for reliability analysis and intervention design in black-box LLM pipelines. Specifically, each intermediate state carries an unobserved semantic error load that evolves through inherited-error amplification, context dependence, and fresh error injection. The resulting recurrence decomposes final error, identifies stable, critical, and unstable regimes, and quantifies system-level risk under parameter uncertainty. We prove that the checkpoint-placement objective is monotone submodular, yielding near-optimal budgeted interventions, and provide a proxy-based partial-identification procedure for black-box estimation. Experiments show that proxy-scale propagation estimates predict relative downstream risk and select checkpoints that reduce accumulated proxy load and directionally lower failure rates under fixed budgets.