Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation
Abstract
Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at the same time. We study how instruction-following degrades as such constraints accumulate. We introduce a benchmark that stacks 24 verifier-checked instructions, one to twenty at a time, and evaluate three production-tier LLMs (Claude Sonnet 4.6, GPT-5-mini, Gemini 2.5 Flash). Instruction-following degrades non-linearly: the follow rate falls from ~96% at one instruction to 0.60, 0.43 and 0.20 at twenty, driven by a structured and reproducible set of pairwise conflicts across all 231 pairs. A single "output JSON" constraint, for example, is jointly unsatisfiable with nine others, and the interaction landscape correlates across models and tasks. We then evaluate a training-free remedy: an instruction compiler that rewrites the stacked prompt in a single LLM call and is reused across queries at no per-query cost. Its benefit is capability-graded. It recovers +11.0 points of follow rate for the weakest target, which is also the tier most often deployed at scale, while leaving the strongest essentially unchanged. Cluster-robust tests, same-baseline controls, and a within-family scaling ladder over nine targets (Spearman -0.85 between raw following rate and recovery) attribute the gain to the rewrite itself rather than to additional tokens, reordering, or measurement headroom. We release the benchmark, verifiers, and cached runs for full reproduction.