AppCIF-Bench: An Application-Level Complex Instruction Following Benchmark for Large Language Models
Jiahao Xu ⋅ Xuefang Zhao ⋅ Xinhua Feng ⋅ Hongrui Yang ⋅ Yixian Liu ⋅ Zhichao Hu ⋅ Lliu Yuhong
Abstract
Complex Instruction Following (CIF) is a foundational capability for deploying Large Language Models (LLMs) in production-grade applications. It demands stable, precise outputs under increasing horizon length, deeply nested rules, and multi-turn interactions. Yet existing benchmarks are predominantly atomic and short-horizon, failing to capture the logical depth and distractor pressure of real-world settings. We introduce **AppCIF**, the first application-level CIF benchmark. AppCIF stratifies difficulty along two orthogonal axes: *Horizon Length* ($H$), the span over which distractors accumulate and critical signals must be tracked; and *Logical Depth* ($L$), the nesting depth of rules to be unfolded. Together, $\langle H, L\rangle$ defines four hierarchical complexity levels, scaling from clean few-step execution (L1) to extreme pressure deduction (L4). AppCIF comprises 466 curated samples across five application scenarios. It adopts a "**Prompt-as-Environment**" paradigm, pairing Complex System Prompts (averaging ~14k words) with executable Code Simulation Engines (averaging ~3.5k lines of Python) to sustain rigorous state-machine deduction. The benchmark is constructed via a **Labor-Divided Human-AI pipeline** with scenario-specific collaboration strategies, and evaluated on 12 representative LLMs. Results reveal sharp performance stratification: Hard Satisfaction Rate ranges from 1.1% to 45.9%, exposing severe bottlenecks in long-horizon, distractor-laden instruction execution.
Chat is not available.
Successful Page Load