The Complexity Kink: LLM Rubric Instruments for Causal Inference on Code Generation Reliability
Abstract
Large language models (LLMs) are increasingly evaluated and deployed as code generators, but benchmarks still lack a reliable way to measure when a programming task becomes structurally too complex for a model to solve. Many code evaluations estimate task complexity from the code a model produces. This makes the key variable endogenous: when a model fails on a difficult prompt, it may emit a short stub, partial solution, or broken program whose measured complexity is low, causing hard failures to be reclassified as easy cases. This can hide whether reliability degrades smoothly or changes regime at a sharp complexity threshold, a distinction that matters for model selection, benchmark design, and deployment guardrails. We present the Complexity Kink benchmark, a prompt-side experiment that measures intended solution complexity before generation and separates it from generated-output complexity and functional correctness. We construct a 5,000-prompt Python benchmark from OpenCodeInstruct, stratified to cover the range of structural complexity, and score each prompt with a fixed six-dimensional rubric covering branching, iteration, state, data structures, edge cases, and algorithmic composition. Four out-of-panel LLM judges score the full prompt set, yielding complete coverage and high composite inter-rater reliability (ICC = 0.865). We then evaluate a 21-model panel with unit-test execution and estimate the relationship between prompt-side complexity and pass rate using instrumental variables and threshold regression. By making task complexity observable before generation, this framework tests whether the apparent complexity kink is a real structural break or a measurement artifact, and gives researchers and practitioners a more defensible way to compare models, diagnose reliability limits, and design evaluations that do not erase the failures that matter most.