Internal Safety Collapse in Frontier Large Language Models
Abstract
This work identifies a critical failure mode in frontier large language models (LLMs), which we term \textbf{Internal Safety Collapse} (ISC): \textit{under certain task conditions, models enter a state in which they continuously generate large volumes of harmful content while executing otherwise benign tasks}. To systematically study ISC, we introduce \TVD{} (Task, Validator, Data), a framework that instantiates controlled workflows around domain tools, where valid completion requires filling harmful content as structured data. We collect 53 representative workflows across 8 disciplines, from toxicity evaluation to molecular docking and pathogen genome analysis, showing that ISC reproduces across all eight disciplines. In the worst case over three interaction settings, the four frontier LLMs average a \textbf{95.3\%} safety-failure rate. Frontier LLMs with stronger task-completion capability show higher unsafe-completion rates: long-horizon execution skill becomes a liability when workflow completion requires harmful content. We even observe extremely severe harmful content closely resembling outputs from early-generation, unaligned LLMs in 2023. Despite substantial safety alignment efforts, frontier LLMs continue to retain inherently unsafe internal capabilities: alignment reshapes observable outputs but does not eliminate the underlying risk profile. These findings underscore the need for caution when deploying LLMs in high-stakes settings, including scientific pipelines and autonomous agents. Complementary source code is provided at \url{https://anonymous.4open.science/r/NIPS-ISC-Code-Share-BE11}.