Position: A Safe LLM and a Safe Harness Do Not Make a Safe Agent
Abstract
AI agents combine a learned neural component, such as a language model, with symbolic components, such as memory, tools, and environments. Existing safety alignment work improves both sides independently: neural alignment reduces unsafe model behavior, while symbolic alignment constrains memory, tool use, permissions, and action. This position paper argues that safe neural and symbolic components do not necessarily compose into a safe agent. Agent safety is often contextual: whether an answer, retrieval, or tool call is safe can depend on broader safety context, e.g., prior interactions, provenance, user boundaries, and authorization. For example, a biosafety AI research assistant might answer separate, benign-looking questions about viral delivery systems, protocol troubleshooting, and relevant literature, while the accumulated interaction begins to support the design of a dangerous pathogen. We propose contextual agent alignment as a research agenda for identifying safety-relevant context, designing safer agents that use this context, and evaluating contextual safety. Several documented agent safety failures, including cumulative dual-use risk in research assistants, peer-preservation in oversight settings, and contextual privacy violations, can be viewed as instances of contextual safety failure. This agenda is important for the broader AI safety community because frontier systems are increasingly deployed as agents in real-world settings.