What Is a Morally Aligned AI Agent? A Philosophical Coverage Map and Its Failure Under Composition
Syed A Haider
Abstract
Alignment research has a vocabulary problem: we ask systems to be safe, reliable, and trustworthy without specifying what those words require of a system philosophically, then optimize precisely for targets we have not defined. We decompose agency into eight dimensions and morality into three, cross them into an $8 \times 3$ coverage map, and, unlike prior framework papers, supply the coding protocol and survey standing behind every cell. Coverage is uneven in a specific way: deliberation reaches only two of the eight agency dimensions, against four for normativity and six for judgment. We then ask whether these properties survive composition into agentic systems. Largely they do not: two are blocked by standing impossibility results and one attenuates along delegation chains. Per-agent alignment does not lift to system-level alignment, which reframes several observed multi-agent failures as structural rather than incidental.
Chat is not available.
Successful Page Load