Instrumental Convergence Revisited: Towards a Broader Characterization of Runaway Optimization
Abstract
Instrumental convergence and runaway optimization constitute a major source of concern in AI risk discourse. We believe that their conceptions present in the discourse are both too broad and too narrow. They are too broad because they make use of some notions -- such as "agency", "power", or "generality" -- that conflate several concepts that should be tracked separately. They are also too narrow because they tend to assume -- both in canonical sources and in live discourse -- that they straightforwardly rely on properties such as vNM rationality, coherence, human-level intelligence, or pursuing goals. Going beyond those limitations of the concepts is needed for us to apprehend the full scope of risks from runaway optimization. We also make explicit what we believe to be the key generative intuition behind optimization risks and call it the "Central Dogma of AI Risk".