The Watchtower Imperative: Lifespan-Calibrated Evaluation for Frontier AI
Abstract
The gap between AI capability deployment and safety assurance is structural and cannot be closed through internal evaluations alone. Analyzing documented harms (2017-2026) across the human lifespan reveals that disparate incidents — from voice assistants endangering children to AI companions implicated in adolescent suicide and automated care denials for the elderly — reflect a single failure: evaluating systems against an implicit adult user model. We introduce five developmental moderators structured into a Lifespan Vulnerability Matrix across six life stages. The matrix reveals a non-monotonic risk distribution: vulnerability peaks at both ends of the lifespan, governance coverage is asymmetrical, and authority to convert model errors into real-world impact is scrutinized least. We propose lifespan design principles generalizing existing child-protective instruments, and an institution-agnostic, three-tier external evaluation framework — automated simulation, expert red teaming, and post-deployment monitoring — indexed directly to matrix risk cells.