Who Holds the Guardrails? Measuring Safety-Control Provenance in Publicly Exposed Generative AI
Abstract
Most child-safety evaluations focus on provider-managed AI services. We instead study publicly reachable, unauthenticated self-hosted deployments, where model weights, system instructions, moderation, and administrative control may have different provenance. From search-index data, we identify 123,745 endpoints across 31 service fingerprints without submitting prompts or invoking models. For 11,506 endpoints with visible inventories, we attribute system instructions to the publisher, operator, or a third party. 55.9% have none, 25.9% retain the publisher default, only 6.8% are operator-authored, and 11.5% contain instructions from two injection campaigns that exploit unauthenticated model-creation interfaces. We further identify 1,590 deployments of refusal-removed weights and show that the model image-generation stack ships no active safety component by default. Our results show that child-safety-relevant controls belong to deployments and may have different provenance from the upstream model.