Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
Abstract
Recent AI regulations increasingly emphasize the need for mechanisms that preserve the utility of data for AI innovation while preventing misuse, particularly by enforcing purpose limitation in downstream AI applications. In practice, enforcing this principle remains challenging, as released data can be trivially fed into arbitrary models beyond its declared intent. Existing approaches attempt to mitigate this risk by either perturbing data or retraining models to limit unintended use. These strategies, however, offer no protection against inference by unknown or externally trained models, or fundamentally rely on control over the training or deployment. In this work, we introduce non-transferable examples (NTEs), recoded data that act as a task-level "ciphertext" decodable only by a designated model. Where adversarial examples exploit sensitive input directions, NTEs use the complementary insensitive subspace: a training-free, data-agnostic recoding within a model-specific low-sensitivity subspace preserves the authorized model's outputs while degrading unauthorized ones through subspace misalignment. We establish formal bounds certifying output fidelity for the authorized model and showing unauthorized degradation scales with measurable spectral misalignment between models. Empirically, NTEs preserve authorized-model performance across diverse vision backbones and vision-language models, while unauthorized models collapse even under adaptive reconstruction attacks. These results establish NTEs as a practical means to preserve intended data utility while preventing unauthorized exploitation. Our source code and visual demos are available at: https://github.com/model-specific/non-transferable-examples.