SkillXray: Detecting Privacy Harms in Agent Skills
Mariana Fernandez-Espinosa ⋅ Kieleh Ngong Ivoline Clarisse ⋅ Keerthiram Murugesan ⋅ Karthikeyan Natesan Ramamurthy
Abstract
Agent skills extend AI assistants through installable instructions and code that can access users’ files, accounts, and data. Even when performing their advertised tasks, skills may infer sensitive attributes, retain unnecessary information, or disclose data to unexpected recipients. We study these potential privacy-of-use harms through SkillXray, a static audit framework that analyzes skills before installation without executing them. SkillXray extracts data operations from code and natural-language instructions, grounds them in source evidence, and adapts Contextual Integrity to assess their appropriateness. Applied to 100 real skills from public marketplaces, 54 of which handle personal data, SkillXray flags at least one potential violation in 72% of them, and its judgments agree with a blind human annotator on 80% of moves ($\kappa=0.58$). The violations that turn up are mundane ones, such as sensitive inferences the user never saw, uses beyond the stated task, and disclosures to unexpected recipients. These findings show that privacy harms in agent skills can arise from ordinary functionality.
Chat is not available.
Successful Page Load