Evaluating Relational Privacy in Paediatric LLM Workflows
Abstract
Language-model systems face privacy risks beyond the memorisation of training data. Such systems increasingly transform records containing information about multiple people, yet privacy evaluations often focus only on whether sensitive information is disclosed, rather than whether it is disclosed to the appropriate recipient for the appropriate purpose. This distinction matters because the same third-party fact may be inappropriate in one context but necessary in another. We introduce a paired-recipient evaluation that holds a synthetic paediatric record fixed while varying recipient and purpose. Across five LLMs, relational-privacy prompting reduces prohibited third-party disclosure from 24.4\% to 1.5\%, but paired disclosure accuracy on 26 required/prohibited recipient pairs from 13 records falls from 51.5\% at baseline to 0\%. Low leakage alone can therefore mask a failure to use third-party information appropriately across contexts. Thus, evaluating relational privacy requires measuring both inappropriate disclosure and appropriate use.