Generating Diverse and Verified Synthetic Clinical Notes with Labels by Construction
Abstract
Under strict privacy regulations in the clinical domain, de-identification is necessary to enable the use of clinical data such as for natural language processing. However, we currently lack sophisticated ways to tell which de-identification models actually are suitable. Specifically, we need benchmarks that are challenging, diverse, and clinically plausible, all while ensuring label correctness. To this end, we present SPARCLING, an LLM-based pipeline to generate synthetic discharge letters with character-exact PHI span annotations. Our pipeline combines structured attribute sampling, template generation, and red teaming to obtain diverse PHI-labeled discharge letters. In our experiments, we demonstrate this superior content diversity over existing public datasets. Further, we show SPARCLING's downstream utility as a benchmark to uncover failure modes of de-identifiers in the German clinical domain. Finally, our experiments show that training on our synthetic data can yield more capable de-identifiers than training on the scarce public data alone. We will publish our code and the generated benchmark dataset.