Persona-Aware Medical Text Simplification with Prefix-Tuning and Reinforcement Learning from AI Feedback
Abstract
Persona-Aware Medical Text Simplification with Prefix-Tuning and Reinforcement Learning from AI Feedback Biomedical text simplification has gained increasing attention with the advancement of large language models (LLMs), which demonstrate strong capabilities in adapting complex medical information for improved healthcare accessibility. However, existing simplification approaches often target a generic audience, despite healthcare communication serving readers with diverse levels of medical knowledge, ranging from patients and the general public to researchers and clinical experts [1]. This can result in outputs that remain too complex for lay readers or oversimplify information intended for more knowledgeable readers. We investigate persona-aware biomedical text simplification across four reader profiles: Layman, Pre-medical Student, Researcher, and Domain Expert. We propose a persona-conditioned prefix-tuning approach to control linguistic complexity and information presentation according to the target reader. Our approach learns persona-specific continuous prefix representations that are prepended to the input and optimised while keeping the underlying LLM parameters frozen, enabling each persona profile to induce distinct linguistic complexity and simplification [2]. We further investigate LLM-as-a-Judge-guided reinforcement learning using Proximal Policy Optimisation (PPO) for persona alignment [3]. We evaluate these approaches on the PERCS dataset [4], across LLaMA and Mistral model families, including their biomedical variants, and compare them with zero-shot, few-shot, and chain-of-thought prompting baselines using SARI, BERTScore, ROUGE, and FKGL. Prefix-tuning improves persona-specific control over prompting-based approaches, achieving a SARI score of 53.19% and BERTScore of 96.41% for the Layman persona. Improvements are strongest for Layman and Pre-medical Student personas, while Researcher and Domain Expert outputs retain greater linguistic complexity. However, reinforcement learning does not consistently improve the prefix-tuned models, with reward-based alignment sometimes degrading simplification performance, highlighting limitations in using automated reward signals to capture persona-specific simplification quality. Overall, our findings demonstrate that persona-conditioned adaptation provides more controllable biomedical text simplification while preserving semantic content and also highlighting the importance of reward design for effective alignment. Cunha, R., Ferreira, T.C., Pagano, A., Alves, F.: A persona-based corpus in the diabetes self-care domain-applying a human-centered approach to a low-resource context. In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). pp. 1353 Li, X.L., Liang, P.: Prefix-tuning: Optimizing continuous prompts for generation. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. pp. 4582–4597(2021) Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O.: Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017) Salvi, R.C., Chawla, C., Jain, D., Panigrahi, S., Akhtar, M.S., Yadav, S.: Percs: Persona-guided controllable biomedical summarization dataset. arXiv preprint arXiv:2512.03 (2025)