Discovering Programmatic Policies from Reinforcement Learning-Based Traffic Signal Controllers
Abstract
Traffic signal control (TSC) is critical for mitigating urban congestion. Deep reinforcement learning (DRL) has achieved strong performance by learning adaptive policies from traffic observations, but learned neural controllers are often opaque, structurally complex, and difficult to generalize across scenarios. We propose PPL-TSC, a programmatic policy learning framework that distills compact and human-readable TSC policies from DRL teachers. PPL-TSC formulates policy learning as an evolutionary search over phase utility functions (PUFs), which score candidate signal phases and select the highest-utility phase for execution. Each generation consists of four steps: policy generation, policy review, policy calibration, and policy evaluation. In the generation step, a large language model (LLM) acts as a semantics-aware evolutionary operator to propose and refine candidate PUFs. Generated PUFs are screened through explicit checks for syntactic validity, mathematical soundness, and consistency with TSC principles, with invalid ones repaired through iterative feedback. Valid PUFs are then calibrated to improve behavioral fidelity to the DRL teacher, and evaluated using a joint score combining teacher fidelity and traffic-control performance to select candidates for the next generation. Extensive experiments across diverse traffic scenarios and representative DRL-based controllers show that PPL-TSC produces compact and human-readable policies, matches teacher performance, improves transfer generalization, and substantially reduces computational cost.