Hardware-Aware Sim-to-Real Reinforcement Learning for Low-Cost Intelligent Traffic Control in Sub-Saharan Africa
Eunice Adebusayo Adewusi ⋅ Marvin Ogore
Abstract
Rapid urbanization in Sub-Saharan Africa has intensified a mobility crisis, with cities like Lagos experiencing 70-minute average commutes [1] and \\$4B in annual economic losses [2]. Traditional Adaptive Traffic Control Systems (ATCS) are cost-prohibitive, up to $339k per intersection [3], and rely on rigid, lane-based assumptions that fail to accommodate the non-lane-based, heterogeneous traffic patterns, characterized by a complex mix of cars, motorcycles, and high-occupancy informal minibuses, typical of African cities. To address these, we developed a low-cost, retrofittable, edge-deployed traffic control system using a simulation-trained Proximal Policy Optimization (PPO) agent deployed on Raspberry Pi hardware based on the Mobility Justice framework that prioritizes multi-modal equity, safety, and public transit over private cars, shifting the optimization paradigm from absolute vehicle-centric metrics to person-throughput and affordability. Transferring simulation-trained policies to the physical world or edge devices introduces a severe sim-to-real gap [4] from sensor noise, processing jitter, and latency mismatches. To enable resource-aware edge deployment, we model traffic control as a Markov Decision Process (MDP) and apply a 96.5% state-space reduction, compressing a high-dimensional 113D simulation state to a lightweight 4D queue vector: $S = [q_N, q_S, q_E, q_W]$. The action space is a binary choice between green phases. We train a compact, three-layer policy network ($128 \rightarrow 64 \rightarrow 32$) using PPO to beat a classical Longest-Queue-First (LQF) heuristic simulated at each timestep. The comparative reward function is $R = 3.0C + R_{\text{comp}} + R_{\text{strat}} + R_{\text{cong}}$, where $C$ is cleared vehicles, $R_{\text{strat}}$ rewards clearing the longest queue, $R_{\text{cong}}$ penalizes congestion, and $R_{\text{comp}}$ scales dynamically based on the performance differential $\Delta = C_{\text{PPO}} - C_{\text{LQF}}$ ($+8.0 \times \Delta$ for outperforming the baseline and $+5.0 \times \Delta$ for underperformance). While our long-term paradigm targets multi-modal person-throughput, this initial deployment optimizes vehicular clearance to validate the system’s feasibility under physical constraints. Multi-seed validation across five independent training runs confirmed exceptional reproducibility with a Coefficient of Variation (CV) of 1.3\% across seeds. We apply domain randomization, simulating queue capacities, arrival rates, mechanical debounce, sensor noise, and GPIO processing latency. On-device, a hybrid rule-based safety wrapper monitors and overrides PPO actions that select empty phases while conflicting roads are congested, guaranteeing 100\% clearing efficiency and zero crashes. We construct a physical prototype utilizing a Raspberry Pi 4 Model B connected via GPIO to 12 LEDs (signals) and 4 debounced tactile push buttons (arrivals). To guarantee fair benchmarking under stochastic conditions, we implement a Record & Replay methodology: a fixed-timing baseline first controls traffic while recording the exact timestamps of vehicle arrivals; the PPO agent then replays the identical sequence, isolating the control strategy as the sole experimental variable. Under rigorous hardware-in-the-loop (HIL) evaluation across 25 diverse scenarios, the PPO agent achieved statistically significant improvements over the baseline. In simulation, the agent secured a 72% win rate, yielding an 8.9% reduction in average vehicle delay ($p=0.018$, Wilcoxon signed-rank test) and an 8.8% reduction in queue length ($p=0.025$), while boosting cars-cleared-per-switch efficiency by 233%. On physical hardware, the agent maintained a 100% win rate across comparison tests, improving throughput by 2.8% and reducing phase changes by 20% by dynamically extending green times. Edge inference executed with a mean latency of 6.84 ms (max 11.16 ms), providing a 14.6$\times$ safety margin below the 100 ms real-time threshold. This work is the first hardware-validated, low-cost Reinforcement Learning (RL) traffic control system tailored for Sub-Saharan Africa. Deploying it as an intersection retrofit offers a 1.8$\times$ annual ROI and a 7-month investment payback period for developing municipalities.
Chat is not available.
Successful Page Load