Deep Reinforcement Learning for Energy Efficient 5G Base Stations Using Semi Markov Modeling
Abstract
5G networks consume substantially more energy than earlier generations, and base stations are responsible for most of this consumption. Operators address this through the Advanced Sleep Mechanism (ASM) strategy, which moves an idle base station through progressively deeper sleep states to conserve power, and Semi Markov Process (SMP) models have long been used to analyze this strategy. However, existing SMP models rely on sleep transition parameters that are fixed before deployment, and since real User Request (UR) traffic fluctuates continuously, these static parameters cannot adapt to changing conditions, which limits the energy savings achievable in practice. To address this, the objective of this work is to propose a framework that first models the base station analytically using an SMP and then augments this model with an adaptive Proximal Policy Optimization (PPO) agent for online control. The base station is modeled as an SMP with six states, comprising Active, Inactive, and four progressive sleep modes, and steady state probabilities are derived through a two stage embedded Markov chain analysis. Closed form expressions are obtained for the power saving factor, average power consumption, and energy efficiency, and these expressions are evaluated over a wide grid of arrival and service rates to characterize the analytical performance landscape of the ASM strategy. Since this SMP model operates under fixed parameters and cannot respond to dynamic traffic conditions, a PPO based reinforcement learning agent is developed to serve as a discrete event simulation controller over the same ASM strategy. The SMP expressions are evaluated once to establish a steady state performance reference, after which the PPO agent takes over all sleep management decisions. At each decision step the agent observes ten state variables, including the current sleep mode, timer progress, traffic intensity, and queue occupancy, and selects one of five actions that control sleep depth and timer duration. It is trained with a power saving reward that directly tracks time spent in a sleep mode, and a queue based safety override together with an auto wake mechanism jointly protect service quality without requiring an explicit penalty term. Training runs for 2000 episodes across five random seeds, after which the trained policy is evaluated across a wide range of arrival rates and compared directly against the SMP analytical baseline. Results show that the PPO based policy consistently outperforms the SMP baseline, with the performance gap widening as traffic intensity increases. At peak load, the power saving factor improves by up to 26 percentage points over the baseline, average power consumption is reduced by up to 37.4 \%, and energy efficiency improves by up to 58.5 \%, with the greatest benefit observed precisely under the high traffic conditions that are hardest to manage efficiently.