Uppsats

Reinforcement Learning based Centralized Joint Orchestration of Radio, Processing, and Fronthaul Resources for Energy Efficient Cell-Free Massive MIMO Networks

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2026

Språk: Engelska

Sammanfattning

Future wireless networks face growing energy demands due to dense infrastructure and computationally intensive processing. Cell-Free Massive Multiple-Input Multiple-Output (CF-mMIMO) has been recognized as a promising architecture to enhance user fairness and spectral efficiency through joint transmission among distributed Access Pointss (APs). However, largescale deployment of CF-MIMO with optical fronthaul introduces high endto- end power consumption across radio, fronthaul, and cloud domains. More recently, Open Radio Access Network (O-RAN) allows controlling radio and cloud resources jointly, allowing potential energy savings. To reduce the energy consumption of CF-mMIMO, this thesis considers CF-mMIMO on top of O-RAN architecture, and proposes a Reinforcement Learning (RL)-based framework for energy-efficient joint radio, fronthaul, and processing resource allocation in centralized CF-MIMO networks. This study first models the end-to-end power consumption of centralized CF-mMIMO, demonstrating that the power consumption in both cloud and radio dominantly depends on active antenna ports and the number of active radio units. Then, Proximal Policy Optimization (PPO) algorithm is designed to learn near-real-time control policies at the near-real-time RAN Intelligent Controller (near-RT RIC), dynamically adjusting active antennas based on long-term channel characteristics to minimize total network power while maintaining users’ Spectral Efficiency (SE) requirements. To enhance the performance, a greedy post-processing stage is developed to further prune redundant activations without compromising service quality. Simulation results show that the proposed RL-based controller achieves over 50% total power reduction compared to the full activation scheme, and around 15% reduction over the heuristic baseline. Specifically, by integrating a greedy strategy with the RL-based control, the total power consumption can be further reduced by an additional 50%. Moreover, the trained model performs well to different user densities and SE targets, demonstrating high robustness and scalability. Its inference time is at the millisecond (ms) level, offering a significant speed advantage over traditional optimization methods. The joint orchestration of radio and processing resources reduces the O-Cloud power consumption by 60%, demonstrating the significant energy-saving potential of the proposed end-to-end control framework.

Information

Författare
Ge, Zilin
Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.