SearcharxivSearch

arXiv subjects

Mohamed Shalma

Publications and source records attributed to Mohamed Shalma.

4 recordsLinked to original sources

Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage

Deep reinforcement learning (DRL) has recently gained a great attention due to its real-time adaptation and effectiveness in complex optimization problems. This paper investigates the optimal deployment of millimeter-wave (mmWave) base stations (BSs) in a realistic, non-convex campus topology. The optimization problem is NP-hard, due to the non-convex, non-smooth nature of the max-min fairness objective. To overcome these constraints, we formulate the BS placement as a Markov Decision Process (MDP) and systematically benchmark four DRL schemes: a discrete single-agent Deep Q-Network (DQN), a spatially partitioned Multi-Agent DQN, a continuous single-agent Deep Deterministic Policy Gradient (DDPG), and a geographically partitioned multi-agent DDPG framework. Numerical evaluations reveal that the multi-agent DDPG approach substantially outperforms single-agent in dense scenarios. Additionally full coverage is achieved, and a fairness Jain's index of 0.94 is obtained. Finally, the multi-agent demonstrates highly efficient computational convergence of dense scenarios with $400$ users.

cs.LG

Optimal Base Station Placement for Beyond 5G Networks with Non-Convex Topology

This paper investigates the optimal placement of a millimeter-wave (mmWave) base station (BS) within a realistic U-shaped environment with non-convex topology. The problem is challenging and NP-hard due to the non-convex topology and the non-convex objective functions which are the sum-rate maximization and max-min fairness, the latter being additionally non-smooth. To address this challenge, the BS placement is formulated as a Markov Decision Process (MDP). Then, we propose two deep reinforcement learning (DRL) techniques: First, the deployment area is discretized into a grid and optimized using a Deep Q-Network (DQN). Second, the U-shaped region is partitioned into continuous subspaces, where a Deep Deterministic Policy Gradient (DDPG) agent is dedicated to each subspace then the best BS placement is selected among partitions. Results demonstrate that optimal placement achieves full coverage and yields a Jain index of 0.99. Furthermore, the proposed partitioned multi-space DDPG achieves better solution than DQN with lower complexity.

eess.SP

On the Sensitivity of Active RIS Systems to CSI Errors: Joint Optimization and Performance Trade-off

In this paper, the problem of maximizing the sum-rate is addressed for a multi-user uplink scenario that is assisted by an active reconfigurable intelligent surface (RIS). The maximization is achieved by optimizing the beamforming at the base station, the users' transmit power, active RIS elements phase shifts, and active gains in presence of imperfect channel state information (CSI). The non-convex maximization problem is decomposed into sub-problems and solved via iterative approaches including the Lagrangian method, the projected gradient descent, multi-variate Taylor expansion and fractional programming. Numerical results show that the active RIS is more sensitive to CSI imperfections than passive one at high error variances.

eess.SP

Hybrid Deep Reinforcement Learning for Joint Resource Allocation in Multi-Active RIS-Aided Uplink Communications

Active Reconfigurable Intelligent Surfaces (RIS) are a promising technology for 6G wireless networks. This paper investigates a novel hybrid deep reinforcement learning (DRL) framework for resource allocation in a multi-user uplink system assisted by multiple active RISs. The objective is to maximize the minimum user rate by jointly optimizing user transmit powers, active RIS configurations, and base station (BS) beamforming. We derive a closed-form solution for optimal beamforming and employ DRL algorithms: Soft actor-critic (SAC), deep deterministic policy gradient (DDPG), and twin delayed DDPG (TD3) to solve the high-dimensional, non-convex power and RIS optimization problem. Simulation results demonstrate that SAC achieves superior performance with high learning rate leading to faster convergence and lower computational cost compared to DDPG and TD3. Furthermore, the closed-form of optimally beamforming enhances the minimum rate effectively.

eess.SP