SearcharxivSearch

arXiv subjects

Aditya Chaudhari

Publications and source records attributed to Aditya Chaudhari.

2 recordsLinked to original sources

Optimal Control of District Cooling Energy Plant with Reinforcement Learning and MPC

We consider the problem of optimal control of district cooling energy plants (DCEPs) consisting of multiple chillers, a cooling tower, and a thermal energy storage (TES), in the presence of time-varying electricity price. A straightforward application of model predictive control (MPC) requires solving a challenging mixed-integer nonlinear program (MINLP) because of the on/off of chillers and the complexity of the DCEP model. Reinforcement learning (RL) is an attractive alternative since its real-time control computation is much simpler. But designing an RL controller is challenging due to myriad design choices and computationally intensive training. In this paper, we propose an RL controller and an MPC controller for minimizing the electricity cost of a DCEP, and compare them via simulations. The two controllers are designed to be comparable in terms of objective and information requirements. The RL controller uses a novel Q-learning algorithm that is based on least-squares policy iteration. We describe the design choices for the RL controller, including the choice of state space and basis functions, that are found to be effective. The proposed MPC controller does not need a mixed integer solver for implementation, but only a nonlinear program (NLP) solver. A rule-based baseline controller is also proposed to aid in comparison. Simulation results show that the proposed RL and MPC controllers achieve similar savings over the baseline controller, about 17%.

eess.SY

A time-optimal feedback control for a particular case of the game of two cars

In this paper, a time-optimal feedback solution to the game of two cars, for the case where the pursuer is faster and more agile than the evader, is presented. The concept of continuous subsets of the reachable set is introduced to characterize the time-optimal pursuit-evasion game under feedback strategies. Using these subsets it is shown that, if initially the pursuer is distant enough from the evader, then the feedback saddle point strategies for both the pursuer and the evader are coincident with one of the common tangents from the minimum radius turning circles of the pursuer to the minimum radius turning circles of the evader. Using geometry, four feasible tangents are identified and the feedback min-max strategy for the pursuer and the max-min strategy for the evader are derived by solving a $2 \times 2$ matrix game at each instant. Insignificant computational effort is involved in evaluating the pursuer and evader inputs using the proposed feedback control law and hence it is suitable for real-time implementation. The proposed law is validated further by comparing the resulting trajectories with those obtained by solving the differential game using numerical techniques.

eess.SY