SearcharxivSearch

arXiv subjects

Mohammad S. Ramadan

Publications and source records attributed to Mohammad S. Ramadan.

8 recordsLinked to original sources

Simultaneous improvement of control and estimation for battery management systems

Standard battery management systems treat the control and state estimation problems as decoupled objectives, relying on certainty equivalence controllers that are blind to the varying observability induced by nonlinear open-circuit voltage models. In this paper, we show that for a broad class of objectives, including the peak shaving and valley filling scenarios common in grid-connected energy storage, the expected cost of a stochastic battery system can be exactly parametrized by the conditional mean and covariance of the state of charge. This reformulation reveals a direct coupling between the control input and estimation quality, a coupling that certainty equivalence controllers ignore, and motivates a dual-control approach in which the controller actively reduces estimation uncertainty by driving the state to high observability regions without compromising the control objective. We derive a deterministic surrogate to this stochastic cost and pose the dual-control problem as a computationally tractable model predictive control problem. We validate our approach on a nine-battery system tracking a time-varying power/demand reference trajectory. We report simultaneous improvements in control cost (up to 20\% reduction) and state estimation error (up to 30\% reduction). The estimation improvement is reported across different state estimators: extended Kalman filter, unscented Kalman filter, and a moving horizon estimator, confirming that the estimation improvement of our approach is not restricted to a specific state observer.

eess.SY

Actively Learning Reinforcement Learning: A Stochastic Optimal Control Approach

In this paper we propose a framework towards achieving two intertwined objectives: (i) equipping reinforcement learning with active exploration and deliberate information gathering, such that it regulates state and parameter uncertainties resulting from modeling mismatches and noisy sensory; and (ii) overcoming the computational intractability of stochastic optimal control. We approach both objectives by using reinforcement learning to compute the stochastic optimal control law. On one hand, we avoid the curse of dimensionality prohibiting the direct solution of the stochastic dynamic programming equation. On the other hand, the resulting stochastic optimal control reinforcement learning agent admits caution and probing, that is, optimal online exploration and exploitation. Unlike fixed exploration and exploitation balance, caution and probing are employed automatically by the controller in real-time, even after the learning process is terminated. We conclude the paper with a numerical simulation, illustrating how a Linear Quadratic Regulator with the certainty equivalence assumption may lead to poor performance and filter divergence, while our proposed approach is stabilizing, of an acceptable performance, and computationally convenient.

cs.LG

Monte Carlo Grid Dynamic Programming: Almost Sure Convergence and Probability Constraints

Dynamic Programming (DP) suffers from the well-known ``curse of dimensionality'', further exacerbated by the need to compute expectations over process noise in stochastic models. This paper presents a Monte Carlo-based sampling approach for the state space and an interpolation procedure for the resulting value function, dependent on the process noise density, in a "self-approximating" fashion, eliminating the need for ordering or set-membership tests. We provide proof of almost sure convergence for the value iteration (and consequently, policy iteration) procedure. The proposed meshless sampling and interpolation algorithm alleviates the burden of gridding the state space, traditionally required in DP, and avoids constructing a piecewise constant value function over a grid. Moreover, we demonstrate that the proposed interpolation procedure is well-suited for handling probabilistic constraints by sampling both infeasible and feasible regions. The curse of dimensionality cannot be avoided, however, this approach offers a practical framework for addressing lower-order stochastic nonlinear systems with probabilistic constraints, while eliminating the need for linear interpolations and set membership tests. Numerical examples are presented to further explain and illustrate the convenience of the proposed algorithms.

eess.SY

Extended Kalman filter -- Koopman operator for tractable stochastic optimal control

The theory of dual control was introduced more than seven decades ago. Although it has provided rich insights to the fields of control, estimation, and system identification, dual control is generally computationally prohibitive. In recent years, however, the use of Koopman operator theory for control applications has been emerging. This paper presents a new reformulation of the stochastic optimal control problem that, employing the Koopman operator, yields a standard LQR problem with the dual control as its solution. We provide a numerical example that demonstrates the effectiveness of the proposed approach compared with certainty equivalence control, when applied to systems with varying observability.

eess.SY

A Control Approach for Nonlinear Stochastic State Uncertain Systems with Probabilistic Safety Guarantees

This paper presents an algorithm to apply nonlinear control design approaches in the case of stochastic systems with partial state observation. Deterministic nonlinear control approaches are formulated under the assumption of full state access and, often, relative degree one. We propose a control design approach that first generates a control policy for nonlinear deterministic models with full state observation. The resulting control policy is then used to build an importance-like probability distribution over the space of control sequences which are to be evaluated for the true stochastic and state-uncertain dynamics. This distribution serves in the sampling step within a random search control optimization procedure, to focus the exploration effort on certain regions of the control space. The sampled control sequences are assigned costs determined by a prescribed finite-horizon performance and safety measure, which is based on the stochastic dynamics. This sampling algorithm is parallelizable and shown to have computational complexity indifferent to the state dimension, and to be able to guarantee safety over the prescribed prediction horizon. A numerical simulation is provided to test the applicability and effectiveness of the presented approach and compare it to a certainty equivalence controller.

eess.SY

Optimal Trajectories for Propellant-Free Rendezvous Missions

The paper provides a new approach to utilizing space environmental forces in time- and energy-optimal, propellant-less spacecraft rendezvous missions. Considering the nonlinear form of the relative dynamic equations, rendezvous missions are posed as optimal control problems subject to input saturation. We conduct a direct optimal control approach to obtain optimal trajectories and control inputs. Initially, we consider the differential drag only and conduct a comprehensive analysis of the effect of altitude on the required control input and achieved cost function. Lorentz forces are then utilized with the differential drag, reducing the time required for time-optimal missions. For energy-optimal missions with combined differential drag and Lorentz forces, a weighting matrix in the cost function is introduced to adjust the relative contributions of these forces.

physics.space-ph

State estimation for control: an approach for output-feedback stochastic MPC

The paper provides a new approach to the determination of a single state value for stochastic output feedback problems using paradigms from Model Predictive Control, particularly the distinction between open-loop and closed-loop control and between deterministic optimal control and stochastic optimal control. The State Selection Algorithm is presented and relies on given dynamics and constraints, a nominal deterministic state-feedback controller, and a sampling based method to select the best state value, based on optimizing a prescribed finite-horizon performance function, over the available candidates provided by a particle filter. The cost function is minimized over the horizon with controls determined by the nominal controller and the selected states. So, the minimization is performed not over the selection of the control other than through the choice of state value to use. The algorithm applies generally to nonlinear stochastic systems and relies on Monte Carlo sampling and averaging. However, in linear quadratic polyhedrally constrained cases the technique reduces to a quadratic program for the state value. The algorithm is evaluated in a set of computational examples, which illustrate its efficacy and limitations. Numerical aspects and the opportunity for parallelization are discussed. The examples demonstrate the algorithm operating, in closed-loop with its attendant particle filter, over the long horizon.

math.OC

Maximum Likelihood Recursive State Estimation using the Expectation Maximization Algorithm

A Maximum Likelihood recursive state estimator is derived for non-linear and non-Gaussian state-space models. The estimator combines a particle filter to generate the conditional density and the Expectation Maximization algorithm to compute the maximum likelihood state estimate iteratively. Algorithms for maximum likelihood state filtering, prediction and smoothing are presented. The convergence properties of these algorithms, which are inherited from the Expectation Maximization algorithm, are proven and examined in two examples. It is shown that, with randomized reinitialization, which is feasible because of the algorithm simplicity, these methods are able to converge to the Maximum Likelihood Estimate (MLE) of multimodal, truncated and skewed densities, as well as those of disjoint support.

stat.ME