Searcharxiv⌕ Search

arXiv subjects

Zhaorong Zhang

Publications and source records attributed to Zhaorong Zhang.

14 recordsLinked to original sources

Relaxed Control with Entropy Regularization for Itô Stochastic Systems with Input Delay

This paper investigates the infinite-horizon classical stochastic optimal control problem with input delay under an entropy-regularized relaxed control framework. In particular, by constructing a relaxed system and introducing an entropy regularization term, we reformulate the classical optimal control problem into an entropy regularized formulation, and derive the optimal controller that follows a Gaussian distribution. Furthermore, we show that the optimal Gaussian control distribution converges to the optimal Dirac measure as the exploration weight tends to zero. Numerical simulation is provided to validate the effectiveness of the proposed method.

math.OC↗

OCP-GN: A Scalable Second-order Optimizer for Stochastic Optimization

This paper proposes a novel second-order optimization algorithm based on the Optimal Control Principle (OCP), applicable to large-scale optimization problems in neural network training. The algorithm has a computational complexity of O(d) and strong robustness. Extensive experiments on multiple benchmarks demonstrate the significant superiority of the proposed method.

cs.CV↗

Distributed Load Frequency Control of Multi-Area Smart Grid

In this paper, we investigate the distributed load frequency control problem in a multi-area smart grid under external load disturbances and measurement noise. The novelty lies in that the information privacy is fully taken into account, that is, the internal structural parameters and operational states of each area are not shared with non-neighboring areas, which makes traditional distributed optimal control methods ineffective. The main contribution is to propose a distributed algorithm for the global optimal power regulation command under information privacy constraints via distributed approximation of the control Riccati equation, the estimation Riccati equation, and the state estimation. Simulation results show that the proposed algorithm can approximate the performance of centralized optimal control, and the performance index under the proposed distributed controller is smaller than that under the commonly used distributed control.

math.OC↗

Distributed Algorithm for the Global Optimal Controller of Nonlinear Multi-Agent Systems

In this paper, we investigate the distributed optimal control problem for a kind of nonlinear multi-agent systems. In particular,both the state and the system dynamic structures of each agent are private and can only be shared among communicating agents.This type of information structure is inevitable in fields such as collaborative control for industrial confidentiality, and renders traditional distributed control methods using all systems' dynamic structures ineffective. The primary contribution is the proposal of a distributed algorithm for the global optimal controller under such practical information structure via distributed approximation of the Hamilton-Jacobi-Bellman equation. Practical numerical simulation demonstrates the effectiveness of the proposed algorithm.

math.OC↗

Distributed Solving of Linear Quadratic Optimal Controller with Terminal State Constraint

This paper is concerned with the linear quadratic (LQ) optimal control of continuous-time system with terminal state constraint. In particular, multiple agents exist in the system which can only access partial information of the matrix parameters. This makes the classical solving method based on Riccati equation with global information suffering. The main contribution is to present a distributed algorithm to derive the optimal controller which is consisting of the distributed iterations for the Riccati equation, a backward differential equation driven by the optimal Lagrange multiplier and the optimal state. Furthermore, the proposed distributed iteration method is extended to solve the consensus control problem for heterogeneous multi-agent systems, achieving the globally optimal performance of the system. The effectiveness of the proposed algorithm is verified by two numerical examples, where the performance index under the proposed distributed controller is smaller than that under the commonly used consensus control.

math.OC↗

Reinforcement Learning for Stochastic LQ Control of Discrete-Time Systems with Multiplicative Noises

This paper considers a stochastic linear quadratic problem for discrete-time systems with multiplicative noises over an infinite horizon. To obtain the optimal solution, we propose an online iterative algorithm of reinforcement learning based on Bellman dynamic programming principle. The algorithm avoids the direct calculation of algebra Riccati equations. It merely takes advantage of state trajectories over a short interval instead of all iterations, significantly simplifying the calculation process. Under the stabilizable initial values, numerical examples shed light on our theoretical results.

math.OC↗

Q-Learning for Linear Quadratic Optimal Control with Terminal State Constraint

This paper is concerned with the linear quadratic optimal control of discrete-time time-varying system with terminal state constraint. The main contribution is to propose a Q-learning algorithm for the optimal controller when the time-varying system matrices and input matrices are both unknown. Different from the existing Q-learning algorithms in the literature which are mainly for the unconstrained optimal control problem, the novelty of the proposed algorithm is available to deal with the case with terminal state constraints. A numerical example is illustrated to verify the effectiveness of the proposed algorithm.

math.OC↗

Reinforcement Learning-Based Optimal Control for Multiplicative-Noise Systems with Input Delay

In this paper, the reinforcement learning (RL)-based optimal control problem is studied for multiplicative-noise systems, where input delay is involved and partial system dynamics is unknown. To solve a variant of Riccati-ZXL equations, which is a counterpart of standard Riccati equation and determines the optimal controller, we first develop a necessary and sufficient stabilizing condition in form of several Lyapunov-type equations, a parallelism of the classical Lyapunov theory. Based on the condition, we provide an offline and convergent algorithm for the variant of Riccati-ZXL equations. According to the convergent algorithm, we propose a RL-based optimal control design approach for solving linear quadratic regulation problem with partially unknown system dynamics. Finally, a numerical example is used to evaluate the proposed algorithm.

math.OC↗

Distributed Q-Learning for Stochastic LQ Control with Unknown Uncertainty

This paper studies a discrete-time stochastic control problem with linear quadratic criteria over an infinite-time horizon. We focus on a class of control systems whose system matrices are associated with random parameters involving unknown statistical properties. In particular, we design a distributed Q-learning algorithm to tackle the Riccati equation and derive the optimal controller stabilizing the system. The key technique is that we convert the problem of solving the Riccati equation into deriving the zero point of a matrix equation and devise a distributed stochastic approximation method to compute the estimates of the zero point. The convergence analysis proves that the distributed Q-learning algorithm converges to the correct value eventually. A numerical example sheds light on that the distributed Q-learning algorithm converges asymptotically.

math.OC↗

Convergence Rate of a Message-passing Algorithm for Solving Linear Systems

This paper studies the convergence rate of a message-passing distributed algorithm for solving a large-scale linear system. This problem is generalised from the celebrated Gaussian Belief Propagation (BP) problem for statistical learning and distributed signal processing, and this message-passing algorithm is generalised from the well-celebrated Gaussian BP algorithm. Under the assumption of generalised diagonal dominance, we reveal, through painstaking derivations, several bounds on the convergence rate of the message-passing algorithm. In particular, we show clearly how the convergence rate of the algorithm can be explicitly bounded using the diagonal dominance properties of the system. When specialised to the Gaussian BP problem, our work also offers new theoretical insight into the behaviour of the BP algorithm because we use a purely linear algebraic approach for convergence analysis.

eess.SY↗

Distributed Weighted Least-squares Estimation for Networked Systems with Edge Measurements

This paper studies the problem of distributed weighted least-squares (WLS) estimation for an interconnected linear measurement network with additive noise. Two types of measurements are considered: self measurements for individual nodes, and edge measurements for the connecting nodes. Each node in the network carries out distributed estimation by using its own measurement and information transmitted from its neighbours. We study two distributed estimation algorithms: a recently proposed distributed WLS algorithm and the so-called Gaussian Belief Propagation (BP) algorithm. We first establish the equivalence of the two algorithms. We then prove a key result which shows that the information matrix is always generalised diagonally dominant, under some very mild condition. Using these two results and some known convergence properties of the Gaussian BP algorithm, we show that the aforementioned distributed WLS algorithm gives the globally optimal WLS estimate asymptotically. A bound on its convergence rate is also presented.

eess.SY↗

A Fast Converging Distributed Solver for Linear Systems with Generalised Diagonal Dominance

This paper proposes a new distributed algorithm for solving linear systems associated with a sparse graph under a generalised diagonal dominance assumption. The algorithm runs iteratively on each node of the graph, with low complexities on local information exchange between neighbouring nodes, local computation and local storage. For an acyclic graph under the condition of diagonal dominance, the algorithm is shown to converge to the correct solution in a finite number of iterations, equalling the diameter of the graph. For a loopy graph, the algorithm is shown to converge to the correct solution asymptotically. Simulations verify that the proposed algorithm significantly outperforms the classical Jacobi method and a recent distributed linear system solver based on average consensus and orthogonal projection.

eess.SP↗

Convergence of Message-Passing for Distributed Convex Optimisation with Scaled Diagonal Dominance

This paper studies the convergence properties the well-known message-passing algorithm for convex optimisation. Under the assumption of pairwise separability and scaled diagonal dominance, asymptotic convergence is established and a simple bound for the convergence rate is provided for message-passing. In comparison with previous results, our results do not require the given convex program to have known convex pairwise components and that our bound for the convergence rate is tighter and simpler. When specialised to quadratic optimisation, we generalise known results by providing a very simple bound for the convergence rate.

math.OC↗

On Convergence Rate of the Gaussian Belief Propagation Algorithm for Markov Networks

Gaussian Belief Propagation (BP) algorithm is one of the most important distributed algorithms in signal processing and statistical learning involving Markov networks. It is well known that the algorithm correctly computes marginal density functions from a high dimensional joint density function over a Markov network in a finite number of iterations when the underlying Gaussian graph is acyclic. It is also known more recently that the algorithm produces correct marginal means asymptotically for cyclic Gaussian graphs under the condition of walk summability. This paper extends this convergence result further by showing that the convergence is exponential under the walk summability condition, and provides a simple bound for the convergence rate.

stat.ML↗