SearcharxivSearch

arXiv subjects

Minyue Fu

Publications and source records attributed to Minyue Fu.

At least 19 recordsLinked to original sources

Reinforcement Learning-Based Output Feedback LQR for Continuous-Time MIMO Systems

This article studies model-free output feedback linear quadratic regulation (LQR) for continuous-time linear systems with an $n$-dimensional state, an $m$-dimensional input, and a $p$-dimensional output, using filtered input--output data. Since the system state is unavailable, existing methods rely on dynamic filters to parameterize the hidden state using measurable input--output signals. However, the intrinsic dimension of the resulting filter-based parametrization can be smaller than the dimension of the complete filtered vector, and this deterministic redundancy can make the Bellman regressions rank deficient. We characterize this intrinsic dimension and show that the conventional filtered vector contains only $2n$ independent components for single-input multi-output (SIMO) systems and $n(m+1)$ independent components for general multi-input multi-output (MIMO) systems. Based on this characterization, a reduced filtered vector is extracted directly from data and used to develop reduced model-free output feedback policy iteration and value iteration equations, eliminating the redundant directions and decreasing the number of unknown parameters while retaining a fully input--output data-based implementation. A numerical example illustrates the rank reduction and the effectiveness of the learned controller.

eess.SY

Comparative Study of Q-Learning for State-Feedback LQG Control with an Unknown Model

We study the problem of designing a state feedback linear quadratic Gaussian (LQG) controller for a system in which the system matrices as well as the process noise covariance are unknown. We do a rigorous comparison between two approaches. The first is the classic one in which a system identification stage is used to estimate the unknown parameters, which are then used in a state-feedback LQG (SF-LQG) controller design. The second approach is a recently proposed one using a reinforcement learning paradigm called Q-learning. We do the comparison in terms of complexity and accuracy of the resulting controller. We show that the classic approach asymptotically efficient, giving virtually no room for improvement in terms of accuracy. We also propose a novel Q-learning-based method which we show asymptotically achieves the optimal controller design. We complement our proposed method with a numerically efficient algorithmic implementation aiming at making it competitive in terms of computations. Nevertheless, our complexity analysis shows that the classic approach is still numerically more efficient than this Q-learning-based alternative. We then conclude that the classic approach remains being the best choice for addressing the SF-LQG design in the case of unknown parameters.

eess.SY

Nanometer Scanning with Micrometer Sensing: Beating Quantization Constraints in Lissajous Trajectory Tracking

This paper addresses the task of tracking Lissajous trajectories in the presence of quantized positioning sensors. To do so, theoretical results on tracking of continuous time periodic signals in the presence of output quantization are provided. With these results in hand, the application to Lissajous tracking is explored. The method proposed relies on the internal model principle and dispenses perfect knowledge of the system equations. Numerical results show that an arbitrary small scanning resolution is achievable despite large sensor quantization intervals.

eess.SY

Decentralized Strategies for Finite Population Linear-Quadratic-Gaussian Games and Teams

This paper is concerned with a new class of mean-field games which involve a finite number of agents. Necessary and sufficient conditions are obtained for the existence of the decentralized open-loop Nash equilibrium in terms of non-standard forward-backward stochastic differential equations (FBSDEs). By solving the FBSDEs, we design a set of decentralized strategies by virtue of two differential Riccati equations. Instead of the $\varepsilon$-Nash equilibrium in classical mean-field games, the set of decentralized strategies is shown to be a Nash equilibrium. For the infinite-horizon problem, a simple condition is given for the solvability of the algebraic Riccati equation arising from consensus. Furthermore, the social optimal control problem is studied. Under a mild condition, the decentralized social optimal control and the corresponding social cost are given.

math.OC

LQG Differential Stackelberg Game under Nested Observation Information Pattern

We investigate the linear quadratic Gaussian Stackelberg game under a class of nested observation information pattern. Two decision makers implement control strategies relying on different information sets: The follower uses its observation data to design its strategy, whereas the leader implements its strategy using global observation data. We show that the solution requires solving a new type of forward-backward stochastic differential equations whose drift terms contain two types of conditional expectation terms associated to the adjoint variables. We then propose a method to find the functional relations between each adjoint pair, i.e., each pair formed by an adjoint variable and the conditional expectation of its associated state. The proposed method follows a layered pattern. More precisely, in the inner layer, we seek the functional relation for the adjoint pair under the sigma-sub-algebra generated by follower's observation information; and in the outer layer, we look for the functional relation for the adjoint pair under the sigma-sub-algebra generated by leader's observation information. Our result shows that the optimal open-loop solution admits an explicit feedback type representation. More precisely, the feedback coefficient matrices satisfy tuples of coupled forward-backward differential Riccati equations, and feedback variables are computed by Kalman-Bucy filtering.

math.OC

Reinforcement Learning Approach to Estimation in Linear Systems

This paper addresses two important estimation problems for linear systems, namely system identification and model-free state estimation. Our focus is on ARMAX models with unknown parameters. We first provide a reinforcement learning algorithm for system identification with guaranteed consistency. This algorithm is then used to provide a novel solution to model-free state estimation. These results are then applied to solving the model-free LQG control problem in the reinforcement learning setting.

eess.SY

Distributed Newton Optimization with Maximized Convergence Rate

The distributed optimization problem is set up in a collection of nodes interconnected via a communication network. The goal is to find the minimizer of a global objective function formed by the addition of partial functions locally known at each node. A number of methods are available for addressing this problem, having different advantages. The goal of this work is to achieve the maximum possible convergence rate. As the first step towards this end, we propose a new method which we show converges faster than other available options. As with most distributed optimization methods, convergence rate depends on a step size parameter. As the second step towards our goal we complement the proposed method with a fully distributed method for estimating the optimal step size that maximizes convergence speed. We provide theoretical guarantees for the convergence of the resulting method in a neighborhood of the solution. Also, for the case in which the global objective function has a single local minimum, we provide a different step size selection criterion together with theoretical guarantees for convergence. We present numerical experiments showing that, when using the same step size, our method converges significantly faster than its rivals. Experiments also show that the distributed step size estimation method achieves an asymptotic convergence rate very close to the theoretical maximum.

math.OC

Distributed Kalman Estimation with Decoupled Local Filters

We study a distributed Kalman filtering problem in which a number of nodes cooperate without central coordination to estimate a common state based on local measurements and data received from neighbors. This is typically done by running a local filter at each node using information obtained through some procedure for fusing data across the network. A common problem with existing methods is that the outcome of local filters at each time step depends on the data fused at the previous step. We propose an alternative approach to eliminate this error propagation. The proposed local filters are guaranteed to be stable under some mild conditions on certain global structural data, and their fusion yields the centralized Kalman estimate. The main feature of the new approach is that fusion errors introduced at a given time step do not carry over to subsequent steps. This offers advantages in many situations including when a global estimate in only needed at a rate slower than that of measurements or when there are network interruptions. If the global structural data can be fused correctly asymptotically, the stability of local filters is equivalent to that of the centralized Kalman filter. Otherwise, we provide conditions to guarantee stability and bound the resulting estimation error. Numerical experiments are given to show the advantage of our method over other existing alternatives.

eess.SY

Convergence Rate of a Message-passing Algorithm for Solving Linear Systems

This paper studies the convergence rate of a message-passing distributed algorithm for solving a large-scale linear system. This problem is generalised from the celebrated Gaussian Belief Propagation (BP) problem for statistical learning and distributed signal processing, and this message-passing algorithm is generalised from the well-celebrated Gaussian BP algorithm. Under the assumption of generalised diagonal dominance, we reveal, through painstaking derivations, several bounds on the convergence rate of the message-passing algorithm. In particular, we show clearly how the convergence rate of the algorithm can be explicitly bounded using the diagonal dominance properties of the system. When specialised to the Gaussian BP problem, our work also offers new theoretical insight into the behaviour of the BP algorithm because we use a purely linear algebraic approach for convergence analysis.

eess.SY

Convergence and Accuracy Analysis for A Distributed Static State Estimator based on Gaussian Belief Propagation

This paper focuses on the distributed static estimation problem and a Belief Propagation (BP) based estimation algorithm is proposed. We provide a complete analysis for convergence and accuracy of it. More precisely, we offer conditions under which the proposed distributed estimator is guaranteed to converge and we give concrete characterizations of its accuracy. Our results not only give a new algorithm with good performance but also provide a useful analysis framework to learn the properties of a distributed algorithm. It yields better theoretical understanding of the static distributed state estimator and may generate more applications in the future.

eess.SY

Distributed Weighted Least-squares Estimation for Networked Systems with Edge Measurements

This paper studies the problem of distributed weighted least-squares (WLS) estimation for an interconnected linear measurement network with additive noise. Two types of measurements are considered: self measurements for individual nodes, and edge measurements for the connecting nodes. Each node in the network carries out distributed estimation by using its own measurement and information transmitted from its neighbours. We study two distributed estimation algorithms: a recently proposed distributed WLS algorithm and the so-called Gaussian Belief Propagation (BP) algorithm. We first establish the equivalence of the two algorithms. We then prove a key result which shows that the information matrix is always generalised diagonally dominant, under some very mild condition. Using these two results and some known convergence properties of the Gaussian BP algorithm, we show that the aforementioned distributed WLS algorithm gives the globally optimal WLS estimate asymptotically. A bound on its convergence rate is also presented.

eess.SY

The Vulnerability of Cyber-Physical System under Stealthy Attacks

In this paper, we study the impact of stealthy attacks on the Cyber-Physical System (CPS) modeled as a stochastic linear system. An attack is characterised by a malicious injection into the system through input, output or both, and it is called stealthy (resp.~strictly stealthy) if it produces bounded changes (resp.~no changes) in the detection residue. Correspondingly, a CPS is called vulnerable (resp.~strictly vulnerable) if it can be destabilized by a stealthy attack (resp.~strictly stealthy attack). We provide necessary and sufficient conditions for the vulnerability and strictly vulnerability. For the invulnerable case, we also provide a performance bound for the difference between healthy and attacked system. Numerical examples are provided to illustrate the theoretical results.

eess.SY

Statistical Approach to Detection of Attacks for Stochastic Cyber-Physical Systems

We study the problem of detecting an attack on a stochastic cyber-physical system. We aim to treat the problem in its most general form. We start by introducing the notion of asymptotically detectable attacks, as those attacks introducing changes to the system's output statistics which persist asymptotically. We then provide a necessary and sufficient condition for asymptotic detectability. This condition preserves generality as it holds under no restrictive assumption on the system and attacking scheme. To show the importance of this condition, we apply it to detect certain attacking schemes which are undetectable using simple statistics. Our necessary and sufficient condition naturally leads to an algorithm which gives a confidence level for attack detection. We present simulation results to illustrate the performance of this algorithm.

eess.SY

Multi-sensor State Estimation over Lossy Channels using Coded Measurements

This paper focuses on a networked state estimation problem for a spatially large linear system with a distributed array of sensors, each of which offers partial state measurements, and the transmission is lossy. We propose a measurement coding scheme with two goals. Firstly, it permits adjusting the communication requirements by controlling the dimension of the vector transmitted by each sensor to the central estimator. Secondly, for a given communication requirement, the scheme is optimal, within the family of linear causal coders, in the sense that the weakest channel condition is required to guarantee the stability of the estimator. For this coding scheme, we derive the minimum mean-square error (MMSE) state estimator, and state a necessary and sufficient condition with a trivial gap, for its stability. We also derive a sufficient but easily verifiable stability condition, and quantify the advantage offered by the proposed coding scheme. Finally, simulations results are presented to confirm our claims.

eess.SY

A Fast Converging Distributed Solver for Linear Systems with Generalised Diagonal Dominance

This paper proposes a new distributed algorithm for solving linear systems associated with a sparse graph under a generalised diagonal dominance assumption. The algorithm runs iteratively on each node of the graph, with low complexities on local information exchange between neighbouring nodes, local computation and local storage. For an acyclic graph under the condition of diagonal dominance, the algorithm is shown to converge to the correct solution in a finite number of iterations, equalling the diameter of the graph. For a loopy graph, the algorithm is shown to converge to the correct solution asymptotically. Simulations verify that the proposed algorithm significantly outperforms the classical Jacobi method and a recent distributed linear system solver based on average consensus and orthogonal projection.

eess.SP

A Distributed Adaptive Scheme for Multi-Agent Systems

In traditional adaptive control, the certainty equivalence principle suggests a two-step design scheme. A controller is first designed for the ideal situation assuming the uncertain parameter was known and it renders a Lyapunov function. Then, the uncertain parameter in the controller is replaced by its estimation that is updated by an adaptive law along the gradient of Lyapunov function. This principle does not generally work for a multi-agent system as an adaptive law based on the gradient of (centrally constructed) Lyapunov function cannot be implemented in a distributed fashion, except for limited situations. In this paper, we propose a novel distributed adaptive scheme, not relying on gradient of Lyapunov function, for general multi-agent systems. In this scheme, asymptotic consensus of a second-order uncertain multi-agent system is achieved in a network of directed graph.

eess.SY

Convergence of Message-Passing for Distributed Convex Optimisation with Scaled Diagonal Dominance

This paper studies the convergence properties the well-known message-passing algorithm for convex optimisation. Under the assumption of pairwise separability and scaled diagonal dominance, asymptotic convergence is established and a simple bound for the convergence rate is provided for message-passing. In comparison with previous results, our results do not require the given convex program to have known convex pairwise components and that our bound for the convergence rate is tighter and simpler. When specialised to quadratic optimisation, we generalise known results by providing a very simple bound for the convergence rate.

math.OC

On Convergence Rate of the Gaussian Belief Propagation Algorithm for Markov Networks

Gaussian Belief Propagation (BP) algorithm is one of the most important distributed algorithms in signal processing and statistical learning involving Markov networks. It is well known that the algorithm correctly computes marginal density functions from a high dimensional joint density function over a Markov network in a finite number of iterations when the underlying Gaussian graph is acyclic. It is also known more recently that the algorithm produces correct marginal means asymptotically for cyclic Gaussian graphs under the condition of walk summability. This paper extends this convergence result further by showing that the convergence is exponential under the walk summability condition, and provides a simple bound for the convergence rate.

stat.ML