SearcharxivSearch

arXiv subjects

Frank L. Lewis

Publications and source records attributed to Frank L. Lewis.

14 recordsLinked to original sources

Provable Reinforcement Learning for Networked Control Systems with Stochastic Packet Disordering

This paper formulates a stochastic optimal control problem for linear networked control systems featuring stochastic packet disordering with a unique stabilizing solution certified. The problem is solved by proposing reinforcement learning algorithms. A measurement method is first presented to deal with PD and calculate the newest control input. The NCSs with stochastic PD are modeled as stochastic NCSs. Then, given a cost function, a modified algebraic Riccati equation is derived within the formulation. We propose offline policy iteration and value iteration algorithms to solve the MARE associated with provable convergence. These two algorithms require knowledge of NCS dynamics and PD probabilities. To release that, we further design online model-free off-policy and Q-learning algorithms with an online estimation method for PD probability. Both model-free algorithms solve the optimal control problem using real-time system states, control inputs, and PD probability estimates. Simulation results verify the proposed formulation and algorithms at last.

eess.SY

Structural Balance of Complex Weighted Graphs and Multi-partite Consensus

The structural balance of a signed graph is known to be necessary and sufficient to obtain a bipartite consensus among agents with friend-foe relationships. In the real world, relationships are multifarious, and the coexistence of different opinions is ubiquitous. We are therefore motivated to study the multi-partite consensus problem of multi-agent systems, for which we extend the concept of structural balance to graphs with complex edge weights. It is shown that the generalized structural balance property is necessary and sufficient for achieving multi-partite consensus.

eess.SY

Data-Driven Inverse Reinforcement Learning for Expert-Learner Zero-Sum Games

In this paper, we formulate inverse reinforcement learning (IRL) as an expert-learner interaction whereby the optimal performance intent of an expert or target agent is unknown to a learner agent. The learner observes the states and controls of the expert and hence seeks to reconstruct the expert's cost function intent and thus mimics the expert's optimal response. Next, we add non-cooperative disturbances that seek to disrupt the learning and stability of the learner agent. This leads to the formulation of a new interaction we call zero-sum game IRL. We develop a framework to solve the zero-sum game IRL problem that is a modified extension of RL policy iteration (PI) to allow unknown expert performance intentions to be computed and non-cooperative disturbances to be rejected. The framework has two parts: a value function and control action update based on an extension of PI, and a cost function update based on standard inverse optimal control. Then, we eventually develop an off-policy IRL algorithm that does not require knowledge of the expert and learner agent dynamics and performs single-loop learning. Rigorous proofs and analyses are given. Finally, simulation experiments are presented to show the effectiveness of the new approach.

cs.LG

Prescribed-Time Control and Its Latest Developments

Prescribed-time (PT) control, originated from \textit{Song et al.}, has gained increasing attention among control community. The salient feature of PT control lies in its ability to achieve system stability within a finite settling time user-assignable in advance irrespective of initial conditions. It is such a unique feature that has enticed many follow-up studies on this technically important area, motivating numerous research advancements. In this article, we provide a comprehensive survey on the recent developments in PT control. Through a concise introduction to the concept of PT control, and a unique taxonomy covering: 1) from robust PT control to adaptive PT control; 2) from PT control for single-input-single-output (SISO) systems to multi-input-multi-output (MIMO) systems; and 3) from PT control for single systems to multi-agent systems, we present an accessible review of this interesting topic. We highlight key techniques, fundamental assumptions adopted in various developments as well as some new design ideas. We also discuss several possibles future research directions towards PT control.

eess.SY

Sequential Convex Programming for Collaboration of Connected and Automated Vehicles

This paper investigates the collaboration of multiple connected and automated vehicles (CAVs) in different scenarios. In general, the collaboration of CAVs can be formulated as a nonlinear and nonconvex model predictive control (MPC) problem. Most of the existing approaches available for utilization to solve such an optimization problem suffer from the drawback of considerable computational burden, which hinders the practical implementation in real time. This paper proposes the use of sequential convex programming (SCP), which is a powerful approach to solving the nonlinear and nonconvex MPC problem in real time. To appropriately deploy the methodology, as a first stage, SCP requires linearization and discretization when addressing the nonlinear dynamics of the system model adequately. Based on the linearization and discretization, the original MPC problem can be transformed into a quadratically constrained quadratic programming (QCQP) problem. Besides, SCP also involves convexification to handle the associated nonconvex constraints. Thus, the nonconvex QCQP can be reduced to a quadratic programming (QP) problem that can be solved rather quickly. Therefore, the computational efficiency is suitably improved despite the existence of nonlinear and nonconvex characteristics, whereby the implementation is realized in real time. Furthermore, simulation results in three different scenarios of autonomous driving are presented to validate the effectiveness and efficiency of our proposed approach.

math.OC

Learning nonlinear dynamics in synchronization of knowledge-based leader-following networks

Knowledge-based leader-following synchronization of heterogeneous nonlinear multi-agent systems is a challenging problem since the leader's dynamic information is unknown to any follower node. This paper proposes a learning-based fully distributed observer for a class of nonlinear leader systems, which can simultaneously learn the leader's dynamics and states. This class of leader dynamics is rather general and does not require a bounded Jacobian matrix. Based on this learning-based distributed observer, we further synthesize an adaptive distributed control law for solving the leader-following synchronization problem of multiple Euler-Lagrange systems subject to an uncertain nonlinear leader system. The results are illustrated by a simulation example.

eess.SY

Neuro-adaptive Cooperative Tracking Control with Prescribed Performance of Unknown Higher-order Nonlinear Multi-agent Systems

This paper is concerned with the design of a distributed cooperative synchronization controller for a class of higher-order nonlinear multi-agent systems. The objective is to achieve synchronization and satisfy a predefined time-based performance. Dynamics of the agents (also called the nodes) are assumed to be unknown to the controller and are estimated using Neural Networks. The proposed robust neuro-adaptive controller drives different states of nodes systematically to synchronize with the state of the leader node within the constraints of the prescribed performance. The nodes are connected through a weighted directed graph with a time-invariant topology. Only few nodes have access to the leader. Lyapunov-based stability proofs demonstrate that the multi-agent system is uniformly ultimately bounded stable. Highly nonlinear heterogeneous networked systems with uncertain parameters and external disturbances were used to validate the robustness and performance of the new novel approach. Simulation results considered two different examples: single-input single-output and multi-input multi-output, which demonstrate the effectiveness of the proposed controller. Keywords: Prescribed performance, Transformed error, Multi-agents, Neuro-Adaptive, Distributed adaptive control, Consensus, Transient, Steady-state error, Communication graph, Networked Systems, Synchronization, Robustness, Estimation, Estimator, Observer, Filter, operator, small, error, dynamics, kinematics, equilibrium, asymptotic, zero, unknown, time-varying, neighborhood, global, node, agent, Neural Networks, semi-global, stable, stability, uncertain, noise, bias, singular value, matrix, bounded, origin, comparison, rigid body, 3D, space, mapping, Laplacian matrix, directed graph, disturbance, Theory, undirected graph, Inertial measurement units, IMUs, single-input single-output, multi-input multi-output, SISO, MIMO.

math.OC

Semi-Definite Relaxation Based ADMM for Cooperative Planning and Control of Connected Autonomous Vehicles

This paper investigates the cooperative planning and control problem for multiple connected autonomous vehicles (CAVs) in different scenarios. In the existing literature, most of the methods suffer from significant problems in computational efficiency. Besides, as the optimization problem is nonlinear and nonconvex, it typically poses great difficultly in determining the optimal solution. To address this issue, this work proposes a novel and completely parallel computation framework by leveraging the alternating direction method of multipliers (ADMM). The nonlinear and nonconvex optimization problem in the autonomous driving problem can be divided into two manageable subproblems; and the resulting subproblems can be solved by using effective optimization methods in a parallel framework. Here, the differential dynamic programming (DDP) algorithm is capable of addressing the nonlinearity of the system dynamics rather effectively; and the nonconvex coupling constraints with small dimensions can be approximated by invoking the notion of semi-definite relaxation (SDR), which can also be solved in a very short time. Due to the parallel computation and efficient relaxation of nonconvex constraints, our proposed approach effectively realizes real-time implementation and thus also extra assurance of driving safety is provided. In addition, two transportation scenarios for multiple CAVs are used to illustrate the effectiveness and efficiency of the proposed method.

cs.MA

Local Policy Optimization for Trajectory-Centric Reinforcement Learning

The goal of this paper is to present a method for simultaneous trajectory and local stabilizing policy optimization to generate local policies for trajectory-centric model-based reinforcement learning (MBRL). This is motivated by the fact that global policy optimization for non-linear systems could be a very challenging problem both algorithmically and numerically. However, a lot of robotic manipulation tasks are trajectory-centric, and thus do not require a global model or policy. Due to inaccuracies in the learned model estimates, an open-loop trajectory optimization process mostly results in very poor performance when used on the real system. Motivated by these problems, we try to formulate the problem of trajectory optimization and local policy synthesis as a single optimization problem. It is then solved simultaneously as an instance of nonlinear programming. We provide some results for analysis as well as achieved performance of the proposed technique under some simplifying assumptions.

cs.LG

A Unified Strategy for Solution Seeking in Graphical N-coalition Noncooperative Games

This paper aims to reduce the communication and computation costs of the Nash equilibrium seeking strategy for the $N$-coalition noncooperative games proposed in [1]. The objective is achieved in two manners: 1. An interference graph is introduced to describe the interactions among the agents in each coalition. 2. The Nash equilibrium seeking strategy is designed with the interference graphs considered. The convergence property of the proposed Nash equilibrium seeking strategy is analytically investigated. It is shown that the agents' actions generated by the proposed method converge to a neighborhood of the Nash equilibrium of the graphical $N$-coalition noncooperative games under certain conditions. Several special cases where there is only one coalition and/or there are coalitions with only one agent are considered. The results for the special cases demonstrate that the proposed seeking strategy achieves the solution seeking for noncooperative games, social cost minimization problems and single-agent optimization problems in a unified framework. Numerical examples are presented to support the theoretical results.

math.OC

Neuro-adaptive distributed control with prescribed performance for the synchronization of unknown nonlinear networked systems

This paper proposes a neuro-adaptive distributive cooperative tracking control with prescribed performance function (PPF) for highly nonlinear multi-agent systems. PPF allows error tracking from a predefined large set to be trapped into a predefined small set. The key idea is to transform the constrained system into unconstrained one through transformation of the output error. Agents' dynamics are assumed to be completely unknown, and the controller is developed for strongly connected structured network. The proposed controller allows all agents to follow the trajectory of the leader node, while satisfying necessary dynamic requirements. The proposed approach guarantees uniform ultimate boundedness of the transformed error and the adaptive neural network weights. Simulations include two examples to validate the robustness and smoothness of the proposed controller against highly nonlinear heterogeneous networked system with time varying uncertain parameters and external disturbances.

math.OC

On the Identifiability of the Influence Model for Stochastic Spatiotemporal Spread Processes

The influence model is a discrete-time stochastic model that succinctly captures the interactions of a network of Markov chains. The model produces a reduced-order representation of the stochastic network, and can be used to describe and tractably analyze probabilistic spatiotemporal spread dynamics, and hence has found broad usage in network applications such as social networks, traffic management, and failure cascades in power systems. This paper provides sufficient and necessary conditions for the identifiability of the influence model, and also develops estimators for the model structure through exploiting the model's special properties. In addition, we analyze conditions for the identifiability of the partially observed influence model (POIM), for which not all of the sites can be measured.

eess.SY

Adaptive synchronisation of unknown nonlinear networked systems with prescribed performance

This paper proposes an adaptive tracking control with prescribed performance function for distributive cooperative control of highly nonlinear multi-agent systems. The use of such approach confines the tracking error within a large predefined set to a predefined smaller set. The key idea is to transform the constrained system into unconstrained one through the transformation of the output error. Agents' dynamics are assumed unknown, and the controller is developed for a strongly connected structured network. The proposed controller allows all agents to follow the trajectory of the leader node, while satisfying the necessary dynamic requirements. The proposed approach guarantees uniform ultimate boundedness for the transformed error as well as a bounded adaptive estimate of the unknown parameters and dynamics. Simulations include two examples to validate the robustness and smoothness of the proposed controller against highly nonlinear heterogeneous multi-agent system with uncertain time-variant parameters and external disturbances. Keywords: Prescribed performance, Transformed error, Multi-agents, Distributed adaptive control, Adaptive Consensus, Transient, Steady-state error, Semi-global asymptotic stability, uniformly ultimately bounded, Nonlinear Networked Systems, Distributed Control, Robustness.

math.OC

Distributed Nonlinear MPC of Multi-Agent Systems with Data Compression and Random Delays - Extended Version

This is an extended version of a technical note accepted for publication in IEEE Transactions on Automatic Control. The note proposes an Input to State practically Stable (ISpS) formulation of distributed nonlinear model predictive controller (NMPC) for formation control of constrained autonomous vehicles in presence of communication bandwidth limitation and transmission delays. Planned trajectories are compressed using neural networks resulting in considerable reduction of data packet size, while being robust to propagation delays and uncertainty in neighbors' trajectories. Collision avoidance is achieved by means of spatially filtered potential field. Analytical results proving ISpS and generalized small gain conditions are presented for both strongly- and weakly-connected networks, and illustrated by simulations.

eess.SY