SearcharxivSearch

arXiv subjects

Majid Mazouchi

Publications and source records attributed to Majid Mazouchi.

10 recordsLinked to original sources

Finite-time Koopman Identifier: A Unified Batch-online Learning Framework for Joint Learning of Koopman Structure and Parameters

In this paper, a unified batch-online learning approach is introduced to learn a linear representation of nonlinear system dynamics using the Koopman operator. The presented system modeling approach leverages a novel incremental Koopman-based update law that retrieves a mini-batch of samples stored in a memory to not only minimizes the instantaneous Koopman operator's identification errors but also the identification errors for the batch of retrieved samples. Discontinuous modifications of gradient flows are presented for the online update law to assure finite-time convergence under easy-to-verify conditions defined on the batch of data. Therefore, this unified online-batch framework allows performing joint sample- and time-domain analysis for converging the Koopman operator's parameters. More specifically, it is shown that if the collected mini-batch of samples guarantees a rank condition, then finite-time guarantee in the time domain can be certified and the settling time depends on the quality of collected samples being reused in the update law. Moreover, the efficiency of the proposed Koopman-based update law is further analyzed by showing that the identification regret in continuous time grows sub-linearly with time. Furthermore, to avoid learning corrupted dynamics due to the selection of an inappropriate set of Koopman observables, a higher-layer meta learner employs a discrete Bayesian optimization algorithm to obtain the best library of observable functions for the operator. Since finite-time convergence of the Koopman model for each set of observable is guaranteed under a rank condition on stored data, the fitness of each set of observables can be obtained based on the identification error on the stored samples in the proposed framework and even without implementing any controller based on the learned system.

eess.SY

A Risk-Averse Preview-based $Q$-Learning Algorithm: Application to Highway Driving of Autonomous Vehicles

A risk-averse preview-based $Q$-learning planner is presented for navigation of autonomous vehicles. To this end, the multi-lane road ahead of a vehicle is represented by a finite-state non-stationary Markov decision process (MDP). A risk assessment unit module is then presented that leverages the preview information provided by sensors along with a stochastic reachability module to assign reward values to the MDP states and update them as scenarios develop. A sampling-based risk-averse preview-based Q-learning algorithm is finally developed that generates samples using the preview information and reward function to learn risk-averse optimal planning strategies without actual interaction with the environment. The risk factor is imposed on the objective function to avoid fluctuation of the Q values, which can jeopardize the vehicle's safety and/or performance. The overall hybrid automaton model of the system is leveraged to develop a feasibility check unit module that detects unfeasible plans and enables the planner system to react proactively to the changes of the environment. Finally, to verify the efficiency of the presented algorithm, its implementation on two highway driving scenarios of an autonomous vehicle in a varying traffic density is considered.

eess.SY

A Convex Optimization Approach for Control of Linear Quadratic Systems with Multiplicative Noise via System Level Synthesis

This paper presents a convex optimization-based solution to the design of state-feedback controllers for solving the linear quadratic regulator (LQR) problem of uncertain discrete-time systems with multiplicative noise. To synthesize a tractable solution, the recently developed system level synthesis (SLS) framework is leveraged. It is shown that SLS shifts the controller synthesis task from the design of a robust controller to the design of the entire set-valued closed-loop system responses. To this end, the closed-loop system response is entirely characterized by probabilistic set-valued maps from the additive noise to control actions and states. A bi-level convex optimization over the achievable set-valued closed-loop responses is then developed to optimize the expected value of the LQR cost against the worst-case closed-loop system response. The solution to this robust optimization problem may be too conservative since it aims at enforcing the design constraints for all possible system realizations. To deal with this issue, the presented optimization problem is next reformulated as a chance-constrained program (CCP) in which the guarantees are not intended in a deterministic sense of satisfaction against all possible closed-loop system responses, but are instead intended in a probabilistic sense of satisfaction against all but a small fraction of the system responses. To approximately solve the CCP without the requirement of knowing the probabilistic description of the uncertainty in the system matrices, the so-called scenario optimization approach is employed, which provides probabilistic guarantees based on a finite number of system realizations and results in a convex optimization program with moderate computational complexity. Finally, numerical simulations are presented to illustrate the theoretical findings.

eess.SY

Performance Analysis of Event-Triggered Consensus Control for Multi-agent Systems under Cyber-Physical Attacks

This work presents a rigorous analysis of the adverse effects of cyber-physical attacks on the performance of multi-agent consensus with event-triggered control protocols. It is shown how a strategic malicious attack on sensors and actuators can deceive the triggering condition of both state-based event-triggered mechanism and combinational state-based event-triggered mechanism, which are commonplace and widely used in the literature. More precisely, it is first shown that a deception attack in the case of combinational state-based event-triggered mechanism can result in a non-triggering misbehavior, in the sense that the compromised agent does not trigger any event and consequently results in partial feedback disconnectivity by preventing information from reaching the local neighbors of the compromised agent. This indicates that the combinational state-based event-triggered mechanism can be leveraged by the attacker to harm the network connectivity by rendering the recent data unavailable to agents. It is then shown that the deception attack in the case of state-based event-triggered mechanism can result in a continuous-triggering misbehavior in the sense that the event-triggered mechanism continuously generates triggering events resulting in undesirable phenomenon of Zeno behavior. Finally, numerical simulations are presented to illustrate the theoretical findings.

eess.SY

Secure Event-Triggered Distributed Kalman Filters for State Estimation over Wireless Sensor Networks

In this paper, we analyze the adverse effects of cyber-physical attacks as well as mitigate their impacts on the event-triggered distributed Kalman filter (DKF). We first show that although event-triggered mechanisms are highly desirable, the attacker can leverage the event-triggered mechanism to cause non-triggering misbehavior which significantly harms the network connectivity and its collective observability. We also show that an attacker can mislead the event-triggered mechanism to achieve continuous-triggering misbehavior which not only drains the communication resources but also harms the network's performance. An information-theoretic approach is presented next to detect attacks on both sensors and communication channels. In contrast to the existing results, the restrictive Gaussian assumption on the attack signal's probability distribution is not required. To mitigate attacks, a meta-Bayesian approach is presented that incorporates the outcome of the attack detection mechanism to perform second-order inference. The proposed second-order inference forms confidence and trust values about the truthfulness or legitimacy of sensors' own estimates and those of their neighbors, respectively. Each sensor communicates its confidence to its neighbors. Sensors then incorporate the confidence they receive from their neighbors and the trust they formed about their neighbors into their posterior update laws to successfully discard corrupted information. Finally, the simulation result validates the effectiveness of the presented resilient event-triggered DKF.

eess.SY

Fully-Heterogeneous Containment Control of a Network of Leader-Follower Systems

This paper develops a distributed solution to the fully-heterogeneous containment control problem (CCP), for which not only the followers' dynamics but also the leaders' dynamics are non-identical. A novel formulation of the fully-heterogeneous CCP is first presented in which each follower constructs its virtual exo-system. To build these virtual exo-systems by followers, a novel distributed algorithm is developed to calculate the so-called normalized level of influences (NLIs) of all leaders on each follower and a novel adaptive distributed observer is designed to estimate the dynamics and states of all leaders that have an influence on each follower. Then, a distributed control protocol is proposed based on the cooperative output regulation framework, utilizing this virtual exo-system. Based on estimations of leaders' dynamics and states and NLIs of leaders on each follower, the solutions of the so-called linear regulator equations are calculated in a distributed manner, and consequently, a distributed control protocol is designed for solving the output containment problem. Finally, theoretical results are verified by performing numerical simulations.

cs.MA

Assured Learning-enabled Autonomy: A Metacognitive Reinforcement Learning Framework

Reinforcement learning (RL) agents with pre-specified reward functions cannot provide guaranteed safety across variety of circumstances that an uncertain system might encounter. To guarantee performance while assuring satisfaction of safety constraints across variety of circumstances, an assured autonomous control framework is presented in this paper by empowering RL algorithms with metacognitive learning capabilities. More specifically, adapting the reward function parameters of the RL agent is performed in a metacognitive decision-making layer to assure the feasibility of RL agent. That is, to assure that the learned policy by the RL agent satisfies safety constraints specified by signal temporal logic while achieving as much performance as possible. The metacognitive layer monitors any possible future safety violation under the actions of the RL agent and employs a higher-layer Bayesian RL algorithm to proactively adapt the reward function for the lower-layer RL agent. To minimize the higher-layer Bayesian RL intervention, a fitness function is leveraged by the metacognitive layer as a metric to evaluate success of the lower-layer RL agent in satisfaction of safety and liveness specifications, and the higher-layer Bayesian RL intervenes only if there is a risk of lower-layer RL failure. Finally, a simulation example is provided to validate the effectiveness of the proposed approach.

cs.AI

A Convex Programming Approach to Data-Driven Risk-Averse Reinforcement Learning

This paper presents a model-free reinforcement learning (RL) algorithm to solve the risk-averse optimal control (RAOC) problem for discrete-time nonlinear systems. While successful RL algorithms have been presented to learn optimal control solutions under epistemic uncertainties (i.e., lack of knowledge of system dynamics), they do so by optimizing the expected utility of outcomes, which ignores the variance of cost under aleatory uncertainties (i.e., randomness). Performance-critical systems, however, must not only optimize the expected performance, but also reduce its variance to avoid performance fluctuation during RL's course of operation. To solve the RAOC problem, this paper presents the following three variants of RL algorithms and analyze their advantages and preferences for different situations/systems: 1) a one-shot static convex program -based RL, 2) an iterative value iteration (VI) algorithm that solves a linear programming (LP) optimization at each iteration, and 3) an iterative policy iteration (PI) algorithm that solves a convex optimization at each iteration and guarantees the stability of the consecutive control policies. Convergence of the exact optimization problems, which are infinite-dimensional in all three cases, to the optimal risk-averse value function is shown. To turn these optimization problems into standard optimization problems with finite decision variables and constraints, function approximation for value estimations as well as constraint sampling are leveraged. Data-driven implementations of these algorithms are provided based on Q-function which enables learning the optimal value without any knowledge of the system dynamics. The performance of the approximated solutions is also verified through a weighted sup-norm bound and the Lyapunov bound. A simulation example is provided to verify the effectiveness of the presented approach.

eess.SY

Data-driven Dynamic Multi-objective Optimal Control: An Aspiration-satisfying Reinforcement Learning Approach

This paper presents an iterative data-driven algorithm for solving dynamic multi-objective (MO) optimal control problems arising in control of nonlinear continuous-time systems. It is first shown that the Hamiltonian functional corresponding to each objective can be leveraged to compare the performance of admissible policies. Hamiltonian-inequalities are then used for which their satisfaction guarantees satisfying the objectives' aspirations. An aspiration-satisfying dynamic optimization framework is then presented to optimize the main objective while satisfying the aspiration of other objectives. Relation to satisficing (good enough) decision-making framework is shown. A Sum-of-Square (SOS) based iterative algorithm is developed to solve the formulated aspiration-satisfying MO optimization. To obviate the requirement of complete knowledge of the system dynamics, a data-driven satisficing reinforcement learning approach is proposed to solve the SOS optimization problem in real-time using only the information of the system trajectories measured during a time interval without having full knowledge of the system dynamics. Finally, two simulation examples are provided to show the effectiveness of the proposed algorithm.

eess.SY

Observer-based Adaptive Optimal Output Containment Control problem of Linear Heterogeneous Multi-agent Systems with Relative Output Measurements

This paper develops an optimal relative output-feedback based solution to the containment control problem of linear heterogeneous multi-agent systems. A distributed optimal control protocol is presented for the followers to not only assure that their outputs fall into the convex hull of the leaders' output (i.e., the desired or safe region), but also optimizes their transient performance. The proposed optimal control solution is composed of a feedback part, depending of the followers' state, and a feed-forward part, depending on the convex hull of the leaders' state. To comply with most real-world applications, the feedback and feed-forward states are assumed to be unavailable and are estimated using two distributed observers. That is, since the followers cannot directly sense their absolute states, a distributed observer is designed that uses only relative output measurements with respect to their neighbors (measured for example by using range sensors in robotic) and the information which is broadcasted by their neighbors to estimate their states. Moreover, another adaptive distributed observer is designed that uses exchange of information between followers over a communication network to estimate the convex hull of the leaders' state. The proposed observer relaxes the restrictive requirement of knowing the complete knowledge of the leaders' dynamics by all followers. An off-policy reinforcement learning algorithm on an actor-critic structure is next developed to solve the optimal containment control problem online, using relative output measurements and without requirement of knowing the leaders' dynamics by all followers. Finally, the theoretical results are verified by numerical simulations.

eess.SY