SearcharxivSearch

arXiv subjects

Antoine Aspeel

Publications and source records attributed to Antoine Aspeel.

14 recordsLinked to original sources

Exploiting Over-Approximation Errors as Preview Information for Nonlinear Control

We study the control of nonlinear constrained systems via over-approximations. Our key observation is that the over-approximation error, rather than being an unknown disturbance, can be exploited as input-dependent preview information. This leads to the notion of informed policies, which depend on both the state and the error. We formulate the concretization problem -- recovering a valid input for the true system from a preview-based policy -- as a fixed-point equation. Existence of solutions follows from the Brouwer fixed-point theorem, while efficient computation is enabled through closed-form, linear, or convex programs for input-affine systems, and through an iterative method based on the Banach fixed-point theorem for nonlinear systems.

math.OC

Active Reward Machine Inference From Raw State Trajectories

Reward machines are automaton-like structures that capture the memory required to accomplish a multi-stage task. When combined with reinforcement learning or optimal control methods, they can be used to synthesize robot policies to achieve such tasks. However, specifying a reward machine by hand, including a labeling function capturing high-level features that the decisions are based on, can be a daunting task. This paper deals with the problem of learning reward machines directly from raw state and policy information. As opposed to existing works, we assume no access to observations of rewards, labels, or machine nodes, and show what trajectory data is sufficient for learning the reward machine in this information-scarce regime. We then extend the result to an active learning setting where we incrementally query trajectory extensions to improve data (and indirectly computational) efficiency. Results are demonstrated with several grid world examples.

cs.RO

Learning Reward Machines from Partially Observed Policies

Inverse reinforcement learning is the problem of inferring a reward function from an optimal policy or demonstrations by an expert. In this work, it is assumed that the reward is expressed as a reward machine whose transitions depend on atomic propositions associated with the state of a Markov Decision Process (MDP). Our goal is to identify the true reward machine using finite information. To this end, we first introduce the notion of a prefix tree policy which associates a distribution of actions to each state of the MDP and each attainable finite sequence of atomic propositions. Then, we characterize an equivalence class of reward machines that can be identified given the prefix tree policy. Finally, we propose a SAT-based algorithm that uses information extracted from the prefix tree policy to solve for a reward machine. It is proved that if the prefix tree policy is known up to a sufficient (but finite) depth, our algorithm recovers the exact reward machine up to the equivalence class. This sufficient depth is derived as a function of the number of MDP states and (an upper bound on) the number of states of the reward machine. These results are further extended to the case where we only have access to demonstrations from an optimal policy. Several examples, including discrete grid and block worlds, a continuous state-space robotic arm, and real data from experiments with mice, are used to demonstrate the effectiveness and generality of the approach.

cs.LG

Mode-Prefix-Based Control of Switched Linear Systems with Applications to Fault Tolerance

In this paper, we consider the problem of designing prefix-based optimal controllers for switched linear systems over finite horizons. This problem arises in fault-tolerant control, when system faults result in abrupt changes in dynamics. We consider a class of mode-prefix-based linear controllers that depend only on the history of the switching signal. The proposed optimal control problems seek to minimize both expected performance and worst-case performance over switching signals. We show that this problem can be reduced to a convex optimization problem. To this end, we synthesize one controller for each switching signal under a prefix constraint that ensures consistency between controllers. Then, system level synthesis is used to obtain a convex program in terms of the system-level parameters. In particular, it is shown that the prefix constraints are linear in terms of the system-level parameters. Finally, we apply this framework for optimal control of a fighter jet model suffering from system faults, illustrating how fault tolerance is ensured.

math.OC

Minimal L2-Consistent Data-Transmission

In this work, we consider non-collocated sensors and actuators, and we address the problem of minimizing the number of sensor-to-actuator transmissions while ensuring that the L2 gain of the system remains under a threshold. By using causal factorization and system level synthesis, we reformulate this problem as a rank minimization problem over a convex set. When heuristics like nuclear norm minimization are used for rank minimization, the resulting matrix is only numerically low rank and must be truncated, which can lead to an infeasible solution. To address this issue, we introduce approximate causal factorization to control the factorization error and provide a bound on the degradation of the L2 gain in terms of the factorization error. The effectiveness of our method is demonstrated using a benchmark.

eess.SY

A Simulation Preorder for Koopman-like Lifted Control Systems

This paper introduces a simulation preorder among lifted systems, a generalization of finite-dimensional Koopman approximations (also known as approximate immersions) to systems with inputs. It is proved that this simulation relation implies the containment of both the open- and closed-loop behaviors. Optimization-based sufficient conditions are derived to verify the simulation relation in two special cases: i) a nonlinear (unlifted) system and an affine lifted system and, ii) two affine lifted systems. Numerical examples demonstrate the approach.

eess.SY

Ellipsoidal embeddings of graphs

Due to their flexibility to represent almost any kind of relational data, graph-based models have enjoyed a tremendous success over the past decades. While graphs are inherently only combinatorial objects, however, many prominent analysis tools are based on the algebraic representation of graphs via matrices such as the graph Laplacian, or on associated graph embeddings. Such embeddings associate to each node a set of coordinates in a vector space, a representation which can then be employed for learning tasks such as the classification or alignment of the nodes of the graph. As the geometric picture provided by embedding methods enables the use of a multitude of methods developed for vector space data, embeddings have thus gained interest both from a theoretical as well as a practical perspective. Inspired by trace-optimization problems, often encountered in the analysis of graph-based data, here we present a method to derive ellipsoidal embeddings of the nodes of a graph, in which each node is assigned a set of coordinates on the surface of a hyperellipsoid. Our method may be seen as an alternative to popular spectral embedding techniques, to which it shares certain similarities we discuss. To illustrate the utility of the embedding we conduct a case study in which we analyse synthetic and real world networks with modular structure, and compare the results obtained with known methods in the literature.

cs.SI

A Low Rank Approach to Minimize Sensor-to-Actuator Communication in Finite Horizon Output Feedback

Many modern controllers are composed of different components that communicate in real-time over some network with limited resources. In this work, we are interested in designing a controller that can be implemented with a minimum number of sensor-to-actuator messages, while satisfying safety constraints over a finite horizon. For finite horizon problems, a linear time-varying controller with memory can be represented as a block-lower-triangular matrix. We show that the rank of this matrix exactly captures the minimum number of messages needed to be sent from the sensors to actuators to implement such a controller. Moreover, we introduce a novel matrix factorization called causal factorization that gives the required implementation. Finally, we show that the rank of the controller is the same as the rank of the Youla parameter, enabling the Youla parametrization (or analogous parametrizations) to be used to design the controller, which reduces the overall design problem into a rank minimization one over a convex set. Finally, convex relaxations for rank are used to demonstrate that our approach leads to 20-50% less messages on a simulation than a benchmark method.

eess.SY

Koopman-inspired Implicit Backward Reachable Sets for Unknown Nonlinear Systems

Koopman liftings have been successfully used to learn high dimensional linear approximations for autonomous systems for prediction purposes, or for control systems for leveraging linear control techniques to control nonlinear dynamics. In this paper, we show how learned Koopman approximations can be used for state-feedback correct-by-construction control. To this end, we introduce the Koopman over-approximation, a (possibly hybrid) lifted representation that has a simulation-like relation with the underlying dynamics. Then, we prove how successive application of controlled predecessor operation in the lifted space leads to an implicit backward reachable set for the actual dynamics. Finally, we demonstrate the approach on two nonlinear control examples with unknown dynamics.

eess.SY

Optimal Intermittent Particle Filter

The problem of the optimal allocation (in the expected mean square error sense) of a measurement budget for particle filtering is addressed. We propose three different optimal intermittent filters, whose optimality criteria depend on the information available at the time of decision making. For the first, the stochastic program filter, the measurement times are given by a policy that determines whether a measurement should be taken based on the measurements already acquired. The second, called the offline filter, determines all measurement times at once by solving a combinatorial optimization program before any measurement acquisition. For the third one, which we call online filter, each time a new measurement is received, the next measurement time is recomputed to take all the information that is then available into account. We prove that in terms of expected mean square error, the stochastic program filter outperforms the online filter, which itself outperforms the offline filter. However, these filters are generally intractable. For this reason, the filter estimate is approximated by a particle filter. Moreover, the mean square error is approximated using a Monte-Carlo approach, and different optimization algorithms are compared to approximately solve the combinatorial programs (a random trial algorithm, greedy forward and backward algorithms, a simulated annealing algorithm, and a genetic algorithm). Finally, the performance of the proposed methods is illustrated on two examples: a tumor motion model and a common benchmark for particle filtering.

eess.SY

Positive semi-definite embedding for dimensionality reduction and out-of-sample extensions

In machine learning or statistics, it is often desirable to reduce the dimensionality of a sample of data points in a high dimensional space $\mathbb{R}^d$. This paper introduces a dimensionality reduction method where the embedding coordinates are the eigenvectors of a positive semi-definite kernel obtained as the solution of an infinite dimensional analogue of a semi-definite program. This embedding is adaptive and non-linear. We discuss this problem both with weak and strong smoothness assumptions about the learned kernel. A main feature of our approach is the existence of an out-of-sample extension formula of the embedding coordinates in both cases. This extrapolation formula yields an extension of the kernel matrix to a data-dependent Mercer kernel function. Our empirical results indicate that this embedding method is more robust with respect to the influence of outliers, compared with a spectral embedding method.

cs.LG

Optimal Control for Linear Networked Control Systems with Information Transmission Constraints

This paper addresses the problem of robust control of a linear discrete-time system subject to bounded disturbances and to measurement and control budget constraints. Using Q-parameterization and a polytope containment method, we prove that the co-design of an affine feedback controller, a measurement schedule and a control schedule can be exactly formulated as a mixed integer linear program with 2 binary variables per time step. As a consequence, this problem can be solved efficiently, even when an exhaustive search for measurement and control times would have been impossible in a reasonable amount of time.

math.OC

Optimal measurement budget allocation for particle filtering

Particle filtering is a powerful tool for target tracking. When the budget for observations is restricted, it is necessary to reduce the measurements to a limited amount of samples carefully selected. A discrete stochastic nonlinear dynamical system is studied over a finite time horizon. The problem of selecting the optimal measurement times for particle filtering is formalized as a combinatorial optimization problem. We propose an approximated solution based on the nesting of a genetic algorithm, a Monte Carlo algorithm and a particle filter. Firstly, an example demonstrates that the genetic algorithm outperforms a random trial optimization. Then, the interest of non-regular measurements versus measurements performed at regular time intervals is illustrated and the efficiency of our proposed solution is quantified: better filtering performances are obtained in 87.5% of the cases and on average, the relative improvement is 27.7%.

eess.SY

Optimal Intermittent Measurements for Tumor Tracking in X-ray Guided Radiotherapy

In radiation therapy, tumor tracking is a challenging task that allows a better dose delivery. One practice is to acquire X-ray images in real-time during treatment, that are used to estimate the tumor location. These informations are used to predict the close future tumor trajectory. Kalman prediction is a classical approach for this task. The main drawback of X-ray acquisition is that it irradiates the patient, including its healthy tissues. In the classical Kalman framework, X-ray measurements are taken regularly, i.e. at a constant rate. In this paper, we propose a new approach which relaxes this constraint in order to take measurements when they are the most useful. Our aim is for a given budget of measurements to optimize the tracking process. This idea naturally brings to an optimal intermittent Kalman predictor for which measurement times are selected to minimize the mean squared prediction error over the complete fraction. This optimization problem can be solved directly when the respiratory model has been identified and the optimal sampling times can be computed at once. These optimal measurement times are obtained by solving a combinatorial optimization problem using a genetic algorithm. We created a test benchmark on trajectories validated on one patient. This new prediction method is compared to the regular Kalman predictor and a relative improvement of 9:8% is observed on the root mean square position estimation error.

stat.AP