SearcharxivSearch

arXiv subjects

Sebastian Peitz

Publications and source records attributed to Sebastian Peitz.

At least 37 records · Page 2Linked to original sources

Equivariance and partial observations in Koopman operator theory for partial differential equations

The Koopman operator has become an essential tool for data-driven analysis, prediction and control of complex systems. The main reason is the enormous potential of identifying linear function space representations of nonlinear dynamics from measurements. This equally applies to ordinary, stochastic, and partial differential equations (PDEs). Until now, with a few exceptions only, the PDE case is mostly treated rather superficially, and the specific structure of the underlying dynamics is largely ignored. In this paper, we show that symmetries in the system dynamics can be carried over to the Koopman operator, which allows us to significantly increase the model efficacy. Moreover, the situation where we only have access to partial observations (i.e., measurements, as is very common for experimental data) has not been treated to its full extent, either. Moreover, we address the highly-relevant case where we cannot measure the full state, where alternative approaches (e.g., delay coordinates) have to be considered. We derive rigorous statements on the required number of observables in this situation, based on embedding theory. We present numerical evidence using various numerical examples including the wave equation and the Kuramoto-Sivashinsky equation.

math.DS

Accelerating the analysis of optical quantum systems using the Koopman operator

The prediction of photon echoes is a crucial technique for understanding optical quantum systems. However, it typically requires numerous simulations with varying parameters and input pulses, rendering numerical studies computationally expensive. This article investigates the use of data-driven surrogate models based on the Koopman operator to accelerate this process while maintaining accuracy over many time steps. To this end, we employ a bilinear Koopman model using extended dynamic mode decomposition to simulate the optical Bloch equations for an ensemble of inhomogeneously broadened two-level systems. These systems are well suited to describe the excitation of excitonic resonances in semiconductor nanostructures, such as ensembles of semiconductor quantum dots. We conduct a detailed study to determine the number of system simulations required for the resulting data-driven Koopman model to achieve sufficient accuracy across a wide range of parameter settings. We analyze the L2 error and the relative error of the photon echo peak and investigate how the control positions relate to stabilization. After proper training, our methods can predict the dynamics of the quantum ensemble accurately and with numerical efficiency.

quant-ph

Solving Partial Differential Equations with Equivariant Extreme Learning Machines

We utilize extreme-learning machines for the prediction of partial differential equations (PDEs). Our method splits the state space into multiple windows that are predicted individually using a single model. Despite requiring only few data points (in some cases, our method can learn from a single full-state snapshot), it still achieves high accuracy and can predict the flow of PDEs over long time horizons. Moreover, we show how additional symmetries can be exploited to increase sample efficiency and to enforce equivariance.

cs.LG

Variance representations and convergence rates for data-driven approximations of Koopman operators

We rigorously derive novel error bounds for extended dynamic mode decomposition (EDMD) to approximate the Koopman operator for discrete- and continuous time (stochastic) systems; both for i.i.d. and ergodic sampling under non-restrictive assumptions. We show exponential convergence rates for i.i.d. sampling and provide the first superlinear convergence rates for ergodic sampling of deterministic systems. The proofs are based on novel exact variance representations for the empirical estimators of mass and stiffness matrix. Moreover, we verify the accuracy of the derived error bounds and convergence rates by means of numerical simulations for highly-complex dynamical systems including a nonlinear partial differential equation.

math.DS

Invariance of Gaussian RKHSs under Koopman operators of stochastic differential equations with constant matrix coefficients

We consider the Koopman operator semigroup $(K^t)_{t\ge 0}$ associated with stochastic differential equations of the form $dX_t = AX_t\,dt + B\,dW_t$ with constant matrices $A$ and $B$ and Brownian motion $W_t$. We prove that the reproducing kernel Hilbert space $\bH_C$ generated by a Gaussian kernel with a positive definite covariance matrix $C$ is invariant under each Koopman operator $K^t$ if the matrices $A$, $B$, and $C$ satisfy the following Lyapunov-like matrix inequality: $AC^2 + C^2A^\top\le 2BB^\top$. In this course, we prove a characterization concerning the inclusion $\bH_{C_1}\subset\bH_{C_2}$ of Gaussian RKHSs for two positive definite matrices $C_1$ and $C_2$. The question of whether the sufficient Lyapunov-condition is also necessary is left as an open problem.

math.PR

A multiobjective continuation method to compute the regularization path of deep neural networks

Sparsity is a highly desired feature in deep neural networks (DNNs) since it ensures numerical efficiency, improves the interpretability of models (due to the smaller number of relevant features), and robustness. For linear models, it is well known that there exists a \emph{regularization path} connecting the sparsest solution in terms of the $\ell^1$ norm, i.e., zero weights and the non-regularized solution. Very recently, there was a first attempt to extend the concept of regularization paths to DNNs by means of treating the empirical loss and sparsity ($\ell^1$ norm) as two conflicting criteria and solving the resulting multiobjective optimization problem for low-dimensional DNN. However, due to the non-smoothness of the $\ell^1$ norm and the high number of parameters, this approach is not very efficient from a computational perspective for high-dimensional DNNs. To overcome this limitation, we present an algorithm that allows for the approximation of the entire Pareto front for the above-mentioned objectives in a very efficient manner for high-dimensional DNNs with millions of parameters. We present numerical examples using both deterministic and stochastic gradients. We furthermore demonstrate that knowledge of the regularization path allows for a well-generalizing network parametrization. To the best of our knowledge, this is the first algorithm to compute the regularization path for non-convex multiobjective optimization problems (MOPs) with millions of degrees of freedom.

cs.LG

Fast Convergence of Inertial Multiobjective Gradient-like Systems with Asymptotic Vanishing Damping

We present a new gradient-like dynamical system related to unconstrained convex smooth multiobjective optimization which involves inertial effects and asymptotic vanishing damping. To the best of our knowledge, this system is the first inertial gradient-like system for multiobjective optimization problems including asymptotic vanishing damping, expanding the ideas laid out in [H. Attouch and G. Garrigos, Multiobjective optimization: an inertial approach to Pareto optima, preprint, arXiv:1506.02823, 201]. We prove existence of solutions to this system in finite dimensions and further prove that its bounded solutions converge weakly to weakly Pareto optimal points. In addition, we obtain a convergence rate of order $O(t^{-2})$ for the function values measured with a merit function. This approach presents a good basis for the development of fast gradient methods for multiobjective optimization.

math.OC

On the continuity and smoothness of the value function in reinforcement learning and optimal control

The value function plays a crucial role as a measure for the cumulative future reward an agent receives in both reinforcement learning and optimal control. It is therefore of interest to study how similar the values of neighboring states are, i.e., to investigate the continuity of the value function. We do so by providing and verifying upper bounds on the value function's modulus of continuity. Additionally, we show that the value function is always Hölder continuous under relatively weak assumptions on the underlying system and that non-differentiable value functions can be made differentiable by slightly "disturbing" the system.

eess.SY

A Descent Method for Nonsmooth Multiobjective Optimization in Hilbert Spaces

The efficient optimization method for locally Lipschitz continuous multiobjective optimization problems from [1] is extended from finite-dimensional problems to general Hilbert spaces. The method iteratively computes Pareto critical points, where in each iteration, an approximation of the subdifferential is computed in an efficient manner and then used to compute a common descent direction for all objective functions. To prove convergence, we present some new optimality results for nonsmooth multiobjective optimization problems in Hilbert spaces. Using these, we can show that every accumulation point of the sequence generated by our algorithm is Pareto critical under common assumptions. Computational efficiency for finding Pareto critical points is numerically demonstrated for multiobjective optimal control of an obstacle problem.

math.OC

Distributed Control of Partial Differential Equations Using Convolutional Reinforcement Learning

We present a convolutional framework which significantly reduces the complexity and thus, the computational effort for distributed reinforcement learning control of dynamical systems governed by partial differential equations (PDEs). Exploiting translational invariances, the high-dimensional distributed control problem can be transformed into a multi-agent control problem with many identical, uncoupled agents. Furthermore, using the fact that information is transported with finite velocity in many cases, the dimension of the agents' environment can be drastically reduced using a convolution operation over the state space of the PDE. In this setting, the complexity can be flexibly adjusted via the kernel width or by using a stride greater than one. Moreover, scaling from smaller to larger systems -- or the transfer between different domains -- becomes a straightforward task requiring little effort. We demonstrate the performance of the proposed framework using several PDE examples with increasing complexity, where stabilization is achieved by training a low-dimensional deep deterministic policy gradient agent using minimal computing resources.

cs.LG

Error bounds for kernel-based approximations of the Koopman operator

We consider the data-driven approximation of the Koopman operator for stochastic differential equations on reproducing kernel Hilbert spaces (RKHS). Our focus is on the estimation error if the data are collected from long-term ergodic simulations. We derive both an exact expression for the variance of the kernel cross-covariance operator, measured in the Hilbert-Schmidt norm, and probabilistic bounds for the finite-data estimation error. Moreover, we derive a bound on the prediction error of observables in the RKHS using a finite Mercer series expansion. Further, assuming Koopman-invariance of the RKHS, we provide bounds on the full approximation error. Numerical experiments using the Ornstein-Uhlenbeck process illustrate our results.

math.DS

Error analysis of kernel EDMD for prediction and control in the Koopman framework

Extended Dynamic Mode Decomposition (EDMD) is a popular data-driven method to approximate the Koopman operator for deterministic and stochastic (control) systems. This operator is linear and encompasses full information on the (expected stochastic) dynamics. In this paper, we analyze a kernel-based EDMD algorithm, known as kEDMD, where the dictionary consists of the canonical kernel features at the data points. The latter are acquired by i.i.d. samples from a user-defined and application-driven distribution on a compact set. We prove bounds on the prediction error of the kEDMD estimator when sampling from this (not necessarily ergodic) distribution. The error analysis is further extended to control-affine systems, where the considered invariance of the Reproducing Kernel Hilbert Space is significantly less restrictive in comparison to invariance assumptions on an a-priori chosen dictionary.

math.DS

Learning Bilinear Models of Actuated Koopman Generators from Partially-Observed Trajectories

Data-driven models for nonlinear dynamical systems based on approximating the underlying Koopman operator or generator have proven to be successful tools for forecasting, feature learning, state estimation, and control. It has become well known that the Koopman generators for control-affine systems also have affine dependence on the input, leading to convenient finite-dimensional bilinear approximations of the dynamics. Yet there are still two main obstacles that limit the scope of current approaches for approximating the Koopman generators of systems with actuation. First, the performance of existing methods depends heavily on the choice of basis functions over which the Koopman generator is to be approximated; and there is currently no universal way to choose them for systems that are not measure preserving. Secondly, if we do not observe the full state, then it becomes necessary to account for the dependence of the output time series on the sequence of supplied inputs when constructing observables to approximate Koopman operators. To address these issues, we write the dynamics of observables governed by the Koopman generator as a bilinear hidden Markov model, and determine the model parameters using the expectation-maximization (EM) algorithm. The E-step involves a standard Kalman filter and smoother, while the M-step resembles control-affine dynamic mode decomposition for the generator. We demonstrate the performance of this method on three examples, including recovery of a finite-dimensional Koopman-invariant subspace for an actuated system with a slow manifold; estimation of Koopman eigenfunctions for the unforced Duffing equation; and model-predictive control of a fluidic pinball system based only on noisy observations of lift and drag.

math.DS

Multiobjective Optimization of Non-Smooth PDE-Constrained Problems

Multiobjective optimization plays an increasingly important role in modern applications, where several criteria are often of equal importance. The task in multiobjective optimization and multiobjective optimal control is therefore to compute the set of optimal compromises (the Pareto set) between the conflicting objectives. The advances in algorithms and the increasing interest in Pareto-optimal solutions have led to a wide range of new applications related to optimal and feedback control - potentially with non-smoothness both on the level of the objectives or in the system dynamics. This results in new challenges such as dealing with expensive models (e.g., governed by partial differential equations (PDEs)) and developing dedicated algorithms handling the non-smoothness. Since in contrast to single-objective optimization, the Pareto set generally consists of an infinite number of solutions, the computational effort can quickly become challenging, which is particularly problematic when the objectives are costly to evaluate or when a solution has to be presented very quickly. This article gives an overview of recent developments in the field of multiobjective optimization of non-smooth PDE-constrained problems. In particular we report on the advances achieved within Project 2 "Multiobjective Optimization of Non-Smooth PDE-Constrained Problems - Switches, State Constraints and Model Order Reduction" of the DFG Priority Programm 1962 "Non-smooth and Complementarity-based Distributed Parameter Systems: Simulation and Hierarchical Optimization".

math.OC

Multi-Objective Trust-Region Filter Method for Nonlinear Constraints using Inexact Gradients

In this article, we build on previous work to present an optimization algorithm for nonlinearly constrained multi-objective optimization problems. The algorithm combines a surrogate-assisted derivative-free trust-region approach with the filter method known from single-objective optimization. Instead of the true objective and constraint functions, so-called fully linear models are employed, and we show how to deal with the gradient inexactness in the composite step setting, adapted from single-objective optimization as well. Under standard assumptions, we prove convergence of a subset of iterates to a quasi-stationary point and if constraint qualifications hold, then the limit point is also a KKT-point of the multi-objective problem.

math.OC

Learning a model is paramount for sample efficiency in reinforcement learning control of PDEs

The goal of this paper is to make a strong point for the usage of dynamical models when using reinforcement learning (RL) for feedback control of dynamical systems governed by partial differential equations (PDEs). To breach the gap between the immense promises we see in RL and the applicability in complex engineering systems, the main challenges are the massive requirements in terms of the training data, as well as the lack of performance guarantees. We present a solution for the first issue using a data-driven surrogate model in the form of a convolutional LSTM with actuation. We demonstrate that learning an actuated model in parallel to training the RL agent significantly reduces the total amount of required data sampled from the real system. Furthermore, we show that iteratively updating the model is of major importance to avoid biases in the RL training. Detailed ablation studies reveal the most important ingredients of the modeling process. We use the chaotic Kuramoto-Sivashinsky equation do demonstarte our findings.

cs.LG

Towards reliable data-based optimal and predictive control using extended DMD

While Koopman-based techniques like extended Dynamic Mode Decomposition are nowadays ubiquitous in the data-driven approximation of dynamical systems, quantitative error estimates were only recently established. To this end, both sources of error resulting from a finite dictionary and only finitely-many data points in the generation of the surrogate model have to be taken into account. We generalize the rigorous analysis of the approximation error to the control setting while simultaneously reducing the impact of the curse of dimensionality by using a recently proposed bilinear approach. In particular, we establish uniform bounds on the approximation error of state-dependent quantities like constraints or a performance index enabling data-based optimal and predictive control with guarantees.

math.OC

On the Universal Transformation of Data-Driven Models to Control Systems

The advances in data science and machine learning have resulted in significant improvements regarding the modeling and simulation of nonlinear dynamical systems. It is nowadays possible to make accurate predictions of complex systems such as the weather, disease models or the stock market. Predictive methods are often advertised to be useful for control, but the specifics are frequently left unanswered due to the higher system complexity, the requirement of larger data sets and an increased modeling effort. In other words, surrogate modeling for autonomous systems is much easier than for control systems. In this paper we present the framework QuaSiModO (Quantization-Simulation-Modeling-Optimization) to transform arbitrary predictive models into control systems and thus render the tremendous advances in data-driven surrogate modeling accessible for control. Our main contribution is that we trade control efficiency by autonomizing the dynamics - which yields mixed-integer control problems - to gain access to arbitrary, ready-to-use autonomous surrogate modeling techniques. We then recover the complexity of the original problem by leveraging recent results from mixed-integer optimization. The advantages of QuaSiModO are a linear increase in data requirements with respect to the control dimension, performance guarantees that rely exclusively on the accuracy of the predictive model in use, and little prior knowledge requirements in control theory to solve complex control problems.

math.OC