SearcharxivSearch

arXiv subjects

Christian Fiedler

Publications and source records attributed to Christian Fiedler.

At least 19 recordsLinked to original sources

Consensus-based optimization for linearly separable functions

Consensus-based optimization (CBO) is an efficient metaheuristic for global optimisation with attractive mathematical properties, allowing global convergence results even in non-convex settings. In practice it suffers greatly from the curse of dimensionality, as do most particle-based optimisers. Different strategies have been proposed to apply CBO even for high-dimensional optimisation problems, the most prominent being the so-called anisotropic noise model. However, a recent work by Bonandin et al. highlighted that this method performs well primarily on separable objective functions, and even small coordinate rotations severely degrade the performance. Motivated by this observation, in this work we study the case where the objective function is separable only in a linearly transformed coordinate system. To leverage the efficiency of anisotropic CBO we propose a reparametrisation scheme, which provably estimates such a coordinate transformation and then applies CBO in the new variables. Furthermore, we show how our scheme can be interpreted as noise adaptation in CBO. Numerical examples highlight the efficacy of the method, demonstrating significant performance improvements on challenging benchmark objectives.

math.OC

Reliable sampling-based RKHS norm estimation via superconvergence

Kernel methods are one of the cornerstones of learning-based control, modern system identification, surrogate modelling, and related fields. A key advantage of this class of learning and function approximation methods is the availability of quantitative error bounds, which in turn play a key role in guaranteeing the safety of learned controllers and related learning-based algorithms. However, these error bounds rely on a particular property of the target function -- its reproducing kernel Hilbert space (RKHS) norm -- which is usually impossible to obtain in practice. Motivated by this severe shortcoming, we present a novel sampling-based RKHS norm estimation approach with a solid theoretical foundation, leveraging very recent advances in the theory of superconvergence in kernel methods. Our method is applicable to a broad range of practically relevant function classes and requires only reasonable prior knowledge about the target function. Extensive numerical experiments demonstrate the efficacy and practical applicability of the proposed method. By providing a reliable RKHS norm estimation approach, we remove a major obstacle to the practical deployment of learning-based control algorithms.

math.NA

The Mean-Field Limit of Online Stochastic Vector Balancing

We study an online vector balancing problem, in which $n$ independent Gaussian random vectors $\boldsymbol{\zeta}(1),\dots,\boldsymbol{\zeta}(n) \sim \mathcal{N}(0, I_n)$, each of dimension $n$, arrive one at a time. The goal is to choose signs $\varepsilon(1),\dots,\varepsilon(n) \in \{\pm 1\}$ with $\varepsilon(k)$ depending only on $\boldsymbol{\zeta}(1),\dots,\boldsymbol{\zeta}(k)$, so as to minimize the expected $\ell^{\infty}$ norm of the signed sum $\frac{1}{\sqrt{n}}\sum_{k = 1}^n \varepsilon(k) \boldsymbol{\zeta}(k)$. Prior work showed that the optimal value $V^n$ is $O(1)$, at least for Rademacher $\boldsymbol{\zeta}(k)$'s, by constructing specific algorithms. Our main contribution is to determine the exact limit $V^{\infty} = \lim_{n\to\infty} V^n$ as the value of a nonstandard stochastic control problem of mean-field type: find the narrowest terminal interval into which a Brownian motion can be adaptively steered under a uniform-in-time $L^2$ constraint on the drift. The proof of the lower bound $V^{\infty} \leq \liminf_{n \to \infty} V^n$ uses probabilistic compactness arguments, and is very flexible. In fact, we show that the lower bound is universal, in that it holds as long as the entries of the $\boldsymbol{\zeta}(k)$ vectors are i.i.d. with mean zero, variance 1, and finite fourth moment. The proof of the upper bound $\limsup_{n \to \infty} V^n \leq V^{\infty}$ is more delicate, relying on dynamic programming principles and a priori bounds obtained from a coupling procedure involving the F\"ollmer drift, which makes explicit use of the Gaussian structure. In addition to our main convergence result, we provide some analysis and asymptotics for the limiting mean-field control problem.

math.PR

Statistical Learning Theory for Distributional Classification

In supervised learning with distributional inputs in the two-stage sampling setup, relevant to applications like learning-based medical screening or causal learning, the inputs (which are probability distributions) are not accessible in the learning phase, but only samples thereof. This problem is particularly amenable to kernel-based learning methods, where the distributions or samples are first embedded into a Hilbert space, often using kernel mean embeddings (KMEs), and then a standard kernel method like Support Vector Machines (SVMs) is applied, using a kernel defined on the embedding Hilbert space. In this work, we contribute to the theoretical analysis of this latter approach, with a particular focus on classification with distributional inputs using SVMs. We establish a new oracle inequality and derive consistency and learning rate results. Furthermore, for SVMs using the hinge loss and Gaussian kernels, we formulate a novel variant of an established noise assumption from the binary classification literature, under which we can establish learning rates. Finally, some of our technical tools like a new feature space for Gaussian kernels on Hilbert spaces are of independent interest.

cs.LG

Fuk-Nagaev inequality in smooth Banach spaces: Optimum bounds for distributions of heavy-tailed martingales

We derive a Fuk-Nagaev inequality for the maxima of norms of martingale sequences in smooth Banach spaces which allow for a finite number of higher conditional moments. The bound is obtained by combining an optimization approach for a Chernoff bound due to Rio (2017) with a classical bound for moment generating functions of smooth Banach space norms by Pinelis (1994). Our result improves comparable infinite-dimensional bounds in the literature by removing unnecessary centering terms and giving precise constants. As an application, we propose a McDiarmid-type bound for vector-valued functions which allow for a uniform bound on their conditional higher moments.

math.PR

Towards optimal control of ensembles of discrete-time systems

The control of ensembles of dynamical systems is an intriguing and challenging problem, arising for example in quantum control. We initiate the investigation of optimal control of ensembles of discrete-time systems, focusing on minimising the average finite horizon cost over the ensemble. For very general nonlinear control systems and stage and terminal costs, we establish existence of minimisers under mild assumptions. Furthermore, we provide a $\Gamma$-convergence result which enables consistent approximation of the challenging ensemble optimal control problem, for example, by using empirical probability measures over the ensemble. Our results form a solid foundation for discrete-time optimal control of ensembles, with many interesting avenues for future research.

math.OC

Kernel conditional tests from learning-theoretic bounds

We propose a framework for hypothesis testing on conditional probability distributions, which we then use to construct statistical tests of functionals of conditional distributions. These tests identify the inputs where the functionals differ with high probability, and include tests of conditional moments or two-sample tests. Our key idea is to transform confidence bounds of a learning method into a test of conditional expectations. We instantiate this principle for kernel ridge regression (KRR) with subgaussian noise. An intermediate data embedding then enables more general tests -- including conditional two-sample tests -- via kernel mean embeddings of distributions. To have guarantees in this setting, we generalize existing pointwise-in-time or time-uniform confidence bounds for KRR to previously-inaccessible yet essential cases such as infinite-dimensional outputs with non-trace-class kernels. These bounds also circumvent the need for independent data, allowing for instance online sampling. To make our tests readily applicable in practice, we introduce bootstrapping schemes leveraging the parametric form of testing thresholds identified in theory to avoid tuning inaccessible parameters. We illustrate the tests on examples, including one in process monitoring and comparison of dynamical systems. Overall, our results establish a comprehensive foundation for conditional testing on functionals, from theoretical guarantees to an algorithmic implementation, and advance the state of the art on confidence bounds for vector-valued least squares estimation.

cs.LG

Safety in safe Bayesian optimization and its ramifications for control

A recurring and important task in control engineering is parameter tuning under constraints, which conceptually amounts to optimization of a blackbox function accessible only through noisy evaluations. For example, in control practice parameters of a pre-designed controller are often tuned online in feedback with a plant, and only safe parameter values should be tried, avoiding for example instability. Recently, machine learning methods have been deployed for this important problem, in particular, Bayesian optimization (BO). To handle safety constraints, algorithms from safe BO have been utilized, especially SafeOpt-type algorithms, which enjoy considerable popularity in learning-based control, robotics, and adjacent fields. However, we identify two significant obstacles to practical safety. First, SafeOpt-type algorithms rely on quantitative uncertainty bounds, and most implementations replace these by theoretically unsupported heuristics. Second, the theoretically valid uncertainty bounds crucially depend on a quantity - the reproducing kernel Hilbert space norm of the target function - that at present is impossible to reliably bound using established prior engineering knowledge. By careful numerical experiments we show that these issues can indeed cause safety violations. To overcome these problems, we propose Lipschitz-only Safe Bayesian Optimization (LoSBO), a safe BO algorithm that relies only on a known Lipschitz bound for its safety. Furthermore, we propose a variant (LoS-GP-UCB) that avoids gridding of the search space and is therefore applicable even for moderately high-dimensional problems.

eess.SY

Automatic nonlinear MPC approximation with closed-loop guarantees

Safety guarantees are vital in many control applications, such as robotics. Model predictive control (MPC) provides a constructive framework for controlling safety-critical systems, but is limited by its computational complexity. We address this problem by presenting a novel algorithm that automatically computes an explicit approximation to nonlinear MPC schemes while retaining closed-loop guarantees. Specifically, the problem can be reduced to a function approximation problem, which we then tackle by proposing ALKIA-X, the Adaptive and Localized Kernel Interpolation Algorithm with eXtrapolated reproducing kernel Hilbert space norm. ALKIA-X is a non-iterative algorithm that ensures numerically well-conditioned computations, a fast-to-evaluate approximating function, and the guaranteed satisfaction of any desired bound on the approximation error. Hence, ALKIA-X automatically computes an explicit function that approximates the MPC, yielding a controller suitable for safety-critical systems and high sampling rates. We apply ALKIA-X to approximate two nonlinear MPC schemes, demonstrating reduced computational demand and applicability to realistic problems.

eess.SY

On Safety in Safe Bayesian Optimization

Optimizing an unknown function under safety constraints is a central task in robotics, biomedical engineering, and many other disciplines, and increasingly safe Bayesian Optimization (BO) is used for this. Due to the safety critical nature of these applications, it is of utmost importance that theoretical safety guarantees for these algorithms translate into the real world. In this work, we investigate three safety-related issues of the popular class of SafeOpt-type algorithms. First, these algorithms critically rely on frequentist uncertainty bounds for Gaussian Process (GP) regression, but concrete implementations typically utilize heuristics that invalidate all safety guarantees. We provide a detailed analysis of this problem and introduce Real-\b{eta}-SafeOpt, a variant of the SafeOpt algorithm that leverages recent GP bounds and thus retains all theoretical guarantees. Second, we identify assuming an upper bound on the reproducing kernel Hilbert space (RKHS) norm of the target function, a key technical assumption in SafeOpt-like algorithms, as a central obstacle to real-world usage. To overcome this challenge, we introduce the Lipschitz-only Safe Bayesian Optimization (LoSBO) algorithm, which guarantees safety without an assumption on the RKHS bound, and empirically show that this algorithm is not only safe, but also exhibits superior performance compared to the state-of-the-art on several function classes. Third, SafeOpt and derived algorithms rely on a discrete search space, making them difficult to apply to higher-dimensional problems. To widen the applicability of these algorithms, we introduce Lipschitz-only GP-UCB (LoS-GP-UCB), a variant of LoSBO applicable to moderately high-dimensional problems, while retaining safety.

cs.LG

Mean field limits for discrete-time dynamical systems via kernel mean embeddings

Mean field limits are an important tool in the context of large-scale dynamical systems, in particular, when studying multiagent and interacting particle systems. While the continuous-time theory is well-developed, few works have considered mean field limits for deterministic discrete-time systems, which are relevant for the analysis and control of large-scale discrete-time multiagent system. We prove existence results for the mean field limit of very general discrete-time control systems, for which we utilize kernel mean embeddings. These results are then applied in a typical optimal control setup, where we establish the mean field limit of the relaxed dynamic programming principle. Our results can serve as a rigorous foundation for many applications of mean field approaches for discrete-time dynamical systems.

eess.SY

On kernel-based statistical learning in the mean field limit

In many applications of machine learning, a large number of variables are considered. Motivated by machine learning of interacting particle systems, we consider the situation when the number of input variables goes to infinity. First, we continue the recent investigation of the mean field limit of kernels and their reproducing kernel Hilbert spaces, completing the existing theory. Next, we provide results relevant for approximation with such kernels in the mean field limit, including a representer theorem. Finally, we use these kernels in the context of statistical learning in the mean field limit, focusing on Support Vector Machines. In particular, we show mean field convergence of empirical and infinite-sample solutions as well as the convergence of the corresponding risks. On the one hand, our results establish rigorous mean field limits in the context of kernel methods, providing new theoretical tools and insights for large-scale problems. On the other hand, our setting corresponds to a new form of limit of learning problems, which seems to have not been investigated yet in the statistical learning theory literature.

cs.LG

Lipschitz and Hölder Continuity in Reproducing Kernel Hilbert Spaces

Reproducing kernel Hilbert spaces (RKHSs) are very important function spaces, playing an important role in machine learning, statistics, numerical analysis and pure mathematics. Since Lipschitz and Hölder continuity are important regularity properties, with many applications in interpolation, approximation and optimization problems, in this work we investigate these continuity notion in RKHSs. We provide several sufficient conditions as well as an in depth investigation of reproducing kernels inducing prescribed Lipschitz or Hölder continuity. Apart from new results, we also collect related known results from the literature, making the present work also a convenient reference on this topic.

math.FA

Practical and Rigorous Uncertainty Bounds for Gaussian Process Regression

Gaussian Process Regression is a popular nonparametric regression method based on Bayesian principles that provides uncertainty estimates for its predictions. However, these estimates are of a Bayesian nature, whereas for some important applications, like learning-based control with safety guarantees, frequentist uncertainty bounds are required. Although such rigorous bounds are available for Gaussian Processes, they are too conservative to be useful in applications. This often leads practitioners to replacing these bounds by heuristics, thus breaking all theoretical guarantees. To address this problem, we introduce new uncertainty bounds that are rigorous, yet practically useful at the same time. In particular, the bounds can be explicitly evaluated and are much less conservative than state of the art results. Furthermore, we show that certain model misspecifications lead to only graceful degradation. We demonstrate these advantages and the usefulness of our results for learning-based control with numerical examples.

cs.LG

Reproducing kernel Hilbert spaces in the mean field limit

Kernel methods, being supported by a well-developed theory and coming with efficient algorithms, are among the most popular and successful machine learning techniques. From a mathematical point of view, these methods rest on the concept of kernels and function spaces generated by kernels, so called reproducing kernel Hilbert spaces. Motivated by recent developments of learning approaches in the context of interacting particle systems, we investigate kernel methods acting on data with many measurement variables. We show the rigorous mean field limit of kernels and provide a detailed analysis of the limiting reproducing kernel Hilbert space. Furthermore, several examples of kernels, that allow a rigorous mean field limit, are presented.

stat.ML

A Kernel Two-sample Test for Dynamical Systems

Evaluating whether data streams are drawn from the same distribution is at the heart of various machine learning problems. This is particularly relevant for data generated by dynamical systems since such systems are essential for many real-world processes in biomedical, economic, or engineering systems. While kernel two-sample tests are powerful for comparing independent and identically distributed random variables, no established method exists for comparing dynamical systems. The main problem is the inherently violated independence assumption. We propose a two-sample test for dynamical systems by addressing three core challenges: we (i) introduce a novel notion of mixing that captures autocorrelations in a relevant metric, (ii) propose an efficient way to estimate the speed of mixing relying purely on data, and (iii) integrate these into established kernel two-sample tests. The result is a data-driven method that is straightforward to use in practice and comes with sound theoretical guarantees. In an example application to anomaly detection from human walking data, we show that the test is readily applicable without any human expert knowledge and feature engineering.

stat.ML

Revisiting the derivation of stage costs in infinite horizon discrete-time optimal control

In many applications of optimal control, the stage cost is not fixed, but rather a design choice with considerable impact on the control performance. In infinite horizon optimal control, the choice of stage cost is often restricted by the requirement of uniform cost controllability, which is nontrivial to satisfy. Here we revisit a previously proposed constructive technique for stage cost design. We generalize its setting, weaken the required assumptions and add additional flexibility. Furthermore, we show that the required assumptions essentially cannot be weakened anymore. By providing improved design options for stage costs, this work contributes to expanding the applicability of optimization-based control methodologies, in particular, model predictive control.

math.OC

Learning-enhanced robust controller synthesis with rigorous statistical and control-theoretic guarantees

The combination of machine learning with control offers many opportunities, in particular for robust control. However, due to strong safety and reliability requirements in many real-world applications, providing rigorous statistical and control-theoretic guarantees is of utmost importance, yet difficult to achieve for learning-based control schemes. We present a general framework for learning-enhanced robust control that allows for systematic integration of prior engineering knowledge, is fully compatible with modern robust control and still comes with rigorous and practically meaningful guarantees. Building on the established Linear Fractional Representation and Integral Quadratic Constraints framework, we integrate Gaussian Process Regression as a learning component and state-of-the-art robust controller synthesis. In a concrete robust control example, our approach is demonstrated to yield improved performance with more data, while guarantees are maintained throughout.

eess.SY