SearcharxivSearch

arXiv subjects

Tizian Wenzel

Publications and source records attributed to Tizian Wenzel.

At least 19 recordsLinked to original sources

Reliable sampling-based RKHS norm estimation via superconvergence

Kernel methods are one of the cornerstones of learning-based control, modern system identification, surrogate modelling, and related fields. A key advantage of this class of learning and function approximation methods is the availability of quantitative error bounds, which in turn play a key role in guaranteeing the safety of learned controllers and related learning-based algorithms. However, these error bounds rely on a particular property of the target function -- its reproducing kernel Hilbert space (RKHS) norm -- which is usually impossible to obtain in practice. Motivated by this severe shortcoming, we present a novel sampling-based RKHS norm estimation approach with a solid theoretical foundation, leveraging very recent advances in the theory of superconvergence in kernel methods. Our method is applicable to a broad range of practically relevant function classes and requires only reasonable prior knowledge about the target function. Extensive numerical experiments demonstrate the efficacy and practical applicability of the proposed method. By providing a reliable RKHS norm estimation approach, we remove a major obstacle to the practical deployment of learning-based control algorithms.

math.NA

ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules

Tabular foundation models such as TabPFN and TabICL already produce full predictive distributions, yet prevailing regression benchmarks evaluate them almost exclusively via point-estimate metrics (RMSE, $R^2$). This discards precisely the distributional information these models are designed to provide - a critical gap for high-stakes domains where not all kinds of errors are equally costly. We introduce ScoringBench, an open and extensible benchmark that evaluates tabular regression models under a comprehensive suite of proper scoring rules - including CRPS, CRLS, interval score, energy score, and weighted CRPS - alongside standard point metrics. ScoringBench covers 97 regression datasets from diverse domains, supports transparent community contributions via a git-based leaderboard, and provides two complementary ranking protocols: an ordinal Demsar/autorank approach and a magnitude-preserving z-score ranking approach. Evaluating several models - spanning in-context learners, fine-tuned foundation models, gradient-boosted trees, and MLPs - we find that model rankings shift substantially depending on the scoring rule: models that excel on point-estimate metrics can rank poorly on probabilistic ones, and the top-performing model under one proper scoring rule may rank noticeably lower under another. These results demonstrate that the choice of evaluation metric is not a technicality but a modelling decision - and, for applications where e.g. tail errors are disproportionately costly, a domain-specific requirement with direct consequences for model deployment.

cs.AI

Distributional Regression with Tabular Foundation Models: Evaluating Probabilistic Predictions via Proper Scoring Rules

Modern tabular foundation models such as TabPFN and TabICL naturally produce full predictive distributions, while the benchmarks used to evaluate them (TabArena, TALENT, and others) still rely almost exclusively on point-estimate metrics (RMSE, $R^2$). This mismatch implicitly rewards machine learning models or pipelines that elicit a good conditional mean while ignoring the quality of the predictive distribution. We make the case for using proper scoring rules for training, fine-tuning, and benchmarking (ranking) of tabular foundation models. Although all strictly proper scoring rules are theoretically equivalent at the population level, they may differ on finite data: We demonstrate analytically and empirically that different scoring rules can induce different inductive biases during finite-sample optimization, leading to different model performance. We validate this finding by running fine-tuning experiments with TabPFN and TabICL using different scoring rules for various data sets, revealing non-trivial interactions between training objectives and evaluation metrics. Our results show that practitioners can adapt tabular foundation models to task-specific scoring objectives, and that the choice of scoring rule can influence model behavior in practice.

cs.LG

Piecewise linear interpolation via kernels

We consider piecewise linear interpolation from the perspective of kernel interpolation and quadrature. If the Sobolev space $W_2^1(0, 1)$ is equipped with a suitable inner product, its reproducing kernel is piecewise linear and gives rise to piecewise linear interpolation. We show that such kernels are Green kernels for certain second-order partial differential equations and use kernel-based superconvergence theory to obtain rates of convergence for approximation of functions lying in $W_2^s(0, 1)$ for $s \in [1, 2]$. The rates coincide with classical rates for linear splines.

math.NA

Refined rates of convergence for target-data dependent greedy generalized interpolation with Sobolev kernels

Greedy methods have recently been successfully applied to generalized kernel interpolation, or the recovery of a function from data stemming from the evaluation of linear functionals, including the approximation of solutions of linear PDEs by symmetric collocation. When applied to kernels generating Sobolev spaces as their native Hilbert spaces, some of these greedy methods can provide the same error guarantee of generalized interpolation on quasi-uniform points. More importantly, certain target-data-adaptive methods even give a dimension- and smoothness-independent improvement in the speed of convergence over quasi-uniform points, thus offering advantages for high-dimensional problems. These convergence rates however contain a spurious logarithmic term that limits this beneficial effect. The goal of this note is to remove this factor, and this is possible by using estimates on metric entropy numbers.

math.NA

On the optimal shape parameter for kernel methods: Sharp direct and inverse statements

The search for the optimal shape parameter for Radial Basis Function (RBF) kernel approximation has been an outstanding research problem for decades. In this work, we establish a theoretical framework for this problem by leveraging a recently established theory on sharp direct, inverse and saturation statements for kernel based approximation. In particular, we link the search for the optimal shape parameter to superconvergence phenomena. Our analysis is carried out for finitely smooth Sobolev kernels, thereby covering large classes of radial kernels used in practice, including those emerging from current machine-learning methodologies. Our results elucidate how approximation regimes, kernel regularity, and parameter choices interact, thereby clarifying a question that has remained unresolved for decades.

math.NA

Sharp inverse statements for kernel approximation: Superconvergence and saturation

This article establishes sharp inverse and saturation statements for kernel-based approximation using finitely smooth Sobolev kernels on bounded Lipschitz regions. The analysis focuses on the superconvergence regime, for which direct statements have only recently been obtained. The resulting theory yields a one-to-one correspondence between the smoothness of a target function - quantified in terms of power spaces - and the achievable approximation rates by kernel-based approximation. In this way, we extend existing results beyond the escaping-the-native-space regime and provide a unified characterization covering the full scale of admissible smoothness spaces.

math.NA

Sobolev Algorithm for Local Smoothness Analysis (SALSA) via Sharp Direct and Inverse Statements

We extend sharp direct and inverse approximation statements for kernel-based methods for finitely smooth kernels, i.e. those whose native spaces are norm-equivalent to Sobolev spaces. In particular, our inverse results are now formulated for a broad class of approximation schemes beyond interpolation, extending existing theory. Building on these results, we propose a novel Sobolev Algorithm for Local Smoothness Analysis (SALSA) for detecting local smoothness properties of target data, including their degree of smoothness and non-smoothness. The method is rigorously grounded based on the sharp direct and inverse statements. Numerical experiments in various settings highlight the effectiveness of the proposed algorithm.

math.NA

Spectral equivalence of unsymmetric kernel matrices and applications

Symmetric kernel matrices are a well-researched topic in the literature of kernel based approximation. In particular stability properties in terms of lower bounds on the smallest eigenvalue of such symmetric kernel matrices are thoroughly investigated, as they play a fundamental role in theory and practice. In this work, we focus on unsymmetric kernel matrices and derive stability properties under small shifts by establishing a spectral equivalence to their unshifted, symmetric versions. This extends and generalizes results for translational invariant kernels upon Quak et al [SIAM Journal on Math. Analysis, 1993] and Sivakumar and Ward [Numerische Mathematik, 1993], however focussing instead on finitely smooth kernels. As applications, we consider convolutional kernels over domains, which are no longer translational invariant, but which are still an important class of kernels for applications. For these, we derive novel lower bounds for the smallest eigenvalue of the kernel matrices in terms of the separation distance of the data points, and thus derive stability bounds in terms of the condition number.

math.NA

Kernel-based Greedy Approximation of Parametric Elliptic Boundary Value Problems

We recently introduced a scale of kernel-based greedy schemes for approximating the solutions of elliptic boundary value problems. The procedure is based on a generalized interpolation framework in reproducing kernel Hilbert spaces and was coined PDE-$\beta$-greedy procedure, where the parameter $\beta \geq 0$ is used in a greedy selection criterion and steers the degree of function adaptivity. Algebraic convergence rates have been obtained for Sobolev-space kernels and solutions of finite smoothness. We now report a result of exponential convergence rates for the case of infinitely smooth kernels and solutions. We furthermore extend the approximation scheme to the case of parametric PDEs by the use of state-parameter product kernels. In the surrogate modelling context, the resulting approach can be interpreted as an a priori model reduction approach, as no solution snapshots need to be precomputed. Numerical results show the efficiency of the approximation procedure for problems which occur as challenges for other parametric MOR procedures: non-affine geometry parametrizations, moving sources or high-dimensional domains.

math.NA

General superconvergence for kernel-based approximation

Kernel interpolation is a fundamental technique for approximating functions from scattered data, with a well-understood convergence theory when interpolating elements of a reproducing kernel Hilbert space. Beyond this classical setting, research has focused on two regimes: misspecified interpolation, where the kernel smoothness exceeds that of the target function, and superconvergence, where the target is smoother than the Hilbert space. This work addresses the latter, where smoother target functions yield improved convergence rates, and extends existing results by characterizing superconvergence for projections in general Hilbert spaces. We show that functions lying in ranges of certain operators, including adjoint of embeddings, exhibit accelerated convergence, which we extend across interpolation scales between these ranges and the full Hilbert space. In particular, we analyze Mercer operators and embeddings into $L_p$ spaces, linking the images of adjoint operators to Mercer power spaces. Applications to Sobolev spaces are discussed in detail, highlighting how superconvergence depends critically on boundary conditions. Our findings generalize and refine previous results, offering a broader framework for understanding and exploiting superconvergence. The results are supported by numerical experiments.

math.NA

Sharp inverse statements for kernel interpolation

While direct statements for kernel based interpolation on regions $Ω\subset \mathbb{R}^d$ are well researched, far less is known about corresponding inverse statements. The available inverse statements for kernel based interpolation so far are not sharp. In this paper, we derive sharp inverse statements for interpolation using finitely smooth kernels, such as popular radial basis function (RBF) kernels like the class of Matérn or Wendland kernels. In particular, the results show that there is a one-to-one correspondence between the smoothness of a function and its approximation rate via kernel interpolation: If a function can be approximated with a given rate, it has a corresponding smoothness and vice versa.

math.NA

Analysis of Structured Deep Kernel Networks

In this paper, we leverage a recent deep kernel representer theorem to connect kernel based learning and (deep) neural networks in order to understand their interplay. In particular, we show that the use of special types of kernels yields models reminiscent of neural networks that are founded in the same theoretical framework of classical kernel methods, while benefiting from the computational advantages of deep neural networks. Especially the introduced Structured Deep Kernel Networks (SDKNs) can be viewed as neural networks (NNs) with optimizable activation functions obeying a representer theorem. This link allows us to analyze also NNs within the framework of kernel networks. We prove analytic properties of the SDKNs which show their universal approximation properties in three different asymptotic regimes of unbounded number of centers, width and depth. Especially in the case of unbounded depth, more accurate constructions can be achieved using fewer layers compared to corresponding constructions for ReLU neural networks. This is made possible by leveraging properties of kernel approximation.

cs.LG

Adaptive meshfree approximation for linear elliptic partial differential equations with PDE-greedy kernel methods

We consider meshless approximation for solutions of boundary value problems (BVPs) of elliptic Partial Differential Equations (PDEs) via symmetric kernel collocation. We discuss the importance of the choice of the collocation points, in particular by using greedy kernel methods. We introduce a scale of PDE-greedy selection criteria that generalizes existing techniques, such as the PDE-$P$-greedy and the PDE-$f$-greedy rules for collocation point selection. For these greedy selection criteria we provide bounds on the approximation error in terms of the number of greedily selected points and analyze the corresponding convergence rates. This is achieved by a novel analysis of Kolmogorov widths of special sets of BVP point-evaluation functionals. Especially, we prove that target-data dependent algorithms that make use of the right hand side functions of the BVP exhibit faster convergence rates than the target-data independent PDE-$P$-greedy. The convergence rate of the PDE-$f$-greedy possesses a dimension independent rate, which makes it amenable to mitigate the curse of dimensionality. The advantages of these greedy algorithms are highlighted by numerical examples.

math.NA

Kernel Methods in the Deep Ritz framework: Theory and practice

In this contribution, kernel approximations are applied as ansatz functions within the Deep Ritz method. This allows to approximate weak solutions of elliptic partial differential equations with weak enforcement of boundary conditions using Nitsche's method. A priori error estimates are proven in different norms leveraging both standard results for weak solutions of elliptic equations and well-established convergence results for kernel methods. This availability of a priori error estimates renders the method useful for practical purposes. The procedure is described in detail, meanwhile providing practical hints and implementation details. By means of numerical examples, the performance of the proposed approach is evaluated numerically and the results agree with the theoretical findings.

math.NA

Greedy Kernel Methods for Approximating Breakthrough Curves for Reactive Flow from 3D Porous Geometry Data

We address the challenging application of 3D pore scale reactive flow under varying geometry parameters. The task is to predict time-dependent integral quantities, i.e., breakthrough curves, from the given geometries. As the 3D reactive flow simulation is highly complex and computationally expensive, we are interested in data-based surrogates that can give a rapid prediction of the target quantities of interest. This setting is an example of an application with scarce data, i.e., only having available few data samples, while the input and output dimensions are high. In this scarce data setting, standard machine learning methods are likely to ail. Therefore, we resort to greedy kernel approximation schemes that have shown to be efficient meshless approximation techniques for multivariate functions. We demonstrate that such methods can efficiently be used in the high-dimensional input/output case under scarce data. Especially, we show that the vectorial kernel orthogonal greedy approximation (VKOGA) procedure with a data-adapted two-layer kernel yields excellent predictors for learning from 3D geometry voxel data via both morphological descriptors or principal component analysis.

math.NA

Finetuning greedy kernel models by exchange algorithms

Kernel based approximation offers versatile tools for high-dimensional approximation, which can especially be leveraged for surrogate modeling. For this purpose, both "knot insertion" and "knot removal" approaches aim at choosing a suitable subset of the data, in order to obtain a sparse but nevertheless accurate kernel model. In the present work, focussing on kernel based interpolation, we aim at combining these two approaches to further improve the accuracy of kernel models, without increasing the computational complexity of the final kernel model. For this, we introduce a class of kernel exchange algorithms (KEA). The resulting KEA algorithm can be used for finetuning greedy kernel surrogate models, allowing for an reduction of the error up to 86.4% (17.2% on average) in our experiments.

cs.LG

A comparison study of supervised learning techniques for the approximation of high dimensional functions and feedback control

Approximation of high dimensional functions is in the focus of machine learning and data-based scientific computing. In many applications, empirical risk minimisation techniques over nonlinear model classes are employed. Neural networks, kernel methods and tensor decomposition techniques are among the most popular model classes. We provide a numerical study comparing the performance of these methods on various high-dimensional functions with focus on optimal control problems, where the collection of the dataset is based on the application of the State-Dependent Riccati Equation.

math.NA