SearcharxivSearch

arXiv subjects

Fabio Nobile

Publications and source records attributed to Fabio Nobile.

At least 19 recordsLinked to original sources

On the sample complexity of the active subspace method

Active subspaces identify low-dimensional linear structure in high-dimensional parameter-to-output maps by estimating the dominant eigenspace of a gradient covariance operator. In practice this covariance is replaced by a Monte Carlo estimator built from a limited number of gradient evaluations. Classical analyses based on controlling the covariance error in operator norm lead to sample-complexity estimates that can be substantially more pessimistic than the sampling rules commonly used in computations. This paper studies the empirical active subspace method directly in the projection-error metric relevant for ridge approximation. We derive non-asymptotic quasi-optimality bounds governed by a regularized inverse Christoffel function associated with the gradient field. Under a bounded-gradient assumption, the resulting estimates already improve the sample-complexity estimates obtained from operator-norm covariance bounds. We then show that additional smoothness of the gradient map, expressed through membership in a reproducing kernel Hilbert space, yields sharper coherence estimates and motivates tractable importance sampling from kernel diagonal measures. Furthermore, the same smoothness assumption yields a priori decay bounds for the population active subspace tail energy, which can be combined with our finite-sample estimate to prescribe rank, regularization scale, and sample size, allowing to fully characterize the a priori sample complexity. The abstract assumptions are verified for lognormal Gaussian and affine uniform parametric elliptic PDEs using weighted summability of Hermite and Legendre series expansions.

math.NA

Dynamical Low-Rank Filters for Data Assimilation

We propose dynamical low-rank (DLR) type filters for data-assimilation problems based on stochastic differential equations (SDEs). In detail, first we derive a DLRA filter for minimizing jointly the mean and covariance error, as well as a strategy to efficiently include the relevant orthogonal directions. This last approach allows the main subspace to evolve also according to the observation operator. Those procedures naturally extend to a Kalman-Bucy type filter when dealing with linear drift, and to ensemble methods, too, resulting also suitable for problems described by nonlinear drift and possible non-Gaussian distribution. Moreover, we further propose a preliminary particle-type DLRA filter that shows potentiality in nonlinear settings. Numerical simulations show the efficacy of these procedures in relevant applications, opening up to further studies in these filtering directions.

math.NA

Dynamical Low-Rank Smoothing

Computational costs often make smoothing procedures prohibitive for high-dimensional data assimilation problems. To address this challenge, we propose a dynamical low-rank approximation (DLRA) methodology for smoothing concerning frameworks based on stochastic differential equations. We extend the previously developed joint mean-and-covariance optimization (JMCO) filtering setting to derive a reduced-order smoother via the Rauch--Tung--Striebel recursion and establish the corresponding Kalman--Bucy smoothing for affine drift dynamics. The resulting algorithms retain the adaptive nature of DLRA while significantly reducing the computational time and storage of the whole smoothing procedure.

math.NA

A transition-density-based operator learning method for Fokker-Planck equations with various initial conditions

Solving Fokker-Planck equations (FPEs) for multiple initial conditions typically requires repeated computations, leading to substantial computational costs. In this work, we propose a transition-density-based operator learning method to efficiently approximate the solution operator of FPEs with various initial conditions. The core idea is to learn the transition probability density function (PDF) of the underlying stochastic differential equation (SDE), from which the solution associated with a new initial distribution can be obtained through the Chapman-Kolmogorov equation without retraining the model. A major challenge in learning the transition PDF lies in the singular behavior induced by the Dirac initial condition. To address it, we introduce a conditional normalizing flow whose base distribution is given by the explicit transition PDF of a linearized SDE. This base distribution captures the short-time behavior of the target transition PDF and allows the normalizing flow to learn a near-identity transformation at small times. We further incorporate a time-weighted loss function to stabilize training near the initial time and develop an importance-sampling strategy for evaluating solutions associated with general initial conditions. A variety of numerical experiments are presented to illustrate the effectiveness and robustness of the proposed method.

cs.LG

Neural Galerkin Normalizing Flows for Bayesian Inference of Diffusions with Inaccessible Boundaries

One of the primary challenges in Bayesian inference on the parameters of a diffusion model from discrete observations is the unavailability of an analytical expression for the transition density function between consecutive observation times, which is needed to derive the likelihood function. Extending previous studies that solve Fokker-Planck (FP) type partial differential equations with Normalizing Flows, we propose a new Normalizing Flow architecture to learn the transition density function of the diffusion process between two observation times. We do so by solving in a Neural Galerkin framework the associated FP equation with a Dirac mass as initial condition, over a specified training distribution of the initial datum and the coefficients of the diffusion. We specifically focus on processes whose diffusion matrix vanishes in certain inaccessible boundary regions, such as Stochastic Volatility models that satisfy a Feller condition. The product of the obtained transition densities evaluated along the observed trajectory approximates the likelihood function, thereby enabling cheap posterior sampling via Markov chain Monte Carlo (MCMC). After the offline training phase, inference becomes significantly more efficient, as it avoids the need to solve the FP equation in real time for each parameter proposed by the MCMC sampler or to rely on other likelihood-free methods for Bayesian inference that involve repeated simulation of diffusion bridges.

cs.LG

Optimized multilevel Monte Carlo methods in Banach spaces

We present a theoretical and numerical analysis of Monte Carlo methods for the estimation of statistical moments of random variables $X:\Omega\rightarrow E$ taking values in a Banach space $E$. For practical computation, we consider finite-dimensional approximation subspaces ${(E_\ell)_{\ell\in\mathbb{N}}\subset E}$ of increasing dimension. We develop a refined error analysis that explicitly accounts for a dependence of the Rademacher type constants on the dimension of $E_\ell$, leading to novel complexity results for single- and multilevel Monte Carlo methods to estimate the mean and injective moments of arbitrary order, which are, in certain cases, sharper than those derived in [Kirchner, Schwab, J. Funct. Anal, 2024]. Moreover, we show that, in favorable cases, the resulting error-vs.-work bounds are independent of the Rademacher type of $E$. We then focus on $L^p(S)$-valued random variables for a $\sigma$-finite measure space satisfying certain approximation properties, and prove that for a random variable $X\in L^q(\Omega;L^p(S))\cap L^p(S;L^q(\Omega))$, with $q\in (1,\infty)$ and $p\in [1,\infty)$, the $L^q$-convergence rate of a Monte Carlo estimator is determined exclusively by the integrability parameter $\min\{q,2\}$, with no dependence on the Rademacher type $\min\{p,2\}$ of $L^p(S)$. We further investigate the impact of measuring the (multilevel) Monte Carlo error in the $L^q(\Omega;L^p(S))$-norm while $X$ possesses additional regularity, $X\in L^{\tilde{q}}(\Omega;L^p(S))\cap L^p(S;L^{\tilde{q}}(\Omega))$ with $\tilde{q}\in [q,\infty)$. This analysis reveals an interplay between the sampling error and the strong approximation error, and leads to optimized error-vs.-work bounds for both single- and multilevel Monte Carlo methods. Numerical experiments confirm the sharpness of the analyses presented.

math.NA

Neural Galerkin Normalizing Flow for Transition Probability Density Functions of Diffusion Models

We propose a new Neural Galerkin Normalizing Flow framework to approximate the transition probability density function of a diffusion process by solving the corresponding Fokker-Planck equation with an atomic initial distribution, parametrically with respect to the location of the initial mass. By using Normalizing Flows, we look for the solution as a transformation of the transition probability density function of a reference stochastic process, ensuring that our approximation is structure-preserving and automatically satisfies positivity and mass conservation constraints. By extending Neural Galerkin schemes to the context of Normalizing Flows, we derive a system of ODEs for the time evolution of the Normalizing Flow's parameters. Adaptive sampling routines are used to evaluate the Fokker-Planck residual in meaningful locations, which is of vital importance to address high-dimensional PDEs. Numerical results show that this strategy captures key features of the true solution and enforces the causal relationship between the initial datum and the density function at subsequent times. After completing an offline training phase, online evaluation becomes significantly more cost-effective than solving the PDE from scratch. The proposed method serves as a promising surrogate model, which could be deployed in many-query problems associated with stochastic differential equations, like Bayesian inference, simulation, and diffusion bridge generation.

cs.LG

LAGO: A Local-Global Optimization Framework Combining Trust Region Methods and Bayesian Optimization

We introduce LAGO, a LocAl-Global Optimization framework coupling Bayesian Optimization (BO) and gradient-based trust region local refinement through an adaptive competition mechanism for smooth expensive-to-evaluate objective functions with available gradients. At each iteration, global and local optimization strategies independently propose candidate points, and the next evaluation is selected based on predicted improvement. LAGO separates global exploration from local refinement at the proposal level: the BO acquisition function is optimized outside the active trust region, while local candidates are proposed within the trust region. Points in the vicinity of the accepted local step are incorporated in the global GP dataset only when satisfying a lengthscale-based minimum-distance criterion, hence reducing the risk of numerical instability during local exploitation. LAGO enhances BO with efficient local refinement when reaching promising regions, and reverts to exploratory behavior when local steps are not competitive.

cs.LG

Dynamical Low-Rank Ensemble Kalman filter for State/Parameter estimation

We propose a Dynamical Low-Rank Ensemble Kalman Filter (DLR-ENKF) for efficient joint state-parameter estimation in high-dimensional dynamical systems. The method extends the DLR-ENKF formulation of arXiv:2509.11210 to the augmented state-parameter framework, tracking the filtering density within a dynamically evolving low-dimensional subspace. Key developments include a time-integration strategy that combines the Basis Update & Galerkin scheme with forecast/analysis discretisation, and a DEIM-based hyper-reduction technique for efficient evaluation of nonlinear terms. We demonstrate the effectiveness, robustness, and computational advantages of the proposed approach on benchmark problems. The results highlight the potential of dynamically evolving reduced bases to achieve accurate filtering and parameter estimation at reduced computational cost.

math.NA

Numerical Methods for Dynamical Low-Rank Approximations of Stochastic Differential Equations -- Part I: Time discretization

In this work (Part I), we study three time-discretization schemes for the Dynamical Low-Rank Approximation (DLRA) of high-dimensional stochastic differential equations (SDEs). Specifically, we consider the Dynamically Orthogonal (DO) method for DLRA proposed and analyzed in arXiv:2308.11581v4, which approximates the true solution by a linear combination of few products between deterministic orthonormal modes and stochastic modes, both time-dependent. The first scheme considered consists in a forward discretization in time of both deterministic and stochastic components, in a Euler-Maruyama style. Its convergence is proven subject to a time-step restriction dependent on the smallest singular value of the Gram matrix associated to the stochastic modes, which, on its turn, is shown to be always positive, provided that the SDE under study is driven by a non-degenerate noise. The second and the third schemes, on the other hand, are staggered ones, alternating updates of the deterministic and the stochastic modes in half steps, and have a projector splitting nature. We show stability of the second scheme and prove convergence with constants independent of the smallest singular value. The third scheme works better in practice, although our theoretical convergence bounds are worse than those for the second one. Computational experiments support our theoretical results. In this work we do not consider the discretization in probability, which will be the topic of Part II.

math.NA

A function approximation algorithm using multilevel active subspaces

The Active Subspace (AS) method is a widely used technique for identifying the most influential directions in high-dimensional input spaces that affect the output of a computational model. The standard AS algorithm requires a sufficient number of gradient evaluations (samples) of the input output map to achieve quasi-optimal reconstruction of the active subspace, which can lead to a significant computational cost if the samples include numerical discretization errors which have to be kept sufficiently small. To address this issue, we propose a multilevel version of the Active Subspace method (MLAS) that utilizes samples computed with different accuracies and yields different active subspaces across accuracy levels, which can match the accuracy of single-level AS with reduced computational cost, making it suitable for downstream tasks such as function approximation. In particular, we propose to perform the latter via optimally-weighted least-squares polynomial approximation in the different active subspaces, and we present an adaptive algorithm to choose dynamically the dimensions of the active subspaces and polynomial spaces. We demonstrate the practical viability of the MLAS method with polynomial approximation through numerical experiments based on random partial differential equations (PDEs).

math.NA

Dynamical Low-Rank Approximations for Kalman Filtering

We propose a dynamical low rank approximation of the Kalman-Bucy process (DLR-KBP), which evolves the filtering distribution of a partially continuously observed linear SDE on a small time-varying subspace at reduced computational cost. This reduction is valid in presence of small noise and when the filtering distribution concentrates around a low dimensional subspace. We further extend this approach to a DLR-ENKF process, where particles are evolved in a low dimensional time-varying subspace at reduced cost. This allows for a significantly larger ensemble size compared to standard EnKF at equivalent cost, thereby lowering the Monte Carlo error and improving filter accuracy. Theoretical properties of the DLR-KBP and DLR-ENKF are investigated, including a propagation of chaos property. Numerical experiments demonstrate the effectiveness of the technique.

math.NA

Stochastic gradient with least-squares control variates

The stochastic gradient descent (SGD) method is a widely used approach for solving stochastic optimization problems, but its convergence is typically slow. Existing variance reduction techniques, such as SAGA, improve convergence by leveraging stored gradient information; however, they are restricted to settings where the objective functional is a finite sum, and their performance degrades when the number of terms in the sum is large. In this work, we propose a novel approach which is well suited when the objective is given by an expectation over random variables with a continuous probability distribution. Our method constructs a control variate by fitting a linear model to past gradient evaluations using weighted discrete least-squares, effectively reducing variance while preserving computational efficiency. We establish theoretical sublinear convergence guarantees for strongly convex objectives and demonstrate the method's effectiveness through numerical experiments on random PDE-constrained optimization problems.

math.OC

Multilevel quadrature formulae for the optimal control of random PDEs

This manuscript presents a framework for using multilevel quadrature formulae to compute the solution of optimal control problems constrained by random partial differential equations. Our approach consists in solving a sequence of optimal control problems discretized with different levels of accuracy of the physical and probability discretizations. The final approximation of the control is then obtained in a postprocessing step, by suitably combining the adjoint variables computed on the different levels. We present a general convergence and complexity analysis for an unconstrained linear quadratic problem under abstract assumptions on the spatial discretization and on the quadrature formulae. We detail our framework for the specific case of a MultiLevel Monte Carlo (MLMC) quadrature formula, and numerical experiments confirm the better computational complexity of our MLMC approach compared to a standard Monte Carlo sample average approximation, even beyond the theoretical assumptions.

math.NA

Robust high-order low-rank BUG integrators based on explicit Runge--Kutta methods

In this work, we introduce high-order Basis-Update & Galerkin (BUG) integrators based on explicit Runge-Kutta methods for large-scale matrix differential equations. These dynamical low-rank integrators extend the BUG integrator to arbitrary explicit Runge-Kutta schemes by performing a BUG step at each stage of the method. The resulting Runge-Kutta BUG (RK-BUG) integrators are robust with respect to small singular values, fully forward in time, and high-order accurate, while enabling conservation and rank adaptivity. We prove that RK-BUG integrators retain the order of convergence of the underlying Runge-Kutta method until the error reaches a plateau corresponding to the low-rank truncation error, which vanishes as the rank becomes full. This theoretical analysis is supported by several numerical experiments. The results demonstrate the high-order convergence of the RK-BUG integrator and its superior accuracy compared to other existing dynamical low-rank integrators.

math.NA

Probabilistic Load Forecasting of Distribution Power Systems based on Empirical Copulas

Accurate and reliable electricity load forecasts are becoming increasingly important as the share of intermittent resources in the system increases. Distribution System Operators (DSOs) are called to accurately forecast their production and consumption to place optimal bids in the day-ahead market. Forecasts must account for the volatility of weather-parameters that impacts both the production and consumption of electricity. If DSO-loads are small or lower-granularity forecasts are needed, parametric statistical methods may fail to provide reliable performance since they rely on a priori statistical distributions of the variables to forecast. In this paper, we introduce a Probabilistic Load Forecast (PLF) method based on Empirical Copulas (ECs). The model is datadriven, does not need a priori assumption on parametric distribution for variables, nor the dependence structure (copula). It employs a kernel density estimate of the underlying distribution using beta kernels that have bounded support on the unit hypercube. The method naturally supports variables with widely different distributions, such as weather data (including forecasted ones) and historic electricity consumption, and produces a conditional probability distribution for every time step in the forecast, which allows inferring the quantiles of interest. The proposed non-parametric approach differs significantly from previous forecasting methods based on copulas, which typically uses copulas to model hierarchical dependence. The bandwidth of the beta kernel density estimators is optimized using Integrated Square Error (ISE). We present results from an open dataset and showcase the strength of the model with respect to Quantile Regression (QR) using standard probabilistic evaluation metrics.

eess.SY

Dynamical Low-Rank Approximation for Stochastic Differential Equations

In this paper, we set the mathematical foundations of the Dynamical Low-Rank Approximation (DLRA) method for stochastic differential equations (SDEs). DLRA aims at approximating the solution as a linear combination of a small number of basis vectors with random coefficients (low rank format) with the peculiarity that both the basis vectors and the random coefficients vary in time. While the formulation and properties of DLRA are now well understood for random/parametric equations, the same cannot be said for SDEs and this work aims to fill this gap. We start by rigorously formulating a Dynamically Orthogonal (DO) approximation (an instance of DLRA successfully used in applications) for SDEs, which we then generalize to define a parametrization independent DLRA for SDEs. We show local well-posedness of the DO equations and their equivalence with the DLRA formulation. We also characterize the explosion time of the DO solution by a loss of linear independence of the random coefficients defining the solution expansion and give sufficient conditions for global existence.

math.NA

Petrov-Galerkin Dynamical Low Rank Approximation:SUPG stabilisation of advection-dominated problems

We propose a novel framework of generalised Petrov-Galerkin Dynamical Low Rank Approximations (DLR) in the context of random PDEs. It builds on the standard Dynamical Low Rank Approximations in their Dynamically Orthogonal formulation. It allows to seamlessly build-in many standard and well-studied stabilisation techniques that can be framed as either generalised Galerkin methods, or Petrov-Galerkin methods. The framework is subsequently applied to the case of Streamine Upwind/Petrov Galerkin (SUPG) stabilisation of advection-dominated problems with small stochastic perturbations of the transport field. The norm-stability properties of two time discretisations are analysed. Numerical experiments confirm that the stabilising properties of the SUPG method naturally carry over to the DLR framework.

math.NA