SearcharxivSearch

arXiv subjects

Dirk Nuyens

Publications and source records attributed to Dirk Nuyens.

At least 19 recordsLinked to original sources

A machine-learned probability distribution in the phase space of turbulent channel flow for synthetic turbulence and flow reconstruction

Although a complete characterisation of the probability distribution in the phase space of turbulent flows remains elusive, accurately sampling this distribution is essential for both synthetic turbulence generation and turbulent flow reconstruction. Motivated by these applications, we examine to what extent a machine-learned distribution can approximate the physical invariant distribution of turbulent channel flow at $\mathrm{Re}_\tau=180$. We assess three important properties of the approximation: physical ensemble statistics, consistent conditional sampling, and dynamical invariance. To this end, a flow-based generative model is trained on a minimal conditional flow unit, which we define as the smallest domain outside which conditional fields, given a single observation at the domain centre, are indistinguishable from unconditional fields in terms of mean-square discrepancy to other conditional fields. We also introduce a consistent procedure for sampling from the conditional learned distribution. Comparisons with direct numerical simulation show that synthetic turbulent fields reproduce key statistical and dynamical features of turbulence, including intermittency and nonlinear energy transfer. The consistency of conditional sampling is demonstrated in a flow reconstruction problem, and subsequently used to generate synthetic turbulent velocity fields on a large domain. When adopted as initial conditions in direct numerical simulations, these fields yield physical and statistically stationary ensemble statistics, indicating that the learned distribution provides a good approximation to the natural distribution of the turbulent dynamical system.

physics.flu-dyn

Fourier Neural Operators with rank-1 lattice points and hyperbolic cross

The \emph{Fourier neural operator} (FNO) is a neural network architecture that learns mappings between function spaces. Its efficient implementation is based on the multi-dimensional Fourier transform. By deriving general regularity bounds for the FNO with respect to both the spatial and parametric variables, we prove that the generalization error of the FNO can be improved by replacing spatial tensor product grids with purpose-built rank-1 lattice points, and by using a second lattice carefully constructed as training points in the parametric space. We achieve more accurate and efficient approximations from fewer network parameters, fewer spatial points, and fewer training samples. In addition, the architecture is simplified, because the high-dimensional Fourier transform on rank-1 lattices requires only a \emph{one-dimensional fast Fourier transform}, and we can use a \emph{hyperbolic cross} frequency index set with lattice points. We demonstrate the benefits of our \emph{lattice-based hyperbolic-cross FNOs} for an elliptic PDE on the torus.

math.NA

Scrap Composition Estimation in EAF and BOF: State-Space Models, Hyperparameters, and Validation

Accurate knowledge of scrap composition can increase the usage of recycled material to produce steel, reducing the need for raw ore extraction and minimizing environmental impact by conserving natural resources and lowering carbon emissions. First, we introduce two state-space models for the elemental composition of scrap in Electric Arc Furnaces (EAF) and Basic Oxygen Furnaces (BOF): a linear model for elements that transfer entirely into steel, and a non-linear model for elements that partition between steel and slag. The models are fitted with the Kalman filter and the unscented Kalman filter, respectively, using only data already collected in the standard steel production process. Crucially, the resulting scrap composition estimates can in turn be used to predict the elemental composition of future steel production. Second, we analyze how key hyperparameters affect estimation accuracy and stability, and we provide practical guidelines for tuning them from expert knowledge and historical data. Third, we validate the models on real BOF data from ArcelorMittal, using Cu and Cr as representative elements. Both filters outperform windowed non-negative least squares regression, a strong baseline method for scrap composition estimation, yielding reliable real-time estimates of scrap composition.

eess.SY

Regularity and tailored regularization of Deep Neural Networks, with application to parametric PDEs in uncertainty quantification

In this paper we consider Deep Neural Networks (DNNs) with a smooth activation function as surrogates for high-dimensional functions that are somewhat smooth but costly to evaluate. We consider the standard (non-periodic) DNNs as well as propose a new model of periodic DNNs which are especially suited for a class of periodic target functions when Quasi-Monte Carlo lattice points are used as training points. The primary contribution of this paper is the derivation of explicit bounds for all mixed derivatives of DNNs with respect to their input parameters. The bounds depend on the neural network parameters as well as the choice of activation function, with explicit constants. These bounds are fully general and remain independent of both the target function and the training data. By imposing restrictions on the network parameters to match the regularity features of the target functions, we prove that DNNs with $N$ tailor-constructed lattice training points can achieve the generalization error (or $L_2$ approximation error) bound ${\tt tol} + \mathcal{O}(N^{-r/2})$, where ${\tt tol}\in (0,1)$ is the tolerance achieved by the training error in practice, and $r = 1/p^*$, with $p^*$ being the ``summability exponent'' of a sequence that characterises the decay of the input variables in the target functions, and with the implied constant independent of the dimensionality of the input data. We apply our analysis to popular models of parametric elliptic PDEs in uncertainty quantification. In our numerical experiments, we restrict the network parameters during training by adding tailored regularization terms, and we show that for an algebraic equation mimicking the parametric PDE problems the DNNs trained with tailored regularization perform significantly better.

math.NA

Lattice-based Deep Neural Networks: Regularity and Tailored Regularization

This survey article is concerned with the application of lattice rules to Deep Neural Networks (DNNs), lattice rules being a family of quasi-Monte Carlo methods. They have demonstrated effectiveness in various contexts for high-dimensional integration and function approximation. They are extremely easy to implement thanks to their very simple formulation -- all that is required is a good integer generating vector of length matching the dimensionality of the problem. In recent years there has been a burst of research activities on the application and theory of DNNs. We review our recent article on using lattice rules as training points for DNNs with a smooth activation function, where we obtained explicit regularity bounds of the DNNs. By imposing restrictions on the network parameters to match the regularity features of the target function, we prove that DNNs with tailored lattice training points can achieve good theoretical generalization error bounds, with implied constants independent of the input dimension. We also demonstrate numerically that DNNs trained with our tailored regularization perform significantly better than with standard $\ell_2$ regularization.

cs.LG

Quasi-Monte Carlo methods for uncertainty quantification of wave propagation and scattering problems modelled by the Helmholtz equation

We analyse and implement a quasi-Monte Carlo (QMC) finite element method (FEM) for the forward problem of uncertainty quantification (UQ) for the Helmholtz equation with random coefficients, both in the second-order and zero-order terms of the equation, thus modelling wave scattering in random media. The problem is formulated on the infinite propagation domain, after scattering by the heterogeneity, and also (possibly) a bounded impenetrable scatterer. The spatial discretization scheme includes truncation to a bounded domain via a perfectly matched layer (PML) technique and then FEM approximation. A special case is the problem of an incident plane wave being scattered by a bounded sound-soft impenetrable obstacle surrounded by a random heterogeneous medium, or more simply, just scattering by the random medium. The random coefficients are assumed to be affine separable expansions with infinitely many independent uniformly distributed and bounded random parameters. As quantities of interest for the UQ, we consider the expectation of general linear functionals of the solution, with a special case being the far-field pattern of the scattered field. The numerical method consists of (a) dimension truncation in parameter space, (b) application of an adapted QMC method to compute expected values, and (c) computation of samples of the PDE solution via PML truncation and FEM approximation. Our error estimates are explicit in $s$ (the dimension truncation parameter), $N$ (the number of QMC points), $h$ (the FEM grid size) and (most importantly), $k$ (the Helmholtz wavenumber). The method is also exponentially accurate with respect to the PML truncation radius. Illustrative numerical experiments are given.

math.NA

Quasi-Monte Carlo methods for uncertainty quantification of tumor growth modeled by a parametric semi-linear parabolic reaction-diffusion equation

We study the application of a quasi-Monte Carlo (QMC) method to a class of semi-linear parabolic reaction-diffusion partial differential equations used to model tumor growth. Mathematical models of tumor growth are largely phenomenological in nature, capturing infiltration of the tumor into surrounding healthy tissue, proliferation of the existing tumor, and patient response to therapies, such as chemotherapy and radiotherapy. Considerable inter-patient variability, inherent heterogeneity of the disease, sparse and noisy data collection, and model inadequacy all contribute to significant uncertainty in the model parameters. It is crucial that these uncertainties can be efficiently propagated through the model to compute quantities of interest (QoIs), which in turn may be used to inform clinical decisions. We show that QMC methods can be successful in computing expectations of meaningful QoIs. Well-posedness results are developed for the model and used to show a theoretical error bound for the case of uniform random fields. The theoretical linear error rate, which is superior to that of standard Monte Carlo, is verified numerically. Encouraging computational results are also provided for lognormal random fields, prompting further theoretical development.

math.NA

A localized consensus-based sampling algorithm

We propose a localized consensus-based method for sampling from non-Gaussian distributions, a task that frequently arises when solving Bayesian inverse problems. Our method arises from an alternative derivation of consensus-based sampling (CBS). Starting from ensemble-preconditioned Langevin dynamics, we replace the potential by its Moreau envelope -- a smoother approximation -- in order to replace the gradient in the Langevin equation with a proximal operator. We then approximate this operator by a weighted mean. In the limit of infinitely smoothing the potential to a quadratic function, this procedure recovers the standard CBS dynamics. In addition, outside this limit, we retrieve a refined variant of polarized CBS. We call the resulting algorithm localized consensus-based sampling, since particles interact more with nearby particles than with faraway ones. Our method is affine-invariant, exact for Gaussian targets in the mean-field limit, and demonstrates improved robustness over polarized CBS in numerical experiments. Like other consensus-based methods, localized CBS is gradient-free and easily parallelizable.

math.NA

Error bounds for function approximation using generated sets

This paper explores the use of "generated sets" $\{ \{ k \boldsymbolζ \} : k = 1, \ldots, n \}$ for function approximation in reproducing kernel Hilbert spaces which consist of multi-dimensional functions with an absolutely convergent Fourier series. The algorithm is a least squares algorithm that samples the function at the points of a generated set. We show that there exist $\boldsymbolζ \in [0,1]^d$ for which the worst-case $L_2$ error has the optimal order of convergence if the space has polynomially converging approximation numbers. In fact, this holds for a significant portion of the generators. Additionally we show that a restriction to rational generators is possible with a slight increase of the bound. Furthermore, we specialise the results to the weighted Korobov space, where we derive a bound applicable to low values of sample points, and state tractability results.

math.NA

Sensitivity Analysis of State Space Models for Scrap Composition Estimation in EAF and BOF

This study develops and analyzes linear and nonlinear state space models for estimating the elemental composition of scrap steel used in steelmaking, with applications to Electric Arc Furnace (EAF) and Basic Oxygen Furnace (BOF) processes. The models incorporate mass balance equations and are fitted using a modified Kalman filter for linear cases and the Unscented Kalman Filter (UKF) for nonlinear cases. Using Cu and Cr as representative elements, we assess the sensitivity of model predictions to measurement noise in key process variables, including steel mass, steel composition, scrap input mass, slag mass, and iron oxide fraction in slag. Results show that the models are robust to moderate noise levels in most variables, particularly when errors are below $10\%$. However, accuracy significantly deteriorates with noise in slag mass estimation. These findings highlight the practical feasibility and limitations of applying state space models for real-time scrap composition estimation in industrial settings.

eess.SY

A randomised lattice rule algorithm with pre-determined generating vector and random number of points for Korobov spaces with $0 < α\le 1/2$

In previous work (Kuo, Nuyens, Wilkes, 2023), we showed that a lattice rule with a pre-determined generating vector but random number of points can achieve the near optimal convergence of $O(n^{-α-1/2+ε})$, $ε> 0$, for the worst case expected error, commonly referred to as the randomised error, for numerical integration of high-dimensional functions in the Korobov space with smoothness $α> 1/2$. Compared to the optimal deterministic rate of $O(n^{-α+ε})$, $ε> 0$, such a randomised algorithm is capable of an extra half in the rate of convergence. In this paper, we show that a pre-determined generating vector also exists in the case of $0 < α\le 1/2$. Also here we obtain the near optimal convergence of $O(n^{-α-1/2+ε})$, $ε> 0$; or in more detail, we obtain $O(\sqrt{r} \, n^{-α-1/2+1/(2r)+ε'})$ which holds for any choices of $ε' > 0$ and $r \in \mathbb{N}$ with $r > 1/(2α)$.

math.NA

Random-prime--fixed-vector randomised lattice-based algorithm for high-dimensional integration

We show that a very simple randomised algorithm for numerical integration can produce a near optimal rate of convergence for integrals of functions in the $d$-dimensional weighted Korobov space. This algorithm uses a lattice rule with a fixed generating vector and the only random element is the choice of the number of function evaluations. For a given computational budget $n$ of a maximum allowed number of function evaluations, we uniformly pick a prime $p$ in the range $n/2 < p \le n$. We show error bounds for the randomised error, which is defined as the worst case expected error, of the form $O(n^{-α- 1/2 + δ})$, with $δ> 0$, for a Korobov space with smoothness $α> 1/2$ and general weights. The implied constant in the bound is dimension-independent given the usual conditions on the weights. We present an algorithm that can construct suitable generating vectors \emph{offline} ahead of time at cost $O(d n^4 / \ln n)$ when the weight parameters defining the Korobov spaces are so-called product weights. For this case, numerical experiments confirm our theory that the new randomised algorithm achieves the near optimal rate of the randomised error.

math.NA

Comparison of Two Search Criteria for Lattice-based Kernel Approximation

The kernel interpolant in a reproducing kernel Hilbert space is optimal in the worst-case sense among all approximations of a function using the same set of function values. In this paper, we compare two search criteria to construct lattice point sets for use in lattice-based kernel approximation. The first candidate, $\calP_n^*$, is based on the power function that appears in machine learning literature. The second, $\calS_n^*$, is a search criterion used for generating lattices for approximation using truncated Fourier series. We find that the empirical difference in error between the lattices constructed using $\calP_n^*$ and $\calS_n^*$ is marginal. The criterion $\calS_n^*$ is preferred as it is computationally more efficient and has a proven error bound.

math.NA

Constructing Embedded Lattice-based Algorithms for Multivariate Function Approximation with a Composite Number of Points

We approximate $d$-variate periodic functions in weighted Korobov spaces with general weight parameters using $n$ function values at lattice points. We do not limit $n$ to be a prime number, as in currently available literature, but allow any number of points, including powers of $2$, thus providing the fundamental theory for construction of embedded lattice sequences. Our results are constructive in that we provide a component-by-component algorithm which constructs a suitable generating vector for a given number of points or even a range of numbers of points. It does so without needing to construct the index set on which the functions will be represented. The resulting generating vector can then be used to approximate functions in the underlying weighted Korobov space. We analyse the approximation error in the worst-case setting under both the $L_2$ and $L_{\infty}$ norms. Our component-by-component construction under the $L_2$ norm achieves the best possible rate of convergence for lattice-based algorithms, and the theory can be applied to lattice-based kernel methods and splines. Depending on the value of the smoothness parameter $α$, we propose two variants of the search criterion in the construction under the $L_{\infty}$ norm, extending previous results which hold only for product-type weight parameters and prime $n$. We also provide a theoretical upper bound showing that embedded lattice sequences are essentially as good as lattice rules with a fixed value of $n$. Under some standard assumptions on the weight parameters, the worst-case error bound is independent of $d$.

math.NA

Scaled lattice rules for integration on $\mathbb{R}^d$ achieving higher-order convergence with error analysis in terms of orthogonal projections onto periodic spaces

We introduce a new method to approximate integrals $\int_{\mathbb{R}^d} f(\boldsymbol{x}) \, \mathrm{d} \boldsymbol{x}$ which simply scales lattice rules from the unit cube $[0,1]^d$ to properly sized boxes on $\mathbb{R}^d$, hereby achieving higher-order convergence that matches the smoothness of the integrand function $f$ in a certain Sobolev space of dominating mixed smoothness. Our method only assumes that we can evaluate the integrand function $f$ and does not assume a particular density nor the ability to sample from it. In particular, for the theoretical analysis we show a new result that the method of adding Bernoulli polynomials to a function to make it "periodic" on a box without changing its integral value over the box, is equivalent to an orthogonal projection from a well chosen Sobolev space of dominating mixed smoothness to an associated periodic Sobolev space of the same dominating mixed smoothness, which we call a Korobov space. We note that the Bernoulli polynomial method is often not used because of its excessive computational complexity and also here we only make use of it in our theoretical analysis. We show that our new method of applying scaled lattice rules to increasing boxes can be interpreted as orthogonal projections with decreasing projection error. Such a method would not work on the unit cube since then the committed error caused by non-periodicity of the integrand would be constant, but for integration on the Euclidean space we can use the certain decay towards zero when the boxes grow. Hence we can bound the truncation error as well as the projection error and show higher-order convergence in applying scaled lattice rules for integration on Euclidean space. We illustrate our theoretical analysis by numerical experiments which confirm our findings.

math.NA

Lattice field computations via recursive numerical integration

We investigate the application of efficient recursive numerical integration strategies to models in lattice gauge theory from quantum field theory. Given the coupling structure of the physics problems and the group structure within lattice cubature rules for numerical integration, we show how to approach these problems efficiently by means of Fast Fourier Transform techniques. In particular, we consider applications to the quantum mechanical rotor and compact U(1) lattice gauge theory, where the physical dimensions are two and three. This proceedings article reviews our results presented in J. Comput. Phys 443 (2021) 110527.

hep-lat

MDFEM: Multivariate decomposition finite element method for elliptic PDEs with lognormal diffusion coefficients using higher-order QMC and FEM

We introduce the multivariate decomposition finite element method for elliptic PDEs with lognormal diffusion coefficient $a=\exp(Z)$ where $Z$ is a Gaussian random field defined by an infinite series expansion $Z(\boldsymbol{y}) = \sum_{j\ge1} y_j\,ϕ_j$ with $y_j\sim\mathcal{N}(0,1)$ and a given sequence of functions $\{ϕ_j\}_{j\ge1}$. We use the MDFEM to approximate the expected value of a linear functional of the solution of the PDE which is an infinite-dimensional integral over the parameter space. The proposed algorithm uses the multivariate decomposition method (MDM) to compute the infinite-dimensional integral by a decomposition into finite-dimensional integrals, which we resolve using quasi-Monte Carlo (QMC) methods, and for which we use the finite element method (FEM) to solve different instances of the PDE. We develop higher-order quasi-Monte Carlo rules for integration over the finite-dimensional Euclidean space with respect to the Gaussian distribution by use of a truncation strategy. By linear transformations of interlaced polynomial lattice rules from the unit cube to a multivariate box of the Euclidean space we achieve higher-order convergence rates for functions belonging to a class of anchored Gaussian Sobolev spaces, taking into account the truncation error. Under appropriate conditions, the MDFEM achieves higher-order convergence rates in term of error versus cost, i.e., to achieve an accuracy of $O(ε)$ the computational cost is $O(ε^{-1/λ-d'/λ}) = O(ε^{-(p^*+d'/τ)/(1-p^*)})$ where $ε^{-1/λ}$ and $ε^{-d'/λ}$ are respectively the cost of the quasi-Monte Carlo cubature and the finite element approximations, with $d' = d \, (1+δ')$ for some $δ' \ge 0$ and $d$ the physical dimension, and $0 < p^* \le (2+d'/τ)^{-1}$ is a parameter representing the sparsity of $\{ϕ_j\}_{j\ge1}$.

math.NA

MDFEM: Multivariate decomposition finite element method for elliptic PDEs with uniform random diffusion coefficients using higher-order QMC and FEM

We introduce the multivariate decomposition finite element method (MDFEM) for solving elliptic PDEs with uniform random diffusion coefficients. We show that the MDFEM can be used to reduce the computational complexity of estimating the expected value of a linear functional of the solution of the PDE. The proposed algorithm combines the multivariate decomposition method (MDM), to compute infinite dimensional integrals, with the finite element method (FEM), to solve different instances of the PDE. The strategy of the MDFEM is to decompose the infinite-dimensional problem into multiple finite-dimensional ones which lends itself to easier parallelization than to solve a single large dimensional problem. Our first result adjusts the analysis of the multivariate decomposition method to incorporate the log-factor which typically appears in error bounds for multivariate quadrature, i.e., cubature, methods; and we take care of the fact that the number of points $n$ needs to come, e.g., in powers of 2 for higher order approximations. For the further analysis we specialize the cubature methods to be two types of quasi-Monte Carlo (QMC) rules, being digitally shifted polynomial lattice rules and interlaced polynomial lattice rules. The second and main contribution then presents a bound on the error of the MDFEM and shows higher-order convergence w.r.t. the total computational cost in case of the interlaced polynomial lattice rules in combination with a higher-order finite element method.

math.NA