SearcharxivSearch

arXiv subjects

Björn Engquist

Publications and source records attributed to Björn Engquist.

At least 19 recordsLinked to original sources

Convergence of a Randomized Newton Method in Nonconvex Optimization

We analyze a stochastic Newton optimization scheme for locating the unique global minimizer of a general nonconvex objective function. The method couples a Newton algorithm to additive Gaussian noise with state-dependent variance. In the bounded domain setting, we prove global almost sure convergence. The proof is based on two features of the algorithm: a nondegenerate exploratory property that ensures entrance into a neighborhood of the minimizer after a finite number of steps, and a decaying-noise property that yields contraction with high probability and prevents infinitely many exits from the neighborhood of the minimum.

math.OC

Diffusion-guided optimization for full waveform inversion

We present a diffusion-guided full waveform inversion (FWI) study in which pretrained diffusion generative models are used as learned regularizers inside a PDE-constrained seismic inversion loop. We compare three training-free guidance strategies: Manifold-Preserving Guided Diffusion (MPGD), SDEdit-based initialization, and Split Gibbs Diffusion Sampling (SGDS), which alternates between FWI likelihood updates and diffusion-prior denoising. The proposed workflow keeps wave-equation modeling in the inversion loop and uses a geological prior to stabilize model components that are weakly constrained by the seismic data. Controlled GeoFWI experiments, benchmark-scale Marmousi and Overthrust tests, a difficult Sigsbee2A salt test, and noise-degradation studies show that SGDS improves reconstruction quality relative to conventional L2 and total-variation regularized FWI in clean and moderately noisy synthetic settings. Overall, these experiments demonstrate that diffusion-guided optimization can serve as a practical learned regularization strategy for synthetic FWI benchmarks while preserving the wave-equation modeling loop.

physics.geo-ph

Sampling Using Hybrid Stochastic Dynamics

This work proposes a framework for sampling from the Gibbs distribution of a given potential using hybrid stochastic dynamics. In this framework, two distinct sampling dynamics are run in different regions of the state space. The two dynamics are coupled across the interface through natural transmission conditions that preserve the target distribution. Using a specially constructed regularization scheme, we establish an exponential rate of convergence for the hybrid dynamics to equilibrium. We also analyze the metastability properties of the hybrid dynamics in a radially symmetric landscape, showing that the hybrid scheme can improve the mean exit time. This advantage is further confirmed by the numerical experiments.

math.NA

A Dilation-based Seamless Multiscale Method For Elliptic Problems

Many numerical methods for multiscale differential equations require a scale separation between the larger and the smaller scales to achieve accuracy and computational efficiency. In the area of multiscale dynamical systems, so-called, seamless methods have been introduced to reduce the requirement of scale separation. We will translate these methods to numerical homogenization problems and extend the technique to multiple dimensions. The initial step is to prove that a one-dimensional \sepia{second-order} elliptic operator with oscillatory coefficients can be rewritten as a multiscale dynamical system. Inspired by this, multiscale elliptic operators in higher dimensions are approximated by a novel approach based on local dilation, which provides a middle ground for balancing intractability and accuracy without the need for full resolution. The dilation operator can be further generalized to preserve important structures by properly decomposing the coefficient field. Error estimates are developed and promising numerical results of different examples are included.

math.NA

An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization

We propose a new gradient descent algorithm with added stochastic terms for finding the global optimizers of nonconvex optimization problems. A key component in the algorithm is the adaptive tuning of the randomness based on the value of the objective function. In the language of simulated annealing, the temperature is state-dependent. With this, we prove the global convergence of the algorithm with an algebraic rate both in probability and in the parameter space. This is a significant improvement over the classical rate from using a more straightforward control of the noise term. The convergence proof is based on the actual discrete setup of the algorithm, not just its continuous limit as often done in the literature. We also present several numerical examples to demonstrate the efficiency and robustness of the algorithm for reasonably complex objective functions.

math.OC

Sampling with Adaptive Variance for Multimodal Distributions

We propose and analyze a class of adaptive sampling algorithms for multimodal distributions on a bounded domain, which share a structural resemblance to the classic overdamped Langevin dynamics. We first demonstrate that this class of linear dynamics with adaptive diffusion coefficients and vector fields can be interpreted and analyzed as weighted Wasserstein gradient flows of the Kullback--Leibler (KL) divergence between the current distribution and the target Gibbs distribution, which directly leads to the exponential convergence of both the KL and $χ^2$ divergences, with rates depending on the weighted Wasserstein metric and the Gibbs potential. We then show that a derivative-free version of the dynamics can be used for sampling without gradient information of the Gibbs potential and that for Gibbs distributions with nonconvex potentials, this approach could achieve significantly faster convergence than the classical overdamped Langevin dynamics. A comparison of the mean transition times between local minima of a nonconvex potential further highlights the better efficiency of the derivative-free dynamics in sampling.

cs.LG

Neural Inverse Operators for Solving PDE Inverse Problems

A large class of inverse problems for PDEs are only well-defined as mappings from operators to functions. Existing operator learning frameworks map functions to functions and need to be modified to learn inverse maps from data. We propose a novel architecture termed Neural Inverse Operators (NIOs) to solve these PDE inverse problems. Motivated by the underlying mathematical structure, NIO is based on a suitable composition of DeepONets and FNOs to approximate mappings from operators to functions. A variety of experiments are presented to demonstrate that NIOs significantly outperform baselines and solve PDE inverse problems robustly, accurately and are several orders of magnitude faster than existing direct and PDE-constrained optimization methods.

cs.LG

Adaptive State-Dependent Diffusion for Derivative-Free Optimization

This paper develops and analyzes a stochastic derivative-free optimization strategy. A key feature is the state-dependent adaptive variance. We prove global convergence in probability with algebraic rate and give the quantitative results in numerical examples. A striking fact is that convergence is achieved without explicit information of the gradient and even without comparing different objective function values as in established methods such as the simplex method and simulated annealing. It can otherwise be compared to annealing with state-dependent temperature.

math.OC

A Generalized Weighted Optimization Method for Computational Learning and Inversion

The generalization capacity of various machine learning models exhibits different phenomena in the under- and over-parameterized regimes. In this paper, we focus on regression models such as feature regression and kernel regression and analyze a generalized weighted least-squares optimization method for computational learning and inversion with noisy data. The highlight of the proposed framework is that we allow weighting in both the parameter space and the data space. The weighting scheme encodes both a priori knowledge on the object to be learned and a strategy to weight the contribution of different data points in the loss function. Here, we characterize the impact of the weighting scheme on the generalization error of the learning method, where we derive explicit generalization errors for the random Fourier feature model in both the under- and over-parameterized regimes. For more general feature maps, error bounds are provided based on the singular values of the feature matrix. We demonstrate that appropriate weighting from prior knowledge can improve the generalization capability of the learned model.

cs.LG

Heterogeneous multiscale methods for rough-wall laminar viscous flow

We develop numerical multiscale methods for viscous boundary layer flow. The goal is to derive effective boundary conditions, or wall laws, through high resolution simulations localized to the boundary coupled to a coarser simulation in the domain interior. The multiscale framework is analyzed in the context of laminar flow over a rough boundary. Asymptotic analysis shows that, up to a small perturbation, the method recovers the slip constant in the wall law derived from periodic homogenization theory. Numerical experiments illustrate the utility of the method for more general roughness patterns and fair field flow conditions.

math.NA

Effects of resolution inhomogeneity in large-eddy simulation

Large Eddy Simulation (LES) of turbulence in complex geometries is often conducted using strongly inhomogeneous resolution. The issues associated with resolution inhomogeneity are related to the noncommutativity of the filtering and differentiation operators, which introduces a commutation term into the governing equations. Neglect of this commutation term gives rise to commutation error. While the commutation error is well recognized, it is often ignored in practice. Moreover, the commutation error arising from the implicit part of the filter (i.e., projection onto the underlying discretization) has not been well investigated. Modeling the commutator between numerical projection and differentiation is crucial for correcting errors induced by resolution inhomogeneity in practical LES settings, which typically rely solely on implicit filtering. Here, we employ a multiscale asymptotic analysis to investigate the characteristics of the commutator. This provides a statistical description of the commutator, which can serve as a target for the statistical characteristics of a commutator model. Further, we investigate how commutation error manifests in simulation and demonstrate its impact on the convection of a packet of homogeneous isotropic turbulence through an inhomogeneous grid. A connection is made between the commutation error and the propagation properties of the underlying numerics. A modeling approach for the commutator is proposed that is applicable to LES with filters that include projections to the discrete solution space and that respects the numerical properties of the LES evolution equation. It may also be useful in addressing other LES modeling issues such as discretization error.

physics.flu-dyn

Near wall patch representation of wall bounded turbulence

Recent experimental and computational studies indicate that near wall turbulent flows can be characterized by universal small scale autonomous dynamics that are modulated by large scale structures. We formulate numerical simulations of near wall turbulence in a small domain localized to the boundary, whose size scales in viscous units. To mimic the environment in which the near wall turbulence evolves, our formulation accounts for the flux of mean momentum through the upper boundary of the domain. Comparisons of the model's two dimensional energy spectra and low order single-point statistics with the corresponding quantities computed from zero and mild favorable pressure gradient direct numerical simulations indicate it successfully captures the dynamics of the small scale near wall turbulence.

physics.flu-dyn

Seismic Inversion and the Data Normalization for Optimal Transport

Full waveform inversion (FWI) has recently become a favorite technique for the inverse problem of finding properties in the earth from measurements of vibrations of seismic waves on the surface. Mathematically, FWI is PDE constrained optimization where model parameters in a wave equation are adjusted such that the misfit between the computed and the measured dataset is minimized. In a sequence of papers, we have shown that the quadratic Wasserstein distance from optimal transport is to prefer as misfit functional over the standard $L^2$ norm. Datasets need however first to be normalized since seismic signals do not satisfy the requirements of optimal transport. There has been a puzzling contradiction in the results. Normalization methods that satisfy theorems pointing to ideal properties for FWI have not performed well in practical computations, and other scaling methods that do not satisfy these theorems have performed much better in practice. In this paper, we will shed light on this issue and resolve this contradiction.

math.NA

Seismic Imaging and Optimal Transport

Seismology has been an active science for a long time. It changed character about 50 years ago when the earth's vibrations could be measured on the surface more accurately and more frequently in space and time. The full wave field could be determined, and partial differential equations (PDE) started to be used in the inverse process of finding properties of the interior of the earth. We will briefly review earlier techniques but mainly focus on Full Waveform Inversion (FWI) for the acoustic formulation. FWI is a PDE constrained optimization in which the variable velocity in a forward wave equation is adjusted such that the solution matches measured data on the surface. The minimization of the mismatch is usually coupled with the adjoint state method, which also includes the solution to an adjoint wave equation. The least-squares norm is the conventional objective function measuring the difference between simulated and measured data, but it often results in the minimization trapped in local minima. One way to mitigate this is by selecting another misfit function with better convexity properties. Here we propose using the quadratic Wasserstein metric as a new misfit function in FWI. The optimal map defining the quadratic Wasserstein metric can be computed by solving a Monge-Ampere equation. Theorems pointing to the advantages of using optimal transport over the least-squares norm will be discussed, and a number of large-scale computational examples will be presented.

math.NA

Application of Optimal Transport and the Quadratic Wasserstein Metric to Full-Waveform Inversion

Conventional full-waveform inversion (FWI) using the least-squares norm ($L^2$) as a misfit function is known to suffer from cycle skipping. This increases the risk of computing a local rather than the global minimum of the misfit. In our previous work, we proposed the quadratic Wasserstein metric ($W_2$) as a new misfit function for FWI. The $W_2$ metric has been proved to have many ideal properties with regards to convexity and insensitivity to noise. When the observed and predicted seismic data are regarded as two density functions, the quadratic Wasserstein metric corresponds to the optimal cost of rearranging one density into the other, where the transportation cost is quadratic in distance. The difficulty of transforming seismic signals into nonnegative density functions is discussed. Unlike the $L^2$ norm, $W_2$ measures not only amplitude differences, but also global phase shifts, which helps to avoid cycle skipping issues. In this work, we build on our earlier method to cover more realistic high-resolution applications by embedding the $W_2$ technique into the framework of the adjoint-state method and applying it to seismic relevant 2D examples: the Camembert, the Marmousi, and the 2004 BP models. We propose a new way of using the $W_2$ metric trace-by-trace in FWI and compare it to global $W_2$ via the solution of the Monge-Ampère equation. With corresponding adjoint source, the velocity model can be updated using the l-BFGS method. Numerical results show the effectiveness of $W_2$ for alleviating cycle skipping issues and sensitivity to noise. Both mathematical theory and numerical examples demonstrate that the quadratic Wasserstein metric is a good candidate for a misfit function in seismic inversion.

physics.geo-ph

Nonuniform sampling and multiscale computation

In homogenization theory and multiscale modeling, typical functions satisfy the scaling law $f^ε(x) = f(x,x/ε)$, where $f$ is periodic in the second variable and $ε$ is the smallest relevant wavelength, $0<ε\ll1$. Our main result is a new $L^{2}$-stability estimate for the reconstruction of such bandlimited multiscale functions $f^ε$ from periodic nonuniform samples. The goal of this paper is to demonstrate the close relation between and sampling strategies developed in information theory and computational grids in multiscale modeling. This connection is of much interest because numerical simulations often involve discretizations by means of sampling, and meshes are routinely designed using tools from information theory. The proposed sampling sets are of optimal rate according to the minimal sampling requirements of Landau \cite{Landau}.

math.NA

A parallel directional Fast Multipole Method

This paper introduces a parallel directional fast multipole method (FMM) for solving N-body problems with highly oscillatory kernels, with a focus on the Helmholtz kernel in three dimensions. This class of oscillatory kernels requires a more restrictive low-rank criterion than that of the low-frequency regime, and thus effective parallelizations must adapt to the modified data dependencies. We propose a simple partition at a fixed level of the octree and show that, if the partitions are properly balanced between p processes, the overall runtime is essentially O(N log N/p+ p). By the structure of the low-rank criterion, we are able to avoid communication at the top of the octree. We demonstrate the effectiveness of our parallelization on several challenging models.

math.NA

A parallel sweeping preconditioner for heterogeneous 3D Helmholtz equations

A parallelization of a sweeping preconditioner for 3D Helmholtz equations without large cavities is introduced and benchmarked for several challenging velocity models. The setup and application costs of the sequential preconditioner are shown to be O(γ^2 N^{4/3}) and O(γ N log N), where γ(ω) denotes the modestly frequency-dependent number of grid points per Perfectly Matched Layer. Several computational and memory improvements are introduced relative to using black-box sparse-direct solvers for the auxiliary problems, and competitive runtimes and iteration counts are reported for high-frequency problems distributed over thousands of cores. Two open-source packages are released along with this paper: "Parallel Sweeping Preconditioner (PSP)" and the underlying distributed multifrontal solver, "Clique".

math.NA