SearcharxivSearch

arXiv subjects

Mohammad Motamed

Publications and source records attributed to Mohammad Motamed.

17 recordsLinked to original sources

Symmetrized Sinkhorn-Gibbs Inference for Oscillatory Inverse Problems

Oscillatory inverse problems often exhibit highly nonconvex discrepancy landscapes due to signal misalignment and cycle-skipping phenomena, posing significant challenges for uncertainty quantification. While Gibbs posteriors provide a flexible framework for incorporating problem-specific discrepancy measures, their performance depends strongly on the empirical risk landscape induced by the chosen discrepancy measure. Optimal transport accounts for the spatial and temporal structure of the underlying feature domain when comparing oscillatory signals. However, the direct application of classical optimal transport is precluded because observed and predicted signals are typically signed and therefore do not satisfy the positivity requirements of classical transport formulations. We introduce a symmetrized Sinkhorn-Gibbs inference framework for oscillatory inverse problems. The proposed approach combines a normalization procedure for signed signals with a symmetrized Sinkhorn loss that exploits complementary transport information from the normalized signals and their normalized negations. The resulting loss is incorporated into the Gibbs posterior framework, yielding a Gibbs inference methodology tailored to oscillatory data. We establish smoothness properties of the proposed loss, prove well-definedness of the resulting Gibbs posterior, derive robustness guarantees, and develop an adaptive sampling strategy for posterior computation. Numerical experiments demonstrate less oscillatory empirical risk landscapes with fewer spurious local minima, more accurate posterior inference, greater robustness to observational noise, and improved population-level recovery than Gibbs posteriors based on Euclidean and trace-wise Wasserstein losses.

math.OC

Fourier Residual Networks Achieve Spectral Accuracy for Discontinuous Functions

We present a constructive approximation framework for analyzing the expressive power of Fourier residual networks in approximating a broad class of one-dimensional functions. Our study covers both piecewise continuous functions -- including those with jump discontinuities in the function and its derivatives -- and fully smooth functions. We show that Fourier residual networks achieve spectral convergence without requiring periodicity or continuity, thereby overcoming key limitations of classical linear Fourier approximation and nonlinear methods, without being restricted to Barron-type function spaces. Our approach builds on classical techniques from approximation theory, including fixed-point iteration and Hermite interpolation by trigonometric polynomials. We support our theoretical results with numerical experiments based on both the constructed approximations and a randomized algorithm developed in our earlier work.

math.NA

Deep Learning without Global Optimization by Random Fourier Neural Networks

We introduce a new training algorithm for deep neural networks that utilize random complex exponential activation functions. Our approach employs a Markov Chain Monte Carlo sampling procedure to iteratively train network layers, avoiding global and gradient-based optimization while maintaining error control. It consistently attains the theoretical approximation rate for residual networks with complex exponential activation functions, determined by network complexity. Additionally, it enables efficient learning of multiscale and high-frequency features, producing interpretable parameter distributions. Despite using sinusoidal basis functions, we do not observe Gibbs phenomena in approximating discontinuous target functions.

cs.LG

Approximation Error and Complexity Bounds for ReLU Networks on Low-Regular Function Spaces

In this work, we consider the approximation of a large class of bounded functions, with minimal regularity assumptions, by ReLU neural networks. We show that the approximation error can be bounded from above by a quantity proportional to the uniform norm of the target function and inversely proportional to the product of network width and depth. We inherit this approximation error bound from Fourier features residual networks, a type of neural network that uses complex exponential activation functions. Our proof is constructive and proceeds by conducting a careful complexity analysis associated with the approximation of a Fourier features residual network by a ReLU network.

stat.ML

Residual Multi-Fidelity Neural Network Computing

In this work, we consider the general problem of constructing a neural network surrogate model using multi-fidelity information. Motivated by error-complexity estimates for ReLU neural networks, we formulate the correlation between an inexpensive low-fidelity model and an expensive high-fidelity model as a possibly non-linear residual function. This function defines a mapping between 1) the shared input space of the models along with the low-fidelity model output, and 2) the discrepancy between the outputs of the two models. The computational framework proceeds by training two neural networks to work in concert. The first network learns the residual function on a small set of high- and low-fidelity data. Once trained, this network is used to generate additional synthetic high-fidelity data, which is used in the training of the second network. The trained second network then acts as our surrogate for the high-fidelity quantity of interest. We present four numerical examples to demonstrate the power of the proposed framework, showing that significant savings in computational cost may be achieved when the output predictions are desired to be accurate within small tolerances.

cs.LG

Non-degenerate Marginal-Likelihood Calibration with Application to Quantum Characterization

We propose a marginal likelihood strategy within the Kennedy-O'Hagan (KOH) Bayesian framework, where a Gaussian process (GP) models the discrepancy between a physical system and its simulator. Our approach introduces a novel marginalized likelihood by integrating out the degenerate eigenspace of the covariance matrix, rather than approximating the original likelihood. Unlike approximation methods that compromise accuracy for computational efficiency, our method defines an exact likelihood -- distinct from the original but preserving all relevant information. This formulation achieves computational efficiency and stability, even for large datasets where the covariance matrix nears degeneracy. Applied to the characterization of a superconducting quantum device at Lawrence Livermore National Laboratory, the approach enhances the predictive accuracy of the Lindblad master equations for modeling Ramsey measurement data by effectively quantifying uncertainties consistent with the quantum data.

quant-ph

Deterministic and Bayesian Characterization of Quantum Computing Devices

Motivated by the noisy and fluctuating behavior of current quantum computing devices, this paper presents a data-driven characterization approach for estimating transition frequencies and decay times in a Lindbladian dynamical model of a superconducting quantum device. The data includes parity events in the transition frequency between the first and second excited states. A simple but effective mathematical model, based upon averaging solutions of two Lindbladian models, is demonstrated to accurately capture the experimental observations. A deterministic point estimate of the device parameters is first performed to minimize the misfit between data and Lindbladian simulations. These estimates are used to make an informed choice of prior distributions for the subsequent Bayesian inference. An additive Gaussian noise model is developed for the likelihood function, which includes two hyper-parameters to capture the noise structure of the data. The outcome of the Bayesian inference are posterior probability distributions of the transition frequencies, which for example can be utilized to design risk neutral optimal control pulses. The applicability of our approach is demonstrated on experimental data from the Quantum Device and Integration Testbed (QuDIT) at Lawrence Livermore National Laboratory, using a tantalum-based superconducting transmon device.

quant-ph

Approximation Power of Deep Neural Networks: an explanatory mathematical survey

This survey provides an in-depth and explanatory review of the approximation properties of deep neural networks, with a focus on feed-forward and residual architectures. The primary objective is to examine how effectively neural networks approximate target functions and to identify conditions under which they outperform traditional approximation methods. Key topics include the nonlinear, compositional structure of deep networks and the formalization of neural network tasks as optimization problems in regression and classification settings. The survey also addresses the training process, emphasizing the role of stochastic gradient descent and backpropagation in solving these optimization problems, and highlights practical considerations such as activation functions, overfitting, and regularization techniques. Additionally, the survey explores the density of neural networks in the space of continuous functions, comparing the approximation capabilities of deep ReLU networks with those of other approximation methods. It discusses recent theoretical advancements in understanding the expressiveness and limitations of these networks. A detailed error-complexity analysis is also presented, focusing on error rates and computational complexity for neural networks with ReLU and Fourier-type activation functions in the context of bounded target functions with minimal regularity assumptions. Alongside recent known results, the survey introduces new findings, offering a valuable resource for understanding the theoretical foundations of neural network approximation. Concluding remarks and further reading suggestions are provided.

cs.LG

A multi-fidelity neural network surrogate sampling method for uncertainty quantification

We propose a multi-fidelity neural network surrogate sampling method for the uncertainty quantification of physical/biological systems described by ordinary or partial differential equations. We first generate a set of low/high-fidelity data by low/high-fidelity computational models, e.g. using coarser/finer discretizations of the governing differential equations. We then construct a two-level neural network, where a large set of low-fidelity data are utilized in order to accelerate the construction of a high-fidelity surrogate model with a small set of high-fidelity data. We then embed the constructed high-fidelity surrogate model in the framework of Monte Carlo sampling. The proposed algorithm combines the approximation power of neural networks with the advantages of Monte Carlo sampling within a multi-fidelity framework. We present two numerical examples to demonstrate the accuracy and efficiency of the proposed method. We show that dramatic savings in computational cost may be achieved when the output predictions are desired to be accurate within small tolerances.

math.NA

Hierarchical Low-Rank Approximation of Regularized Wasserstein Distance

Sinkhorn divergence is a measure of dissimilarity between two probability measures. It is obtained through adding an entropic regularization term to Kantorovich's optimal transport problem and can hence be viewed as an entropically regularized Wasserstein distance. Given two discrete probability vectors in the $n$-simplex and supported on two bounded spaces in ${\mathbb R}^d$, we present a fast method for computing Sinkhorn divergence when the cost matrix can be decomposed into a $d$-term sum of asymptotically smooth Kronecker product factors. The method combines Sinkhorn's matrix scaling iteration with a low-rank hierarchical representation of the scaling matrices to achieve a near-linear complexity ${\mathcal O}(n \log^3 n)$. This provides a fast and easy-to-implement algorithm for computing Sinkhorn divergence, enabling its applicability to large-scale optimization problems, where the computation of classical Wasserstein metric is not feasible. We present a numerical example related to signal processing to demonstrate the applicability of quadratic Sinkhorn divergence in comparison with quadratic Wasserstein distance and to verify the accuracy and efficiency of the proposed method.

math.NA

Fuzzy-Stochastic Partial Differential Equations

We introduce and study a new class of partial differential equations (PDEs) with hybrid fuzzy-stochastic parameters, coined fuzzy-stochastic PDEs. Compared to purely stochastic PDEs or purely fuzzy PDEs, fuzzy-stochastic PDEs offer powerful models for accurate representation and propagation of hybrid aleatoric-epistemic uncertainties inevitable in many real-world problems. We will use the level-set representation of fuzzy functions and define the solution to fuzzy-stochastic PDE problems through a corresponding parametric problem, and further present theoretical results on the well-posedness and regularity of such problems. We also propose a numerical strategy for computing output fuzzy-stochastic quantities, such as fuzzy failure probabilities and fuzzy probability distributions. We present two numerical examples to compute various fuzzy-stochastic quantities and to demonstrate the applicability of fuzzy-stochastic PDEs to complex engineering problems.

math.AP

Wasserstein metric-driven Bayesian inversion with applications to signal processing

We present a Bayesian framework based on a new exponential likelihood function driven by the quadratic Wasserstien metric. Compared to conventional Bayesian models based on Gaussian likelihood functions driven by the least-squares norm ($L_2$ norm), the new framework features several advantages. First, the new framework does not rely on the likelihood of the measurement noise and hence can treat complicated noise structures such as combined additive and multiplicative noise. Secondly, unlike the normal likelihood function, the Wasserstein-based exponential likelihood function does not usually generate multiple local extrema. As a result, the new framework features better convergence to correct posteriors when a Markov Chain Monte Carlo sampling algorithm is employed. Thirdly, in the particular case of signal processing problems, while a normal likelihood function measures only the amplitude differences between the observed and simulated signals, the new likelihood function can capture both the amplitude and the phase differences. We apply the new framework to a class of signal processing problems, that is, the inverse uncertainty quantification of waveforms, and demonstrate its advantages compared to Bayesian models with normal likelihood functions.

math.NA

A Fuzzy-Stochastic Multiscale Model for Fiber Composites: A one-dimensional study

We study mathematical and computational models for computing the deformation of fiber-reinforced cross-plied laminates due to external forces. This requires an understanding of both micro-structural effects and different sources of uncertainty in the problem. We first show that the uncertainties in the problem are of both statistical (aleatoric) and systematic (epistemic) types and that current multiscale stochastic models, such as stationary random fields, which are based on precise probability theory, are not capable of correctly characterizing uncertainty in fiber composites. Next, we motivate the applicability of models based on imprecise uncertainty theory and present a novel fuzzy-stochastic model, which can more accurately describe uncertainties in fiber composites. The new model is constructed by combining stochastic fields and fuzzy variables through a simple calibration-validation approach. Finally, we construct a global-local multiscale algorithm for efficiently computing output quantities of interest. The method aims at approximating required quantities, such as displacements and stresses, in regions of relatively small size, e.g. hot spots or zones. The algorithm uses the concept of representative volume elements and computes a global solution to construct a local approximation that captures the microscale features of the solution. The results are based on and backed by real experimental data.

math.NA

A Sparse Stochastic Collocation Technique for High-Frequency Wave Propagation with Uncertainty

We consider the wave equation with highly oscillatory initial data, where there is uncertainty in the wave speed, initial phase and/or initial amplitude. To estimate quantities of interest related to the solution and their statistics, we combine a high-frequency method based on Gaussian beams with sparse stochastic collocation. Although the wave solution, $u^\varepsilon$, is highly oscillatory in both physical and stochastic spaces, we provide theoretical arguments and numerical evidence that quantities of interest based on local averages of $|u^\varepsilon|^2$ are smooth, with derivatives in the stochastic space uniformly bounded in $\varepsilon$, where $\varepsilon$ denotes the short wavelength. This observable related regularity makes the sparse stochastic collocation approach more efficient than Monte Carlo methods. We present numerical tests that demonstrate this advantage.

math.NA

Fast Bayesian Optimal Experimental Design for Seismic Source Inversion

We develop a fast method for optimally designing experiments in the context of statistical seismic source inversion. In particular, we efficiently compute the optimal number and locations of the receivers or seismographs. The seismic source is modeled by a point moment tensor multiplied by a time-dependent function. The parameters include the source location, moment tensor components, and start time and frequency in the time function. The forward problem is modeled by elastodynamic wave equations. We show that the Hessian of the cost functional, which is usually defined as the square of the weighted L2 norm of the difference between the experimental data and the simulated data, is proportional to the measurement time and the number of receivers. Consequently, the posterior distribution of the parameters, in a Bayesian setting, concentrates around the "true" parameters, and we can employ Laplace approximation and speed up the estimation of the expected Kullback-Leibler divergence (expected information gain), the optimality criterion in the experimental design procedure. Since the source parameters span several magnitudes, we use a scaling matrix for efficient control of the condition number of the original Hessian matrix. We use a second-order accurate finite difference method to compute the Hessian matrix and either sparse quadrature or Monte Carlo sampling to carry out numerical integration. We demonstrate the efficiency, accuracy, and applicability of our method on a two-dimensional seismic source inversion problem.

stat.CO

Taylor Expansion and Discretization Errors in Gaussian Beam Superposition

The Gaussian beam superposition method is an asymptotic method for computing high frequency wave fields in smoothly varying inhomogeneous media. In this paper we study the accuracy of the Gaussian beam superposition method and derive error estimates related to the discretization of the superposition integral and the Taylor expansion of the phase and amplitude off the center of the beam. We show that in the case of odd order beams, the error is smaller than a simple analysis would indicate because of error cancellation effects between the beams. Since the cancellation happens only when odd order beams are used, there is no remarkable gain in using even order beams. Moreover, applying the error estimate to the problem with constant speed of propagation, we show that in this case the local beam width is not a good indicator of accuracy, and there is no direct relation between the error and the beam width. We present numerical examples to verify the error estimates.

math.NA

Finite difference schemes for second order systems describing black holes

In the harmonic description of general relativity, the principle part of Einstein's equations reduces to 10 curved space wave equations for the componenets of the space-time metric. We present theorems regarding the stability of several evolution-boundary algorithms for such equations when treated in second order differential form. The theorems apply to a model black hole space-time consisting of a spacelike inner boundary excising the singularity, a timelike outer boundary and a horizon in between. These algorithms are implemented as stable, convergent numerical codes and their performance is compared in a 2-dimensional excision problem.

gr-qc