SearcharxivSearch

arXiv subjects

Yimin Zhong

Publications and source records attributed to Yimin Zhong.

At least 19 recordsLinked to original sources

A Fourier approach to Gromov's filling area conjecture

We prove that every compact connected Riemannian isometric filling $M$ of a circle of length $2π$ satisfies $\operatorname{Area}(M) \geq \frac{14ζ(3)}π \approx 5.35677$, regardless of orientability or topological types. Our new approach uses the odd Fourier coefficients of the distance functions from boundary points. For orientable fillings, we use a cubic resonant perturbation to obtain $\operatorname{Area}(M)>5.40154$.

math.DG

A Multi-Level Machine Learning Framework for Inverse Scattering Problems with Multi-Frequency Data

In this work, we propose a multi-level machine learning framework for solving inverse scattering problems with multi-frequency data. The multi-level neural network is built along the frequency axis of the scattering problem, wherein at each fixed frequency, a new level of network is added to the existing architecture to update the reconstruction. By marching through the frequency levels, the proposed multi-level computational framework is able to obtain higher-order Fourier modes of the imaging target as the depth of the neural network grows and higher-frequency data are used. Furthermore, the overall learning problem is decomposed into a sequence of simpler local tasks, each associated with a single frequency. This decomposition significantly reduces the complexity of the optimization problem and mitigates the risk of convergence to undesirable local minima, resulting in a robust and reliable training procedure for solving inverse scattering problems. We conduct various numerical experiments for the inverse source scattering problem and the inverse medium scattering problem to illustrate the effectiveness and robustness of the proposed machine learning framework. In addition, theoretical analysis in the neural tangent kernel regime shows that the proposed multi-level architecture progressively recovers the higher-order Fourier components of the imaging target.

math.NA

Fourier Multi-Component and Multi-Layer Neural Networks: Unlocking High-Frequency Potential

The architecture of a neural network and the choice of its activation function are both fundamental to its performance. Equally important is ensuring that these two elements are well matched, as their alignment is key to effective representation and learning. In this paper, we introduce the Fourier Multi-Component and Multi-Layer Neural Network (FMMNN), a model that combines sine-type activations with the multi-component and multi-layer structure of MMNNs. In an FMMNN, each component is represented as a trainable linear combination of fixed random sine-type basis functions, while multi-layer composition generates more complex and adaptive high-frequency features. We establish that FMMNNs retain exponential expressive power for function approximation even under a low-rank architectural structure. We also analyze the optimization landscape of FMMNNs and find it to be substantially more favorable than that of standard fully connected neural networks, especially for high-frequency targets. In addition, we propose a scaled random initialization method for the first-layer weights in FMMNNs, which accelerates training and improves final performance when sufficient samples are available. Extensive numerical experiments support our theoretical insights, showing that FMMNNs achieve strong accuracy and favorable convergence behavior on oscillatory function-approximation benchmarks.

cs.LG

Forward and inverse problems of a semilinear transport equation

We study forward and inverse problems for a semilinear radiative transport model where the absorption coefficient depends on the angular average of the transport solution. Our first result is the well-posedness theory for the transport model with general boundary data, which significantly improves previous theories for small boundary data. For the inverse problem of reconstructing the nonlinear absorption coefficient from internal data, we develop stability results for the reconstructions and unify an $L^1$ stability theory for both the diffusion and transport regimes by introducing a weighted norm that penalizes the contribution from the boundary region. The problems studied here are motivated by applications such as photoacoustic imaging of multi-photon absorption of heterogeneous media.

math.AP

From Frequency Bias to Spectral Balance: Operator-Aware Preconditioners for PINNs

When neural networks (NNs) are used as a type of nonlinear parametric representation to solve partial differential equations (PDEs), they often display frequency-dependent learning dynamics that can differ from those seen in direct function approximation tasks, resulting from a balance between the frequency bias of the NN representation and that of the underlying differential operator. Although many commonly used NNs exhibit a bias towards low-frequency modes in representation, the presence of differential operators in the loss function, which amplifies high-frequency components, can lead to high frequency bias. In this work, using second order elliptic PDEs as an example, we show how these two factors compete and lead to an overall frequency bias in different situations. Once the balance is determined, it is important to design computational strategies to counter the resulting bias to improve training efficiency. We propose a simple operator-aware preconditioning strategy that rebalances the optimization landscape and the learning dynamics by applying an auxiliary integral operator to the residual. The integral kernel can be the Green's function of a reference elliptic operator or an approximation, and integrates easily with common NN solvers for PDEs. Extensive experiments, including multiscale and variable-coefficient problems, show that the approach restores more balanced learning dynamics across modes and substantially improves both convergency and accuracy.

math.NA

What Can One Expect When Solving PDEs Using Shallow Neural Networks?

We use elliptic partial differential equations (PDEs) as examples to show various properties and behaviors when shallow neural networks (SNNs) are used to represent the solutions. In particular, we study the numerical ill-conditioning, frequency bias, and the balance between the differential operator and the shallow network representation for different formulations of the PDEs and with various activation functions. Our study shows that the performance of Physics-Informed Neural Networks (PINNs) or Deep Ritz Method (DRM) using linear SNNs with power ReLU activation is dominated by their inherent ill-conditioning and spectral bias against high frequencies. Although this can be alleviated by using non-homogeneous activation functions with proper scaling, achieving such adaptivity for nonlinear SNNs remains costly due to ill-conditioning.

math.NA

Robustness of data-driven approaches in limited angle tomography

The limited angle Radon transform is notoriously difficult to invert due to its ill-posedness. In this work, we give a mathematical explanation that data-driven approaches can stably reconstruct more information compared to traditional methods like filtered backprojection. In addition, we use experiments based on the U-Net neural network to validate our theory.

math.NA

Structured and Balanced Multi-Component and Multi-Layer Neural Networks

In this work, we propose a balanced multi-component and multi-layer neural network (MMNN) structure to accurately and efficiently approximate functions with complex features, in terms of both degrees of freedom and computational cost. The main idea is inspired by a multi-component approach, in which each component can be effectively approximated by a single-layer network, combined with a multi-layer decomposition strategy to capture the complexity of the target function. Although MMNNs can be viewed as a simple modification of fully connected neural networks (FCNNs) or multi-layer perceptrons (MLPs) by introducing balanced multi-component structures, they achieve a significant reduction in training parameters, a much more efficient training process, and improved accuracy compared to FCNNs or MLPs. Extensive numerical experiments demonstrate the effectiveness of MMNNs in approximating highly oscillatory functions and their ability to automatically adapt to localized features.

cs.LG

Why Shallow Networks Struggle to Approximate and Learn High Frequencies

In this work, we present a comprehensive study combining mathematical and computational analysis to explain why a two-layer neural network struggles to handle high frequencies in both approximation and learning, especially when machine precision, numerical noise, and computational cost are significant factors in practice. Specifically, we investigate the following fundamental computational issues: (1) the minimal numerical error achievable under finite precision, (2) the computational cost required to attain a given accuracy, and (3) the stability of the method with respect to perturbations. The core of our analysis lies in the conditioning of the representation and its learning dynamics. Explicit answers to these questions are provided, along with supporting numerical evidence.

cs.LG

Transport models for wave propagation in scattering media with nonlinear absorption

This work considers the propagation of high-frequency waves in highly-scattering media where physical absorption of a nonlinear nature occurs. Using the classical tools of the Wigner transform and multiscale analysis, we derive semilinear radiative transport models for the phase-space intensity and the diffusive limits of such transport models. As an application, we consider an inverse problem for the semilinear transport equation, where we reconstruct the absorption coefficients of the equation from a functional of its solution. We obtain a uniqueness result on the inverse problem.

math.AP

Error Analysis for the Implicit Boundary Integral Method

The implicit boundary integral method (IBIM) provides a framework to construct quadrature rules on regular lattices for integrals over irregular domain boundaries. This work provides a systematic error analysis for IBIMs on uniform Cartesian grids for boundaries with different degree of regularities. We first show that the quadrature error gains an addition order of $\frac{d-1}{2}$ from the curvature for a strongly convex smooth boundary due to the ``randomness'' in the signed distances. This gain is discounted for degenerated convex surfaces. We then extend the error estimate to general boundaries under some special circumstances, including how quadrature error depends on the boundary's local geometry relative to the underlying grid. Bounds on the variance of the quadrature error under random shifts and rotations of the lattices are also derived.

math.NA

How much can one learn a partial differential equation from its solution?

In this work we study the problem about learning a partial differential equation (PDE) from its solution data. PDEs of various types are used as examples to illustrate how much the solution data can reveal the PDE operator depending on the underlying operator and initial data. A data driven and data adaptive approach based on local regression and global consistency is proposed for stable PDE identification. Numerical experiments are provided to verify our analysis and demonstrate the performance of the proposed algorithms.

math.NA

Corrected Trapezoidal Rule-IBIM for linearized Poisson-Boltzmann equation

In this paper, we solve the linearized Poisson-Boltzmann equation, used to model the electric potential of macromolecules in a solvent. We derive a corrected trapezoidal rule with improved accuracy for a boundary integral formulation of the linearized Poisson-Boltzmann equation. More specifically, in contrast to the typical boundary integral formulations, the corrected trapezoidal rule is applied to integrate a system of compacted supported singular integrals using uniform Cartesian grids in $\mathbb{R}^3$, without explicit surface parameterization. A Krylov method, accelerated by a fast multipole method, is used to invert the resulting linear system. We study the efficacy of the proposed method, and compare it to an existing, lower order method. We then apply the method to the computation of electrostatic potential of macromolecules immersed in solvent. The solvent excluded surfaces, defined by a common approach, are merely piecewise smooth, and we study the effectiveness of the method for such surfaces.

math.NA

Instability of an inverse problem for the stationary radiative transport near the diffusion limit

In this work, we study the instability of an inverse problem of radiative transport equation with angularly averaged measurement near the diffusion limit, i.e. the normalized mean free path (the Knudsen number) $0 < \eps \ll 1$. It is well-known that there is a transition of stability from Hölder type to logarithmic type with $\eps\to 0$, the theory of this transition of stability is still an open problem. In this study, we show the transition of stability by establishing the balance of two different regimes depending on the relative sizes of $\eps$ and the perturbation in measurements. When $\eps$ is sufficiently small, we obtain exponential instability, which stands for the diffusive regime, and otherwise we obtain Hölder instability instead, which stands for the transport regime.

math-ph

How much can one learn from a single solution of a PDE?

Linear evolution PDE $\partial_t u(x,t) = -\mathcal{L} u$, where $\mathcal{L}$ is a strongly elliptic operator independent of time, is studied as an example to show if one can superpose snapshots of a single (or a finite number of) solution(s) to construct an arbitrary solution. Our study shows that it depends on the growth rate of the eigenvalues, $μ_n$, of $\mathcal{L}$ in terms of $n$. When the statement is true, a simple data-driven approach for model reduction and approximation of an arbitrary solution of a PDE without knowing the underlying PDE is designed. Numerical experiments are presented to corroborate our analysis.

math.NA

Inverse Source Problem for Acoustically-Modulated Electromagnetic Waves

We propose a method to reconstruct the electrical current density from acoustically-modulated boundary measurements of time-harmonic electromagnetic fields. We show that the current can be uniquely reconstructed with Lipschitz stability. We also report numerical simulations to illustrate the analytical results.

math.AP

Inverse Boundary Problem for the Two Photon Absorption Transport Equation

This work studies the inverse boundary problem for the two photon absorption radiative transport equation. We show that the absorption coefficients and scattering coefficients can be uniquely determined from the \emph{albedo} operator. If scattering is absent, we do not require smallness of the incoming source and the reconstructions of the absorption coefficients are explicit.

math.AP