SearcharxivSearch

arXiv subjects

Yunan Yang

Publications and source records attributed to Yunan Yang.

At least 19 recordsLinked to original sources

Two Adjoint Perspectives on Fokker-Planck Optimization: A Microscopic-Macroscopic Correspondence

The Fokker-Planck equation admits both a macroscopic Eulerian description through probability densities and a microscopic Lagrangian description through stochastic trajectories. Consequently, optimization problems constrained by the Fokker-Planck equation can be formulated from either perspective. Surprisingly, the corresponding adjoint equations appear to be fundamentally different: the macroscopic adjoint is governed by the backward Kolmogorov equation, whereas the microscopic adjoint evolves pathwise along stochastic trajectories. In this note, we reconcile these two formulations by establishing their correspondence in the continuum setting. We further show that, although their discrete gradients no longer coincide after discretization, both provide consistent numerical approximations of the continuum gradient. Explicit convergence rates are established for both discretization strategies.

math.NA

Reduced order model for parametric Boltzmann equation and its application to inverse problems

The Boltzmann equation plays an important role in modeling mesoscopic behavior in a wide range of scientific and engineering applications. However, its numerical solution is computationally expensive due to the high dimensionality of the model and the nonlinear nonlocal collision operator, especially for steady-state problems that require iterative solvers. This cost becomes prohibitive for inverse problems, where the induced optimization problem requires repeated forward solves. In this work, we propose a reduced-order model (ROM) for the parametric Boltzmann equation to address this computational challenge. The ROM constructs a low-dimensional approximation space for the parameter-induced solution manifold through a residual-based greedy strategy, and the reduced solution is then obtained via residual minimization over the reduced space, subject to mass conservation. The overall efficiency of the ROM is achieved by exploiting the quadratic structure of the collision operator and a precomputed separable approximation of the collision kernel. The resulting ROM is further applied to a thermally-driven inverse problem for reconstructing collision parameters from the observed macroscopic temperature data. This is accomplished either by directly replacing the PDE constraint with the ROM, leading to a bilevel optimization formulation, or by reformulating the task as a single-level optimization problem through the Karush--Kuhn--Tucker (KKT) conditions. Numerical experiments in both collision-dominated and transport-dominated cases are performed to demonstrate the efficiency and accuracy of the proposed ROM and its effectiveness in inverse problems. In particular, the resulting inverse problem is computationally much more tractable, achieving speedups of several orders of magnitude over that based on the full-order model while maintaining comparable accuracy.

math.NA

Structure-Preserving Reduced-Order Modeling via Low-Rank Transport Signatures

Parametrized PDEs with density-valued solutions are often difficult to approximate with classical linear reduced-order models, especially in transport-dominated regimes. We introduce an optimal-transport-based reduced-order modeling that represents each density by the Kantorovich potential transporting a fixed reference density to the target density, and then maps these potentials to transport signatures using a weighted Laplacian associated with the reference measure. This embeds the density-valued solution map in a Hilbert space while preserving control of the induced transport maps and Wasserstein error. We treat the signature map as a continuous matrix indexed by parameters and space, construct a low-rank skeleton decomposition using a maximal-volume criterion, and learn the parameter-to-coefficient map with a neural network for efficient non-intrusive online evaluation. The reconstructed solution is obtained by pushing forward the reference density, so mass preservation is built into the method. We prove a mean-squared Wasserstein error bound separating low-rank approximation, discretization, sampling, and learning errors, and demonstrate the method on a two-dimensional continuity equation, where transport signatures yield substantially lower-rank structure than the original density snapshots.

math.NA

Data-Adaptive Learning of Dynamical Systems by Matching Transfer Operators and Invariant Measures

Trajectory-based learning of dynamical systems is often fragile in the presence of noise, chaos, or sparse observations, as small pointwise errors can rapidly amplify. We introduce a transition-statistics approach to system identification that learns dynamics by matching the induced motion of probability mass across a data-adaptive mesh. Given trajectory data, we build an unstructured partition of state space and approximate the Perron--Frobenius operator with a regularized Ulam transition matrix. We replace hard cell indicators with continuous, piecewise-smooth partition-of-unity weights, yielding a Markov matrix supporting gradient-based optimization with respect to the parameters of a learned vector field. This enables two related training objectives: matching invariant measures through the stationary eigenvectors of the transition matrices, and matching the full transition matrices to capture transport between regions of state space. Numerical experiments on Lorenz-63, Lorenz-96, and a reduced-order NOAA sea surface temperature forecasting problem show that transition-statistics matching gives more reliable long-time dynamics than pointwise trajectory matching, particularly under measurement noise and sparse sampling. The approach provides a robust operator-theoretic alternative to trajectory-level losses for learning chaotic and partially observed dynamical systems.

math.NA

Sampling Using Hybrid Stochastic Dynamics

This work proposes a framework for sampling from the Gibbs distribution of a given potential using hybrid stochastic dynamics. In this framework, two distinct sampling dynamics are run in different regions of the state space. The two dynamics are coupled across the interface through natural transmission conditions that preserve the target distribution. Using a specially constructed regularization scheme, we establish an exponential rate of convergence for the hybrid dynamics to equilibrium. We also analyze the metastability properties of the hybrid dynamics in a radially symmetric landscape, showing that the hybrid scheme can improve the mean exit time. This advantage is further confirmed by the numerical experiments.

math.NA

Well-Posedness and Efficient Algorithms for Inverse Optimal Transport with Bregman Regularization

This work analyzes the inverse optimal transport (IOT) problem under Bregman regularization. We establish well-posedness results, including existence, uniqueness (up to equivalence classes of solutions), and stability, under several structural assumptions on the cost matrix. On the computational side, we investigate the existence of solutions to the optimization problem with general constraints on the cost matrix and provide a sufficient condition guaranteeing existence. In addition, we propose an inexact block coordinate descent (BCD) method for the problem with a strongly convex penalty term. In particular, when the penalty is quadratic, the subproblems admit a diagonal Hessian structure, which enables highly efficient element-wise Newton updates. We establish a linear convergence rate for the algorithm and demonstrate its practical performance through numerical experiments, including the validation of stability bounds, the investigation of regularization effects, and the application to a marriage matching dataset.

math.OC

On the Unique Recovery of Transport Maps and Vector Fields from Finite Measure-Valued Data

We establish guarantees for the unique recovery of vector fields and transport maps from finite measure-valued data, yielding new insights into generative models, data-driven dynamical systems, and PDE inverse problems. In particular, we provide general conditions under which a diffeomorphism can be uniquely identified from its pushforward action on finitely many densities, i.e., when the data $\{(ρ_j,f_\#ρ_j)\}_{j=1}^m$ uniquely determines $f$. As a corollary, we introduce a new metric which compares diffeomorphisms by measuring the discrepancy between finitely many pushforward densities in the space of probability measures. We also prove analogous results in an infinitesimal setting, where derivatives of the densities along a smooth vector field are observed, i.e., when $\{(ρ_j,\text{div} (ρ_j v))\}_{j=1}^m$ uniquely determines $v$. Our analysis makes use of the Whitney and Takens embedding theorems, which provide estimates on the required number of densities $m$, depending only on the intrinsic dimension of the problem. We additionally interpret our results through the lens of Perron--Frobenius and Koopman operators and demonstrate how our techniques lead to new guarantees for the well-posedness of certain PDE inverse problems related to continuity, advection, Fokker--Planck, and advection-diffusion-reaction equations. Finally, we present illustrative numerical experiments demonstrating the unique identification of transport maps from finitely many pushforward densities, and of vector fields from finitely many weighted divergence observations.

stat.ML

Inference of interacting kernel in the mean-field regime

We study the problem of reconstructing interaction kernels in systems of interacting agents from macroscopic measurements when posed as an optimization problem. The reconstruction procedure depends on the formulation of the forward model, which may be given either by a finite-dimensional coupled ODE system tracking individual agent trajectories or by a mean-field PDE describing the evolution of the agent density. We investigate the similarities and differences between these two formulations in the mean-field regime. While the first variation derived from the particle system does not provide an unbiased estimator of the first variation associated with the limiting PDE, we prove that, under mild assumptions, the two are close in a weak sense with a convergence rate $\mathcal{O}(N^{-1/2})$. This rate is further confirmed by numerical evidences.

math.NA

Adjoint DSMC Method for Spatially Inhomogeneous Boltzmann Equation with General Boundary Conditions

We develop adjoint Direct Simulation Monte Carlo (DSMC) formulations for the spatially inhomogeneous Boltzmann equation with periodic, specular reflecting, diffuse thermal, and prescribed inflow boundary conditions. Periodic and specular boundaries are treated using a pathwise particle adjoint conditional on the realized event history. For diffuse thermal boundaries, we introduce a randomized-time regularization of wall-crossing events and use score-function terms to differentiate the resulting boundary probabilities. Reparameterization of the outgoing half-Maxwellian samples provides sensitivities with respect to wall temperatures and tangential wall velocities. Prescribed inflow requires a different construction because perturbations of incoming particles affect subsequent cell populations, local collision frequencies, and collision schedules. We therefore derive an ensemble adjoint based on the locally linearized Boltzmann collision operator and evaluate boundary sensitivities using local inflow-source scores, with injection counts held fixed. For a scalar objective, the dominant adjoint cost is largely independent of the number of parameters. Numerical experiments for Maxwell molecules validate the formulations against centered finite differences for thermal, mixed thermal-specular, two-sided inflow, and high-Mach Couette-flow configurations. The results demonstrate consistent gradient estimates, Monte Carlo convergence, stability with respect to the thermal regularization parameter, and accurate sensitivity calculation in a regime with limited relative statistical noise.

math.NA

Velocity Reconstruction from Flow-Induced Magnetic Fields

We study the inverse problem of reconstructing an incompressible velocity field $\boldsymbol{v}$ from observations of the induced magnetic field $\boldsymbol{b}$. In the presence of a strong, constant background field $\mathbf{F}$, the evolution of the magnetic perturbation $\boldsymbol{b}$ is governed by the linearized induction equation. We analyze the system on both the entire space $\Omega = \mathbb{R}^d$ and a periodic domain $\Omega = \prod_{i=1}^d [0, L_i)$, which models a homogeneous medium with side lengths $L_i > 0$. We analyze this problem by decomposing it into the injectivity of a parabolic forward map and the solvability of a divergence-free transport sub-problem. On the whole space $\mathbb{R}^d$, we show that the transport sub-problem is well-posed when data is prescribed on a non-characteristic hypersurface transverse to $\mathbf{F}$. On the torus, we establish a sharp uniqueness criterion based on the rational dependence of the ratios $\{F_i/L_i\}_{i=1}^d$ between the background-field components and the corresponding domain periods. Furthermore, we show that for the reconstructed velocity to belong to $L^2$, a sufficient condition is that the background field must satisfy a Diophantine condition. The proof combines injectivity of the parabolic forward map with uniqueness for a steady transport equation along $\mathbf{F}$.

math.AP

Learning where to learn: Training data distribution optimization for scientific machine learning

In scientific machine learning, models are routinely deployed with parameter values or boundary conditions far from those used in training. This paper studies the learning-where-to-learn problem of designing a training data distribution that minimizes average prediction error across a family of deployment regimes. A theoretical analysis shows how the training distribution shapes deployment accuracy. This motivates two adaptive algorithms based on bilevel or alternating optimization in the space of probability measures. Discretized implementations using parametric distribution classes or nonparametric particle-based gradient flows deliver optimized training distributions that outperform nonadaptive designs. Once trained, the resulting models exhibit improved sample complexity and robustness to distribution shift. This framework unlocks the potential of principled data acquisition for learning functions and solution operators of partial differential equations.

cs.LG

Operator learning meets inverse problems: A probabilistic perspective

Operator learning offers a robust framework for approximating mappings between infinite-dimensional function spaces. It has also become a powerful tool for solving inverse problems in the computational sciences. This chapter surveys methodological and theoretical developments at the intersection of operator learning and inverse problems. It begins by summarizing the probabilistic and deterministic approaches to inverse problems, and pays special attention to emerging measure-centric formulations that treat observed data or unknown parameters as probability distributions. The discussion then turns to operator learning by covering essential components such as data generation, loss functions, and widely used architectures for representing function-to-function maps. The core of the chapter centers on the end-to-end inverse operator learning paradigm, which aims to directly map observed data to the solution of the inverse problem without requiring explicit knowledge of the forward map. It highlights the unique challenge that noise plays in this data-driven inversion setting, presents structure-aware architectures for both point predictions and posterior estimates, and surveys relevant theory for linear and nonlinear inverse problems. The chapter also discusses the estimation of priors and regularizers, where operator learning is used more selectively within classical inversion algorithms.

math.NA

Measure-Theoretic Time-Delay Embedding

The celebrated Takens' embedding theorem provides a theoretical foundation for reconstructing the full state of a dynamical system from partial observations. However, the classical theorem assumes that the underlying system is deterministic and that observations are noise-free, limiting its applicability in real-world scenarios. Motivated by these limitations, we formulate a measure-theoretic generalization that adopts an Eulerian description of the dynamics and recasts the embedding as a pushforward map between spaces of probability measures. Our mathematical results leverage recent advances in optimal transport. Building on the proposed measure-theoretic time-delay embedding theory, we develop a computational procedure that aims to reconstruct the full state of a dynamical system from time-lagged partial observations, engineered with robustness to handle sparse and noisy data. We evaluate our measure-based approach across several numerical examples, ranging from the classic Lorenz-63 system to real-world applications such as NOAA sea surface temperature reconstruction and ERA5 wind field reconstruction.

math.DS

Robust Frequency Domain Full-Waveform Inversion via HV-Geometry

Conventional frequency-domain full-waveform inversion (FWI) is typically implemented with an $L^2$ misfit function, which suffers from challenges such as cycle skipping and sensitivity to noise. While the Wasserstein metric has proven effective in addressing these issues in time-domain FWI, its applicability in frequency-domain FWI is limited due to the complex-valued nature of the data and reduced transport-like dependency on wave speed. To mitigate these challenges, we introduce the HV metric ($d_{\text{HV}}$), inspired by optimal transport theory, which compares signals based on horizontal and vertical changes without requiring the normalization of data. We implement $d_{\text{HV}}$ as the misfit function in frequency-domain FWI and evaluate its performance on synthetic and real-world datasets from seismic imaging and ultrasound computed tomography (USCT). Numerical experiments demonstrate that $d_{\text{HV}}$ outperforms the $L^2$ and Wasserstein metrics in scenarios with limited prior model information and high noise while robustly improving inversion results on clinical USCT data.

math.OC

Invariant Measures in Time-Delay Coordinates for Unique Dynamical System Identification

While invariant measures are widely employed to analyze physical systems when a direct study of pointwise trajectories is intractable, e.g., due to chaos or noise, they cannot uniquely identify the underlying dynamics. Our first result shows that, in contrast to invariant measures in state coordinates, e.g., $[x(t), y(t), z(t)]$, the invariant measure expressed in time-delay coordinates, e.g., $[x(t), x(t-τ),\ldots, x(t-(m-1)τ)]$, can identify the dynamics up to a topological conjugacy. Our second result resolves the remaining ambiguity: by combining invariant measures constructed from multiple delay frames with distinct observables, the system is uniquely identifiable, provided that a suitable initial condition is satisfied. These guarantees require informative observables and appropriate delay parameters ($m,τ$), which can be limiting in certain settings. We support our theoretical contributions through a series of physical examples demonstrating how invariant measures expressed in delay-coordinates can be used to perform robust system identification in practice.

math.DS

Cryo-EM as a Stochastic Inverse Problem

Cryo-electron microscopy (Cryo-EM) enables high-resolution imaging of biomolecules, but structural heterogeneity remains a major challenge in 3D reconstruction. Traditional methods assume a discrete set of conformations, limiting their ability to recover continuous structural variability. In this work, we formulate cryo-EM reconstruction as a stochastic inverse problem (SIP) over probability measures, where the observed images are modeled as the push-forward of an unknown distribution over molecular structures via a random forward operator. We pose the reconstruction problem as the minimization of a variational discrepancy between observed and simulated image distributions, using statistical distances such as the KL divergence and the Maximum Mean Discrepancy. The resulting optimization is performed over the space of probability measures via a Wasserstein gradient flow, which we numerically solve using particles to represent and evolve conformational ensembles. We validate our approach using synthetic examples, including a realistic protein model, which demonstrates its ability to recover continuous distributions over structural states. We analyze the connection between our formulation and Maximum A Posteriori (MAP) approaches, which can be interpreted as instances of the discretize-then-optimize (DTO) framework. We further provide a consistency analysis, establishing conditions under which DTO methods, such as MAP estimation, converge to the solution of the underlying infinite-dimensional continuous problem. Beyond cryo-EM, the framework provides a general methodology for solving SIPs involving random forward operators.

stat.ML

HV Metric For Time-Domain Full Waveform Inversion

Full-waveform inversion (FWI) is a powerful technique for reconstructing high-resolution material parameters from seismic or ultrasound data. The conventional least-squares (\(L^{2}\)) misfit suffers from pronounced non-convexity that leads to \emph{cycle skipping}. Optimal-transport misfits, such as the Wasserstein distance, alleviate this issue; however, their use requires artificially converting the wavefields into probability measures, a preprocessing step that can modify critical amplitude and phase information of time-dependent wave data. We propose the \emph{HV metric}, a transport-based distance that acts naturally on signed signals, as an alternative metric for the \(L^{2}\) and Wasserstein objectives in time-domain FWI. After reviewing the metric's definition and its relationship to optimal transport, we derive closed-form expressions for the Fréchet derivative and Hessian of the map \(f \mapsto d_{\text{HV}}^2(f,g)\), enabling efficient adjoint-state implementations. A spectral analysis of the Hessian shows that, by tuning the hyperparameters \((κ,λ,ε)\), the HV misfit seamlessly interpolates between \(L^{2}\), \(H^{-1}\), and \(H^{-2}\) norms, offering a tunable trade-off between the local point-wise matching and the global transport-based matching. Synthetic experiments on the Marmousi and BP benchmark models demonstrate that the HV metric-based objective function yields faster convergence and superior tolerance to poor initial models compared to both \(L^{2}\) and Wasserstein misfits. These results demonstrate the HV metric as a robust, geometry-preserving alternative for large-scale waveform inversion.

math.OC

Multi-Window Approaches for Direct and Stable STFT Phase Retrieval

Phase retrieval from phaseless short-time Fourier transform (STFT) measurements is known to be inherently unstable when measurements are taken with respect to a single window. While an explicit inversion formula exists, it is useless in practice due to its instability. In this paper, we overcome this lack of stability by presenting two multi-window approaches that rely on a "good coverage" of the time-frequency plane by the ambiguity functions of the windows. The first is to use the fractional Fourier transform of a dilated Gauss function with various angles as window functions. The essential support of a superposition of the ambiguity function from such window functions is of a "daffodil shape", which converges to a large disc as more angles are used, yielding a much broader coverage in the time-frequency domain. The second approach uses Hermite functions of various degrees as the window functions. The larger the degree, the wider the ambiguity function but with zeros on circles in the time-frequency domain. Combining Hermite functions of different degrees, we can achieve a wide coverage with zeros compensated by the essential support of the ambiguity function from other Hermite windows. Taking advantage of these multi-window procedures, we can stably perform STFT phase retrieval using the direct inversion formula.

math.FA