SearcharxivSearch

arXiv subjects

Xingjie Li

Publications and source records attributed to Xingjie Li.

13 recordsLinked to original sources

The Test and Find Model

We introduce the Test and Find (TF) problem, where a decision maker (DM) faces the following situation: $k$ identical objects are randomly allocated to $n$ distinct boxes (sites) according to some distribution $\pi$, with no more than one object to a box. The DM tests all of the boxes. However, the tests are imperfect: they can give false positive or false negative results. DM has $m$ tags, $1\leq m\leq n$, and, after testing all boxes, she can place a tag on any box that she thinks has a hidden object. She is rewarded $c_i$ for a correct guess and penalized $d_i$ for a wrong guess in box $i$. DM knows all of the parameters of the model and her goal is to maximize the expected reward. We give an explicit solution to this problem. We then turn to the symmetric case, for which we derive more computationally efficient results. We also consider several extensions of the TF model and give detailed solutions. One of these extensions is to the realistic case where $k$, the number of objects, is unknown and random.

math.PR

FlashSinkhorn: IO-Aware Entropic Optimal Transport on GPU

Entropic optimal transport (EOT) via Sinkhorn iterations is widely used in modern machine learning, yet GPU solvers remain inefficient at scale. Tensorized implementations suffer quadratic HBM traffic from dense $n\times m$ interactions, while existing online backends avoid storing dense matrices but still rely on generic tiled map-reduce reduction kernels with limited fusion. We present \textbf{FlashSinkhorn}, an IO-aware EOT solver for squared Euclidean cost that rewrites stabilized log-domain Sinkhorn updates as row-wise LogSumExp reductions of biased dot-product scores, the same normalization as transformer attention. This enables FlashAttention-style fusion and tiling: fused Triton kernels stream tiles through on-chip SRAM and update dual potentials in a single pass, substantially reducing HBM IO per iteration while retaining linear-memory operations. We further provide streaming kernels for transport application, enabling scalable first- and second-order optimization. On A100 GPUs, FlashSinkhorn achieves up to $32\times$ forward-pass and $161\times$ end-to-end speedups over state-of-the-art online baselines on point-cloud OT, improves scalability on OT-based downstream tasks. For reproducibility, we release an open-source implementation at https://github.com/ot-triton-lab/flash-sinkhorn .

cs.LG

Robust and Active Visible-Light Integrated Photonics on Thin-Film Lithium Tantalate for Underwater Optical Wireless Communications

Visible-light integrated photonics enables compact platforms for sensing, precision metrology, and free-space data links at visible wavelengths. However, many applications remain limited by the lack of high-speed and robust modulators in the blue-green band. Here we report, both operating at 532 nm, thin-film lithium tantalate waveguides of propagation losses of dB/cm scale and modulators with a flat frequency response to ~50 GHz. The modulator remains stable when delivering 5 dBm modulated optical power for an hour, which cannot be achieved by thin-film lithium niobate based counterparts under similar conditions and structures. System-level underwater optical wireless communication (UWOC) is validated with 112-Gb/s transmission over 3-m and 64-Gb/s transmission over 9-m underwater links. This represents the first integrated external modulator based UWOC system, overcoming the bandwidth-power-chirp trade-offs of traditional directly modulated laser-based systems. We further demonstrate dual-drive modulators for optical single-sideband and electro-optic frequency-comb generations in the green-wavelength band. These results provide a foundation for complex, robust, and active visible-light photonic integrated circuits for underwater optical applications.

physics.optics

Robust First and Second-Order Differentiation for Regularized Optimal Transport

Applications such as unbalanced and fully shuffled regression can be approached by optimizing regularized optimal transport (OT) distances, such as the entropic OT and Sinkhorn distances. A common approach for this optimization is to use a first-order optimizer, which requires the gradient of the OT distance. For faster convergence, one might also resort to a second-order optimizer, which additionally requires the Hessian. The computations of these derivatives are crucial for efficient and accurate optimization. However, they present significant challenges in terms of memory consumption and numerical instability, especially for large datasets and small regularization strengths. We circumvent these issues by analytically computing the gradients for OT distances and the Hessian for the entropic OT distance, which was not previously used due to intricate tensor-wise calculations and the complex dependency on parameters within the bi-level loss function. Through analytical derivation and spectral analysis, we identify and resolve the numerical instability caused by the singularity and ill-posedness of a key linear system. Consequently, we achieve scalable and stable computation of the Hessian, enabling the implementation of the stochastic gradient descent (SGD)-Newton methods. Tests on shuffled regression examples demonstrate that the second stage of the SGD-Newton method converges orders of magnitude faster than the gradient descent-only method while achieving significantly more accurate parameter estimations.

math.NA

Modeling and Simulation of Traffic on I-485 via Linear Systems and Iterative Methods

Iterative methods such as Jacobi, Gauss-Seidel, and Successive Over-Relaxation (SOR) are fundamental tools in solving large systems of linear equations across various scientific fields, particularly in the field of data science which has become increasingly relevant in the past decade. Iterative methods' use of matrix multiplication rather than matrix inverses makes them ideal for solving large systems quickly. Our research explores the factors of each method that define their respective strengths, limitations, and convergence behaviors to understand how these methods address drawbacks encountered when performing matrix operations by hand, as well as how they can be used in real world applications. After implementing each method by hand to understand how the algorithms work, we developed a Python program that assesses a user-given matrix based on each method's specific convergence criteria. The program compares the spectral radii of all three methods and chooses to execute whichever will yield the fastest convergence rate. Our research revealed the importance of mathematical modeling and understanding specific properties of the coefficient matrix. We observed that Gauss-Seidel is usually the most efficient method because it is faster than Jacobi and doesn't have as strict requirements as SOR, however SOR is ideal in terms of computation speed. We applied the knowledge we gained to create a traffic flow model of the I-485 highway in Charlotte. After creating a program that generates the matrix for this model, we were able to iteratively approximate the flow of cars through neighboring exits using data from the N.C. Department of Transportation. This information identifies which areas are the most congested and can be used to inform future infrastructure development.

cs.DC

NySALT: Nyström-type inference-based schemes adaptive to large time-stepping

Large time-stepping is important for efficient long-time simulations of deterministic and stochastic Hamiltonian dynamical systems. Conventional structure-preserving integrators, while being successful for generic systems, have limited tolerance to time step size due to stability and accuracy constraints. We propose to use data to innovate classical integrators so that they can be adaptive to large time-stepping and are tailored to each specific system. In particular, we introduce NySALT, Nyström-type inference-based schemes adaptive to large time-stepping. The NySALT has optimal parameters for each time step learnt from data by minimizing the one-step prediction error. Thus, it is tailored for each time step size and the specific system to achieve optimal performance and tolerate large time-stepping in an adaptive fashion. We prove and numerically verify the convergence of the estimators as data size increases. Furthermore, analysis and numerical tests on the deterministic and stochastic Fermi-Pasta-Ulam (FPU) models show that NySALT enlarges the maximal admissible step size of linear stability, and quadruples the time step size of the Störmer--Verlet and the BAOAB when maintaining similar levels of accuracy.

math.NA

A Meshfree Peridynamic Model for Brittle Fracture in Randomly Heterogeneous Materials

In this work we aim to develop a unified mathematical framework and a reliable computational approach to model the brittle fracture in heterogeneous materials with variability in material microstructures, and to provide statistic metrics for quantities of interest, such as the fracture toughness. To depict the material responses and naturally describe the nucleation and growth of fractures, we consider the peridynamics model. In particular, a stochastic state-based peridynamic model is developed, where the micromechanical parameters are modeled by a finite-dimensional random vector, or a combination of random variables truncating the Karhunen-Loève decomposition or the principle component analysis (PCA). To solve this stochastic peridynamic problem, probabilistic collocation method (PCM) is employed to sample the random field representing the micromechanical parameters. For each sample, the deterministic peridynamic problem is discretized with an optimization-based meshfree quadrature rule. We present rigorous analysis for the proposed scheme and demonstrate its convergence for a number of benchmark problems, showing that it sustains the asymptotic compatibility spatially and achieves an algebraic or sub-exponential convergence rate in the random space as the number of collocation points grows. Finally, to validate the applicability of this approach on real-world fracture problems, we consider the problem of crystallization toughening in glass-ceramic materials, in which the material at the microstructural scale contains both amorphous glass and crystalline phases. The proposed stochastic peridynamic solver is employed to capture the crack initiation and growth for glass-ceramics with different crystal volume fractions, and the averaged fracture toughness are calculated. The numerical estimates of fracture toughness show good consistency with experimental measurements.

cond-mat.mtrl-sci

An asymptotically compatible probabilistic collocation method for randomly heterogeneous nonlocal problems

In this paper we present an asymptotically compatible meshfree method for solving nonlocal equations with random coefficients, describing diffusion in heterogeneous media. In particular, the random diffusivity coefficient is described by a finite-dimensional random variable or a truncated combination of random variables with the Karhunen-Loève decomposition, then a probabilistic collocation method (PCM) with sparse grids is employed to sample the stochastic process. On each sample, the deterministic nonlocal diffusion problem is discretized with an optimization-based meshfree quadrature rule. We present rigorous analysis for the proposed scheme and demonstrate convergence for a number of benchmark problems, showing that it sustains the asymptotic compatibility spatially and achieves an algebraic or sub-exponential convergence rate in the random coefficients space as the number of collocation points grows. Finally, to validate the applicability of this approach we consider a randomly heterogeneous nonlocal problem with a given spatial correlation structure, demonstrating that the proposed PCM approach achieves substantial speed-up compared to conventional Monte Carlo simulations.

math.NA

ISALT: Inference-based schemes adaptive to large time-stepping for locally Lipschitz ergodic systems

Efficient simulation of SDEs is essential in many applications, particularly for ergodic systems that demand efficient simulation of both short-time dynamics and large-time statistics. However, locally Lipschitz SDEs often require special treatments such as implicit schemes with small time-steps to accurately simulate the ergodic measure. We introduce a framework to construct inference-based schemes adaptive to large time-steps (ISALT) from data, achieving a reduction in time by several orders of magnitudes. The key is the statistical learning of an approximation to the infinite-dimensional discrete-time flow map. We explore the use of numerical schemes (such as the Euler-Maruyama, a hybrid RK4, and an implicit scheme) to derive informed basis functions, leading to a parameter inference problem. We introduce a scalable algorithm to estimate the parameters by least squares, and we prove the convergence of the estimators as data size increases. We test the ISALT on three non-globally Lipschitz SDEs: the 1D double-well potential, a 2D multi-scale gradient system, and the 3D stochastic Lorenz equation with degenerate noise. Numerical results show that ISALT can tolerate time-step magnitudes larger than plain numerical schemes. It reaches optimal accuracy in reproducing the invariant measure when the time-step is medium-large.

math.NA

A review of Local-to-Nonlocal coupling methods in nonlocal diffusion and nonlocal mechanics

Local-to-Nonlocal (LtN) coupling refers to a class of methods aimed at combining nonlocal and local modeling descriptions of a given system into a unified coupled representation. This allows to consolidate the accuracy of nonlocal models with the computational expediency of their local counterparts, while often simultaneously removing additional nonlocal modeling issues such as surface effects. The number and variety of proposed LtN coupling approaches have significantly grown in recent year, yet the field of LtN coupling continues to grow and still has open challenges. This review provides an overview of the state-of-the-art of LtN coupling in the context of nonlocal diffusion and nonlocal mechanics, specifically peridynamics. We present a classification of LtN coupling methods and discuss common features and challenges. The goal of this review is not to provide a preferred way to address LtN coupling but to present a broad perspective of the field, which would serve as guidance for practitioners in the selection of appropriate LtN coupling methods based on the characteristics and needs of the problem under consideration.

math.AP

Force-Based Atomistic/Continuum Blending for Multilattices

We formulate the blended force-based quasicontinuum (BQCF) method for multilattices and develop rigorous error estimates in terms of the approximation parameters: atomistic region, blending region and continuum finite element mesh. Balancing the approximation parameters yields a convergent atomistic/continuum multiscale method for multilattices with point defects, including a rigorous convergence rate in terms of the computational cost. The analysis is illustrated with numerical results for a Stone--Wales defect in graphene.

math.NA

Mean-field Dynamics of Load-Balancing Networks with General Service Distributions

We introduce a general framework for the mean-field analysis of large-scale load-balancing networks with general service distributions. Specifically, we consider a parallel server network that consists of N queues and operates under the $SQ(d)$ load balancing policy, wherein jobs have independent and identical service requirements and each incoming job is routed on arrival to the shortest of $d$ queues that are sampled uniformly at random from $N$ queues. We introduce a novel state representation and, for a large class of arrival processes, including renewal and time-inhomogeneous Poisson arrivals, and mild assumptions on the service distribution, show that the mean-field limit, as $N \rightarrow \infty$, of the state can be characterized as the unique solution of a sequence of coupled partial integro-differential equations, which we refer to as the hydrodynamic PDE. We use a numerical scheme to solve the PDE to obtain approximations to the dynamics of large networks and demonstrate the efficacy of these approximations using Monte Carlo simulations. We also illustrate how the PDE can be used to gain insight into network performance.

math.PR

Coarse graining, dynamic renormalization and the kinetic theory of shock clustering

We demonstrate the utility of the equation free methodology developed by one of the authors (I.G.K) for the study of scalar conservation laws with disordered initial conditions. The numerical scheme is benchmarked on exact solutions in Burgers turbulence corresponding to Levy process initial data. For these initial data, the kinetics of shock clustering is described by Smoluchowski's coagulation equation with additive kernel. The equation free methodology is used to develop a particle scheme that computes self-similar solutions to the coagulation equation, including those with fat tails.

nlin.AO