SearcharxivSearch

arXiv subjects

Jinchao Xu

Publications and source records attributed to Jinchao Xu.

At least 19 recordsLinked to original sources

Configuration-Dependent Lower Bounds for Approximation by Shallow ReLU$^k$ Networks on the Sphere

We establish two related but logically distinct results for shallow ReLU$^k$ neural networks on the unit sphere $\SS^d$. First, for an arbitrary set of inner neural-network parameters, the best $\mathcal{L}^2(\SS^d)$ approximation of a fixed target function with smoothness $r>\tfrac{d+2k+1}{2}$ admits an asymptotic lower bound given by a constant multiple of $n^{-1/2}\underline{h}^{k+1/2}$, where $\underline{h}$ denotes the antipodal separation distance of the normalized inner-parameter set. This lower bound depends explicitly on the parameter configuration through $\underline{h}$ and applies without additional assumptions on the parameters. Second, for antipodally quasi-uniform parameters, $\underline{h}\simeq n^{-1/d}$, and the lower bound establishes the exact saturation order $n^{-\frac{d+2k+1}{2d}}$ for such parameter families: a target function with regularity greater than $\frac{d+2k+1}{2}$ and satisfying the required parity condition can be approximated at this rate, whereas approximation at any strictly faster rate forces the target function to be zero. Our results therefore place linearized neural-network approximation within the classical saturation framework and show that, although ReLU$^k$ network spaces can outperform finite elements of the same degree, this advantage is intrinsically limited.

math.NA

Uniform Inf-Sup Norm Equivalence and Robust Operator Preconditioning for Stokes Flow in tight domains with Periodic Pillars

Many microfluidic and porous-media computations reduce to the same core task: solving a Stokes saddle-point system on a domain perforated by a dense periodic array of pillars, as in deterministic lateral displacement (DLD) particle sorters. After rescaling the device to unit size, the geometry is controlled by a single dimensionless parameter $m$---the number of pillars across the device, equal to the inverse period. In realistic devices $m$ reaches the hundreds or thousands, and as it grows the Stokes inf-sup constant decays like $m^{-1}$, the pressure Schur complement becomes severely ill-conditioned, and standard block solvers slow down in proportion to the pillar density. We remove this bottleneck by identifying the pressure norm that the divergence operator induces on such geometries. For periodic pillar arrays in the proportional-hole regime, we prove that this inf-sup norm is uniformly equivalent to the $L^2+σ_εH^1$ $K$-functional norm at the pore scale $σ_ε\asympε$, with constants independent of the period, the pillar count $m$, and the mesh size $h$. The equivalence identifies the perforated Stokes problem with a Brinkman problem at a homogenized permeability, and its Riesz map reduces to a pressure-mass inverse plus a scaled stiffness inverse. Combined with operator preconditioning, this yields a block preconditioner built from standard algebraic-multigrid solves whose iteration count is essentially independent of both mesh size and pillar density. Two-dimensional Taylor--Hood experiments confirm the predicted robustness in mesh refinement, pillar density, geometric scale, close packing, and time step.

math.NA

Subspace correction as a general framework for convex optimization algorithms

This paper shows that subspace correction methods provide a common algorithmic and theoretical foundation for several classes of convex optimization algorithms, including operator splitting, alternating projection, and multiplier methods. The underlying principle is to decompose a problem into smaller subproblems and combine their solutions, a strategy that appears throughout iterative algorithms. The main tool is an iterate-level formalism, which we call dualization, that abstracts classical primal--dual correspondences and relates a subspace correction method for a dual problem to an algorithm for the corresponding primal problem via a primal--dual consistency relation. At the algorithmic level, dualizing successive subspace correction yields the Peaceman--Rachford and Douglas--Rachford splitting methods; the von Neumann and Dykstra alternating projection algorithms arise as special cases. Dualizing parallel subspace correction yields a parallel splitting method. Multiplier methods, including the alternating direction method of multipliers (ADMM), are connected to subspace correction through the same mechanism together with equivalent block formulations. In particular, a multi-block ADMM-type algorithm is obtained by dualizing a splitting method derived from successive subspace correction. At the theoretical level, the primal--dual consistency relation transfers convergence estimates from subspace correction to the derived algorithms under suitable assumptions. Thus, these operator splitting, alternating projection, and multiplier algorithms can be derived from subspace correction at both the algorithmic and theoretical levels.

math.OC

Sharp Sobolev Approximation on General Domains by Linearized Shallow Networks with Analytic Activations

We study Sobolev approximation on bounded domains by linearized shallow neural networks whose inner parameters are prescribed independently of the target function. Our main step is a one-dimensional construction for analytic activations. We prove that quasi-Chebyshev parameter sets with univariate resolution $m$ generate fixed feature spaces attaining the sharp $H^r$-to-$H^s$ approximation order $m^{-(r-s)}$ for a class of analytic activations satisfying a quantitative non-cancellation condition on their Taylor coefficients. Combining this result with the ridge-function lifting theorem in [SIAM J. Math. Anal. 30 (1998), pp. 155-189] and its extension to arbitrary quasi-uniform direction sets established in this work, we construct tensor-product-type parameter sets that attain the sharp rate $$\|f-f_n\|_{L^2(Ω)}\lesssim n^{-\frac rd}\|f\|_{H^r(Ω)},\quad f\in H^r(Ω)$$ for all $r>0$. In contrast to the finite-difference construction in [Neural Comput. 8 (1996), pp. 164-177], whose explicit admissibility condition may require an extremely small parameter scale, the proposed parameter sets remain distributed over fixed intervals and are therefore more amenable to practical computation.

math.NA

Optimal Neural Network Approximation via Empirical Least Squares with Deterministic Samples

We develop a rigorous theory of discrete residual least-squares approximation for elliptic spectral equations $\mathfrak L_βu=f$ using linearized ReLU$^k$ neural networks on the sphere, where $\mathfrak L_β$ is a positive elliptic spectral multiplier of order $β$. Given a parameter set $Θ_n=\{θ_{j}^*\}_{j=1}^n\subset\mathbb S^d$, we approximate $u$ in the linearized network space $L_n^k(Θ_n)$ by the discrete residual on the collocation points $\{η_i^*\}_{i=1}^m$ \begin{equation*} u_{n,m}\in\arg\min_{v_n\in L_n^k(Θ_n)}\frac1m\sum_{i=1}^m\left(f(η_i^*)-\mathfrak L_βv_n(η_i^*)\right)^2. \end{equation*} With $k>\frac{d-1}{2}+β$, for antipodally quasi-uniform network parameter sets and any quasi-uniform collocation points with $m\gtrsim n$, we prove that \begin{equation*} \|u-u_{n,m}\|_{\mathcal H^β(\mathbb S^d)}\eqsim\|f-\mathfrak L_βu_{n,m}\|_{\mathcal L^2(\mathbb S^d)}\lesssim n^{-\frac{r}{d}} \begin{cases} \|f\|_{\mathcal W^{r,p}(\mathbb S^d)},&\frac{d}{p} 2,\\ \|f\|_{\mathcal H^r(\mathbb S^d)},&r>\frac{d}{2}. \end{cases} \end{equation*} We also establish a high-probability residual estimate, up to a logarithmic factor and an arbitrarily small smoothness loss, for i.i.d.\ uniformly distributed collocation points. The key analytical ingredient is a Bernstein inequality for linearized ReLU$^k$ network spaces. If $\underline h$ denotes the antipodal separation distance of the network parameters, then \begin{equation*} \|v_n\|_{\mathcal H^r(\mathbb S^d)}\lesssim\underline h^{-(r-s)}\|v_n\|_{\mathcal H^s(\mathbb S^d)},\qquad 0\leq s<r<k+\tfrac12. \end{equation*}

math.NA

Self-composing neural operators for high-frequency and multiscale PDE surrogates

Addressing the computational challenges of high-frequency and multiscale partial differential equations (PDEs), this work introduces a self-composing neural operator (SC-NO) framework. Inspired by classical fixed-point iterative solvers (e.g., multigrid, domain decomposition), the proposed architecture constructs a deep operator by repeatedly applying a single, parameter-efficient backbone block. This design mimics the update step of a numerical solver, allowing the model to progressively resolve complex solution features without increasing the parameter count. For practical training, we develop an adaptive ``Train-and-Unroll'' strategy that grows the composition depth during training, acting as a curriculum from shallow to deep self-composed models. We demonstrate the efficacy of this framework on the Helmholtz equation for ultrasound computed tomography (USCT), a problem characterized by high-frequency wave propagation in highly heterogeneous media. By instantiating the backbone with a multigrid-inspired architecture, the SC-NO effectively mitigates the spectral bias often observed in standard operator learning baselines. Numerical experiments show that our method reduces the prediction error significantly compared to Fourier Neural Operators (FNO) and their variants in the 300--500 kHz regime. Furthermore, we provide theoretical analysis linking the self-composition depth to approximation accuracy.

cs.LG

Do Neural Networks Really Beat the Curse of Dimensionality? A Bit-Complexity View

Traditional approximation theory measures convergence rates in terms of the number of parameters or degrees of freedom. However, practical computation operates under finite precision: parameters must be encoded using a finite number of bits. Therefore, approximation efficiency should be evaluated in terms of computational bit complexity, which is intrinsically connected to the metric entropy of the underlying function class. In this work, we develop a unified approximation framework based on binary encoding and metric entropy. We analyze classical methods (including polynomial approximation, sparse grids, and finite elements) as well as shallow and deep neural networks, and compare their approximation rates for function classes with comparable metric entropy. We observe that, when evaluated in terms of bits, most classical methods are in general suboptimal relative to the intrinsic limits dictated by metric entropy, while neural network methods may exhibit different behaviors. We show that when complexity is measured in bits rather than parameters, no method fundamentally exceeds the approximation order achieved by classical approaches. Our results also indicate that many seeming advantages of neural networks, including dimension-independent rates and superconvergence phenomena, stem from differences in function class complexity rather than intrinsic architectural superiority. In this sense, the traditional curse of dimensionality can be misleading; the fundamental limitation is instead a curse of bit complexity, governed by metric entropy.

cs.LG

ReLU$^k$ Neural de Rham Complexes

We construct finite-dimensional de Rham subcomplexes generated by fixed-neuron shallow ReLU$^k$ neural networks, a class of spaces known to provide optimal approximation rates. For neurons of the form $s_i(x)=ω_i\cdot x+b_i$, we introduce spaces of neural differential forms: differential $p$-forms whose coefficients are the ReLU$^k$ ridge functions $σ_{k-p}(s_i)$. These spaces are compatible with the exterior derivative because differentiating a ReLU power lowers its order by one, and for each fixed neuron, differentiation amounts to exterior multiplication by the fixed one-form $ d s_i$. Under a linear independence assumption on the lowest-order family $\{σ_{k-d}(s_i)\}_{i=1}^n$, the global complex decomposes into independent neuron-wise Koszul complexes. We prove exactness in arbitrary dimension and provide a geometric sufficient condition for the required linear independence. Numerical experiments based on the resulting complex provide evidence of stable discretizations and of convergence rates consistent with the underlying approximation theory, and exhibit no spurious modes in eigenvalue problems considered.

math.NA

On the Dimension-Free Approximation of Deep Neural Networks for Symmetric Korobov Functions

Deep neural networks have been widely used as universal approximators for functions with inherent physical structures, including permutation symmetry. In this paper, we construct symmetric deep neural networks to approximate symmetric Korobov functions and prove that both the convergence rate and the constant prefactor scale at most polynomially with respect to the ambient dimension. This represents a substantial improvement over prior approximation guarantees that suffer from the curse of dimensionality. Building on these approximation bounds, we further derive a generalization-error rate for learning symmetric Korobov functions whose leading factors likewise avoid the curse of dimensionality.

cs.LG

McMg: A Learned Phase-Space Multi-channel Multigrid Preconditioner for Helmholtz Equation

Solving heterogeneous Helmholtz equations at high wavenumbers remains challenging because the discretized operator is indefinite, pollution degrades phase accuracy, and scalar coarse-grid correction can discard the local phase and propagation-direction information carried by oscillatory errors. We propose Multi-channel Multigrid (McMg), a learned phase-space multigrid preconditioner for heterogeneous Helmholtz equations. Rather than predicting the solution directly, McMg maps residuals to corrections within an iterative framework. Its central idea is to coarsen physical space while retaining unresolved local wave information in the channel dimension: each coarse node carries a learned packet of amplitude, phase, direction, and scattering coefficients rather than a single scalar unknown. The architecture combines linear multi-channel transfer operators with locally adaptive stencils, neural PDE operators, and medium-dependent smoothers whose coefficients are generated from the wave speed. For a fixed medium, the V-cycle is linear in the residual; nonlinear physical features are computed once in a setup phase and cached, so each online iteration reduces to convolutions with fixed coefficients. We further study generalization across scales. Models trained on small domains transfer directly to larger domains and higher effective wavenumbers, and a Layer-by-Layer Progressive Finetuning (LLPF) strategy improves large-domain scalability by adding new coarse levels while finetuning only the newly introduced parameters. Numerical experiments on high-frequency, high-contrast, and large-scale three-dimensional problems demonstrate that McMg requires substantially fewer iterations and less wall-clock time than strong classical baselines, while consistently outperforming existing neural preconditioners.

math.NA

A polynomial dimension-dependence analysis of Bramble--Pasciak--Xu preconditioners

We investigate the dimension dependence of Bramble--Pasciak--Xu (BPX) preconditioners for high-dimensional partial differential equations and establish that the condition numbers of BPX-preconditioned systems grow only polynomially with the spatial dimension. Our analysis requires a careful derivation of the dimension dependence of several fundamental tools in the theory of finite element methods, including elliptic regularity, the Bramble--Hilbert lemma, trace inequalities, and inverse inequalities. We further analyze an averaged Scott--Zhang-type quasi-interpolation operator, and show that its associated constants scale polynomially with the dimension. Building on these ingredients, we prove a multilevel norm equivalence theorem and derive a BPX preconditioner with explicit polynomial bounds on its dimensional dependence. The analysis is motivated in part by recent tensor and quantum finite element methods, where dimension-explicit conditioning estimates for BPX preconditioners play an important role.

math.NA

A quasi-monolithic localized high-order ALE finite element method for multi-scale fluid-structure interaction problems

This paper presents a quasi-monolithic localized high-order arbitrary Lagrangian-Eulerian (qMLH-ALE) finite element method for multi-scale fluid-structure interaction (FSI) in microfluidic systems. The fluid momentum, the incompressible Neo-Hookean constitutive law, and the left Cauchy-Green tensor $\mathcal{B}$ are assembled into a single implicit system, while the harmonic mesh extension is updated explicitly in a staggered manner. Isoparametric $\mathcal{P}_2$ elements provide third-order geometric approximation of curved fluid-solid interfaces, and a second-order implicit-explicit partitioned Runge-Kutta scheme delivers second-order temporal accuracy without the dissipation of backward Euler. A localized updating strategy confines the moving mesh and the deformation history to a body-fitted sub-domain coupled with a precomputed steady background flow, bridging the scale disparity between local FSI dynamics and the macroscopic microchannel geometry. The Turek-Hron FSI3 benchmark, performed at unit fluid-solid density ratio, reproduces the reference beam-tip amplitude and frequency within $3\%$, confirming stability under the strong added-mass coupling that destabilizes conventional partitioned schemes. Three-dimensional particle-focusing simulations in spiral microchannels further illustrate the framework on long-range multi-scale problems.

math.NA

Sharp inf-sup estimate for the Stokes equation in tight domains with periodic pillars and some numerical implications

The predictive simulation of fluid dynamics in densely packed microfluidic devices, such as Deterministic Lateral Displacement (DLD) arrays, stagnates with standard iterative solvers. We show that this failure is not algorithmic but rooted in the pre-asymptotic degradation of the pressure-velocity coupling stability. For periodic pillar geometries in a generalized lattice framework, we prove that the continuous Ladyzhenskaya-Babuška-Brezzi (LBB) condition, also called the inf-sup constant, deteriorates exactly as $m^{-1}$ up to a positive multiplicative constant, where $m$ is the pillar density (the number of pillars per unit length). This induces a priori error amplification proportional to $m$ and a pressure Schur complement condition number scaling as $\mathcal{O}(m^2)$. To overcome this theoretical limit, we propose a parameter-free, adaptively scaled Augmented Lagrangian (AL) stabilization strategy with penalty $γ\propto m^2$. Numerical experiments on both standard square and asymmetric DLD arrays validate the theoretical bounds: the AL method reduces outer FGMRES iterations from 437 to 22 on a 1.85M-DoF square array and from 687 to 24 on a 1.77M-DoF DLD array.

math.NA

Randomized subspace correction methods for convex optimization

This paper introduces an abstract framework for randomized subspace correction methods for convex optimization, which unifies and generalizes a broad class of existing algorithms, including domain decomposition, multigrid, and block coordinate descent methods. We provide a convergence rate analysis ranging from minimal assumptions to more practical settings, such as sharpness and strong convexity. While most existing studies on block coordinate descent methods focus on nonoverlapping decompositions and smooth or strongly convex problems, our framework extends to more general settings involving arbitrary space decompositions, inexact local solvers, and problems with weaker smoothness or convexity assumptions. The proposed framework is broadly applicable to convex optimization problems arising in areas such as nonlinear partial differential equations, imaging, and data science.

math.OC

High-order DLM-ALE discretizations with robust operator preconditioning for fluid-rigid-body interaction

Motivated by the design of deterministic lateral displacement (DLD) microfluidic devices, we develop a high-order numerical framework for fluid-rigid-body interaction on fitted moving meshes. Rigid-body motion is enforced by a distributed Lagrange multiplier (DLM) formulation, while the moving fluid domain is treated by an arbitrary Lagrangian-Eulerian (ALE) mapping. In space, we use isoparametric Taylor-Hood elements to achieve high-order accuracy and to represent curved boundaries and the fluid-particle interface. In time, we employ a high-order partitioned Runge-Kutta strategy in which the mesh motion is advanced explicitly and the coupled physical fields are advanced implicitly, yielding high-order accuracy for the particle trajectory. The fully coupled system is linearized into a generalized Stokes problem subject to distributed constraints of incompressibility and rigid-body motion. We establish well-posedness of this generalized Stokes formulation at both the continuous and discrete levels, providing the stability foundation for operator preconditioning that is robust with respect to key physical and discretization parameters. Numerical experiments on representative benchmarks, including a DLD case, demonstrate high-order convergence for the fluid solution and rigid-body dynamics, as well as robust iterative convergence of the proposed preconditioners.

math.NA

Solving High-Dimensional PDEs Using Linearized Neural Networks

Linearized shallow neural networks that are constructed by fixing the hidden-layer parameters have recently shown strong performance in solving partial differential equations (PDEs). Such models, widely used in the random feature method (RFM) and extreme learning machines (ELM), transform network training into a linear least-squares problem. In this paper, we conduct a numerical study of the variational (Galerkin) and collocation formulations for these linearized networks. Our numerical results reveal that, in the variational formulation, the associated linear systems are severely ill-conditioned, forming the primary computational bottleneck in scaling the neural network size, even when direct solvers are employed. In contrast, collocation methods combined with robust least-squares solvers exhibit better numerical stability and achieve higher accuracy as we increase neuron numbers. This behavior is consistently observed for both ReLU$^k$ and $\tanh$ activations, with $\tanh$ networks exhibiting even worse conditioning. Furthermore, we demonstrate that random sampling of the hidden layer parameters, commonly used in RFM and ELM, is not necessary for achieving high accuracy. For ReLU$^k$ activations, this follows from existing theory and is verified numerically in this paper, while for $\tanh$ activations, we introduce two deterministic schemes that achieve comparable accuracy.

math.NA

Condition Numbers and Eigenvalue Spectra of Shallow Networks on Spheres

We present an estimation of the condition numbers of the \emph{mass} and \emph{stiffness} matrices arising from shallow ReLU$^k$ neural networks defined on the unit sphere~$\mathbb{S}^d$. In particular, when $\{θ_j^*\}_{j=1}^n \subset \mathbb{S}^d$ is \emph{antipodally quasi-uniform}, the condition number is sharp. Indeed, in this case, we obtain sharp asymptotic estimates for the full spectrum of eigenvalues and characterize the structure of the corresponding eigenspaces, showing that the smallest eigenvalues are associated with an eigenbasis of low-degree polynomials while the largest eigenvalues are linked to high-degree polynomials. This spectral analysis establishes a precise correspondence between the approximation power of the network and its numerical stability.

math.NA

Approximation Rates for Shallow ReLU$^k$ Neural Networks on Sobolev Spaces via the Radon Transform

Let $Ω\subset \mathbb{R}^d$ be a bounded domain. We consider the problem of how efficiently shallow neural networks with the ReLU$^k$ activation function can approximate functions from Sobolev spaces $W^s(L_p(Ω))$ with error measured in the $L_q(Ω)$-norm. Utilizing the Radon transform and recent results from discrepancy theory, we provide a simple proof of nearly optimal approximation rates in a variety of cases, including when $q\leq p$, $p\geq 2$, and $s \leq k + (d+1)/2$. The rates we derive are optimal up to logarithmic factors, and significantly generalize existing results. An interesting consequence is that the adaptivity of shallow ReLU$^k$ neural networks enables them to obtain optimal approximation rates for smoothness up to order $s = k + (d+1)/2$, even though they represent piecewise polynomials of fixed degree $k$.

stat.ML