SearcharxivSearch

FIND YOUR NEXT DISCOVERY

Results for “cs.NA”

Original records, connected by a shared subject.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,600 records · Page 2Linked to original sources

Adaptive multigrid for high-order discontinuous Galerkin methods based on the full approximation scheme

We propose an adaptive multigrid (MG) method for discontinuous Galerkin formulations of elliptic problems using Brandt's full approximation scheme (FAS). Unlike common approaches, this method achieves local $hp$-refinement of hexahedral meshes without the need for hanging nodes. The core component of the FAS-MG method is an overlapping Schwarz smoother, which is optionally accelerated by a Krylov method. This smoother is designed for unstructured curvilinear meshes but maintains a tensor-product structure for fast diagonalization. Numerical experiments demonstrate the exceptional efficiency of the FAS-MG method. Dedicated studies confirm its robustness against high aspect ratios, element deformation, and irregular mesh topology. We also verify its capability for dynamic parallel mesh adaptation using the wave-front benchmark of Červený, Dobrev, and Kolev (SIAM J. Sci. Comp. 41, 2019). Finally, we present preliminary results of extending the method to incompressible Navier-Stokes problems.

math.NA

A multi-class kinetic traffic flow model: discrete-velocity formulation and diffusively-corrected macroscopic limits

This paper introduces a multi-class extension of a discrete-velocity kinetic traffic flow model based on a non-local Prigogine-Herman framework. We derive a hyperbolically scaled system of equations from a continuous kinetic formulation describing interactions between different vehicle classes through braking and relaxation terms. The model is then discretized with respect to the velocity variable for an arbitrary number of vehicle classes, and the structural properties of the resulting formulation are analyzed. In particular, we prove hyperbolicity and total linear degeneracy. Due to the non-conservative structure of the model, we employ a path-conservative finite volume scheme for the numerical approximation of the system. Finally, we derive the corresponding diffusively-corrected macroscopic multi-class model, investigate its stability and present numerical simulations on a single-lane road to illustrate the theoretical findings.

math.NA

Constraint Preserving AFD-WENO Schemes for Relativistic Hydrodynamics with General Equations of State

We develop a high-order physical-constraint-preserving (PCP) alternative finite difference weighted essentially non-oscillatory (AFD-WENO) scheme for the special relativistic hydrodynamics equations with general equations of state. The proposed scheme comprises two key limiters: a state limiter, which acts after the WENO state interpolation step, and a flux limiter, which acts on the final high-order fluxes. The state limiter ensures that the interpolated states are physically admissible, while the flux limiter ensures that the numerical fluxes are physically admissible. The resulting scheme is rigorously proved to satisfy the physical constraints. Incorporating multiple WENO interpolation techniques, including an improved adaptive-order formulation (WENO-AOI), the method is validated through extensive one- and two-dimensional numerical benchmarks with various equations of state. The numerical results demonstrate high-order accuracy, sharp resolution of discontinuities, and robust stability in extreme relativistic regimes.

math.NA

Stein's method for marginals on large graphical models

Many spatial models exhibit locality structures that effectively reduce their intrinsic dimensionality, enabling efficient approximation and sampling of high-dimensional distributions. However, existing approximation techniques primarily focus on joint distributions and do not provide precise accuracy control for low-dimensional marginals, which are of primary interest in many practical scenarios. By leveraging the locality structures, we establish a dimension independent uniform error bound for the marginals of approximate distributions. Inspired by the Stein's method, we introduce a novel $δ$-locality condition that quantifies the locality in distributions, and link it to the structural assumptions such as the sparse graphical models. The theoretical guarantee motivates the localization of existing sampling methods, as we illustrate through the localized likelihood-informed subspace method and localized score matching. We show that by leveraging the locality structure, these methods greatly reduce the sample complexity and computational cost via localized and parallel implementations.

stat.ML

Performance Evaluation of Fast Fourier Transforms on Emerging RISC-V Hardware with Vector Extension Support

This manuscript presents a performance evaluation of Fast Fourier Transform (FFT) implementations on emerging processors supporting the RISC-V Vector Extension (RVV 1.0). By introducing juFFTe, a light-weight high-performance library for discrete Fourier transforms, it is demonstrated how effective vectorization of performance-critical FFT kernels can be achieved on RVV-enabled hardware. Comprehensive benchmarks on three RVV 1.0-ready processors, the SiFive X280, the X100 core of the SpacemiT K3 and the C920v2 core of the Sophon SG2044, reveal substantial performance improvements of juFFTe (https://github.com/FZJ-JSC/juFFTe) over the widely used FFTW3 library. Although RVV-enabled platforms show promising results at this stage of development, a comparison with AMD's Zen 5 architecture indicates that RISC-V needs further maturing to reach the performance of established micro-architectures.

cs.MS

Mathematical and numerical analysis of quantum signal processing

Quantum signal processing (QSP) provides a representation of scalar polynomials of degree $d$ as products of matrices in $\mathrm{SU}(2)$, parameterized by $(d+1)$ real numbers known as phase factors. QSP is the mathematical foundation of quantum singular value transformation (QSVT), which is often regarded as one of the most important quantum algorithms of the past decade, with a wide range of applications in scientific computing, from Hamiltonian simulation to solving linear systems of equations and eigenvalue problems. In this article we survey recent advances in the mathematical and numerical analysis of QSP. In particular, we focus on its generalization beyond polynomials, the computational complexity of algorithms for phase factor evaluation, and the numerical stability of such algorithms. The resolution to some of these problems relies on an unexpected interplay between QSP, nonlinear Fourier analysis on $\mathrm{SU}(2)$, fast polynomial multiplications, and Gaussian elimination for matrices with displacement structure.

quant-ph

Walk based Laplacians for Modeling Diffusion on Complex Networks

We develop a novel framework for modeling diffusion on complex networks by constructing Laplacian-like operators based on walks around a graph. Our approach introduces a parametric family of walk-based Laplacians that naturally incorporate memory effects by excluding or downweighting backtracking trajectories, where walkers immediately revisit nodes. The framework includes: (i) walk-based Laplacians that count all traversals in the network; (ii) nonbacktracking variants that eliminate immediate reversals; and (iii) backtrack-downweighted variants that provide a continuous interpolation between these two regimes. We establish that these operators extend the definition of the standard Laplacian and also preserve some of its properties. We present efficient algorithms using Krylov subspace methods for computing them, ensuring applicability of our proposed framework to large networks. Extensive numerical experiments on real-world networks validate the modeling flexibility of our approach and demonstrate the computational efficiency of the proposed algorithms, including GPU acceleration.

cs.SI

Generalization Error Estimates of Machine Learning Methods for Solving High Dimensional Schrödinger Eigenvalue Problems

We propose a machine learning method for computing eigenvalues and eigenfunctions of the Schrödinger operator on a $d$-dimensional hypercube with Dirichlet boundary conditions. The cut-off function technique is employed to construct trial functions that precisely satisfy the homogeneous boundary conditions. This approach eliminates the error caused by the standard boundary penalty method, improves the overall accuracy of the method, as demonstrated by the typical numerical examples. Under the assumption that the eigenfunctions belong to a spectral Barron space, we derive an explicit convergence rate of the generalization error of the proposed method, which does not suffer from the curse of dimensionality. We verify the assumption by proving a new regularity shift result for the eigenfunctions when the potential function belongs to an appropriate spectral Barron space. Moreover, we extend the generalization error bound to the normalized penalty method, which is widely used in practice.

math.NA

Vanka-smoothed shifted Laplacian multigrid preconditioners for the Helmholtz equations

We present an improved multigrid preconditioner for the acoustic Helmholtz equation with enhanced scalability. Standard multigrid fails to converge for the Helmholtz equation, and the well-known complex shifted Laplacian method overcomes it by adding a complex shift and using the shifted system as a preconditioner. However, the added complex shift grows with the frequency and interferes with the preconditioner's scalability. In this work, we present an additive Vanka smoother that requires a much lower shift than point-wise smoothers, and thereby enhances the scalability. By carefully designing different ingredients of the multigrid cycle, the presented method enables deep V-cycles with a small and bounded shift, even when many levels are used. We validate our method theoretically by local Fourier analysis, and hold numerical experiments for homogeneous and heterogeneous media. We show that our method outperforms plain shifted Laplacian in terms of runtimes and performs well on challenging geophysical media in 2D and 3D.

math.NA

Optimal error estimates for the half-way bounce-back lattice Boltzmann method for the Stokes equations

We give a mathematical proof of the optimal convergence rates for the D2Q9 BGK lattice Boltzmann method with the half-way bounce-back rule for the incompressible Stokes equations in a flat channel. The convergence rates are second-order for the velocity and first-order for the pressure as the lattice spacing $h$ tends to zero, in agreement with formal analyses and numerical experiments, whereas the available rigorous convergence theorems only yield an $O(h^{1/2})$ bound for the velocity error. A key step in the proof is a decomposition of the leading boundary consistency error into macroscopic and kinetic components. These components are absorbed by suitably constructed Stokes and discrete Knudsen layer correctors, respectively. Incorporating these correctors into the prediction function used in previous rigorous analyses, we obtain a refined prediction function with consistency errors of sufficiently high order. Combined with the known weighted $L^2$-stability estimate, this gives the optimal convergence rates.

math.NA

Conformal Uncertainty Quantification Guarantees for Neural Operators

Neural operators provide fast surrogate models for approximating operators between function spaces, but their predictions often lack uncertainty quantification. We develop a split conformal framework to guarantee that a calibrated pointwise band around the neural operator output contains the true solution on at least a $1-γ$ fraction of the evaluation domain, with probability at least $1-α$ over test and calibration inputs, where $α,γ\in(0,1)$. Our method reduces a normalized residual field to its spatial $(1-γ)$-quantile and computes a scaling factor using a held-out calibration dataset. We prove marginal coverage guarantees for measurable residual fields defined on arbitrary probability spaces, covering both continuum domains and fixed discretizations. Under mild assumptions on the data distribution, we show that the coverage conditional on the calibration set follows a Beta distribution, which we verify with numerical experiments on Darcy flow and Navier--Stokes equations, where our calibration yields bands consistently tighter than existing corrections while retaining the target coverage.

math.NA

Energy-Based Physics-Informed Form Finding for Clustered Tensegrity Structures

Tensegrity form-finding and physical property prediction are fundamental problems in structural mechanics, which aim to determine equilibrium configurations and internal force distributions. These problems are challenging due to strong nonlinearity arising from the coupling between geometry and forces, and the need to satisfy equilibrium, stability, and structural constraints. This paper proposes an energy-based learning approach for clustered tensegrity form finding and physical property prediction. The proposed approach incorporates total potential energy minimization and constitutive relations into the training objective, enabling the prediction of equilibrium nodal configurations and the reconstruction of physical quantities such as member forces and force densities. By integrating energy-based physical losses directly into the learning process, the method promotes physical consistency while combining data-driven learning with physics-based constraints. Numerical experiments on tensegrity prism and lander structures demonstrate accurate prediction of equilibrium configurations and internal forces across different training-data ratios, indicating the potential of the proposed approach for nonlinear tensegrity form finding and structural analysis.

cs.LG

A GPU-Accelerated Blocked Adaptive Randomized Range Finder Based on an Implicit Householder QR Decomposition

Low-rank methods can reduce the memory and computational requirements of deep neural network training in approaches such as GaLore. Randomized range finders offer an attractive alternative to singular value decompositions, particularly when the required rank is determined adaptively from a prescribed approximation tolerance. We introduce a blocked adaptive randomized range finder based on an implicit Householder QR decomposition and an optimized hybrid CPU--GPU implementation. The proposed method avoids explicit reorthogonalization. Numerical experiments show that it preserves orthogonality and approximation accuracy in regimes where block Gram--Schmidt without reorthogonalization becomes unstable. The blocked formulation exposes matrix--matrix operations and enables overlap of CPU panel factorization with GPU updates. On an NVIDIA GH200, the overlapped implementation reduces the runtime for the largest tested matrix from 9.91 seconds on the CPU to 0.407 seconds. The method provides a stable and efficient building block for low-rank approximation on heterogeneous systems with applications in computational science and engineering.

math.NA

Selection, Representation, and Execution in Sparse Fourier Neural Operators

Sparse representations are often expected to make models smaller and also reduce inference cost. For Fourier Neural Operators (FNOs), these objectives are not equivalent or do not always align: removing parts of the learned operator can leave the underlying transforms and dense computations unchanged, while changing the grid on which the model is evaluated can introduce overhead of its own. We therefore distinguish sparsity in the representation, in the stored parameters, in the theoretical operation count, and in measured runtime, and present an empirical study of several routes toward sparse FNOs that tests each transition between them separately. Coarsening the execution grid reduces the theoretical cost without reducing measured latency, and adding a correction term recovers accuracy at the cost of making the model slower. Even an 83\% parameter reduction remains slower than the dense baseline under ordinary execution. These results motivate a stricter definition of useful sparsity: the deployed operator must preserve solution accuracy and map its reduced support to a genuinely cheaper execution path.

cs.LG

Multilevel lattice-based kernel approximation for elliptic PDEs with random coefficients

This paper introduces a multilevel kernel-based approximation method to estimate efficiently solutions to elliptic partial differential equations (PDEs) with periodic random coefficients. Building upon the work of Kaarnioja, Kazashi, Kuo, Nobile, Sloan (Numer. Math., 2022) on kernel interpolation with quasi-Monte Carlo (QMC) lattice point sets, we leverage multilevel techniques to enhance computational efficiency while maintaining a given level of accuracy. In the function space setting with product-type weight parameters, the single-level approximation can achieve an accuracy of $\varepsilon>0$ with cost $\mathcal{O}(\varepsilon^{-η-ν-θ})$ for positive constants $η, ν, θ$ depending on the rates of convergence associated with dimension truncation, kernel approximation, and finite element approximation, respectively. Our multilevel approximation can achieve the same $\varepsilon$ accuracy at a reduced cost $\mathcal{O}(\varepsilon^{-η-\max(ν,θ)})$. Full regularity theory and error analysis are provided, followed by numerical experiments that validate the efficacy of the proposed multilevel approximation in comparison to the single-level approach.

math.NA

Wigner-Eckart Factorization of the Polyatomic Boltzmann Collision Operator

We extend the Wigner-Eckart factorization of the spectral Boltzmann collision operator to polyatomic gases with continuous internal energy. Because internal energies are invariant under spatial rotations, the SO(3) reduction survives the Borgnakke-Larsen energy exchange, and the twelve-dimensional collision integral collapses onto a nine-dimensional kinematic core. The core splits into a sparse geometric tensor, evaluated exactly, and a dense physical tensor, integrated by singularity-resolving Gauss rules with an auxiliary Laplace representation of the fractional energy couplings. The quadrature attains near machine precision at the fractional exponents of real gases. The collision invariants are embedded exactly, preserving the translational-internal energy exchange. The factorization compresses the operator by three to nearly four orders of magnitude and accelerates its evaluation 40-fold over dense formulations. The method is validated against the exact monatomic limit, Landau-Teller relaxation, and an analytic frozen-channel Prandtl number, and it matches a published calibration of the same kernel for N2, CO, and H2.

math.NA

D-RADI: A Low-rank ADI Algorithm for Solving Large-scale Discrete-time Algebraic Riccati Equations

The low-rank alternating direction implicit (ADI) method is an efficient numerical technique for solving several types of large-scale matrix equations that admit low-rank solutions. The discrete-time algebraic Riccati equation (DARE) is an important matrix equation with applications in state estimation, controller design, and filter design. In the literature, the low-rank Cholesky factor ADI method for Stein equations has been used within Newton iterations to solve large-scale DAREs. However, no dedicated low-rank ADI solver is available for such DAREs. To address this gap, this paper presents a low-rank ADI solver for large-scale DAREs. We also propose an efficient approach to generate ADI shifts automatically, which makes the proposed solver fully autonomous for solving DAREs. The effectiveness of the proposed solver is compared with MATLAB's \texttt{idare} on a moderate-order problem. Efficiency and accuracy are further demonstrated on large-scale DAREs of order $10^6$. Numerical results confirm that the solver is efficient, accurate, and fully autonomous.

math.NA

Dynamic subscales and unconditional semi-discrete stability for non-residual VMS approximations of generalised Newtonian flows

We analyse a non-residual variational multiscale finite element formulation with dynamic subscales for incompressible generalised Newtonian Navier--Stokes flows. The apparent viscosity is assumed to be bounded and Lipschitz continuous with respect to the shear rate, covering several regularised rheological laws. The method separates, through orthogonal projections, the unresolved pressure-divergence, convective, and pressure-gradient contributions. For the linearised semi-discrete problem, we prove well-posedness and an unconditional stability estimate in an anisotropic VMS norm. The estimate controls viscous dissipation, subscale energies, discrete divergence, and a coupled acceleration-convection-pressure balance. For the non-linear formulation, a fixed-point argument yields existence and optimal-order a priori bounds under suitable regularity and smallness assumptions. The analysis also provides weak-in-time control of the discrete pressure-convection balance. A key point is that the dynamic pressure-related subscale provides the time-derivative contribution needed to control the projected pressure-gradient error, a mechanism unavailable in the corresponding quasi-static setting.

math.NA