SearcharxivSearch

arXiv subjects

Stefan Schnake

Publications and source records attributed to Stefan Schnake.

13 recordsLinked to original sources

Tucker Attention: A generalization of approximate attention mechanisms

The pursuit of reducing the memory footprint of the self-attention mechanism in multi-headed self attention (MHA) spawned a rich portfolio of methods, e.g., group-query attention (GQA) and multi-head latent attention (MLA). The methods leverage specialized low-rank factorizations across embedding dimensions or attention heads. From the point of view of classical low-rank approximation, these methods are unconventional and raise questions of which objects they really approximate and how to interpret the low-rank behavior of the resulting representations. To answer these questions, this work proposes a generalized view on the weight objects in the self-attention layer and a factorization strategy, which allows us to construct a parameter efficient scheme, called Tucker Attention. Tucker Attention requires an order of magnitude fewer parameters for comparable validation metrics, compared to GQA and MLA, as evaluated in LLM and ViT test cases. Additionally, Tucker Attention~encompasses GQA, MLA, MHA as special cases and is fully compatible with flash-attention and rotary position embeddings (RoPE). This generalization strategy yields insights of the actual ranks achieved by MHA, GQA, and MLA, and further enables simplifications for MLA.

cs.LG

Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization

Deploying geospatial foundation models on resource-constrained edge devices demands compact architectures that maintain high downstream performance. However, their large parameter counts and the accuracy loss often induced by compression limit practical adoption. In this work, we leverage manifold-constrained optimization framework DLRT to compress large vision transformer-based geospatial foundation models during transfer learning. By enforcing structured low-dimensional parameterizations aligned with downstream objectives, this approach achieves strong compression while preserving task-specific accuracy. We show that the method outperforms of-the-shelf low-rank methods as LoRA. Experiments on diverse geospatial benchmarks confirm substantial parameter reduction with minimal accuracy loss, enabling high-performing, on-device geospatial models.

cs.CV

Dynamical Low-Rank Compression of Neural Networks with Robustness under Adversarial Attacks

Deployment of neural networks on resource-constrained devices demands models that are both compact and robust to adversarial inputs. However, compression and adversarial robustness often conflict. In this work, we introduce a dynamical low-rank training scheme enhanced with a novel spectral regularizer that controls the condition number of the low-rank core in each layer. This approach mitigates the sensitivity of compressed models to adversarial perturbations without sacrificing accuracy on clean data. The method is model- and data-agnostic, computationally efficient, and supports rank adaptivity to automatically compress the network at hand. Extensive experiments across standard architectures, datasets, and adversarial attacks show the regularized networks can achieve over 94% compression while recovering or improving adversarial accuracy relative to uncompressed baselines.

cs.LG

On high-order/low-order and micro-macro methods for implicit time-stepping of the BGK model

In this paper, a high-order/low-order (HOLO) method is combined with a micro-macro (MM) decomposition to accelerate iterative solvers in fully implicit time-stepping of the BGK equation for gas dynamics. The MM formulation represents a kinetic distribution as the sum of a local Maxwellian and a perturbation. In highly collisional regimes, the perturbation away from initial and boundary layers is small and can be compressed to reduce the overall storage cost of the distribution. The convergence behavior of the MM methods, the usual HOLO method, and the standard source iteration method is analyzed on a linear BGK model. Both the HOLO and MM methods are implemented using a discontinuous Galerkin (DG) discretization in phase space, which naturally preserves the consistency between high- and low-order models required by the HOLO approach. The accuracy and performance of these methods are compared on the Sod shock tube problem and a sudden wall heating boundary layer problem. Overall, the results demonstrate the robustness of the MM and HOLO approaches and illustrate the compression benefits enabled by the MM formulation when the kinetic distribution is near equilibrium.

math.NA

Sparse-grid Discontinuous Galerkin Methods for the Vlasov-Poisson-Lenard-Bernstein Model

Sparse-grid methods have recently gained interest in reducing the computational cost of solving high-dimensional kinetic equations. In this paper, we construct adaptive and hybrid sparse-grid methods for the Vlasov-Poisson-Lenard-Bernstein (VPLB) model. This model has applications to plasma physics and is simulated in two reduced geometries: a 0x3v space homogeneous geometry and a 1x3v slab geometry. We use the discontinuous Galerkin (DG) method as a base discretization due to its high-order accuracy and ability to preserve important structural properties of partial differential equations. We utilize a multiwavelet basis expansion to determine the sparse-grid basis and the adaptive mesh criteria. We analyze the proposed sparse-grid methods on a suite of three test problems by computing the savings afforded by sparse-grids in comparison to standard solutions of the DG method. The results are obtained using the adaptive sparse-grid discretization library ASGarD.

math.NA

Asymptotic Preserving Discontinuous Galerkin Methods for a Linear Boltzmann Semiconductor Model

A key property of the linear Boltzmann semiconductor model is that as the collision frequency tends to infinity, the phase space density $f = f(x,v,t)$ converges to an isotropic function $M(v)ρ(x,t)$, called the drift-diffusion limit, where $M$ is a Maxwellian and the physical density $ρ$ satisfies a second-order parabolic PDE known as the drift-diffusion equation. Numerical approximations that mirror this property are said to be asymptotic preserving. In this paper we build two discontinuous Galerkin methods to the semiconductor model: one with the standard upwinding flux and the other with a $\varepsilon$-scaled Lax-Friedrichs flux, where 1/$\varepsilon$ is the scale of the collision frequency. We show that these schemes are uniformly stable in $\varepsilon$ and are asymptotic preserving. In particular, we discuss what properties the discrete Maxwellian must satisfy in order for the schemes to converge in $\varepsilon$ to an accurate $h$-approximation of the drift diffusion limit. Discrete versions of the drift-diffusion equation and error estimates in several norms with respect to $\varepsilon$ and the spacial resolution are also included.

math.NA

A Predictor-Corrector Strategy for Adaptivity in Dynamical Low-Rank Approximations

In this paper, we present a predictor-corrector strategy for constructing rank-adaptive dynamical low-rank approximations (DLRAs) of matrix-valued ODE systems. The strategy is a compromise between (i) low-rank step-truncation approaches that alternately evolve and compress solutions and (ii) strict DLRA approaches that augment the low-rank manifold using subspaces generated locally in time by the DLRA integrator. The strategy is based on an analysis of the error between a forward temporal update into the ambient full-rank space, which is typically computed in a step-truncation approach before re-compressing, and the standard DLRA update, which is forced to live in a low-rank manifold. We use this error, without requiring its full-rank representation, to correct the DLRA solution. A key ingredient for maintaining a low-rank representation of the error is a randomized singular value decomposition (SVD), which introduces some degree of stochastic variability into the implementation. The strategy is formulated and implemented in the context of discontinuous Galerkin spatial discretizations of partial differential equations and applied to several versions of DLRA methods found in the literature, as well as a new variant. Numerical experiments comparing the predictor-corrector strategy to other methods demonstrate robustness to overcome short-comings of step truncation or strict DLRA approaches: the former may require more memory than is strictly needed while the latter may miss transients solution features that cannot be recovered. The effect of randomization, tolerances, and other implementation parameters is also explored.

math.NA

$H^1$-norm error estimate for a nonstandard finite element approximation of second-order linear elliptic PDEs in non-divergence form

This paper establishes the optimal $H^1$-norm error estimate for a nonstandard finite element method for approximating $H^2$ strong solutions of second order linear elliptic PDEs in non-divergence form with continuous coefficients. To circumvent the difficulty of lacking an effective duality argument for this class of PDEs, a new analysis technique is introduced; the crux of it is to establish an $H^1$-norm stability estimate for the finite element approximation operator which mimics a similar estimate for the underlying PDE operator recently established by the authors and its proof is based on a freezing coefficient technique and a topological argument. Moreover, both the $H^1$-norm stability and error estimate also hold for the linear finite element method.

math.NA

Counting odd numbers in truncations of Pascal's triangle

A "truncation" of Pascal's triangle is a triangular array of numbers that satisfies the usual Pascal recurrence but with a boundary condition that declares some terminal set of numbers along each row of the array to be zero. Presented here is a family of natural truncations of Pascal's triangle that generalize a kind of Catalan triangle. The numbers in each array are realized as differences of binomial coefficients, as counts of certain lattice paths and tableaux, and as entries of representing matrices for certain linear transformations of polynomial spaces. Lucas's theorem is applied to determine precisely those truncations for which the number of odd entries on each row is a power of two.

math.CO

Analysis of the Vanishing Moment Method and its Finite Element Approximations for Second-order Linear Elliptic PDEs in Non-divergence Form

This paper is concerned with continuous and discrete approximations of $W^{2,p}$ strong solutions of second-order linear elliptic partial differential equations (PDEs) in non-divergence form. The continuous approximation of these equations is achieved through the Vanishing Moment Method (VMM) which adds a small biharmonic term to the PDE. The structure of the new fourth-order PDE is a natural fit for Galerkin-type methods unlike the original second order equation since the highest order term is in divergence form. The well-posedness of the weak form of the perturbed fourth order equation is shown as well as error estimates for approximating the strong solution of the original second-order PDE. A $C^1$ finite element method is then proposed for the fourth order equation, and its existence and uniqueness of solutions as well as optimal error estimates in the $H^2$ norm are shown. Lastly, numerical tests are given to show the validity of the method.

math.NA

A Discontinuous Ritz Method for a Class of Calculus of Variations Problems

This paper develops an analogue (or counterpart) to discontinuous Galerkin (DG) methods for approximating a general class of calculus of variations problems. The proposed method, called the discontinuous Ritz (DR) method, constructs a numerical solution by minimizing a discrete energy over DG function spaces. The discrete energy includes standard penalization terms as well as the DG finite element (DG-FE) numerical derivatives developed recently by Feng, Lewis, and Neilan in [Feng2013]. It is proved that the proposed DR method converges and that the DG-FE numerical derivatives exhibit a compactness property which is desirable and crucial for applying the proposed DR method to problems with more complex energy functionals. Numerical tests are provided on the classical $p$-Laplace problem to gauge the performance of the proposed DR method.

math.NA

An enhanced finite element method for a class of variational problems exhibiting the Lavrentiev gap phenomenon

This paper develops an enhanced finite element method for approximating a class of variational problems which exhibit the \textit{Lavrentiev gap phenomenon} in the sense that the minimum values of the energy functional have a nontrivial gap when the functional is minimized on spaces $W^{1,1}$ and $W^{1,\infty}$. To remedy the standard finite element method, which fails to converge for such variational problems, a simple and effective cut-off procedure is utilized to design the (enhanced finite element) discrete energy functional. In essence the proposed discrete energy functional curbs the gap phenomenon by capping the derivatives of its input on a scale of $O(h^{-\alpha})$ (where $h$ denotes the mesh size) for some positive constant $\alpha$. A sufficient condition is proposed for determining the problem-dependent parameter $\a$. Extensive 1-D and 2-D numerical experiment results are provided to show the convergence behavior and the performance of the proposed enhanced finite element method.

math.NA

Interior Penalty Discontinuous Galerkin Methods for Second Order Linear Non-Divergence Form Elliptic PDEs

This paper develops interior penalty discontinuous Galerkin (IP-DG) methods to approximate $W^{2,p}$ strong solutions of second order linear elliptic partial differential equations (PDEs) in non-divergence form with continuous coefficients. The proposed IP-DG methods are closely related to the IP-DG methods for advection-diffusion equations, and they are easy to implement on existing standard IP-DG software platforms. It is proved that the proposed IP-DG methods have unique solutions and converge with optimal rate to the $W^{2,p}$ strong solution in a discrete $W^{2,p}$-norm. The crux of the analysis is to establish a DG discrete counterpart of the Calderon-Zygmund estimate and to adapt a freezing coefficient technique used for the PDE analysis at the discrete level. As a byproduct of our analysis, we also establish broken $W^{1,p}$-norm error estimates for IP-DG approximations of constant coefficient elliptic PDEs. Numerical experiments are provided to gauge the performance of the proposed IP-DG methods and to validate the theoretical convergence results.

math.NA