SearcharxivSearch

arXiv subjects

Volker Schulz

Publications and source records attributed to Volker Schulz.

At least 19 recordsLinked to original sources

A JoLT for the KV cache: Near-lossless KV cache compression via joint Lagrangian allocation of Tucker ranks and a rotated residual for llms

The key-value (KV) cache has become the dominant memory cost of transformer inference: it grows with batch size, context length, and depth, and at long context it, rather than the model weights, sets the throughput ceiling. Existing reductions fall into two families. Low-rank methods factor two-dimensional slices of the cache, either per-head matrices or cross-layer feature blocks, and quantization methods lower the bit-width of every entry. Neither exploits the fact that the cache at a layer is naturally a third-order tensor whose three axes, the heads, the tokens, and the features, carry very different amounts of redundancy. We take this tensor view directly. Our method, JoLT (Joint Lagrangian Tucker), applies a partial Tucker decomposition that compresses only the token and feature axes while leaving the head and layer axes intact, then restores the energy that truncation discards with a rotated low-bit residual: a random orthogonal rotation followed by low-bit quantization. A single Lagrangian dual allocates the Tucker ranks and the residual bit-widths together, per layer group and separately for keys and values, under one byte budget. The result is a near-lossless 2-3x compression. Perplexity stays near-lossless on both a grouped-query-attention model (Mistral-7B-v0.3) and a multi-head-attention model (LLaMA-2-13B), and GSM8K accuracy and needle-in-a-haystack retrieval hold at the uncompressed baseline at 2x on both architectures and through 3x on the GQA model. At 2x, JoLT reconstructs the cache to relative Frobenius error 0.009 (K) and 0.006 (V) on both architectures. A randomized-SVD variant, FlashJoLT, delivers a 5-13x compression-time speedup at 1024-token context and matched quality.

cs.LG

A Probabilistic Approach to Shape Derivatives

We introduce a novel mesh-free and direct method for computing the shape derivative in PDE-constrained shape optimization problems. Our approach is based on a probabilistic representation of the shape derivative and is applicable for second-order semilinear elliptic PDEs with Dirichlet boundary conditions and a general class of target functions. The probabilistic representation derives from an extension of a boundary sensitivity result for diffusion processes due to Costantini, Gobet and El Karoui [14]. Moreover, we present a simulation methodology based on our results that does not necessarily require a mesh of the relevant domain, and provide Taylor tests to verify its numerical accuracy

math.OC

Second Order Shape Optimization for an Interface Identification Problem constrained by Nonlocal Models

Since shape optimization methods have been proven useful for identifying interfaces in models governed by partial differential equations, we show how shape optimization techniques can also be applied to an interface identification problem constrained by a nonlocal Dirichlet problem. Here, we focus on deriving the second shape derivative of the corresponding reduced functional and we further investigate a second order optimization algorithm.

math.OC

Schwarz Methods for Nonlocal Problems

The first domain decomposition methods for partial differential equations were already developed in 1870 by H. A. Schwarz. Here we consider a nonlocal Dirichlet problem with variable coefficients, where a nonlocal diffusion operator is used. We find that domain decomposition methods like the so-called Schwarz methods seem to be a natural way to solve these nonlocal problems. In this work we show the convergence for nonlocal problems, where specific symmetric kernels are employed, and present the implementation of the multiplicative and additive Schwarz algorithms in the above mentioned nonlocal setting.

math.NA

Shape optimization in the space of piecewise-smooth shapes for the Bingham flow variational inequality

This paper sets up an approach for shape optimization problems constrained by variational inequalities (VI) in an appropriate shape space. In contrast to classical VI, where no explicit dependence on the domain is given, VI constrained shape optimization problems are in particular highly challenging because of two main reasons: Firstly, one needs to operate in inherently non-linear, non-convex and infinite-dimensional shape spaces. Secondly, the problem cannot be solved directly without any regularization techniques in general because, e.g., one cannot expect the existence of the shape derivative for an arbitrary shape functional depending on solutions to VI. This paper introduces a specific shape manifold and presents an optimization technique to handle the non-differentiabilities on this shape manifold. In particular, we formulate an optimization system based on G\^ateaux semiderivatives and Eulerian derivatives for a shape optimization problem constrained by the Bingham flow variational inequality. Numerical results show the applicability and efficiency of the proposed approach.

math.OC

Interface Identification constrained by Local-to-Nonlocal Coupling

Models of physical phenomena that use nonlocal operators are better suited for some applications than their classical counterparts that employ partial differential operators. However, the numerical solution of these nonlocal problems can be quite expensive. Therefore, Local-to-Nonlocal couplings have emerged that combine partial differential operators with nonlocal operators. In this work, we make use of an energy-based Local-to-Nonlocal coupling that serves as a constraint for an interface identification problem.

math.OC

Shape Optimization for the Mitigation of Coastal Erosion via Shallow Water Equations

Coastal erosion describes the displacement of land caused by destructive sea waves, currents or tides. Major efforts have been made to mitigate these effects using groins, breakwaters and various other structures. We try to address this problem by applying shape optimization techniques to the obstacles. We model the propagation of waves towards the coastline, using two-dimensional shallow water equations. The obstacle's shape is optimized over an appropriate cost function to minimize the height and velocities of water waves along the shore, without relying on a finite-dimensional design space but based on shape calculus.

math.OC

Shape Optimization for the Mitigation of Coastal Erosion via Smoothed Particle Hydrodynamics

Adjoint-based shape optimization most often relies on Eulerian flow field formulations. However, since Lagrangian particle methods are the natural choice for solving sedimentation problems in oceanography, extensions to the Lagrangian framework are desirable. For the mitigation of coastal erosion, we perform shape optimization for fluid flows, that are described by Lagrangian shallow water equations and discretized via smoothed particle hydrodynamics. The obstacle's shape is hereby optimized over an appropriate cost function to minimize the height of water waves along the shoreline based on shape calculus. Theoretical results will be numerically verified by exploring different scenarios.

physics.flu-dyn

Shape Optimization for the Mitigation of Coastal Erosion via Porous Shallow Water Equations

Coastal erosion describes the displacement of land caused by destructive sea waves, currents or tides. Major efforts have been made to mitigate these effects using groynes, breakwaters and various other structures. We address this problem by applying shape optimization techniques on the obstacles. We model the propagation of waves towards the coastline using two-dimensional porous Shallow Water Equations with artificial viscosity. The obstacle's shape, which is assumed to be permeable, is optimized over an appropriate cost function to minimize the height and velocities of water waves along the shore, without relying on a finite-dimensional design space, but based on shape calculus.

math.OC

Shape optimization for interface identification in nonlocal models

Shape optimization methods have been proven useful for identifying interfaces in models governed by partial differential equations. Here we consider a class of shape optimization problems constrained by nonlocal equations which involve interface-dependent kernels. We derive a novel shape derivative associated to the nonlocal system model and solve the problem by established numerical techniques.

math.OC

nlfem: A flexible 2d Fem Code for Nonlocal Convection-Diffusion and Mechanics

In this work we present the mathematical foundation of an assembly code for finite element approximations of nonlocal models with compactly supported, weakly singular kernels. We demonstrate the code on a nonlocal diffusion model in various configurations and on a two-dimensional bond-based peridynamics model. The code nlfem is published under the MIT License and can be freely downloaded.

math.NA

Shape Optimization for the Mitigation of Coastal Erosion via the Helmholtz Equation

Coastal erosion describes the displacement of land caused by destructive sea waves, currents or tides. Major efforts have been made to mitigate these effects using groins, breakwaters and various other structures. We try to address this problem by applying shape optimization techniques on the obstacles. A first approach models the propagation of waves towards the coastline, using a 2D time-harmonic system based on the famous Helmholtz equation in the form of a scattering problem. The obstacle's shape is optimized over an appropriate cost function to minimize the height of water waves along the shoreline, without relying on a finite-dimensional design space, but based on shape calculus.

math.OC

Pre-Shape Calculus: Foundations and Application to Mesh Quality Optimization

Deformations of the computational mesh arising from optimization routines usually lead to decrease of mesh quality or even destruction of the mesh. We propose a theoretical framework using pre-shapes to generalize classical shape optimization and calculus. We define pre-shape derivatives and derive according structure and calculus theorems. In particular, tangential directions are featured in pre-shape derivatives, in contrast to classical shape derivatives featuring only normal directions. Techniques from classical shape optimization and -calculus are shown to carry over to this framework. An optimization problem class for mesh quality is introduced, which is solvable by use of pre-shape derivatives. This class allows for simultaneous optimization of classical shape objectives and mesh quality without deteriorating the classical shape optimization solution. The new techniques are implemented and numerically tested for 2D and 3D.

math.OC

Tensor numerical method for optimal control problems constrained by an elliptic operator with general rank-structured coefficients

We introduce tensor numerical techniques for solving optimal control problems constrained by elliptic operators in $\mathbb{R}^d$, $d=2,3$, with variable coefficients, which can be represented in a low rank separable form. We construct a preconditioned iterative method with an adaptive rank truncation for solving the equation for the control function, governed by a sum of the elliptic operator and its inverse $M=A + A^{-1}$, both discretized over large $n^{\otimes d}$, $d=2,3$, spatial grids. Two basic solution schemes are proposed and analyzed. In the first approach, one solves iteratively the initial linear system of equations with the matrix $M$ such that the matrix vector multiplication with the elliptic operator inverse, $y=A^{-1} u,$ is performed as an embedded iteration by using a rank-structured solver for the equation of the form $A y=u$. The second numerical scheme avoids the embedded iteration by reducing the initial equation to an equivalent one with the polynomial system matrix of the form $A^2 +I$. For both schemes, a low Kronecker rank spectrally equivalent preconditioner is constructed by using the corresponding matrix valued function of the anisotropic Laplacian diagonalized in the Fourier basis. Numerical tests for control problems in 2D setting confirm the linear-quadratic complexity scaling of the proposed method in the univariate grid size $n$. Further, we numerically demonstrate that for our low rank solution method, a cascadic multigrid approach reduces the number of PCG iterations considerably, however the total CPU time remains merely the same as for the unigrid iteration.

math.NA

Simultaneous Shape and Mesh Quality Optimization using Pre-Shape Calculus

Computational meshes arising from shape optimization routines commonly suffer from decrease of mesh quality or even destruction of the mesh. In this work, we provide an approach to regularize general shape optimization problems to increase both shape and volume mesh quality. For this, we employ pre-shape calculus (cf. arXiv:2012.09124). Existence of regularized solutions is guaranteed. Further, consistency of modified pre-shape gradient systems is established. We present pre-shape gradient system modifications, which permit simultaneous shape optimization with mesh quality improvement. Optimal shapes to the original problem are left invariant under regularization. The computational burden of our approach is limited, since additional solution of possibly larger (non-)linear systems for regularized shape gradients is not necessary. We implement and compare pre-shape gradient regularization approaches for a hard to solve 2D problem. As our approach does not depend on the choice of metrics representing shape gradients, we employ and compare several different metrics.

math.OC

Tensor Method for Optimal Control Problems Constrained by Fractional 3D Elliptic Operator with Variable Coefficients

We introduce the tensor numerical method for solving optimal control problems that are constrained by fractional 2D and 3D elliptic operators with variable coefficients. We solve the governing equation for the control function which includes a sum of the fractional operator and its inverse, both discretized over large 3D $n\times n \times n$ spacial grids. Using the diagonalization of the arising matrix valued functions in the eigenbasis of the 1D Sturm-Liouville operators, we construct the rank-structured tensor approximation with controllable precision for the discretized fractional elliptic operators and the respective preconditioner. The right-hand side in the constraining equation (the optimal design function) is supposed to be represented in a form of a low-rank canonical tensor. Then the equation for the control function is solved in a tensor structured format by using preconditioned CG iteration with the adaptive rank truncation procedure that also ensures the accuracy of calculations, given an $\varepsilon$-threshold. This method reduces the numerical cost for solving the control problem to $O(n \log n)$ (plus the quadratic term $O(n^2)$ with a small weight), which is superior to the approaches based on the traditional linear algebra tools that yield at least $O(n^3 \log n)$ complexity in the 3D case. The storage for the representation of all 3D nonlocal operators and functions involved is also estimated by $O(n \log n)$. This essentially outperforms the traditional methods operating with fully populated $n^3 \times n^3$ matrices and vectors in $\mathbb{R}^{n^3}$. Numerical tests for 2D/3D control problems indicate the almost linear complexity scaling of the rank truncated PCG iteration in the univariate grid size $n$.

math.NA

Tensor product method for fast solution of optimal control problems with fractional multidimensional Laplacian in constraints

We introduce the tensor numerical method for solution of the $d$-dimensional optimal control problems with fractional Laplacian type operators in constraints discretized on large $n^{\otimes d}$ tensor-product Cartesian grids. The approach is based on the rank-structured approximation of the matrix valued functions of the corresponding fractional finite difference Laplacian. We solve the equation for the control function, where the system matrix includes the sum of the fractional $d$-dimensional Laplacian and its inverse. The matrix valued functions of discrete Laplace operator on a tensor grid are diagonalized by using the fast Fourier transform (FFT). Then the low rank approximation of the $d$-dimensional tensors obtained by folding of the corresponding large diagonal matrices of eigenvalues are computed, which allows to solve the governing equation for the control function in a tensor-structured format. The existence of low rank canonical approximation to the class of matrix valued functions involved is justified by using the sinc quadrature approximation method applied to the Laplace transform of the generating function. The linear system of equations for the control function is solved by the PCG iterative method with the rank truncation at each iteration step, where the low Kronecker rank preconditioner is precomputed. The right-hand side, the solution vector, and the governing system matrix are maintained in the rank-structured tensor format which beneficially reduces the numerical cost to $O(n\log n)$, outperforming the standard FFT based methods of complexity $O(n^3\log n)$ for 3D case. Numerical tests for the 2D and 3D control problems confirm the linear complexity scaling of the method in the univariate grid size $n$.

math.NA

A Hybrid Objective Function for Robustness of Artificial Neural Networks -- Estimation of Parameters in a Mechanical System

In several studies, hybrid neural networks have proven to be more robust against noisy input data compared to plain data driven neural networks. We consider the task of estimating parameters of a mechanical vehicle model based on acceleration profiles. We introduce a convolutional neural network architecture that is capable to predict the parameters for a family of vehicle models that differ in the unknown parameters. We introduce a convolutional neural network architecture that given sequential data predicts the parameters of the underlying data's dynamics. This network is trained with two objective functions. The first one constitutes a more naive approach that assumes that the true parameters are known. The second objective incorporates the knowledge of the underlying dynamics and is therefore considered as hybrid approach. We show that in terms of robustness, the latter outperforms the first objective on noisy input data.

cs.LG