SearcharxivSearch

arXiv subjects

Anne Greenbaum

Publications and source records attributed to Anne Greenbaum.

15 recordsLinked to original sources

Linear Systems and Eigenvalue Problems: Open Questions from a Simons Workshop

This document presents a series of open questions arising in matrix computations, i.e., the numerical solution of linear algebra problems. It is a result of working groups at the workshop Linear Systems and Eigenvalue Problems, which was organized at the Simons Institute for the Theory of Computing program on Complexity and Linear Algebra in Fall 2025. The complexity and numerical solution of linear algebra problems is a crosscutting area between theoretical computer science and numerical analysis. The value of the particular problem formulations here is that they were produced via discussions between researchers from both groups. The open questions are organized in five categories: iterative solvers for linear systems, eigenvalue computation, low-rank approximation, randomized sketching, and other areas including tensors, quantum systems, and matrix functions. (Updated to reflect the status of the open problems as of August 20, 2026.)

math.NA

Stable algorithms for general linear systems by preconditioning the normal equations

This paper studies the solution of nonsymmetric linear systems by preconditioned Krylov methods based on the normal equations, LSQR in particular. On some examples, preconditioned LSQR is seen to produce errors many orders of magnitude larger than classical direct methods; this paper demonstrates that the attainable accuracy of preconditioned LSQR can be greatly improved by applying iterative refinement or restarting when the accuracy stalls. This observation is supported by rigorous backward error analysis. This paper also provides a discussion of the relative merits of GMRES and LSQR for solving nonsymmetric linear systems, demonstrates stability for left-preconditioned LSQR without iterative refinement, and shows that iterative refinement can also improve the accuracy of preconditioned conjugate gradient.

math.NA

When is the Resolvent Like a Rank One Matrix?

For a square matrix $A$, the resolvent of $A$ at a point $z \in \mathbb{C}$ is defined as $(A-zI )^{-1}$. We consider the set of points $z \in \mathbb{C}$ where the relative difference in 2-norm between the resolvent and the nearest rank one matrix is less than a given number $\epsilon \in (0,1)$. We establish a relationship between this set and the $\epsilon$-pseudospectrum of $A$, and we derive specific results about this set for Jordan blocks and for a class of large Toeplitz matrices. We also derive disks about the eigenvalues of $A$ that are contained in this set, and this leads to some new results on disks about the eigenvalues that are contained in the $\epsilon$-pseudospectrum of $A$. In addition, we consider the set of points $z \in \mathbb{C}$ where the absolute value of the inner product of the left and right singular vectors corresponding to the largest singular value of the resolvent is less than $\epsilon$. We demonstrate numerically that this set can be almost as large as the one where the relative difference between the resolvent and the nearest rank one matrix is less than $\epsilon$ and we give a partial explanation for this. Some possible applications are discussed.

math.NA

Optimal Polynomial Approximation to Rational Matrix Functions Using the Arnoldi Algorithm

Given an $n$ by $n$ matrix $A$ and an $n$-vector $b$, along with a rational function $R(z) := D(z )^{-1} N(z)$, we show how to find the optimal approximation to $R(A) b$ from the Krylov space, $\mbox{span}( b, Ab, \ldots , A^{k-1} b)$, using the basis vectors produced by the Arnoldi algorithm. To find this optimal approximation requires running $\max \{ \mbox{deg} (D) , \mbox{deg} (N) \} - 1$ extra Arnoldi steps and solving a $k + \max \{ \mbox{deg} (D) , \mbox{deg} (N) \}$ by $k$ least squares problem. Here {\em optimal} is taken to mean optimal in the $D(A )^{*} D(A)$-norm. Similar to the case for linear systems, we show that eigenvalues alone cannot provide information about the convergence behavior of this algorithm and we discuss other possible error bounds for highly nonnormal matrices.

math.NA

Nearly Optimal Approximation of Matrix Functions by the Lanczos Method

Approximating the action of a matrix function $f(\mathbf{A})$ on a vector $\mathbf{b}$ is an increasingly important primitive in machine learning, data science, and statistics, with applications such as sampling high dimensional Gaussians, Gaussian process regression and Bayesian inference, principle component analysis, and approximating Hessian spectral densities. Over the past decade, a number of algorithms enjoying strong theoretical guarantees have been proposed for this task. Many of the most successful belong to a family of algorithms called Krylov subspace methods. Remarkably, a classic Krylov subspace method, called the Lanczos method for matrix functions (Lanczos-FA), frequently outperforms newer methods in practice. Our main result is a theoretical justification for this finding: we show that, for a natural class of rational functions, Lanczos-FA matches the error of the best possible Krylov subspace method up to a multiplicative approximation factor. The approximation factor depends on the degree of $f(x)$'s denominator and the condition number of $\mathbf{A}$, but not on the number of iterations $k$. Our result provides a strong justification for the excellent performance of Lanczos-FA, especially on functions that are well approximated by rationals, such as the matrix square root.

math.NA

GMRES, pseudospectra, and Crouzeix's conjecture for shifted and scaled Ginibre matrices

We study the GMRES algorithm applied to linear systems of equations involving a scaled and shifted $N\times N$ matrix whose entries are independent complex Gaussians. When the right hand side of this linear system is independent of this random matrix, the $N\to\infty$ behavior of the GMRES residual error can be determined exactly. To handle cases where the right hand side depends on the random matrix, we study the pseudospectra and numerical range of Ginibre matrices and prove a restricted version of Crouzeix's conjecture.

math.NA

K-Spectral Sets

We use results in [M. Crouzeix and A. Greenbaum,Spectral sets: numerical range and beyond, SIAM Jour. Matrix Anal. Appl., 40 (2019), pp. 1087-1101] to derive a variety of K-spectral sets and show how they can be used in some applications. We compare the K-values derived here to those that can be derived from a straightforward application of the Cauchy integral formula, by replacing the norm of the integral by the integral of the resolvent norm. While, in some cases, the new upper bounds on the optimal K-value are much tighter than those from the Cauchy integral formula, we show that in many cases of interest, the two values are of the same order of magnitude, with the bounds from the Cauchy integral formula actually being slightly smaller. We give a partial explanation of this in terms of the numerical range of the resolvent at points near an ill-conditioned eigenvalue.

math.NA

Low-memory Krylov subspace methods for optimal rational matrix function approximation

We describe a Lanczos-based algorithm for approximating the product of a rational matrix function with a vector. This algorithm, which we call the Lanczos method for optimal rational matrix function approximation (Lanczos-OR), returns the optimal approximation from a given Krylov subspace in a norm depending on the rational function's denominator, and can be computed using the information from a slightly larger Krylov subspace. We also provide a low-memory implementation which only requires storing a number of vectors proportional to the denominator degree of the rational function. Finally, we show that Lanczos-OR can be used to derive algorithms for computing other matrix functions, including the matrix sign function and quadrature based rational function approximations. In many cases, it improves on the approximation quality of prior approaches, including the standard Lanczos method, with little additional computational overhead.

math.NA

Error bounds for Lanczos-based matrix function approximation

We analyze the Lanczos method for matrix function approximation (Lanczos-FA), an iterative algorithm for computing $f(\mathbf{A}) \mathbf{b}$ when $\mathbf{A}$ is a Hermitian matrix and $\mathbf{b}$ is a given vector. Assuming that $f : \mathbb{C} \rightarrow \mathbb{C}$ is piecewise analytic, we give a framework, based on the Cauchy integral formula, which can be used to derive a priori and a posteriori error bounds for Lanczos-FA in terms of the error of Lanczos used to solve linear systems. Unlike many error bounds for Lanczos-FA, these bounds account for fine-grained properties of the spectrum of $\mathbf{A}$, such as clustered or isolated eigenvalues. Our results are derived assuming exact arithmetic, but we show that they are easily extended to finite precision computations using existing theory about the Lanczos algorithm in finite precision. We also provide generalized bounds for the Lanczos method used to approximate quadratic forms $\mathbf{b}^\textsf{H} f(\mathbf{A}) \mathbf{b}$, and demonstrate the effectiveness of our bounds with numerical experiments.

math.NA

Crouzeix's Conjecture and related problems

In this paper, we establish several results related to Crouzeix's conjecture. We show that the conjecture holds for contractions with eigenvalues that are sufficiently well-separated. This separation is measured by the so-called separation constant, which is defined in terms of the pseudohyperbolic metric. Moreover, we study general properties of related extremal functions and associated vectors. Throughout, compressions of the shift serve as illustrating examples which also allow for refined results.

math.FA

On the Convergence Rate of Variants of the Conjugate Gradient Algorithm in Finite Precision Arithmetic

We consider three mathematically equivalent variants of the conjugate gradient (CG) algorithm and how they perform in finite precision arithmetic. It was shown in [{\em Behavior of slightly perturbed Lanczos and conjugate-gradient recurrences}, Lin.~Alg.~Appl., 113 (1989), pp.~7-63] that under certain conditions the convergence of a slightly perturbed CG computation is like that of exact CG for a matrix with many eigenvalues distributed throughout tiny intervals about the eigenvalues of the given matrix, the size of the intervals being determined by how closely these conditions are satisfied. We determine to what extent each of these variants satisfies the desired conditions, using a set of test problems and show that there is significant correlation between how well these conditions are satisfied and how well the finite precision computation converges before reaching its ultimately attainable accuracy. We show that for problems where the width of the intervals containing the eigenvalues of the associated exact CG matrix makes a significant difference in the behavior of exact CG, the different CG variants behave differently in finite precision arithmetic. For problems where the interval width makes little difference or where the convergence of exact CG is essentially governed by the upper bound based on the square root of the condition number of the matrix, the different CG variants converge similarly in finite precision arithmetic until the ultimate level of accuracy is achieved, although this ultimate level of accuracy may be different for the different variants. This points to the need for testing new CG variants on problems that are especially sensitive to rounding errors.

math.NA

First-order Perturbation Theory for Eigenvalues and Eigenvectors

We present first-order perturbation analysis of a simple eigenvalue and the corresponding right and left eigenvectors of a general square matrix, not assumed to be Hermitian or normal. The eigenvalue result is well known to a broad scientific community. The treatment of eigenvectors is more complicated, with a perturbation theory that is not so well known outside a community of specialists. We give two different proofs of the main eigenvector perturbation theorem. The first, a block-diagonalization technique inspired by the numerical linear algebra research community and based on the implicit function theorem, has apparently not appeared in the literature in this form. The second, based on complex function theory and on eigenprojectors, as is standard in analytic perturbation theory, is a simplified version of well-known results in the literature. The second derivation uses a convenient normalization of the right and left eigenvectors defined in terms of the associated eigenprojector, but although this dates back to the 1950s, it is rarely discussed in the literature. We then show how the eigenvector perturbation theory is easily extended to handle other normalizations that are often used in practice. We also explain how to verify the perturbation results computationally. We conclude with some remarks about difficulties introduced by multiple eigenvalues and give references to work on perturbation of invariant subspaces corresponding to multiple or clustered eigenvalues. Throughout the paper we give extensive bibliographic commentary and references for further reading.

math.NA

Spectral Sets: Numerical Range and Beyond

We extend the proof in [M.~Crouzeix and C.~Palencia, {\em The numerical range is a $(1 + \sqrt{2})$-spectral set}, SIAM Jour.~Matrix Anal.~Appl., 38 (2017), pp.~649-655] to show that other regions in the complex plane are $K$-spectral sets. In particular, we show that various annular regions are $(1 + \sqrt{2} )$-spectral sets and that a more general convex region with a circular hole or cutout is a $(3 + 2 \sqrt{3} )$-spectral set. We demonstrate how these results can be used to give bounds on the convergence rate of the GMRES algorithm for solving linear systems and on that of rational Krylov subspace methods for approximating $f(A)b$, where $A$ is a square matrix, $b$ is a given vector, and $f$ is a function that can be uniformly approximated on such a region by rational functions with poles outside the region.

math.SP

Some Extensions of the Crouzeix-Palencia Result

In [{\em The Numerical Range is a $(1 + \sqrt{2})$-Spectral Set}, SIAM J. Matrix Anal. Appl. 38 (2017), pp.~649-655], Crouzeix and Palencia show that the numerical range of a square matrix or linear operator $A$ is a $(1 + \sqrt{2})$-spectral set for $A$; that is, for any function $f$ analytic in the interior of the numerical range $W(A)$ and continuous on its boundary, the inequality $\| f(A) \| \leq (1 + \sqrt{2} ) \| f \|_{W(A)}$ holds, where the norm on the left is the operator 2-norm and $\| f \|_{W(A)}$ on the right denotes the supremum of $| f(z) |$ over $z \in W(A)$. In this paper, we show how the arguments in their paper can be extended to show that other regions in the complex plane that do {\em not} necessarily contain $W(A)$ are $K$-spectral sets for a value of $K$ that may be close to $1 + \sqrt{2}$. We also find some special cases in which the constant $(1 + \sqrt{2})$ for $W(A)$ can be replaced by $2$, which is the value conjectured by Crouzeix.

math.NA