SearcharxivSearch

arXiv subjects

Michael L. Overton

Publications and source records attributed to Michael L. Overton.

17 recordsLinked to original sources

On the Choice of Sign Defining Householder Transformations

It is well known that, when defining Householder transformations, the correct choice of sign in the standard formula is important to avoid cancellation and hence numerical instability. In this note we point out that when the "wrong" choice of sign is used, the extent of the resulting instability depends in a somewhat subtle way on the data leading to cancellation.

math.NA

An Experimental Comparison of Methods for Computing the Numerical Radius

We make an experimental comparison of methods for computing the numerical radius of an $n\times n$ complex matrix, based on two well-known characterizations, the first a nonconvex optimization problem in one real variable and the second a convex optimization problem in $n^{2}+1$ real variables. We make comparisons with respect to both accuracy and computation time using publicly available software.

math.NA

Multi-fidelity robust controller design with gradient sampling

Robust controllers that stabilize dynamical systems even under disturbances and noise are often formulated as solutions of nonsmooth, nonconvex optimization problems. While methods such as gradient sampling can handle the nonconvexity and nonsmoothness, the costs of evaluating the objective function may be substantial, making robust control challenging for dynamical systems with high-dimensional state spaces. In this work, we introduce multi-fidelity variants of gradient sampling that leverage low-cost, low-fidelity models with low-dimensional state spaces for speeding up the optimization process while nonetheless providing convergence guarantees for a high-fidelity model of the system of interest, which is primarily accessed in the last phase of the optimization process. Our first multi-fidelity method initiates gradient sampling on higher fidelity models with starting points obtained from cheaper, lower fidelity models. Our second multi-fidelity method relies on ensembles of gradients that are computed from low- and high-fidelity models. Numerical experiments with controlling the cooling of a steel rail profile and laminar flow in a cylinder wake demonstrate that our new multi-fidelity gradient sampling methods achieve up to two orders of magnitude speedup compared to the single-fidelity gradient sampling method that relies on the high-fidelity model alone.

math.OC

On Properties of Univariate Max Functions at Local Maximizers

More than three decades ago, Boyd and Balakrishnan established a regularity result for the two-norm of a transfer function at maximizers. Their result extends easily to the statement that the maximum eigenvalue of a univariate real analytic Hermitian matrix family is twice continuously differentiable, with Lipschitz second derivative, at all local maximizers, a property that is useful in several applications that we describe. We also investigate whether this smoothness property extends to max functions more generally. We show that the pointwise maximum of a finite set of $q$-times continuously differentiable univariate functions must have zero derivative at a maximizer for $q=1$, but arbitrarily close to the maximizer, the derivative may not be defined, even when $q=3$ and the maximizer is isolated.

math.OC

Local Minimizers of the Crouzeix Ratio: A Nonsmooth Optimization Case Study

Given a square matrix $A$ and a polynomial $p$, the Crouzeix ratio is the norm of the polynomial on the field of values of $A$ divided by the 2-norm of the matrix $p(A)$. Crouzeix's conjecture states that the globally minimal value of the Crouzeix ratio is 0.5, regardless of the matrix order and polynomial degree, and it is known that 1 is a frequently occurring locally minimal value. Making use of a heavy-tailed distribution to initialize our optimization computations, we demonstrate for the first time that the Crouzeix ratio has many other locally minimal values between 0.5 and 1. Besides showing that the same function values are repeatedly obtained for many different starting points, we also verify that an approximate nonsmooth stationarity condition holds at computed candidate local minimizers. We also find that the same locally minimal values are often obtained both when optimizing over real matrices and polynomials, and over complex matrices and polynomials. We argue that minimization of the Crouzeix ratio makes a very interesting nonsmooth optimization case study, illustrating among other things how effective the BFGS method is for nonsmooth, nonconvex optimization. Our method for verifying approximate nonsmooth stationarity is based on what may be a novel approach to finding approximate subgradients of max functions on an interval. Our extensive computations strongly support Crouzeix's conjecture: in all cases, we find that the smallest locally minimal value is 0.5.

math.OC

Finding the strongest stable weightless column with a follower load and relocatable concentrated masses

We consider the problem of optimal placement of concentrated masses along a massless elastic column that is clamped at one end and loaded by a nonconservative follower force at the free end. The goal is to find the largest possible interval such that the variation in the loading parameter within this interval preserves stability of the structure. The stability constraint is nonconvex and nonsmooth, making the optimization problem quite challenging. We give a detailed analytical treatment for the case of two masses, arguing that the optimal parameter configuration approaches the flutter and divergence boundaries of the stability region simultaneously. Furthermore, we conjecture that this property holds for any number of masses, which in turn suggests a simple formula for the maximal load interval for $n$ masses. This conjecture is strongly supported by extensive computational results, obtained using the recently developed open-source software package GRANSO (GRadient-based Algorithm for Non-Smooth Optimization) to maximize the load interval subject to an appropriate formulation of the nonsmooth stability constraint. We hope that our work will provide a foundation for new approaches to classical long-standing problems of stability optimization for nonconservative elastic systems arising in civil and mechanical engineering.

physics.class-ph

Behavior of Limited Memory BFGS when Applied to Nonsmooth Functions and their Nesterov Smoothings

The motivation to study the behavior of limited-memory BFGS (L-BFGS) on nonsmooth optimization problems is based on two empirical observations: the widespread success of L-BFGS in solving large-scale smooth optimization problems, and the effectiveness of the full BFGS method in solving small to medium-sized nonsmooth optimization problems, based on using a gradient, not a subgradient, oracle paradigm. We first summarize our theoretical results on the behavior of the scaled L-BFGS method with one update applied to a simple convex nonsmooth function that is unbounded below, stating conditions under which the method converges to a non-optimal point regardless of the starting point. We then turn to empirically investigating whether the same phenomenon holds more generally,focusing on a difficult problem of Nesterov, as well as eigenvalue optimization problems arising in semidefinite programming applications. We find that when applied to a nonsmooth function directly, L-BFGS, especially its scaled variant, often breaks down with a poor approximation to an optimal solution, in sharp contrast to full BFGS. Unscaled L-BFGS is less prone to breakdown but conducts far more function evaluations per iteration than scaled L-BFGS does, and thus it is slow. Nonetheless, it is often the case that both variants obtain better results than the provably convergent, but slow, subgradient method. On the other hand, when applied to Nesterov's smooth approximation of a nonsmooth function, scaled L-BFGS is generally much more efficient than unscaled L-BFGS, often obtaining good results even when the problem is quite ill-conditioned. Summarizing, for large-scale nonsmooth optimization problems for which full BFGS and other methods for nonsmooth optimization are not practical, it is often better to apply L-BFGS to a smooth approximation of a nonsmooth problem than to apply it directly to the nonsmooth problem.

math.OC

Fixed-Order H-infinity Controller Design via HIFOO, a Specialized Nonsmooth Optimization Package

We report on our experience with fixed-order H-infinity controller design using the HIFOO toolbox. We applied HIFOO to various benchmark fixed (or reduced) order H-infinity controller design problems in the literature, comparing the results with those published for other methods. The results show that HIFOO can be used as an effective alternative to existing methods for fixed-order HIFOO controller design.

eess.SY

H-infinity Strong Stabilization via HIFOO, a Package for Fixed-Order Controller Design

We report on our experience with strong stabilization using HIFOO, a toolbox for H-infinity fixed-order controller design. We applied HIFOO to 21 fixed-order stable H-infinity controller design problems in the literature, comparing the results with those published for other methods. The results show that HIFOO often achieves good H-infinity performance with low-order stable controllers, unlike other methods in the literature.

eess.SY

First-order Perturbation Theory for Eigenvalues and Eigenvectors

We present first-order perturbation analysis of a simple eigenvalue and the corresponding right and left eigenvectors of a general square matrix, not assumed to be Hermitian or normal. The eigenvalue result is well known to a broad scientific community. The treatment of eigenvectors is more complicated, with a perturbation theory that is not so well known outside a community of specialists. We give two different proofs of the main eigenvector perturbation theorem. The first, a block-diagonalization technique inspired by the numerical linear algebra research community and based on the implicit function theorem, has apparently not appeared in the literature in this form. The second, based on complex function theory and on eigenprojectors, as is standard in analytic perturbation theory, is a simplified version of well-known results in the literature. The second derivation uses a convenient normalization of the right and left eigenvectors defined in terms of the associated eigenprojector, but although this dates back to the 1950s, it is rarely discussed in the literature. We then show how the eigenvector perturbation theory is easily extended to handle other normalizations that are often used in practice. We also explain how to verify the perturbation results computationally. We conclude with some remarks about difficulties introduced by multiple eigenvalues and give references to work on perturbation of invariant subspaces corresponding to multiple or clustered eigenvalues. Throughout the paper we give extensive bibliographic commentary and references for further reading.

math.NA

Analysis of Limited-Memory BFGS on a Class of Nonsmooth Convex Functions

The limited memory BFGS (L-BFGS) method is widely used for large-scale unconstrained optimization, but its behavior on nonsmooth problems has received little attention. L-BFGS can be used with or without "scaling"; the use of scaling is normally recommended. A simple special case, when just one BFGS update is stored and used at every iteration, is sometimes also known as memoryless BFGS. We analyze memoryless BFGS with scaling, using any Armijo-Wolfe line search, on the function $f(x) = a|x^{(1)}| + \sum_{i=2}^{n} x^{(i)}$, initiated at any point $x_0$ with $x_0^{(1)}\not = 0$. We show that if $a\ge 2\sqrt{n-1}$, the absolute value of the normalized search direction generated by this method converges to a constant vector, and if, in addition, $a$ is larger than a quantity that depends on the Armijo parameter, then the iterates converge to a non-optimal point $\bar x$ with $\bar x^{(1)}=0$, although $f$ is unbounded below. As we showed in previous work, the gradient method with any Armijo-Wolfe line search also fails on the same function if $a\geq \sqrt{n-1}$ and $a$ is larger than another quantity depending on the Armijo parameter, but scaled memoryless BFGS fails under a weaker condition relating $a$ to the Armijo parameter than that implying failure of the gradient method. Furthermore, in sharp contrast to the gradient method, if a specific standard Armijo-Wolfe bracketing line search is used, scaled memoryless BFGS fails when $a\ge 2 \sqrt{n-1} $ regardless of the Armijo parameter. Finally, numerical experiments indicate that similar results hold for scaled L-BFGS with any fixed number of updates.

math.OC

Partial smoothness of the numerical radius at matrices whose fields of values are disks

Solutions to optimization problems involving the numerical radius often belong to a special class: the set of matrices having field of values a disk centered at the origin. After illustrating this phenomenon with some examples, we illuminate it by studying matrices around which this set of "disk matrices" is a manifold with respect to which the numerical radius is partly smooth. We then apply our results to matrices whose nonzeros consist of a single superdiagonal, such as Jordan blocks and the Crabb matrix related to a well-known conjecture of Crouzeix. Finally, we consider arbitrary complex three-by-three matrices; even in this case, the details are surprisingly intricate. One of our results is that in this real vector space with dimension 18, the set of disk matrices is a semi-algebraic manifold with dimension 12.

math.NA

Analysis of the Gradient Method with an Armijo-Wolfe Line Search on a Class of Nonsmooth Convex Functions

It has long been known that the gradient (steepest descent) method may fail on nonsmooth problems, but the examples that have appeared in the literature are either devised specifically to defeat a gradient or subgradient method with an exact line search or are unstable with respect to perturbation of the initial point. We give an analysis of the gradient method with steplengths satisfying the Armijo and Wolfe inexact line search conditions on the nonsmooth convex function $f(x) = a|x^{(1)}| + \sum_{i=2}^{n} x^{(i)}$. We show that if $a$ is sufficiently large, satisfying a condition that depends only on the Armijo parameter, then, when the method is initiated at any point $x_0 \in \R^n$ with $x^{(1)}_0\not = 0$, the iterates converge to a point $\bar x$ with $\bar x^{(1)}=0$, although $f$ is unbounded below. We also give conditions under which the iterates $f(x_k)\to-\infty$, using a specific Armijo-Wolfe bracketing line search. Our experimental results demonstrate that our analysis is reasonably tight.

math.OC

Gradient Sampling Methods for Nonsmooth Optimization

This paper reviews the gradient sampling methodology for solving nonsmooth, nonconvex optimization problems. An intuitively straightforward gradient sampling algorithm is stated and its convergence properties are summarized. Throughout this discussion, we emphasize the simplicity of gradient sampling as an extension of the steepest descent method for minimizing smooth objectives. We then provide overviews of various enhancements that have been proposed to improve practical performance, as well as of several extensions that have been made in the literature, such as to solve constrained problems. The paper also includes clarification of certain technical aspects of the analysis of gradient sampling algorithms, most notably related to the assumptions one needs to make about the set of points at which the objective is continuously differentiable. Finally, we discuss possible future research directions.

math.OC

Low-Order Control Design using a Reduced-Order Model with a Stability Constraint on the Full-Order Model

We consider low-order controller design for large-scale linear time-invariant dynamical systems with inputs and outputs. Model order reduction is a popular technique, but controllers designed for reduced-order models may result in unstable closed-loop plants when applied to the full-order system. We introduce a new method to design a fixed-order controller by minimizing the $L_\infty$ norm of a reduced-order closed-loop transfer matrix function subject to stability constraints on the closed-loop systems for both the reduced-order and the full-order models. Since the optimization objective and the constraints are all nonsmooth and nonconvex we use a sequential quadratic programming method based on quasi-Newton updating that is intended for this problem class, available in the open-source software package GRANSO. Using a publicly available test set, the controllers obtained by the new method are compared with those computed by the HIFOO (H-Infinity Fixed-Order Optimization) toolbox applied to a reduced-order model alone, which frequently fail to stabilize the closed-loop system for the associated full-order model.

math.OC

Multiobjective Robust Control with HIFOO 2.0

Multiobjective control design is known to be a difficult problem both in theory and practice. Our approach is to search for locally optimal solutions of a nonsmooth optimization problem that is built to incorporate minimization objectives and constraints for multiple plants. We report on the success of this approach using our public-domain Matlab toolbox HIFOO 2.0, comparing our results with benchmarks in the literature.

math.OC

Maximizing the Closed Loop Asymptotic Decay Rate for the Two-Mass-Spring Control Problem

We consider the following problem: find a fixed-order linear controller that maximizes the closed-loop asymptotic decay rate for the classical two-mass-spring system. This can be formulated as the problem of minimizing the abscissa (maximum of the real parts of the roots) of a polynomial whose coefficients depend linearly on the controller parameters. We show that the only order for which there is a non-trivial solution is 2. In this case, we derive a controller that we prove locally maximizes the asymptotic decay rate, using recently developed techniques from nonsmooth analysis.

math.OC