SearcharxivSearch

arXiv subjects

Bennet Gebken

Publications and source records attributed to Bennet Gebken.

16 recordsLinked to original sources

Superlinear convergence in nonsmooth optimization via higher-order cutting-plane models

A cutting-plane model for a nonsmooth function is the maximum of several first-order expansions centered at different points. Using such a model in a bundle method leads to linear convergence (of serious steps) to a minimum. In smooth optimization, superlinear convergence can be achieved by using higher-order models. We show that the same is true for the nonsmooth case, i.e., we show that cutting-plane models involving higher-order expansions can be used to achieve superlinear convergence in nonsmooth optimization. We first formally define higher-order cutting-plane models for lower-$C^2$ functions and derive an error estimate. Afterwards, we construct a trust-region bundle method based on these models that achieves local superlinear convergence of serious steps, and overall superlinear convergence for certain finite max-type functions. Finally, we verify the superlinear convergence in numerical experiments.

math.OC

Enclosing minima in nonsmooth optimization via trust regions of higher-order cutting-plane models

We propose a globally convergent trust-region bundle method for minimizing lower-$C^2$ functions using higher-order cutting-plane models. Under certain growth assumptions on the objective around its minimum, the method is able to compute infinitely many trust regions of decreasing size that contain the minimum. We show that these growth assumptions are satisfied for certain finite max-type functions with sharp or quadratic growth. Enclosing the minimum in this way can be used to initialize local superlinearly convergent methods, which we demonstrate in numerical experiments.

math.OC

Technical results on the convergence of quasi-Newton methods for nonsmooth optimization

It is well-known by now that the BFGS method is an effective method for minimizing nonsmooth functions. However, despite its popularity, theoretical convergence results are almost non-existent. One of the difficulties when analyzing the nonsmooth case is the fact that the secant equation forces certain eigenvalues of the quasi-Newton matrix to vanish, which is a behavior that has not yet been fully analyzed. In this article, we show what kind of behavior of the eigenvalues would be sufficient to be able to prove the convergence for piecewise differentiable functions. More precisely, we derive assumptions on the behavior from numerical experiments and then prove criticality of the limit under these assumptions. Furthermore, we show how quasi-Newton methods are able to explore the piecewise structure. While we do not prove that the observed behavior of the eigenvalues actually occurs, we believe that these results still give insight, and a certain intuition, for the convergence for nonsmooth functions.

math.OC

Using second-order information in gradient sampling methods for nonsmooth optimization

In this article, we introduce a novel concept for second-order information of a nonsmooth function inspired by the Goldstein eps-subdifferential. It comprises the coefficients of all existing second-order Taylor expansions in an eps-ball around a given point. Based on this concept, we define a model of the objective as the maximum of these Taylor expansions, and derive a sampling scheme for its approximation in practice. Minimization of this model induces a simple descent method, for which we show convergence for the case where the objective is convex or of max-type. While we do not prove any rate of convergence of this method, numerical experiments suggest superlinear behavior with respect to the number of oracle calls of the objective.

math.OC

Analyzing the speed of convergence in nonsmooth optimization via the Goldstein subdifferential with application to descent methods

The Goldstein $\varepsilon$-subdifferential is a relaxed version of the Clarke subdifferential which has recently appeared in several algorithms for nonsmooth optimization. With it comes the notion of $(\varepsilon,\delta)$-critical points, which are points in which the element with the smallest norm in the $\varepsilon$-subdifferential has norm at most $\delta$. To obtain points that are critical in the classical sense, $\varepsilon$ and $\delta$ must vanish. In this article, we analyze at which speed the distance of $(\varepsilon,\delta)$-critical points to the minimum vanishes with respect to $\varepsilon$ and $\delta$. Afterwards, we apply our results to gradient sampling methods and perform numerical experiments. Throughout the article, we put a special emphasis on supporting the theoretical results with simple examples that visualize them.

math.OC

A Descent Method for Nonsmooth Multiobjective Optimization in Hilbert Spaces

The efficient optimization method for locally Lipschitz continuous multiobjective optimization problems from [1] is extended from finite-dimensional problems to general Hilbert spaces. The method iteratively computes Pareto critical points, where in each iteration, an approximation of the subdifferential is computed in an efficient manner and then used to compute a common descent direction for all objective functions. To prove convergence, we present some new optimality results for nonsmooth multiobjective optimization problems in Hilbert spaces. Using these, we can show that every accumulation point of the sequence generated by our algorithm is Pareto critical under common assumptions. Computational efficiency for finding Pareto critical points is numerically demonstrated for multiobjective optimal control of an obstacle problem.

math.OC

A note on the convergence of deterministic gradient sampling in nonsmooth optimization

Approximation of subdifferentials is one of the main tasks when computing descent directions for nonsmooth optimization problems. In this article, we propose a bisection method for weakly lower semismooth functions which is able to compute new subgradients that improve a given approximation in case a direction with insufficient descent was computed. Combined with a recently proposed deterministic gradient sampling approach, this yields a deterministic and provably convergent way to approximate subdifferentials for computing descent directions.

math.OC

Multiobjective Optimization of Non-Smooth PDE-Constrained Problems

Multiobjective optimization plays an increasingly important role in modern applications, where several criteria are often of equal importance. The task in multiobjective optimization and multiobjective optimal control is therefore to compute the set of optimal compromises (the Pareto set) between the conflicting objectives. The advances in algorithms and the increasing interest in Pareto-optimal solutions have led to a wide range of new applications related to optimal and feedback control - potentially with non-smoothness both on the level of the objectives or in the system dynamics. This results in new challenges such as dealing with expensive models (e.g., governed by partial differential equations (PDEs)) and developing dedicated algorithms handling the non-smoothness. Since in contrast to single-objective optimization, the Pareto set generally consists of an infinite number of solutions, the computational effort can quickly become challenging, which is particularly problematic when the objectives are costly to evaluate or when a solution has to be presented very quickly. This article gives an overview of recent developments in the field of multiobjective optimization of non-smooth PDE-constrained problems. In particular we report on the advances achieved within Project 2 "Multiobjective Optimization of Non-Smooth PDE-Constrained Problems - Switches, State Constraints and Model Order Reduction" of the DFG Priority Programm 1962 "Non-smooth and Complementarity-based Distributed Parameter Systems: Simulation and Hierarchical Optimization".

math.OC

On the structure of regularization paths for piecewise differentiable regularization terms

Regularization is used in many different areas of optimization when solutions are sought which not only minimize a given function, but also possess a certain degree of regularity. Popular applications are image denoising, sparse regression and machine learning. Since the choice of the regularization parameter is crucial but often difficult, path-following methods are used to approximate the entire regularization path, i.e., the set of all possible solutions for all regularization parameters. Due to their nature, the development of these methods requires structural results about the regularization path. The goal of this article is to derive these results for the case of a smooth objective function which is penalized by a piecewise differentiable regularization term. We do this by treating regularization as a multiobjective optimization problem. Our results suggest that even in this general case, the regularization path is piecewise smooth. Moreover, our theory allows for a classification of the nonsmooth features that occur in between smooth parts. This is demonstrated in two applications, namely support-vector machines and exact penalty methods.

math.OC

On the Treatment of Optimization Problems with L1 Penalty Terms via Multiobjective Continuation

We present a novel algorithm that allows us to gain detailed insight into the effects of sparsity in linear and nonlinear optimization, which is of great importance in many scientific areas such as image and signal processing, medical imaging, compressed sensing, and machine learning (e.g., for the training of neural networks). Sparsity is an important feature to ensure robustness against noisy data, but also to find models that are interpretable and easy to analyze due to the small number of relevant terms. It is common practice to enforce sparsity by adding the $\ell_1$-norm as a weighted penalty term. In order to gain a better understanding and to allow for an informed model selection, we directly solve the corresponding multiobjective optimization problem (MOP) that arises when we minimize the main objective and the $\ell_1$-norm simultaneously. As this MOP is in general non-convex for nonlinear objectives, the weighting method will fail to provide all optimal compromises. To avoid this issue, we present a continuation method which is specifically tailored to MOPs with two objective functions one of which is the $\ell_1$-norm. Our method can be seen as a generalization of well-known homotopy methods for linear regression problems to the nonlinear case. Several numerical examples - including neural network training - demonstrate our theoretical findings and the additional insight that can be gained by this multiobjective approach.

math.OC

An efficient descent method for locally Lipschitz multiobjective optimization problems

In this article, we present an efficient descent method for locally Lipschitz continuous multiobjective optimization problems (MOPs). The method is realized by combining a theoretical result regarding the computation of descent directions for nonsmooth MOPs with a practical method to approximate the subdifferentials of the objective functions. We show convergence to points which satisfy a necessary condition for Pareto optimality. Using a set of test problems, we compare our method to the multiobjective proximal bundle method by Mäkelä. The results indicate that our method is competitive while being easier to implement. While the number of objective function evaluations is larger, the overall number of subgradient evaluations is lower. Finally, we show that our method can be combined with a subdivision algorithm to compute entire Pareto sets of nonsmooth MOPs.

math.OC

On the Equivariance Properties of Self-adjoint Matrices

We investigate self-adjoint matrices $A\in\mathbb{R}^{n,n}$ with respect to their equivariance properties. We show in particular that a matrix is self-adjoint if and only if it is equivariant with respect to the action of a group $Γ_2(A)\subset \mathbf{O}(n)$ which is isomorphic to $\otimes_{k=1}^n\mathbf{Z}_2$. If the self-adjoint matrix possesses multiple eigenvalues -- this may, for instance, be induced by symmetry properties of an underlying dynamical system -- then $A$ is even equivariant with respect to the action of a group $Γ(A) \simeq \prod_{i = 1}^k \mathbf{O}(m_i)$ where $m_1,\ldots,m_k$ are the multiplicities of the eigenvalues $λ_1,\ldots,λ_k$ of $A$. We discuss implications of this result for equivariant bifurcation problems, and we briefly address further applications for the Procrustes problem, graph symmetries and Taylor expansions.

math.DS

ROM-based multiobjective optimization of elliptic PDEs via numerical continuation

Multiobjective optimization plays an increasingly important role in modern applications, where several objectives are often of equal importance. The task in multiobjective optimization and multiobjective optimal control is therefore to compute the set of optimal compromises (the Pareto set) between the conflicting objectives. Since the Pareto set generally consists of an infinite number of solutions, the computational effort can quickly become challenging which is particularly problematic when the objectives are costly to evaluate as is the case for models governed by partial differential equations (PDEs). To decrease the numerical effort to an affordable amount, surrogate models can be used to replace the expensive PDE evaluations. Existing multiobjective optimization methods using model reduction are limited either to low parameter dimensions or to few (ideally two) objectives. In this article, we present a combination of the reduced basis model reduction method with a continuation approach using inexact gradients. The resulting approach can handle an arbitrary number of objectives while yielding a significant reduction in computing time.

math.OC

Inverse multiobjective optimization: Inferring decision criteria from data

It is a very challenging task to identify the objectives on which a certain decision was based, in particular if several, potentially conflicting criteria are equally important and a continuous set of optimal compromise decisions exists. This task can be understood as the inverse problem of multiobjective optimization, where the goal is to find the objective vector of a given Pareto set. To this end, we present a method to construct the objective vector of a multiobjective optimization problem (MOP) such that the Pareto critical set contains a given set of data points or decision vectors. The key idea is to consider the objective vector in the multiobjective KKT conditions as variable and then search for the objectives that minimize the Euclidean norm of the resulting system of equations. By expressing the objectives in a finite-dimensional basis, we transform this problem into a homogeneous, linear system of equations that can be solved efficiently. There are many important potential applications of this approach. Besides the identification of objectives (both from clean and noisy data), the method can be used for the construction of surrogate models for expensive MOPs, which yields significant speed-ups. Both applications are illustrated using several examples.

math.OC

On the hierarchical structure of Pareto critical sets

In this article we show that the boundary of the Pareto critical set of an unconstrained multiobjective optimization problem (MOP) consists of Pareto critical points of subproblems considering subsets of the objective functions. If the Pareto critical set is completely described by its boundary (e.g. if we have more objective functions than dimensions in the parameter space), this can be used to solve the MOP by solving a number of MOPs with fewer objective functions. If this is not the case, the results can still give insight into the structure of the Pareto critical set. This technique is especially useful for efficiently solving many-objective optimization problems by breaking them down into MOPs with a reduced number of objective functions.

math.OC

A Descent Method for Equality and Inequality Constrained Multiobjective Optimization Problems

In this article we propose a descent method for equality and inequality constrained multiobjective optimization problems (MOPs) which generalizes the steepest descent method for unconstrained MOPs by Fliege and Svaiter to constrained problems by using two active set strategies. Under some regularity assumptions on the problem, we show that accumulation points of our descent method satisfy a necessary condition for local Pareto optimality. Finally, we show the typical behavior of our method in a numerical example.

math.OC