SearcharxivSearch

arXiv subjects

Patrick Mehlitz

Publications and source records attributed to Patrick Mehlitz.

At least 19 recordsLinked to original sources

A penalty-type method for relaxed inverse optimal control problems

This paper is devoted to the introduction and analysis of a penalty-type method for the numerical treatment of a class of bilevel optimization problems arising from inverse optimal control. The algorithm is designed to compute stationary points of the associated relaxed value function reformulation. This is achieved by determining a sequence of stationary points associated with a sequence of surrogate problems where the relaxed value function constraint is penalized, where the updates of upper- and lower-level decision variables are decoupled, and where the penalty parameter is enlarged only in those iterations which do not come along with a sufficient improvement of some feasibility measure. The resulting method does not comprise any linesearch, the lower-level problem has to be evaluated just once per iteration, and the penalty parameter does not need to be driven to infinity. Nevertheless, subsequential convergence results are obtained under reasonable assumptions. Numerical experiments, where the relaxation parameter is also driven to zero, visualize effectiveness of the approach.

math.OC

Elastically safeguarded augmented Lagrangian methods

We investigate, theoretically and numerically, a class of elastically safeguarded augmented Lagrangian methods for nonlinear optimization problems with inequality and equality constraints. Safeguarded augmented Lagrangian methods are known to exhibit substantially stronger global convergence guarantees than their non-safeguarded counterparts, making them attractive in practice. A persistent limitation, however, is that existing methods rely on a fixed safeguard, whose a priori selection can be difficult and inherently limits adaptivity, e.g., with respect to problem scaling. We propose an adaptive safeguarding mechanism that allows the safeguard to grow dynamically, overcoming these drawbacks while preserving the desirable global convergence properties of variants with a fixed safeguard. We further establish complexity bounds in terms of worst-case iteration counts. Numerical experiments comparing ALGENCAN, a state-of-the-art solver with rigid safeguard, against an elastically safeguarded variant thereof confirm that elastic safeguarding yields consistent practical benefits.

math.OC

On directional local minimality and directional optimality conditions in nonsmooth optimization

This paper considers the unconstrained minimization of a lower semicontinuous function. Exploiting first and second subderivatives, directional limiting subdifferentials, and directional proximal subdifferentials, necessary and sufficient first- and second-order optimality conditions are derived that build upon the recently introduced notion of directional local minimality. These results then also yield optimality conditions for conventional nondirectional local minimality which are stated in terms of so-called critical directions and variational objects depending on them. Illustrative examples show that the derived conditions allow for a finer analysis than classical nondirectional optimality conditions.

math.OC

Uniqueness and stability of Lagrange multipliers and associated qualification conditions

This paper is concerned with uniqueness and stability of Lagrange multipliers for constrained optimization problems in abstract spaces. It is well known that validity of the strict Robinson-Zowe-Kurcyusz condition implies the so-called isolated calmness, a one-sided Lipschitz property tailored for set-valued mappings, of some Lagrange multiplier mapping associated with a perturbed version of the original optimization problem, and the latter indeed is enough to guarantee uniqueness of the Lagrange multiplier. The paper studies the isolated calmness of the Lagrange multiplier mapping in detail. Exemplary, it is shown that this condition is sufficient for the Robinson-Zowe-Kurcyusz constraint qualification and, in the presence of additional assumptions, even equivalent to the strict Robinson-Zowe-Kurcyusz condition. Illustrative examples are presented to underline the necessity of postulated assumptions.

math.OC

Approximate directional stationarity and associated qualification conditions

Approximate stationarity conditions provide necessary optimality conditions without requiring additional assumptions by demanding that a perturbed stationarity system possesses solutions as the involved perturbations tend to zero. Together with associated approximate constraint qualifications, which are typically rather mild, they raised much interest in the optimization community during the last decade. In parallel, directional stationarity conditions became quite popular as they sharpen standard stationarity conditions by incorporating data associated with underlying critical directions. The purpose of this paper is twofold. First, we melt the aforementioned concepts of approximate and directional stationarity to formulate and study so-called approximate directional stationarity. For the underlying model problem, an optimization problem with nonsmooth geometric constraints is chosen, which covers diverse practically relevant applications. The role of approximate directional stationarity as a necessary optimality condition is investigated in much detail, complementing results from the literature. Second, we formulate a qualification condition which, based on an approximately directionally stationary point, can be exploited to infer its directional stationarity. The latter condition depends on one particular sequence verifying approximate directional stationarity and merely requires to check a simple condition of Mangasarian--Fromovitz type stated in terms of the directional tools of limiting variational analysis. This contrasts standard approximate constraint qualifications that typically demand a certain stable behavior of all sequences validating approximate stationarity. Throughout, various approaches to verify directional stationarity of local minimizers are established, and illustrative examples are presented to make the theoretical results more accessible.

math.OC

Augmented Lagrangian methods for fully convex composite optimization

This paper is concerned with augmented Lagrangian methods for the treatment of fully convex composite optimization problems. We extend the classical relationship between augmented Lagrangian methods and the proximal point algorithm to the inexact and safeguarded scheme in order to state global primal-dual convergence results. Our analysis distinguishes the regular case, where a stationary minimizer exists, and the irregular case, where all minimizers are nonstationary. Furthermore, we suggest an elastic modification of the standard safeguarding scheme which preserves primal convergence properties while guaranteeing convergence of the dual sequence to a multiplier in the regular situation. Although important for nonconvex problems, the standard safeguarding mechanism leads to weaker convergence guarantees for convex problems than the classical augmented Lagrangian method. Our elastic safeguarding scheme combines the advantages of both while avoiding their shortcomings.

math.OC

Approximate stationarity in disjunctive optimization: concepts, qualification conditions, and application to MPCCs

In this paper, we are concerned with stationarity conditions and qualification conditions for optimization problems with disjunctive constraints. This class covers, among others, optimization problems with complementarity, vanishing, or switching constraints, which are notoriously challenging due to their highly combinatorial structure. The focus of our study is twofold. First, we investigate approximate stationarity conditions and the associated strict constraint qualifications which can be used to infer stationarity of local minimizers. While such concepts are already known in the context of so-called Mordukhovich-stationarity, we introduce suitable extensions associated with strong stationarity. Second, a qualification condition is established which, based on an approximately Mordukhovich- or strongly stationary point, can be used to infer its Mordukhovich- or strong stationarity, respectively. In contrast to the aforementioned strict constraint qualifications, this condition depends on the involved sequences justifying approximate stationarity and, thus, is not a constraint qualification in the narrower sense. However, it is much easier to verify as it merely requires to check the (positive) linear independence of a certain family of gradients. In order to illustrate the obtained findings, they are applied to optimization problems with complementarity constraints, where they can be naturally extended to the well-known concepts of weak and Clarke-stationarity.

math.OC

Revisiting implicit variables in mathematical optimization: simplified modeling and a numerical evidence

Implicit variables of an optimization problem are used to model variationally challenging feasibility conditions in a tractable way while not entering the objective function. Hence, it is a standard approach to treat implicit variables as explicit ones. Recently, it has been shown in terms of a comparatively complex model problem that this approach, generally, is theoretically disadvantageous as the surrogate problem typically suffers from the presence of artificial stationary points and the need for stronger constraint qualifications. The purpose of the present paper is twofold. First, it introduces a much simpler and easier accessible model problem which can be used to recapitulate and even broaden the aforementioned findings. Indeed, we will extend the analysis to two more classes of stationary points and the associated constraint qualifications. These theoretical results are accompanied by illustrative examples from cardinality-constrained, vanishing-constrained, and bilevel optimization. Second, the present paper illustrates, in terms of cardinality-constrained portfolio optimization problems, that treating implicit variables as explicit ones may also be disadvantageous from a numerical point of view.

math.OC

Duality-based single-level reformulations of bilevel optimization problems

Usually, bilevel optimization problems need to be transformed into single-level ones in order to derive optimality conditions and solution algorithms. Among the available approaches, the replacement of the lower-level problem by means of duality relations became popular quite recently. We revisit three realizations of this idea which are based on the lower-level Lagrange, Wolfe, and Mond--Weir dual problem. The resulting single-level surrogate problems are equivalent to the original bilevel optimization problem from the viewpoint of global minimizers under mild assumptions. However, all these reformulations suffer from the appearance of so-called implicit variables, i.e., surrogate variables which do not enter the objective function but appear in the feasible set for modeling purposes. Treating implicit variables as explicit ones has been shown to be problematic when locally optimal solutions, stationary points, and applicable constraint qualifications are compared to the original problem. Indeed, we illustrate that the same difficulties have to be faced when using these duality-based reformulations. Furthermore, we show that the Mangasarian-Fromovitz constraint qualification is likely to be violated at each feasible point of these reformulations, contrasting assertions in some recently published papers.

math.OC

On the directional asymptotic approach in optimization theory

As a starting point of our research, we show that, for a fixed order $\gamma\geq 1$, each local minimizer of a rather general nonsmooth optimization problem in Euclidean spaces is either M-stationary in the classical sense (corresponding to stationarity of order $1$), satisfies stationarity conditions in terms of a coderivative construction of order $\gamma$, or is asymptotically stationary with respect to a critical direction as well as order $\gamma$ in a certain sense. By ruling out the latter case with a constraint qualification not stronger than directional metric subregularity, we end up with new necessary optimality conditions comprising a mixture of limiting variational tools of orders $1$ and $\gamma$. These abstract findings are carved out for the broad class of geometric constraints and $\gamma:=2$, and visualized by examples from complementarity-constrained and nonlinear semidefinite optimization. As a byproduct of the particular setting $\gamma:=1$, our general approach yields new so-called directional asymptotic regularity conditions which serve as constraint qualifications guaranteeing M-stationarity of local minimizers. We compare these new regularity conditions with standard constraint qualifications from nonsmooth optimization. Further, we extend directional concepts of pseudo- and quasi-normality to arbitrary set-valued mappings. It is shown that these properties provide sufficient conditions for the validity of directional asymptotic regularity. Finally, a novel coderivative-like variational tool is used to construct sufficient conditions for the presence of directional asymptotic regularity. For geometric constraints, it is illustrated that all appearing objects can be calculated in terms of initial problem data.

math.OC

Isolated calmness of perturbation mappings in generalized nonlinear programming and local superlinear convergence of Newton-type methods

In this paper, we characterize Lipschitzian properties of different multiplier-free and multiplier-dependent perturbation mappings associated with the stationarity system of a so-called generalized nonlinear program popularized by Rockafellar. Special emphasis is put on the investigation of the isolated calmness property at and around a point. The latter is decisive for the locally fast convergence of the so-called semismooth* Newton-type method by Gfrerer and Outrata. Our central result is the characterization of the isolated calmness at a point of a multiplier-free perturbation mapping via a combination of an explicit condition and a rather mild assumption, automatically satisfied e.g. for standard nonlinear programs. Isolated calmness around a point is characterized analogously by a combination of two stronger conditions. These findings are then related to so-called criticality of Lagrange multipliers, as introduced by Izmailov and extended to generalized nonlinear programming by Mordukhovich and Sarabi. We derive a new sufficient condition (a characterization for some problem classes) of nonexistence of critical multipliers, which has been also used in the literature as an assumption to guarantee local fast convergence of Newton-, SQP-, or multiplier-penalty-type methods. The obtained insights about critical multipliers seem to complement the vast literature on the topic.

math.OC

Local properties and augmented Lagrangians in fully nonconvex composite optimization

A broad class of optimization problems can be cast in composite form, that is, considering the minimization of the composition of a lower semicontinuous function with a differentiable mapping. This paper investigates the versatile template of composite optimization without any convexity assumptions. First- and second-order optimality conditions are discussed. We highlight the difficulties that stem from the lack of convexity when dealing with necessary conditions in a Lagrangian framework and when considering error bounds. Building upon these characterizations, a local convergence analysis is delineated for a recently developed augmented Lagrangian method, deriving rates of convergence in the fully nonconvex setting.

math.OC

Bilevel Optimal Control: Theory, Algorithms, and Applications

In this chapter, we are concerned with inverse optimal control problems, i.e., optimization models which are used to identify parameters in optimal control problems from given measurements. Here, we focus on linear-quadratic optimal control problems with control constraints where the reference control plays the role of the parameter and has to be reconstructed. First, it is shown that pointwise M-stationarity, associated with the reformulation of the hierarchical model as a so-called mathematical problem with complementarity constraints (MPCC) in function spaces, provides a necessary optimality condition under some additional assumptions on the data. Second, we review two recently developed algorithms (an augmented Lagrangian method and a nonsmooth Newton method) for the computational identification of M-stationary points of finite-dimensional MPCCs. Finally, a numerical comparison of these methods, based on instances of the appropriately discretized inverse optimal control problem of our interest, is provided.

math.OC

A fresh look at nonsmooth Levenberg--Marquardt methods with applications to bilevel optimization

In this paper, we revisit the classical problem of solving over-determined systems of nonsmooth equations numerically. We suggest a nonsmooth Levenberg--Marquardt method for its solution which, in contrast to the existing literature, does not require local Lipschitzness of the data functions. This is possible when using Newton-differentiability instead of semismoothness as the underlying tool of generalized differentiation. Conditions for fast local convergence of the method are given. Afterwards, in the context of over-determined mixed nonlinear complementarity systems, our findings are applied, and globalized solution methods, based on a residual induced by the maximum and the Fischer--Burmeister function, respectively, are constructed. The assumptions for fast local convergence are worked out and compared. Finally, these methods are applied for the numerical solution of bilevel optimization problems. We recall the derivation of a stationarity condition taking the shape of an over-determined mixed nonlinear complementarity system involving a penalty parameter, formulate assumptions for local fast convergence of our solution methods explicitly, and present results of numerical experiments. Particularly, we investigate whether the treatment of the appearing penalty parameter as an additional variable is beneficial or not.

math.OC

Fuzzy multiplier, sum and intersection rules in non-Lipschitzian settings: decoupling approach revisited

We revisit the decoupling approach widely used (often intuitively) in nonlinear analysis and optimization and initially formalized about a quarter of a century ago by Borwein & Zhu, Borwein & Ioffe and Lassonde. It allows one to streamline proofs of necessary optimality conditions and calculus relations, unify and simplify the respective statements, clarify and in many cases weaken the assumptions. In this paper we study weaker concepts of quasiuniform infimum, quasiuniform lower semicontinuity and quasiuniform minimum, putting them into the context of the general theory developed by the aforementioned authors. On the way, we unify the terminology and notation and fill in some gaps in the general theory. We establish rather general primal and dual necessary conditions characterizing quasiuniform $\varepsilon$-minima of the sum of two functions. The obtained fuzzy multiplier rules are formulated in general Banach spaces in terms of Clarke subdifferentials and in Asplund spaces in terms of Fr\'echet subdifferentials. The mentioned fuzzy multiplier rules naturally lead to certain fuzzy subdifferential calculus results. An application from sparse optimal control illustrates applicability of the obtained findings.

math.OC

Variational Poisson Denoising via Augmented Lagrangian Methods

In this paper, we denoise a given noisy image by minimizing a smoothness promoting function over a set of local similarity measures which compare the mean of the given image and some candidate image on a large collection of subboxes. The associated convex optimization problem possesses a huge number of constraints which are induced by extended real-valued functions stemming from the Kullback--Leibler divergence. Alternatively, these nonlinear constraints can be reformulated as affine ones, which makes the model seemingly more tractable. For the numerical treatment of both formulations of the model (i.e., the original one as well as the one with affine constraints), we propose a rather general augmented Lagrangian method which is capable of handling the huge amount of constraints. A self-contained, derivative-free, global convergence theory is provided, allowing an extension to other problem classes. For the solution of the resulting subproblems in the setting of our suggested image denoising models, we make use of a suitable stochastic gradient method. Results of several numerical experiments are presented in order to compare both formulations and the associated augmented Lagrangian methods.

math.OC

Notes on the value function approach to multiobjective bilevel optimization

This paper is concerned with the value function approach to multiobjective bilevel optimization which exploits a lower level frontier-type mapping in order to replace the hierarchical model of two interdependent multiobjective optimization problems by a single-level multiobjective optimization problem. As a starting point, different value-function-type reformulations are suggested and their relations are discussed. Here, we focus on the situations where the lower level problem is solved up to efficiency or weak efficiency, and an intermediate solution concept is suggested as well. We study the graph-closedness of the associated efficiency-type and frontier-type mappings. These findings are then used for two purposes. First, we investigate existence results in multiobjective bilevel optimization. Second, for the derivation of necessary optimality conditions via the value function approach, it is inherent to differentiate frontier-type mappings in a generalized way. Here, we are concerned with the computation of upper coderivative estimates for the frontier-type mapping associated with the setting where the lower level problem is solved up to weak efficiency. We proceed in two ways, relying, on the one hand, on a weak domination property and, on the other hand, on a scalarization approach. Throughout the paper, illustrative examples visualize our findings, the necessity of crucial assumptions, and some flaws in the related literature.

math.OC

Convergence Analysis of the Proximal Gradient Method in the Presence of the Kurdyka-{\L}ojasiewicz Property without Global Lipschitz Assumptions

We consider a composite optimization problem where the sum of a continuously differentiable and a merely lower semicontinuous function has to be minimized. The proximal gradient algorithm is the classical method for solving such a problem numerically. The corresponding global convergence and local rate-of-convergence theory typically assumes, besides some technical conditions, that the smooth function has a globally Lipschitz continuous gradient and that the objective function satisfies the Kurdyka-{\L}ojasiewicz property. Though this global Lipschitz assumption is satisfied in several applications where the objective function is, e.g., quadratic, this requirement is very restrictive in the non-quadratic case. Some recent contributions therefore try to overcome this global Lipschitz condition by replacing it with a local one, but, to the best of our knowledge, they still require some extra condition in order to obtain the desired global and rate-of-convergence results. The aim of this paper is to show that the local Lipschitz assumption together with the Kurdyka-{\L}ojasiewicz property is sufficient to recover these convergence results.

math.OC