SearcharxivSearch

arXiv subjects

Wenqing Ouyang

Publications and source records attributed to Wenqing Ouyang.

14 recordsLinked to original sources

Burer-Monteiro factorizability of nuclear norm regularized optimization

This paper studies the relationship between the nuclear norm-regularized minimization problem, which minimizes the sum of a $C^2$ function $h$ and a positive multiple of the nuclear norm, denoted by $f$, and its factorized problem obtained by the Burer-Monteiro technique. We are interested in deriving conditions that ensure every second-order stationary point of the factorized problem corresponds to a global minimizer of $f$, a property we call the $r$-factorizability of $f$ in this paper. Under suitable restricted isometry property (RIP) type assumptions on $h$, we prove the $r$-factorizability of $f$. Moreover, the RIP constant in our paper is tight, in the sense that concrete non-$r$-factorizable $f$ can be constructed when the RIP-type assumption fails to hold. Our technique for constructing such examples is novel and may be of independent interest: specifically, we use a variant of the Von Neumann's trace inequality and relate the existence of such examples to the optimal value of a quadratic program involving the RIP constant, then we explicitly solve this optimization problem to identify the parameter regimes in which such worst-case counterexamples can be constructed.

math.OC

Computing Kurdyka-Łojasiewicz exponents via composition and symmetry

We devise calculus rules for the Kurdyka-Łojasiewicz exponent using the rank theorem and Lie group actions. They apply to a wide class of composite and invariant functions, and are particularly suitable for handling nonisolated local minima. Notably, smoothness plays no role, eschewing gradient and Hessian computations. This provides a unified framework for establishing linear convergence of various algorithms in matrix factorization, $\ell_1$-matrix factorization, matrix sensing, and linear neural networks.

math.OC

Reachability of gradient descent

We show that gradient descent can converge to any local minimum of a smooth semi-algebraic function. This holds if the step sizes are nonsummable and sufficiently small. The same results hold for the subgradient method on locally Lipschitz semi-algebraic functions if the step size is constant.

math.OC

A linesearch-type normal map-based semismooth Newton method for nonsmooth nonconvex composite optimization

We propose a novel linesearch variant of the trust region normal map-based semismooth Newton method developed in [Ouyang and Milzarek, Math. Program. 212(1-2), 389--435 (2025)] for solving a class of nonsmooth, nonconvex composite-type optimization problems. Our approach uses adaptive parameter estimation techniques, which allow us to avoid explicit and potentially expensive Lipschitz constant computations. We provide extensive convergence results including global convergence, convergence of the iterates under the Kurdyka-Łojasiewicz inequality, and transition to fast local q-superlinear convergence. Compared to the original trust region framework, the linesearch-based algorithm is simpler and the overall convergence analysis can be conducted under weaker assumptions -- in particular, without requiring explicit boundedness conditions on the Hessian approximations and iterates. Numerical experiments on sparse logistic regression, image compression, and nonlinear least squares with group penalty terms demonstrate the efficiency of the proposed approach.

math.OC

A MINRES-based Linesearch Algorithm for Nonconvex Optimization with Non-positive Curvature Detection

We propose a MINRES-based Newton-type algorithm for solving unconstrained nonconvex optimization problems. Our approach uses the minimal residual method (MINRES), a well-known solver for indefinite symmetric linear systems, to compute descent directions that leverage second-order and non-positive curvature (NPC) information. Comprehensive asymptotic convergence properties are derived under standard assumptions. In particular, under the Kurdyka-Łojasiewicz inequality and a mild NPC-detectability condition, we prove that our algorithm can avoid strict saddle points and converge to second-order critical points. This is primarily achieved by integrating proper regularization techniques and forward linesearch mechanisms along NPC directions. Furthermore, fast local superlinear convergence to potentially non-isolated minima is established, when the local Polyak-Łojasiewicz condition is satisfied. Numerical experiments on the CUTEst test collection and on a deep auto-encoder problem illustrate the efficiency of the proposed method.

math.OC

Variational Properties of Decomposable Functions. Part II: Strong Second-Order Theory

Local superlinear convergence of the semismooth Newton method usually necessitates assumptions on the uniform invertibility of the utilized, generalized Jacobian matrices, such as, e.g., BD- or CD-regularity. For certain composite-type problems and nonlinear programs (for which explicit representations of the generalized Jacobians of the associated stationarity equations are available), such regularity assumptions are closely connected to strong second-order sufficient conditions. However, general characterizations are not well understood. In this paper, we investigate a strong second-order sufficient condition ($\mathrm{SSOSC}$) for composite problems whose nonsmooth part has a generalized conic-quadratic second subderivative. We discuss the relationship between the $\mathrm{SSOSC}$ and other second order-type conditions that involve the generalized Jacobians of the normal map. In particular, these two conditions are equivalent under certain structural assumptions on the generalized Jacobian matrix of the proximity operator. Leveraging second-order variational techniques and properties, we then verify that the introduced structural conditions hold for a broad class of $C^2$-strictly decomposable functions. Finally, it is shown that the $\mathrm{SSOSC}$ is further equivalent to the strong metric regularity of the subdifferential, the normal map, and the natural residual.

math.OC

Kurdyka-Łojasiewicz exponent via square transformation

We consider one of the most common reparameterization techniques, the square transformation. Assuming the original objective function is the sum of a smooth function and a polyhedral function, we study the variational properties of the objective function after reparameterization. In particular, we first study the minimal norm of the subdifferential of the reparameterized objective function. Second, we compute the second subderivative of the reparameterized objective function on a linear subspace, which allows for fully characterizing the subclass of stationary points of the reparameterized objective function that correspond to stationary points of the original objective function. Finally, utilizing the representation of the minimal norm of the subdifferential, we show that the Kurdyka-Łojasiewicz (KL) exponent of the reparameterized function can be deduced from that of the original function.

math.OC

Variational Properties of Decomposable Functions. Part I: Strict Epi-Calculus and Applications

This work provides a systematic study of the variational properties of decomposable functions which are compositions of an outer support function and an inner smooth mapping under certain constraint qualifications. A particular focus is put on the strict twice epi-differentiability and the associated strict second subderivative of such functions. A lower bound for the strict second subderivative of decomposable functions is derived which allows linking the strict second subderivative of decomposable mappings to the simpler outer support function. Leveraging the variational properties of the support function, we establish the equivalence between the strict twice epi-differentiability of decomposable functions, continuous differentiability of the proximity operator, and the strict complementarity condition. As an application, this allows us to fully characterize the strict saddle point property of decomposable functions. In addition, an explicit formula for the strict second subderivative of decomposable functions is derived if the outer support set is sufficiently regular. This yields an alternative characterization of the strong metric regularity of the subdifferential of decomposable functions at a local minimizer. Finally, we verify that the introduced regularity conditions are satisfied by many practical functions and applications.

math.OC

Kurdyka-Łojasiewicz exponent via Hadamard parametrization

We consider a class of $\ell_1$-regularized optimization problems and the associated smooth "over-parameterized" optimization problems built upon the Hadamard parametrization, or equivalently, the Hadamard difference parametrization (HDP). We characterize the set of second-order stationary points of the HDP-based model and show that they correspond to some stationary points of the corresponding $\ell_1$-regularized model. More importantly, we show that the Kurdyka-Lojasiewicz (KL) exponent of the HDP-based model at a second-order stationary point can be inferred from that of the corresponding $\ell_1$-regularized model under suitable assumptions. Our assumptions are general enough to cover a wide variety of loss functions commonly used in $\ell_1$-regularized models, such as the least squares loss function and the logistic loss function. Since the KL exponents of many $\ell_1$-regularized models are explicitly known in the literature, our results allow us to leverage these known exponents to deduce the KL exponents at second-order stationary points of the corresponding HDP-based models, which were previously unknown. Finally, we demonstrate how these explicit KL exponents at second-order stationary points can be applied to deducing the explicit local convergence rate of a standard gradient descent method for minimizing the HDP-based model.

math.OC

A trust region-type normal map-based semismooth Newton method for nonsmooth nonconvex composite optimization

We propose a novel trust region method for solving a class of nonsmooth, nonconvex composite-type optimization problems. The approach embeds inexact semismooth Newton steps for finding zeros of a normal map-based stationarity measure for the problem in a trust region framework. Based on a new merit function and acceptance mechanism, global convergence and transition to fast local q-superlinear convergence are established under standard conditions. In addition, we verify that the proposed trust region globalization is compatible with the Kurdyka-Łojasiewicz inequality yielding finer convergence results. We further derive new normal map-based representations of the associated second-order optimality conditions that have direct connections to the local assumptions required for fast convergence. Finally, we study the behavior of our algorithm when the Hessian matrix of the smooth part of the objective function is approximated by BFGS updates. We successfully link the KL theory, properties of the BFGS approximations, and a Dennis-Mor{é}-type condition to show superlinear convergence of the quasi-Newton version of our method. Numerical experiments on sparse logistic regression, image compression, and a constrained log-determinant problem illustrate the efficiency of the proposed algorithm.

math.OC

Descent Properties of an Anderson Accelerated Gradient Method With Restarting

Anderson Acceleration (AA) is a popular acceleration technique to enhance the convergence of fixed-point iterations. The analysis of AA approaches typically focuses on the convergence behavior of a corresponding fixed-point residual, while the behavior of the underlying objective function values along the accelerated iterates is currently not well understood. In this paper, we investigate local properties of AA with restarting applied to a basic gradient scheme in terms of function values. Specifically, we show that AA with restarting is a local descent method and that it can decrease the objective function faster than the gradient method. These new results theoretically support the good numerical performance of AA when heuristic descent conditions are used for globalization and they provide a novel perspective on the convergence analysis of AA that is more amenable to nonconvex optimization problems. Numerical experiments are conducted to illustrate our theoretical findings.

math.OC

Nonmonotone Globalization for Anderson Acceleration via Adaptive Regularization

Anderson acceleration (AA) is a popular method for accelerating fixed-point iterations, but may suffer from instability and stagnation. We propose a globalization method for AA to improve stability and achieve unified global and local convergence. Unlike existing AA globalization approaches that rely on safeguarding operations and might hinder fast local convergence, we adopt a nonmonotone trust-region framework and introduce an adaptive quadratic regularization together with a tailored acceptance mechanism. We prove global convergence and show that our algorithm attains the same local convergence as AA under appropriate assumptions. The effectiveness of our method is demonstrated in several numerical experiments.

math.OC

Anderson Acceleration for Nonconvex ADMM Based on Douglas-Rachford Splitting

The alternating direction multiplier method (ADMM) is widely used in computer graphics for solving optimization problems that can be nonsmooth and nonconvex. It converges quickly to an approximate solution, but can take a long time to converge to a solution of high-accuracy. Previously, Anderson acceleration has been applied to ADMM, by treating it as a fixed-point iteration for the concatenation of the dual variables and a subset of the primal variables. In this paper, we note that the equivalence between ADMM and Douglas-Rachford splitting reveals that ADMM is in fact a fixed-point iteration in a lower-dimensional space. By applying Anderson acceleration to such lower-dimensional fixed-point iteration, we obtain a more effective approach for accelerating ADMM. We analyze the convergence of the proposed acceleration method on nonconvex problems, and verify its effectiveness on a variety of computer graphics problems including geometry processing and physical simulation.

math.OC

Accelerating ADMM for Efficient Simulation and Optimization

The alternating direction method of multipliers (ADMM) is a popular approach for solving optimization problems that are potentially non-smooth and with hard constraints. It has been applied to various computer graphics applications, including physical simulation, geometry processing, and image processing. However, ADMM can take a long time to converge to a solution of high accuracy. Moreover, many computer graphics tasks involve non-convex optimization, and there is often no convergence guarantee for ADMM on such problems since it was originally designed for convex optimization. In this paper, we propose a method to speed up ADMM using Anderson acceleration, an established technique for accelerating fixed-point iterations. We show that in the general case, ADMM is a fixed-point iteration of the second primal variable and the dual variable, and Anderson acceleration can be directly applied. Additionally, when the problem has a separable target function and satisfies certain conditions, ADMM becomes a fixed-point iteration of only one variable, which further reduces the computational overhead of Anderson acceleration. Moreover, we analyze a particular non-convex problem structure that is common in computer graphics, and prove the convergence of ADMM on such problems under mild assumptions. We apply our acceleration technique on a variety of optimization problems in computer graphics, with notable improvement on their convergence speed.

cs.GR