SearcharxivSearch

arXiv subjects

Benqi Liu

Publications and source records attributed to Benqi Liu.

6 recordsLinked to original sources

The Minimum Q-Order of BFGS with Exact Line Search Is One

Powell asked whether the smoothness assumptions underlying classical superlinear convergence force a fixed power law between adjacent iterates of exact-line-search variable-metric methods. We answer this question negatively for BFGS: within the smooth strongly convex setting, the smallest possible adjacent-iterate Q-order is one, and this boundary is attained by a single nonterminating run. In every finite dimension at least two, and for any prescribed radius and Hessian tolerance, we construct an infinitely differentiable, globally strongly convex objective that equals the standard quadratic outside the corresponding ball and whose Hessian remains within the prescribed tolerance of the identity in operator norm. The objective has its unique minimizer at the origin and identity Hessian there. Exact-line-search BFGS, initialized with the identity matrix and started inside that ball, converges Q-superlinearly, yet no fixed power greater than one controls all sufficiently late adjacent errors.

math.OC

A counterexample to global convergence of classical DFP under the standard strong Wolfe conditions

A long-standing open question in quasi-Newton optimization asks whether the classical Davidon--Fletcher--Powell (DFP) method converges globally on uniformly convex objectives when all accepted steps satisfy the standard weak Wolfe conditions. We show that the answer is no, even under the standard strong Wolfe conditions. Fix $0<c_1<2/3$ and $2/3\le c_2<1$. We construct a function $f\in C^2(\mathbb{R}^2)$ such that $\frac{1}{2}I\preceq\nabla^2 f(x)\preceq\frac{3}{2}I$ for all $x\in\mathbb{R}^2$. We also choose a fixed positive definite initial inverse Hessian approximation and a sequence of positive step lengths. The classical DFP iteration is well defined, and all accepted steps satisfy the standard strong Wolfe conditions, but $|\nabla f(x_k)|$ converges to a positive constant. The global Hessian condition number is at most three. The construction uses an alternating two-step DFP sequence near a one-dimensional invariant center manifold. Along this sequence, the smaller eigenvalue of the inverse Hessian approximation tends to zero. The changes in the gradient norm between cycle starts are summable, but the total rotation of the associated eigenvectors is unbounded. The accumulation points of the DFP sequence form a circle. A uniform separation bound allows us to interpolate the prescribed function values and gradients. We add smooth functions with pairwise disjoint supports to a quadratic and keep the global Hessian bounds. An affine change of variables gives an identity-initialized example with problem-dependent Hessian bounds. An orthogonal direct sum extends the result to every dimension $n\ge 2$.

math.OC

A Fixed-Penalty Linearized Augmented Lagrangian Method with Classical Multiplier Updates

Augmented Lagrangian methods are effective for nonlinear equality-constrained optimization, but solving their nonlinear primal subproblems can be expensive. For smooth nonconvex problems with deterministic or stochastic objectives, we propose a nonlinear-residual linearized augmented Lagrangian method (NR-LALM) that replaces this subproblem by a regularized Gauss-Newton-type step while retaining the classical multiplier update based on the nonlinear constraint residual. The resulting step is computed from one symmetric positive-definite linear system, but the mismatch between the linearized primal model and the nonlinear-residual update produces a quadratic constraint-linearization error in the multiplier identity. We show that this error can be controlled under local regularity; multiplier boundedness and trajectory localization are derived rather than assumed. With fixed, accuracy-independent parameters, deterministic NR-LALM finds an $\varepsilon$-approximate Karush-Kuhn-Tucker (KKT) pair in $O(\varepsilon^{-2})$ iterations and first-order oracle evaluations. For stochastic objectives, a projected stochastic path-integrated differential estimator with safeguarded restarts requires, in expectation, $O(\varepsilon^{-3})$ stochastic-gradient evaluations and $O(\varepsilon^{-2})$ constraint and Jacobian evaluations. Compactness and a Kurdyka-Lojasiewicz condition further yield finite-length convergence of the deterministic primal-dual sequence. An optional minimum-norm second-order correction reduces the constraint-linearization error from second to fourth order without changing the complexity orders. All theoretical results are formalized in Lean 4. Numerical experiments confirm the predicted error orders and show favorable performance on high-dimensional deterministic and stochastic problems.

math.OC

Restarted Reflected Halpern Acceleration for Augmented Primal-Dual Methods

We study linearly constrained composite convex optimization with a smooth term and a proximable nonsmooth term. We develop a unified augmented primal-dual framework with primal-dual hybrid gradient-type and augmented Chambolle-Pock-type metric choices, including a fully augmented Chambolle-Pock-type family that retains the augmented quadratic term. The exact scheme admits a degenerate proximal-point form; the linearized scheme admits a preconditioned forward-backward form. These representations allow reflected Halpern acceleration to be analyzed directly in primal-dual variables. For the shadow iterates, we prove convergence to Karush-Kuhn-Tucker (KKT) points and nonergodic O(1/k) bounds for the KKT residual and objective gap, with a scalar worst-case example. We show that finite identification belongs to the shadow sequence rather than to the anchored Halpern state. After identification, an affine-face model yields an exact reduced residual identity and a local-sharpness criterion. Finally, we prove linear convergence of restart anchors under fixed-point sharpness on the visited restart set, with local or tail convergence when sharpness follows from local error bounds. Experiments on linear and convex quadratic programs illustrate augmentation and linearization.

math.OC

A Quadratic-Approximation-Based Stochastic Approximation Method for Weakly Convex Stochastic Programming

We propose a novel stochastic approximation algorithm, termed PMQSopt, for solving weakly convex stochastic optimization problems involving expectation-valued functions. The algorithm is constructed by integrating the proximal method of multipliers with quadratic approximations of the original stochastic problem. We analyze the sample complexity of PMQSopt in terms of the total number of stochastic gradient evaluations required. The convergence of the algorithm is characterized by three metrics associated with the $\epsilon$-KKT conditions: the average squared norm of the gradient of the Moreau envelope of the Lagrangian, the average constraint violation, and the average complementarity violation. For each of these metrics, we establish an expected convergence rate of $\mathcal{O}(T^{-1/4})$ after $T$ iterations. Furthermore, we show that with probability at least $1-1/T^{2/3}$, the gradient of the Lagrangian satisfies an $\mathcal{O}(T^{-1/8})$ bound; with probability at least $1-2/T^{2/3}$, the constraint violation achieves an $\mathcal{O}(T^{-1/4})$ bound; and with probability at least $1-3/T^{2/3}$, the complementarity violation attains an $\mathcal{O}(T^{-1/4})$ bound. All results are established under two mild conditions: (i) weak convexity of all problem functions, and (ii) the existence of a strictly feasible point. The proposed PMQSopt algorithm is a sequentially strongly convex programming method that is readily implementable. Numerical experiments illustrate its practical performance.

math.OC

A Proximal Augmented Lagrangian Method Based on Quadratic Approximations for Weakly Convex Optimization

This paper proposes QPALM, a proximal augmented Lagrangian method based on quadratic approximations, for solving nonlinear programming problems with weakly convex objective and constraint functions. The algorithm is constructed by incorporating quadratic approximations of both the objective and constraint functions into a proximal Lagrangian framework. We establish its non-asymptotic convergence rate in terms of the total number of subproblems solved. The convergence of QPALM is characterized by three metrics associated with the $\varepsilon$-KKT conditions: the squared norm of the gradient of the Moreau envelope of the Lagrangian, the average constraint violation, and the average complementarity violation. All three metrics are shown to converge at a rate of $O(T^{-1/3})$ after $T$ iterations. Preliminary numerical results demonstrate the practical efficiency of the proposed method. These results are established under two mild conditions: (i) weak convexity of all problem functions, and (ii) the existence of a strictly feasible point. The proposed QPALM is a sequentially strongly convex programming method that is readily implementable.

math.OC