SearcharxivSearch

arXiv subjects

Liwei Zhang

Publications and source records attributed to Liwei Zhang.

At least 19 recordsLinked to original sources

A counterexample to global convergence of classical DFP under the standard strong Wolfe conditions

A long-standing open question in quasi-Newton optimization asks whether the classical Davidon--Fletcher--Powell (DFP) method converges globally on uniformly convex objectives when all accepted steps satisfy the standard weak Wolfe conditions. We show that the answer is no, even under the standard strong Wolfe conditions. Fix $0<c_1<2/3$ and $2/3\le c_2<1$. We construct a function $f\in C^2(\mathbb{R}^2)$ such that $\frac{1}{2}I\preceq\nabla^2 f(x)\preceq\frac{3}{2}I$ for all $x\in\mathbb{R}^2$. We also choose a fixed positive definite initial inverse Hessian approximation and a sequence of positive step lengths. The classical DFP iteration is well defined, and all accepted steps satisfy the standard strong Wolfe conditions, but $|\nabla f(x_k)|$ converges to a positive constant. The global Hessian condition number is at most three. The construction uses an alternating two-step DFP sequence near a one-dimensional invariant center manifold. Along this sequence, the smaller eigenvalue of the inverse Hessian approximation tends to zero. The changes in the gradient norm between cycle starts are summable, but the total rotation of the associated eigenvectors is unbounded. The accumulation points of the DFP sequence form a circle. A uniform separation bound allows us to interpolate the prescribed function values and gradients. We add smooth functions with pairwise disjoint supports to a quadratic and keep the global Hessian bounds. An affine change of variables gives an identity-initialized example with problem-dependent Hessian bounds. An orthogonal direct sum extends the result to every dimension $n\ge 2$.

math.OC

Primal-Dual Halpern-PAGE Algorithm for Constrained Stochastic Weakly Convex Optimization

We tackle the challenging problem of stochastic weakly convex optimization subject to mixed (equality and inequality) expected-value constraints. While optimal $\mathcal{O}(\epsilon^{-3})$ sample complexity algorithms exist for unconstrained weakly convex problems, dealing with complex functional constraints typically requires cumbersome multi-loop penalty or augmented Lagrangian methods, which suffer from high inner-loop complexity and sensitive parameter tuning. To bridge this fundamental gap, we propose the primal-dual Halpern-PAGE (PD-HP) algorithm. As a purely single-loop method, PD-HP completely bypasses the computational burden of nested iterations. At each step, it merely requires solving a simple strongly convex surrogate subproblem alongside a straightforward dual projection, making it exceptionally efficient and convenient to implement. Crucially, we prove that this computationally lightweight algorithm achieves the optimal $\mathcal{O}(\epsilon^{-3})$ sample complexity for mixed-constrained stochastic weakly convex problems, successfully matching the theoretical lower bounds. Furthermore, when the primal domain is a compact polyhedral convex set, we establish the deterministic stability of the dual multipliers by exploiting the generalized Mangasarian-Fromovitz constraint qualification (MFCQ) alongside Hoffman's error bound. This ensures that our optimal complexity bound holds strictly under the standard, unbounded KKT residual metric without any theoretical gaps or artificial residual truncations.

math.OC

Sample Complexity of Policy Gradient for Log-Growth Control

We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain that optimally stabilizes a scalar linear system driven through a multiplicative-noise actuation channel. The objective $J(K) = \mathbb{E}[\log|1+BK|]$ is the top Lyapunov exponent of the closed loop. This problem carries a structural difficulty we call the cusp obstruction: the optimal gain $K^*$ always places the noise singularity $b_{\rm sing}(K) = -1/K$ in the interior of the support. At this singular optimum the policy gradient exists only as a Cauchy principal value, not as a Lebesgue integral, and the natural single-sample gradient estimator has infinite variance. Standard first-order stochastic-optimization analysis is thus inapplicable at the optimum, and merely smoothing the objective does not resolve the difficulty. The obstruction, however, has an exploitable symmetry: the Cauchy kernel is an odd function of the displacement from the moving pole, so pairing each observation with its reflection through the pole cancels the divergent part. This one cancellation simultaneously controls the population curvature, the gradient-estimator variance, and the bias incurred when the noise density is estimated. Combining these bounds with a closed-form single-transition gradient oracle, we prove that projected mini-batch policy gradient, initialized in any compact subset of the stabilizing region, attains total sample complexity $\tilde{O}(1/\eta)$ when the noise density is known and $\tilde{O}(\eta^{-(2s+1)/(2s)})$ when it must be estimated, for $C^s$ noise densities with $s \geq 2$.

eess.SY

Prox-PEP: A Proximal Partial Exact Penalty Algorithm for Weakly Convex Stochastic Nonlinear Programming

This paper considers stochastic optimization problems with weakly convex objective and constraint functions. We propose Prox-PEP, a proximal method equipped with quadratic subproblems. To handle nonlinear equality constraints, we employ an exact penalty approach, transforming them into inequality constraints with auxiliary slack variables. At each iteration, we construct quadratic approximations for both the objective and the constraint functions to facilitate efficient subproblem computation. By carefully designing the second-order approximation matrices, the subproblem constructed via the augmented Lagrangian function is strictly guaranteed to be strongly convex. Furthermore, we adopt a dynamic strategy for the equality penalty parameter: it monotonically increases up to a predefined threshold and remains constant thereafter. Building upon this algorithmic framework, we establish comprehensive asymptotic complexities. We prove that Prox-PEP achieves an $\mathcal{O}(T^{-1/4})$ average expected oracle complexity for $\epsilon$-KKT stationarity, specifically bounding the squared norm of the gradient of the Moreau envelope of the Lagrangian function, alongside constraint violations and complementarity conditions. Additionally, under standard light-tailed martingale noise assumptions, we derive an $\mathcal{O}(T^{-1/8})$ high-probability convergence bound for the norm of the gradient of the Lagrangian's Moreau envelope, as well as $\mathcal{O}(T^{-1/4})$ high-probability bounds for both constraint violations and complementarity conditions.

math.OC

Stability of Lagrangian Generalized Nash Equilibriums

Lagrangian generalized Nash equilibriums (LGNEs) were introduced by Rockafellar (2024) for a class of generalized Nash equilibrium problems (GNEPs) in which each player's strategy is subject to conic constraints. This paper investigates the stability properties of the LGNE solution set, specifically focusing on the Aubin property, isolated calmness, and Lipschitz continuous single-valued localization. For general conically constrained GNEPs, characterizations of the Aubin property and isolated calmness of the LGNE solution mapping under canonical perturbations are established. These characterizations are formulated using the coderivative and graph derivative of normal cone mappings. Subsequently, these general results are specialized to GNEPs with equality and inequality constraints, yielding explicit characterizations for both the Lipschitz continuous single-valued localization and isolated calmness of the corresponding LGNE solution mapping, which are described by nonsingularity of linear complementarity sytems. For GNEPs with shared conic constraints, the Aubin property and isolated calmness of the consensus LGNE solution mapping--where identical Lagrange multipliers are assigned to the shared constraint--are first characterized. We further analyze the case when the conic constraints are specialized as equalities and inequalities. Finally, for classical conically constrained Nash equilibrium problems, the Aubin property and isolated calmness of the Lagrangian Nash equilibrium solution mapping are also analyzed.

math.OC

A Quadratic-Approximation-Based Stochastic Approximation Method for Weakly Convex Stochastic Programming

We propose a novel stochastic approximation algorithm, termed PMQSopt, for solving weakly convex stochastic optimization problems involving expectation-valued functions. The algorithm is constructed by integrating the proximal method of multipliers with quadratic approximations of the original stochastic problem. We analyze the sample complexity of PMQSopt in terms of the total number of stochastic gradient evaluations required. The convergence of the algorithm is characterized by three metrics associated with the $\epsilon$-KKT conditions: the average squared norm of the gradient of the Moreau envelope of the Lagrangian, the average constraint violation, and the average complementarity violation. For each of these metrics, we establish an expected convergence rate of $\mathcal{O}(T^{-1/4})$ after $T$ iterations. Furthermore, we show that with probability at least $1-1/T^{2/3}$, the gradient of the Lagrangian satisfies an $\mathcal{O}(T^{-1/8})$ bound; with probability at least $1-2/T^{2/3}$, the constraint violation achieves an $\mathcal{O}(T^{-1/4})$ bound; and with probability at least $1-3/T^{2/3}$, the complementarity violation attains an $\mathcal{O}(T^{-1/4})$ bound. All results are established under two mild conditions: (i) weak convexity of all problem functions, and (ii) the existence of a strictly feasible point. The proposed PMQSopt algorithm is a sequentially strongly convex programming method that is readily implementable. Numerical experiments illustrate its practical performance.

math.OC

A Proximal Augmented Lagrangian Method Based on Quadratic Approximations for Weakly Convex Optimization

This paper proposes QPALM, a proximal augmented Lagrangian method based on quadratic approximations, for solving nonlinear programming problems with weakly convex objective and constraint functions. The algorithm is constructed by incorporating quadratic approximations of both the objective and constraint functions into a proximal Lagrangian framework. We establish its non-asymptotic convergence rate in terms of the total number of subproblems solved. The convergence of QPALM is characterized by three metrics associated with the $\varepsilon$-KKT conditions: the squared norm of the gradient of the Moreau envelope of the Lagrangian, the average constraint violation, and the average complementarity violation. All three metrics are shown to converge at a rate of $O(T^{-1/3})$ after $T$ iterations. Preliminary numerical results demonstrate the practical efficiency of the proposed method. These results are established under two mild conditions: (i) weak convexity of all problem functions, and (ii) the existence of a strictly feasible point. The proposed QPALM is a sequentially strongly convex programming method that is readily implementable.

math.OC

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation

Autoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bottleneck. Diffusion VLMs offer a more parallel decoding paradigm, yet directly converting a pretrained autoregressive VLM into a large-block diffusion VLM (dVLM) often leads to substantial quality degradation. In this work, we present BARD, a simple and effective bridging framework that converts a pretrained autoregressive VLM into a same-architecture, decoding-efficient dVLM. Our approach combines progressive supervised block merging, which gradually enlarges the decoding block size, with stage-wise intra-dVLM distillation from a fixed small-block diffusion anchor to recover performance lost at larger blocks. We further incorporate a mixed noise scheduler to improve robustness and token revision during denoising, and memory-friendly training to enable efficient training on long multimodal sequences. A key empirical finding is that direct autoregressive-to-diffusion distillation is poorly aligned and can even hurt performance, whereas distillation within the diffusion regime is consistently effective. Experimental results show that, with $\leq$ 4.4M data, BARD-VL transfers strong multimodal capability from Qwen3-VL to a large-block dVLM. Remarkably, BARD-VL establishes a new SOTA among comparable-scale open dVLMs on our evaluation suite at both 4B and 8B scales. At the same time, BARD-VL achieves up to 3$\times$ decoding throughput speedup compared to the source model. Code is available at https://github.com/fudan-generative-vision/Bard-VL.

cs.CV

Supercell-size scaling of moir\'e band flatness

In moir\'e superlattices, the band flatness governs the degree of wave localization, which is central to harnessing emergent phenomena and designing functional meta-devices. While research has focused on the magic conditions such as magic angle and magic distance for optimal flatness, a fundamental understanding of how flatness changes with the supercell size has remained elusive. Here, we establish a universal scaling between band flatness and supercell size. Theoretically, by recognizing the statistical equivalence between structural perturbations in moir\'e superlattices and disordered systems, we introduce the Thouless number to evaluate the strength of moir\'e localization. This approach allows us to establish a scaling theory for the evolution of band flatness with the supercell size, from which an analytical expression is derived. Our full-wave simulations with one-dimensional and two-dimensional moir\'e superlattices show excellent agreement with the theoretical prediction. Our work reveals a general scaling law for moir\'e band flatness, offering a new perspective for understanding and designing moir\'e-based resonant systems.

physics.optics

Characterize localization length of disordered lattices via critical coupling effect

Light localization by scattering is a fundamental mechanism driving phase transitions of wave transport in disordered systems. Characterizing the localization length in scattering systems is crucial yet challenging. In this Letter, we demonstrate a spatially matched coupling scheme using wavefront shaping to resolve the intrinsic localization length in two-dimensional disordered lattices. By tailoring the incident wavefront, our method facilitates efficient coupling of light to the minimum localized mode. We apply this approach to measure two different self-assembled lattices, and report the first observation of the critical coupling effect, which allows for the direct determination of the characteristic size of minimum localized mode. Our results reveal that for a fixed lattice periodicity, increasing the air-hole diameter significantly reduces this intrinsic localization length. This far-field metrology offers a robust framework for probing wave localization in complex media, which should be useful in various applications such as random lasing and nonlinear optics

physics.optics

Efficient construction of Lie group-equivariant and permutation-invariant spaces

We introduce a practical construction of group-equivariant and permutation-invariant functions of $N$ variables given a finite-dimensional space stable with respect to the group action. The construction applies to any connected linear Lie group and relies on leveraging the Lie algebra to build a matrix $M$ whose kernel is in one-to-one correspondence with the subspace with desired equivariance and invariance properties, removing the need for prior knowledge of Clebsch--Gordan coefficients. A similar construction is proposed for group-equivariant functions alone, without imposing permutation-invariance. For the groups $SO(3)$ and $SU(2)$, we further exploit the structure of the Lie algebra to demonstrate the sparsity pattern and rank of the matrix $M$, which yields the exact dimension of the group-equivariant and permutation-invariant space, as well as the dimension of the group-equivariant space alone. We demonstrate analytically and verify numerically that the proposed method scales linearly with respect to the dimensionality of the basis, offering a high computational gain compared to existing methods in the literature which typically scale exponentially. We finally perform a dimensionality comparison, showing that for large values of~$N$, the dimension of group-equivariant and permutation-invariant spaces is of comparable order as the dimension of permutation-invariant spaces, while pre-asymptotically, the first dimensionality is orders of magnitude lower than the second. Hence a substantial computational gain can be achieved by explicitly enforcing group-equivariance on top of permutation-invariance when approximating such functions.

math.NA

A sequential linear complementarity problem method for generalized Nash equilibrium problems

Generalized Nash equilibrium problems (GNEPs) arise in various applications where multiple players minimize individual cost functions subject to coupled constraints. A relatively unexplored approach to solving such problems is via a sequence of (mixed) linear complementarity problems (LCPs). Compared with the nonlinear equilibrium subproblems arising in recently popular penalty-based methods such as augmented Lagrangian methods, these LCPs are often substantially easier to solve. However, the existing literature on this approach is very limited, largely because of the difficulty of assessing the search directions generated by the subproblems and establishing a principled step-length acceptance criterion. This paper proposes a sequential linear complementarity problem (SLCP) method with a comprehensive convergence analysis. To assess the search directions, we introduce a novel merit function analogous to the $\ell_1$ penalty function in sequential quadratic programming. The merit function is shown to decrease along the search directions generated by the subproblems under suitable assumptions, thereby guaranteeing the global convergence of the SLCP method. We further establish local quadratic convergence and analyze the solvability of the subproblems. Preliminary numerical results demonstrate the effectiveness and competitiveness of the proposed method relative to existing approaches.

math.OC

Halpern Acceleration of the Inexact Proximal Point Method of Rockafellar

This paper investigates a Halpern acceleration of the inexact proximal point method for solving maximal monotone inclusion problems in Hilbert spaces. The proposed Halpern inexact proximal point method (HiPPM) is shown to be globally convergent, and a unified framework is developed to analyze its worst-case convergence behavior. Under mild conditions on the inexactness tolerances, HiPPM achieves an $\mathcal{O}(1/k^{2})$ convergence rate in terms of the squared fixed-point residual. Moreover, under additional well-studied regularity conditions, the method attains a fast linear convergence rate. Building on this framework, we further extend the Halpern acceleration to the inexact augmented Lagrangian method for constrained convex optimization. In the spirit of Rockafellar's classical results, the resulting accelerated inexact augmented Lagrangian method inherits the convergence rate and iteration complexity guarantees of HiPPM. Numerical experiments are provided to support the theoretical findings.

math.OC

A Dual Method for Minimax Quadratic Programming

This paper investigates minimax quadratic programming problems with coupled inequality constraints. By leveraging a duality theorem, we develop a dual algorithm that extends the dual active set method to the minimax setting, transforming the original inequality constrained problem into a sequence of equality constrained subproblems. Under a suitable assumption, we prove that the associated S-pairs do not repeat and that the algorithm terminates in a finite number of iterations, guaranteed by the monotonic decrease of the objective function value. To ensure numerical stability and efficiency, the algorithm is implemented using Cholesky factorization and Givens rotations. Numerical experiments on both randomly generated minimax quadratic programs and illustrative applications demonstrate the accuracy, stability, and computational effectiveness of the proposed algorithm.

math.OC

Second-Order Optimality Conditions for Nonsmooth Constrained Optimization with Applications to Bilevel Programming

Second-order optimality conditions are essential for nonsmooth optimization, where both the objective and constraint functions are Lipschitz continuous and second-order directionally differentiable. This paper provides no-gap second-order necessary and sufficient optimality conditions for such problems without requiring convexity assumptions on the constraint set. We introduce the concept of second-order gph-regularity for constraint functions, which ensures the outer second-order regularity of the feasible region and enables the formulation of comprehensive optimality conditions through the parabolic curve approach. An important application of our results is bilevel optimization, where we derive second-order necessary and sufficient optimality conditions for bi-local optimal solutions, which are based on the local solutions of the lower-level problem. By leveraging the Mangasarian-Fromovitz constraint qualification (MFCQ), strong second-order sufficient condition (SSOSC) and constant rank constraint qualification (CRCQ) of lower-level problem, these second-order conditions are derived without requiring the uniqueness of the lower-level multipliers. In addition, if the linear independence constraint qualification (LICQ) holds, these conditions are expressed solely in terms of the second-order derivatives of the functions defining the bilevel problem, without relying on the second-order information from the solution mapping, which would introduce implicit complexities.

math.OC

Pyramidal Patchification Flow for Visual Generation

Diffusion transformers (DiTs) adopt Patchify, mapping patch representations to token representations through linear projections, to adjust the number of tokens input to DiT blocks and thus the computation cost. Instead of a single patch size for all the timesteps, we introduce a Pyramidal Patchification Flow (PPFlow) approach: Large patch sizes are used for high noise timesteps and small patch sizes for low noise timesteps; Linear projections are learned for each patch size; and Unpatchify is accordingly modified. Unlike Pyramidal Flow, our approach operates over full latent representations other than pyramid representations, and adopts the normal denoising process without requiring the renoising trick. We demonstrate the effectiveness of our approach through two training manners. Training from scratch achieves a $1.6\times$ ($2.0\times$) inference speed over SiT-B/2 for 2-level (3-level) pyramid patchification with slightly lower training FLOPs and similar image generation performance. Training from pretrained normal DiTs achieves even better performance with small training time. The code and checkpoint are at https://github.com/fudan-generative-vision/PPFlow.

cs.CV

An Efficient Smoothing Damped Newton Method for Large-Scale Mathematical Programs with Equilibrium Constraints

Bilevel hyperparameter optimization has received growing attention thanks to the fast development of machine learning. Due to the tremendous size of data sets, the scale of bilevel hyperparameter optimization problem could be extremely large, posing great challenges in designing efficient numerical algorithms. In this paper, we focus on solving the large-scale mathematical programs with equilibrium constraints (MPEC) derived from hyperparameter selection of L1 support vector classification (L1-SVC). We propose a highly efficient smoothing damped Newton method (SDNM) for solving such MPEC. Compared with most existing algorithms where subproblems are solved by packages, our approach fully takes advantage of the structure of MPEC and therefore is package-free. Moreover, the proposed SDNM converges to C-stationary point under MPEC-LICQ with subproblem enjoys a quadratic convergence rate under proper assumptions. Extensive numerical results over LIBSVM dataset show the superior performance of SDNM over other state-of-art algorithms.

math.OC

An Empirical Study: MEMS as a Static Performance Metric

Static performance estimation is essential during compile-time analysis, yet traditional runtime-based methods are costly and platform-dependent. We investigate mems, the number of memory accesses, as a static and architecture-independent performance metric. We develop a Clang-based automated instrumentation tool that rewrites source code to insert path tracing and \textit{mems} counting logic. This allows us to evaluate mems-based performance estimation across ten classical algorithm programs. Experimental results show that within the same program, execution paths with higher mems values consistently exhibit longer runtime. However, this correlation weakens between different programs, suggesting that mems is best suited for comparing performance of different execution paths in a program.

cs.SE