SearcharxivSearch

arXiv subjects

Lexiao Lai

Publications and source records attributed to Lexiao Lai.

16 recordsLinked to original sources

Convergence of difference inclusions: a diameter criterion and step-size conditions

We study bounded realizations of discrete difference inclusions with set-valued increments and additive errors. Our results have two parts. First, we give a general convergence criterion in which changes in a convergent scalar quantity control the diameters of local sequence segments near an accumulation point. We also develop a stratified descent framework for verifying this control. For convergent realizations, the framework yields a stationarity condition for the limit from outer limits of the scaled update map. When applied to first-order methods for minimizing locally Lipschitz objectives definable in polynomially bounded o-minimal structures, the framework yields convergence to critical points for bounded sequences generated by the inexact subgradient, momentum, and stochastic subgradient methods with step sizes of order $k^{-1}$, under the corresponding error and noise conditions. Second, we study first-order methods with polynomial step sizes of order $k^{-a}$. At the square-summability boundary \(a=1/2\), we construct a locally Lipschitz semialgebraic objective for which the exact subgradient method generates a bounded, nonconvergent sequence whose accumulation points satisfy our active-geometry assumptions. Under these assumptions, bounded sequences generated by the momentum method converge for \(1/2<a\leq1\), while bounded sequences generated by the stochastic subgradient method converge almost surely for \(2/3<a\leq1\). The momentum range is sharp for this method class, while the stochastic range matches recent full-sequence convergence results for smooth nonconvex objectives.

math.OC

Certifying optimality in nonconvex robust PCA

Robust principal component analysis seeks to recover a low-rank matrix from fully observed data with sparse corruptions. A scalable approach fits a low-rank factorization by minimizing the sum of entrywise absolute residuals, leading to a nonsmooth and nonconvex objective. Under standard incoherence conditions and a random model for the corruption support, we study factorizations of the ground-truth rank-$r$ matrix with both factors of rank $r$. With high probability, every such factorization is a Clarke critical point. We also characterize the local geometry: when the factorization rank equals $r$, these solutions are sharp local minima; when it exceeds $r$, they are strict saddle points.

math.OC

Manifold constrained steepest descent for smooth and closed-set optimization

We study minimization of smooth functions over feasible sets that have smooth embedded-manifold structure throughout or only on selected regions, using linear minimization oracles (LMOs) to determine search directions under user-chosen norms. Restricting an LMO to a tangent space, however, can require an iterative inner solve. We propose \emph{Manifold Constrained Steepest Descent} (MCSD) and a tangent-projected variant, MCSD-TP, which avoid solving tangent-space LMO subproblems iteratively. The spectral-norm specialization of MCSD on the Stiefel manifold yields \emph{SPEL}, which admits an efficient implementation using matrix-sign computations. Under standard regularity assumptions, both methods attain an \(O(\log T/\sqrt T)\) best-iterate stationarity bound. For closed feasible sets, the hybrid MCSD--PGD combines either smooth method with projected gradient descent; under the corresponding local smoothness condition, every accumulation point is Bouligand stationary. Experiments on Stiefel-constrained PCA, weighted low-rank approximation, and sparse phase retrieval illustrate the proposed methods.

math.OC

Global convergence of the subgradient method for robust signal recovery

We study the subgradient method for factorized robust signal recovery problems, including robust PCA, robust phase retrieval, and robust matrix sensing. The resulting objectives are nonsmooth and nonconvex, and can have unbounded sublevel sets, so standard analyses based on descent and coercivity do not apply. For locally Lipschitz semialgebraic objectives, we develop a convergence framework that replaces these requirements with a boundedness condition on continuous-time subgradient trajectories. Under this condition and sufficiently small step sizes of order $1/k$, we show that iterates of the subgradient method remain bounded and the full sequence converges to a critical point. We then verify the required boundedness property for the three robust objectives by adapting existing trajectory analyses, assuming a mild nondegeneracy condition in the matrix sensing case. Finally, for rank-one symmetric robust PCA, we prove that for almost every initialization, the method cannot converge to spurious critical points; consequently, under the same step-size regime, it converges to a global minimum.

math.OC

Non-Convex Self-Concordant Functions: Practical Algorithms and Complexity Analysis

We extend the standard notion of self-concordance to non-convex optimization and develop a family of second-order algorithms with global convergence guarantees. In particular, two function classes -- \textit{weakly self-concordant} functions and \textit{$F$-based self-concordant} functions -- generalize the self-concordant framework beyond convexity, without assuming the Lipschitz continuity of the gradient or Hessian. For these function classes, we propose a regularized Newton method and an adaptive regularization method that achieve an $\epsilon$-approximate first-order stationary point in $O(\epsilon^{-2})$ iterations. Equipped with an oracle capable of detecting negative curvature, the adaptive algorithm can further attain convergence to an approximate second-order stationary point. Our experimental results demonstrate that the proposed methods offer superior robustness and computational efficiency compared to cubic regularization and trust-region approaches, underscoring the broad potential of self-concordant regularization for large-scale and neural network optimization problems.

math.OC

On the diameter of subgradient sequences in o-minimal structures

We study subgradient sequences of locally Lipschitz functions definable in a polynomially bounded o-minimal structure. We show that the diameter of any subgradient sequence is related to the variation in function values, with error terms dominated by a double summation of step sizes. Consequently, we prove that bounded subgradient sequences converge if the step sizes are of order $1/k$. The proof uses Lipschitz $L$-regular stratifications in o-minimal structures to analyze subgradient sequences via their projections onto different strata.

math.OC

Phase Transitions in Phase-Only Compressed Sensing

The goal of phase-only compressed sensing is to recover a structured signal $\mathbf{x}$ from the phases $\mathbf{z} = {\rm sign}(\mathbf{\Phi}\mathbf{x})$ under some complex-valued sensing matrix $\mathbf{\Phi}$. Exact reconstruction of the signal's direction is possible: we can reformulate it as a linear compressed sensing problem and use basis pursuit (i.e., constrained norm minimization). For $\mathbf{\Phi}$ with i.i.d. complex-valued Gaussian entries, this paper shows that the phase transition is approximately located at the statistical dimension of the descent cone of a signal-dependent norm. Leveraging this insight, we derive asymptotically precise formulas for the phase transition locations in phase-only sensing of both sparse signals and low-rank matrices. Our results prove that the minimum number of measurements required for exact recovery is smaller for phase-only measurements than for traditional linear compressed sensing. For instance, in recovering a 1-sparse signal with sufficiently large dimension, phase-only compressed sensing requires approximately 68% of the measurements needed for linear compressed sensing. This result disproves earlier conjecture suggesting that the two phase transitions coincide. Our proof hinges on the Gaussian min-max theorem and the key observation that, up to a signal-dependent orthogonal transformation, the sensing matrix in the reformulated problem behaves as a nearly Gaussian matrix.

cs.IT

Stability of first-order methods in tame optimization

Modern data science applications demand solving large-scale optimization problems. The prevalent approaches are first-order methods, valued for their scalability. These methods are implemented to tackle highly irregular problems where assumptions of convexity and smoothness are untenable. Seeking to deepen the understanding of these methods, we study first-order methods with constant step size for minimizing locally Lipschitz tame functions. To do so, we propose notions of discrete Lyapunov stability for optimization methods. Concerning common first-order methods, we provide necessary and sufficient conditions for stability. We also show that certain local minima can be unstable, without additional noise in the method. Our analysis relies on the connection between the iterates of the first-order methods and continuous-time dynamics.

math.OC

Proximal random reshuffling under local Lipschitz continuity

We study proximal random reshuffling (PRR) for minimizing the sum of locally Lipschitz or locally smooth functions and a proper lower semicontinuous convex function without assuming coercivity or the existence of limit points. The algorithmic guarantees pertaining to near approximate stationarity rely on a new tracking lemma linking the iterates to trajectories of conservative fields. One of the novelties in the analysis consists in handling set-valued mappings with unbounded values. In the locally smooth case, it improves the known convergence rate from nearly $O(k^{-1/4})$ to nearly $o(k^{-1/2})$.

math.OC

Global stability of first-order methods for coercive tame functions

We consider first-order methods with constant step size for minimizing locally Lipschitz coercive functions that are tame in an o-minimal structure on the real field. We prove that if the method is approximated by subgradient trajectories, then the iterates eventually remain in a neighborhood of a connected component of the set of critical points. Under suitable method-dependent regularity assumptions, this result applies to the subgradient method with momentum, the stochastic subgradient method with random reshuffling and momentum, and the random-permutations cyclic coordinate descent method.

math.OC

Convergence of the momentum method for semialgebraic functions with locally Lipschitz gradients

We propose a new length formula that governs the iterates of the momentum method when minimizing differentiable semialgebraic functions with locally Lipschitz gradients. It enables us to establish local convergence, global convergence, and convergence to local minimizers without assuming global Lipschitz continuity of the gradient, coercivity, and a global growth condition, as is done in the literature. As a result, we provide the first convergence guarantee of the momentum method starting from arbitrary initial points when applied to principal component analysis, matrix sensing, and linear neural networks.

math.OC

Lyapunov stability of the subgradient method with constant step size

We consider the subgradient method with constant step size for minimizing locally Lipschitz semi-algebraic functions. In order to analyze the behavior of its iterates in the vicinity of a local minimum, we introduce a notion of discrete Lyapunov stability and propose necessary and sufficient conditions for stability.

math.OC

Nonsmooth rank-one matrix factorization landscape

We provide the first positive result on the nonsmooth optimization landscape of robust principal component analysis, to the best of our knowledge. It is the object of several conjectures and remains mostly uncharted territory. We identify a necessary and sufficient condition for the absence of spurious local minima in the rank-one case. Our proof exploits the subdifferential regularity of the objective function in order to eliminate the existence quantifier from the first-order optimality condition known as Fermat's rule.

math.OC

Time-Dependent Surveillance-Evasion Games

Surveillance-Evasion (SE) games form an important class of adversarial trajectory-planning problems. We consider time-dependent SE games, in which an Evader is trying to reach its target while minimizing the cumulative exposure to a moving enemy Observer. That Observer is simultaneously aiming to maximize the same exposure by choosing how often to use each of its predefined patrol trajectories. Following the framework introduced in Gilles and Vladimirsky (arXiv:1812.10620), we develop efficient algorithms for finding Nash Equilibrium policies for both players by blending techniques from semi-infinite game theory, convex optimization, and multi-objective dynamic programming on continuous planning spaces. We illustrate our method on several examples with Observers using omnidirectional and angle-restricted sensors on a domain with occluding obstacles.

math.OC