SearcharxivSearch

arXiv subjects

Min Tao

Publications and source records attributed to Min Tao.

17 recordsLinked to original sources

Counterexamples to Whole-Sequence Convergence of Variable-Smoothing Full-Splitting Methods

We study whole-sequence convergence of the smoothing-based full-splitting proximal subgradient method (S-FSPS) for structured nonconvex and nonsmooth fractional programs, introduced by Bo\c{t}, Li, and Tao (SIAM J. Optim., 35(4):2623--2653, 2025) as Algorithm~4.1. Existing theory guarantees only the existence of a subsequence converging to a limiting lifted stationary point. We show that this guarantee is sharp by constructing two admissible instances whose corresponding primal sequences both have cluster set $\{1\}\times\mathbb S^1$ and infinite length, although every cluster point is a limiting lifted stationary point. The construction prescribes a slowly rotating spiral and realizes it exactly through a compatible first-order jet and a $C^{1,1}$ Whitney extension. In the first instance, $A$ has rank one and every smoothing-dual iterate is nonzero. In the second, the feasible set is full-dimensional, $A$ has full row rank, and each of $f\circ K$, $g\circ A$, and the numerator $g\circ A+h$ is nonconstant on the feasible set. The first instance also yields a nonconvergent example for the corresponding variable-smoothing, single-loop, full-splitting method for nonconvex and nonsmooth composite optimization, although that method still admits a subsequence converging to an exact stationary point. Thus, a vanishing but nonsummable smoothing schedule does not imply whole-sequence convergence.

math.OC

Bridging Identification and Second-Order Acceleration: A Fast Alternating Minimization Framework for Composite Optimization

We consider a class of composite optimization problems involving a smooth function and a proper, lower semicontinuous regularizer, which may be nonconvex and nonsmooth. We propose a novel alternating minimization framework that integrates proximal-gradient steps with cubic-regularized Newton updates restricted to a dynamically identified low-dimensional subspace. Under the Kurdyka--{\L}ojasiewicz (KL) property, we establish global convergence of the proposed method to a stationary point. Moreover, by incorporating an adaptive thresholding strategy guided by the KL exponent, we prove a finite identification property without imposing any nondegeneracy assumptions. We further develop a local convergence analysis and show that the proposed method attains a worst-case iteration complexity of $\mathcal{O}(\varepsilon^{-3/2})$ for achieving approximate second-order stationarity. Numerical experiments on both synthetic and real datasets demonstrate the efficiency and effectiveness of the proposed framework.

math.OC

A full splitting algorithm for structured difference-of-convex programs

In this paper, we study a class of nonconvex and nonsmooth structured difference-of-convex (DC) programs, which contain in the convex part the sum of a nonsmooth linearly composed convex function and a differentiable function, and in the concave part another nonsmooth linearly composed convex function. Among the various areas in which such problems occur, we would like to mention in particular the recovery of sparse signals. We propose an adaptive double-proximal, full-splitting algorithm with a moving center approach in the final subproblem, which addresses the challenge of evaluating compositions by decoupling the linear operator from the nonsmooth component. We establish the subsequential convergence of the generated sequence of iterates to an approximate stationary point and prove its global convergence under the Kurdyka-\L ojasiewicz property. We also discuss the tightness of the convergence results and provide insights into the rationale for seeking an approximate KKT point. This is illustrated by constructing a counterexample showing that the algorithm can diverge when seeking exact solutions. Finally, we present a practical version of the algorithm that incorporates a nonmonotone line search, which significantly improves the convergence performance.

math.OC

A Single-loop Proximal Subgradient Algorithm for A Class Structured Fractional Programs

In this paper, we investigate a class of nonconvex and nonsmooth fractional programming problems, where the numerator composed of two parts: a convex, nonsmooth function and a differentiable, nonconvex function, and the denominator consists of a convex, nonsmooth function composed of a linear operator. These structured fractional programming problems have broad applications, including CT reconstruction, sparse signal recovery, the single-period optimal portfolio selection problem and standard Sharpe ratio minimization problem. We develop a single-loop proximal subgradient algorithm that alleviates computational complexity by decoupling the evaluation of the linear operator from the nonsmooth component. We prove the global convergence of the proposed single-loop algorithm to an exact lifted stationary point under the Kurdyka-\L ojasiewicz assumption. Additionally, we present a practical variant incorporating a nonmonotone line search to improve computational efficiency. Finally, through extensive numerical simulations, we showcase the superiority of the proposed approach over the existing state-of-the-art methods for three applications: $L_{1}/S_{\kappa}$ sparse signal recovery, limited-angle CT reconstruction, and optimal portfolio selection.

math.OC

Robust model predictive control for large-scale distributed parameter systems under uncertainty

Control of nonlinear distributed parameter systems (DPS) under uncertainty is a meaningful task for many industrial processes. However, both intrinsic uncertainty and high dimensionality of DPS require intensive computations, while non-convexity of nonlinear systems can inhibit the computation of global optima during the control procedure. In this work, polynomial chaos expansion (PCE) was used to account for the uncertainties in quantities of interest through a systematic data collection from the high-fidelity simulator. Then the proper orthogonal decomposition (POD) method was adopted to project the high-dimensional nonlinear dynamics of the computed statistical moments/bounds onto a low-dimensional subspace, where recurrent neural networks (RNNs) were subsequently built to capture the reduced dynamics. Finally, the reduced RNNs based model predictive control (MPC) would generate a set of sequential optimisation problems, of which near global optima could be computed through the mixed integer linear programming (MILP) reformulation techniques and advanced MILP solver. The effectiveness of the proposed framework is demonstrated through two case studies: a chemical tubular reactor and a cell-immobilisation packed-bed bioreactor for the bioproduction of succinic acid.

math.OC

Multiple Gaussian process models based global sensitivity analysis and efficient optimization of in vitro mRNA transcription process

The in vitro transcription (IVT) process is a critical step in RNA production. To ensure the efficiency of RNA manufacturing, it is essential to optimize and identify its key influencing factors. In this study, multiple Gaussian Process (GP) models are used to perform efficient optimization and global sensitivity analysis (GSA). Firstly, multiple GP models were constructed using the data from multiple experimental replicates, accurately capturing the complexities of the IVT process. Then GSA was conducted to determine the dominant reaction factors, specifically the concentrations of reactants NTP and Mg across all data-driven models. Concurrently, a multi-start optimization algorithm was applied to these GP models to identify optimal operational conditions that maximize RNA yields across all surrogate models. These optimized conditions are subsequently validated through additional experimental data.

q-bio.QM

Model reduction, machine learning based global optimisation for large-scale steady state nonlinear systems

Many engineering processes can be accurately modelled using partial differential equations (PDEs), but high dimensionality and non-convexity of the resulting systems pose limitations on their efficient optimisation. In this work, a model reduction, machine-learning methodology combining principal component analysis (PCA) and artificial neural networks (ANNs) is employed to construct a reduced surrogate model, which can then be utilised by advanced deterministic global optimisation algorithms to compute global optimal solutions with theoretical guarantees. However, such optimisation would still be time-consuming due to the high non-convexity of the activation functions inside the reduced ANN structures. To develop a computationally-efficient optimisation framework, we propose two alternative strategies: The first one is a piecewise-affine reformulation of the nonlinear ANN activation functions, while the second one is based on deep rectifier neural networks with ReLU activation function. The performance of the proposed framework is demonstrated through two illustrative case studies.

math.OC

On NP-Hardness of $L_1/L_2$ Minimization and Bound Theory of Nonzero Entries in Solutions

The \(L_1/L_2\) norm ratio has gained significant attention as a measure of sparsity due to three merits: sharper approximation to the \(L_0\) norm compared to the \(L_1\) norm, being parameter-free and scale-invariant, and exceptional performance with highly coherent matrices. These properties have led to its successful application across a wide range of fields. While several efficient algorithms have been proposed to compute stationary points for \(L_1/L_2\) minimization problems, their computational complexity has remained open. In this paper, we prove that finding the global minimum of both constrained and unconstrained \(L_1/L_2\) models is strongly NP-hard. In addition, we establish uniform upper bounds on the \(L_2\) norm for any local minimizer of both constrained and unconstrained \(L_1/L_2\) minimization models. We also derive upper and lower bounds on the magnitudes of the nonzero entries in any local minimizer of the unconstrained model, aiding in classifying nonzero entries. Finally, we extend our analysis to demonstrate that the constrained and unconstrained \(L_p/L_q\) (\(0 < p \leq 1, 1 < q < +\infty\)) models are also strongly NP-hard.

math.OC

On Partly Smoothness, Activity Identification and Faster Algorithms of $L_1$ over $L_2$ Minimization

The $L_1/L_2$ norm ratio arose as a sparseness measure and attracted a considerable amount of attention due to three merits: (i) sharper approximations of $L_0$ compared to the $L_1$; (ii) parameter-free and scale-invariant; (iii) more attractive than $L_1$ under highly-coherent matrices. In this paper, we first establish the partly smooth property of $L_1$ over $L_2$ minimization relative to an active manifold ${\cal M}$ and also demonstrate its prox-regularity property. Second, we reveal that ADMM$_p$ (or ADMM$^+_p$) can identify the active manifold within a finite iterations. This discovery contributes to a deeper understanding of the optimization landscape associated with $L_1$ over $L_2$ minimization. Third, we propose a novel heuristic algorithm framework that combines ADMM$_p$ (or ADMM$^+_p$) with a globalized semismooth Newton method tailored for the active manifold ${\cal M}$. This hybrid approach leverages the strengths of both methods to enhance convergence. Finally, through extensive numerical simulations, we showcase the superiority of our heuristic algorithm over existing state-of-the-art methods for sparse recovery.

math.OC

SiN-on-SOI Optical Phased Array LiDAR for Ultra-Wide Field of View and 4D Sensing

Three-dimensional (3D) imaging techniques are facilitating the autonomous vehicles to build intelligent system. Optical phased arrays (OPAs) featured by all solid-state configurations are becoming a promising solution for 3D imaging. However, majority of state-of-art OPAs commonly suffer from severe power degradation at the edge of field of view (FoV), resulting in limited effective FoV and deteriorating 3D imaging quality. Here, we synergize chained grating antenna and vernier concept to design a novel OPA for realizing a record wide 160{\deg}-FoV 3D imaging. By virtue of the chained antenna, the OPA exhibits less than 3-dB beam power variation within the 160{\deg} FoV. In addition, two OPAs with different pitch are integrated monolithically to form a quasi-coaxial Vernier OPA transceiver. With the aid of flat beam power profile provided by the chained antennas, the OPA exhibits uniform beam quality at an arbitrary steering angle. The superior beam steering performance enables the OPA to accomplish 160{\deg} wide-FoV 3D imaging based on the frequency-modulated continuous-wave (FMCW) LiDAR scheme. The ranging accuracy is 5.5-mm. Moreover, the OPA is also applied to velocity measurement for 4D sensing. To our best knowledge, it is the first experimental implementation of a Vernier OPA LiDAR on 3D imaging to achieve a remarkable FoV.

physics.optics

A full splitting algorithm for fractional programs with structured numerators and denominators

In this paper, we consider a class of nonconvex and nonsmooth fractional programming problems, that involve the sum of a convex, possibly nonsmooth function composed with a linear operator and a differentiable, possibly nonconvex function in the numerator and a convex, possibly nonsmooth function composed with a linear operator in the denominator. These problems have applications in various fields. We propose an adaptive full-splitting proximal subgradient algorithm that addresses the challenge of decoupling the composition of the nonsmooth component with the linear operator in the numerator. We specifically evaluate the nonsmooth function in the numerator using its proximal operator of its conjugate function. Furthermore, the smooth component in the numerator is evaluated through its gradient, and the nonsmooth in the denominator is managed using its subgradient. We demonstrate subsequential convergence toward an approximate lifted stationary point and ensure global convergence under the Kurdyka-\L ojasiewicz property, all achieved without full-row rank assumptions on the linear operators. We provide further discussions on {\it the tightness of the convergence results of the proposed algorithm and its related variants, and the reasoning behind aiming for an approximate lifted stationary point}. We construct a series of counter-examples to show that the proposed algorithm and its variant might diverge when seeking exact solutions. A practical version incorporating a nonmonotone line search is also developed to enhance its performance significantly. Our theoretical findings are validated through simulations involving limited-angle CT reconstruction and the robust sharp-ratio-type minimization problem.

math.OC

Unified Analysis on L1 over L2 Minimization for signal recovery

In this paper, we carry out a unified study for $L_1$ over $L_2$ sparsity promoting models, which are widely used in the regime of coherent dictionaries for recovering sparse nonnegative/arbitrary signals. First, we provide a unified theoretical analysis on the existence of the global solutions of the constrained and the unconstrained $L_{1}/L_{2}$ models. Second, we analyze the sparse property of any local minimizer of these $L_{1}/L_{2}$ models which serves as a certificate to rule out the nonlocal-minimizer stationary solutions. Third, we derive an analytical solution for the proximal operator of the $L_{1} / L_{2}$ with nonnegative constraint. Equipped with this, we apply the alternating direction method of multipliers to the unconstrained model with nonnegative constraint in a particular splitting way, referred to as ADMM$_p^+$. We establish its global convergence to a d-stationary solution (sharpest stationary) without the Kurdyka-\L ojasiewicz assumption. Extensive numerical simulations confirm the superior of ADMM$_p^+$ over the state-of-the-art methods in sparse recovery. In particular, ADMM$_p^+$ reduces computational time by about $95\%\sim99\%$ while achieving a much higher accuracy than the commonly used scaled gradient projection method for the wavelength misalignment problem.

math.OC

Limited-angle CT reconstruction via the L1/L2 minimization

In this paper, we consider minimizing the L1/L2 term on the gradient for a limited-angle scanning problem in computed tomography (CT) reconstruction. We design a specific splitting framework for an unconstrained optimization model so that the alternating direction method of multipliers (ADMM) has guaranteed convergence under certain conditions. In addition, we incorporate a box constraint that is reasonable for imaging applications, and the convergence for the additional box constraint can also be established. Numerical results on both synthetic and experimental datasets demonstrate the effectiveness and efficiency of our proposed approaches, showing significant improvements over the state-of-the-art methods in the limited-angle CT reconstruction.

math.OC

Minimizing L1 over L2 norms on the gradient

In this paper, we study the L1/L2 minimization on the gradient for imaging applications. Several recent works have demonstrated that L1/L2 is better than the L1 norm when approximating the L0 norm to promote sparsity. Consequently, we postulate that applying L1/L2 on the gradient is better than the classic total variation (the L1 norm on the gradient) to enforce the sparsity of the image gradient. To verify our hypothesis, we consider a constrained formulation to reveal empirical evidence on the superiority of L1/L2 over L1 when recovering piecewise constant signals from low-frequency measurements. Numerically, we design a specific splitting scheme, under which we can prove subsequential and global convergence for the alternating direction method of multipliers (ADMM) under certain conditions. Experimentally, we demonstrate visible improvements of L1/L2 over L1 and other nonconvex regularizations for image recovery from low-frequency measurements and two medical applications of MRI and CT reconstruction. All the numerical results show the efficiency of our proposed approach.

math.NA

Unorganized Malicious Attacks Detection

Recommender system has attracted much attention during the past decade. Many attack detection algorithms have been developed for better recommendations, mostly focusing on shilling attacks, where an attack organizer produces a large number of user profiles by the same strategy to promote or demote an item. This work considers a different attack style: unorganized malicious attacks, where attackers individually utilize a small number of user profiles to attack different items without any organizer. This attack style occurs in many real applications, yet relevant study remains open. We first formulate the unorganized malicious attacks detection as a matrix completion problem, and propose the Unorganized Malicious Attacks detection (UMA) approach, a proximal alternating splitting augmented Lagrangian method. We verify, both theoretically and empirically, the effectiveness of our proposed approach.

cs.IR

Convergence analysis of the direct extension of ADMM for multiple-block separable convex minimization

Recently, the alternating direction method of multipliers (ADMM) has found many efficient applications in various areas; and it has been shown that the convergence is not guaranteed when it is directly extended to the multiple-block case of separable convex minimization problems where there are $m\ge 3$ functions without coupled variables in the objective. This fact has given great impetus to investigate various conditions on both the model and the algorithm's parameter that can ensure the convergence of the direct extension of ADMM (abbreviated as "e-ADMM"). Despite some results under very strong conditions (e.g., at least $(m-1)$ functions should be strongly convex) that are applicable to the generic case with a general $m$, some others concentrate on the special case of $m=3$ under the relatively milder condition that only one function is assumed to be strongly convex. We focus on extending the convergence analysis from the case of $m=3$ to the more general case of $m\ge3$. That is, we show the convergence of e-ADMM for the case of $m\ge 3$ with the assumption of only $(m-2)$ functions being strongly convex; and establish its convergence rates in different scenarios such as the worst-case convergence rates measured by iteration complexity and the asymptotically linear convergence rate under stronger assumptions. Thus the convergence of e-ADMM for the general case of $m\ge 4$ is proved; this result seems to be still unknown even though it is intuitive given the known result of the case of $m=3$. Even for the special case of $m=3$, our convergence results turn out to be more general than the exiting results that are derived specifically for the case of $m=3$.

math.OC

On the Optimal Linear Convergence Rate of a Generalized Proximal Point Algorithm

The proximal point algorithm (PPA) has been well studied in the literature. In particular, its linear convergence rate has been studied by Rockafellar in 1976 under certain condition. We consider a generalized PPA in the generic setting of finding a zero point of a maximal monotone operator, and show that the condition proposed by Rockafellar can also sufficiently ensure the linear convergence rate for this generalized PPA. Indeed we show that these linear convergence rates are optimal. Both the exact and inexact versions of this generalized PPA are discussed. The motivation to consider this generalized PPA is that it includes as special cases the relaxed versions of some splitting methods that are originated from PPA. Thus, linear convergence results of this generalized PPA can be used to better understand the convergence of some widely used algorithms in the literature. We focus on the particular convex minimization context and specify Rockafellar's condition to see how to ensure the linear convergence rate for some efficient numerical schemes, including the classical augmented Lagrangian method proposed by Hensen and Powell in 1969 and its relaxed version, the original alternating direction method of multipliers (ADMM) by Glowinski and Marrocco in 1975 and its relaxed version (i.e., the generalized ADMM by Eckstein and Bertsekas in 1992). Some refined conditions weaker than existing ones are proposed in these particular contexts.

math.OC