SearcharxivSearch

arXiv subjects

Yaohua Hu

Publications and source records attributed to Yaohua Hu.

17 recordsLinked to original sources

Global $o(1/k^2)$ Merit Complexity of Regularized Newton Methods for Convex Multiobjective Optimization

We investigate a regularized Newton method for unconstrained convex multi-objective optimization with twice continuously differentiable objectives whose Hessians are Lipschitz continuous. At each iteration, the method minimizes the quadratically regularized max-envelope of the local quadratic models. Using a Tanabe-type merit function, we prove that this merit decays at the global asymptotic rate $o(1/k^2)$ under the compactness assumption on the initial component-wise lower level set. This result also covers the single-objective case as a special case. Finally, we construct an explicit one-dimensional convex bi-objective family showing that no uniform merit estimate of order $\mathcal O(k^{-(2+\delta)})$ can hold for any fixed $\delta>0$. Thus the exponent $2$ is essentially sharp in the uniform polynomial sense, despite the $o(1/k^2)$ decay on each fixed trajectory.

math.OC

$\ell_{1\text{-}2}$ Regularization for Sparse Optimization: Consistency and Global Convergence

The $\ell_{1\text{-}2}$ regularization method has a strong sparsity promoting capability in approaching sparse solutions of linear inverse problems and gained successful applications in various mathematics and applied science fields. This paper aims to investigate the consistency theory and global convergent algorithms for the $\ell_{1\text{-}2}$ regularization problem. In the theoretical aspect, we introduce a notion of restricted eigenvalue condition relative to the $\ell_{1\text{-}2}$ penalty, and employ it to establish an oracle property and a recovery bound for the global solution of the $\ell_{1\text{-}2}$ regularization problem. In the algorithmic aspect, we propose two types of iterative thresholding algorithms with the truncation technique and the continuation technique, respectively, to solve the $\ell_{1\text{-}2}$ regularization problem. Moreover, under the assumption of the well-known restricted isometry property, we establish the convergence of the proposed algorithms to the ground true sparse solution within a tolerance relevant to the noise level and the recovery bound. Preliminary numerical results show that our proposed algorithms can approach the ground true sparse solution and significantly enhance the sparsity recovery capability, compared with the popular sparse optimization algorithms in the literature.

math.OC

Dynamic Proximal Gradient Algorithms for Schatten-$p$ Quasi-Norm Regularized Problems

This paper investigates numerical solution methods for the Schatten-$p$ quasi-norm regularized problem with $p \in [0,1]$, which has been widely studied for finding low-rank solutions of linear inverse problems and gained successful applications in various mathematics and applied science fields. We propose a dynamic proximal gradient algorithm that, through the use of the Cayley transformation, avoids computationally expensive singular value decompositions at each iteration, thereby significantly reducing the computational complexity. The algorithm incorporates two step size selection strategies: an adaptive backtracking search and an explicit step size rule. We establish the sublinear convergence of the proposed algorithm for all $p \in [0,1]$ within the framework of the Kurdyka-Lojasiewicz property. Notably, under mild assumptions, we show that the generated sequence converges to a stationary point of the objective function of the problem. For the special case when $p=1$, the linear convergence is further proved under the strict complementarity-type regularity condition commonly used in the linear convergence analysis of the forward-backward splitting algorithms. Preliminary numerical results validate the superior computational efficiency of the proposed algorithm.

math.OC

HierLight-YOLO: A Hierarchical and Lightweight Object Detection Network for UAV Photography

The real-time detection of small objects in complex scenes, such as the unmanned aerial vehicle (UAV) photography captured by drones, has dual challenges of detecting small targets (<32 pixels) and maintaining real-time efficiency on resource-constrained platforms. While YOLO-series detectors have achieved remarkable success in real-time large object detection, they suffer from significantly higher false negative rates for drone-based detection where small objects dominate, compared to large object scenarios. This paper proposes HierLight-YOLO, a hierarchical feature fusion and lightweight model that enhances the real-time detection of small objects, based on the YOLOv8 architecture. We propose the Hierarchical Extended Path Aggregation Network (HEPAN), a multi-scale feature fusion method through hierarchical cross-level connections, enhancing the small object detection accuracy. HierLight-YOLO includes two innovative lightweight modules: Inverted Residual Depthwise Convolution Block (IRDCB) and Lightweight Downsample (LDown) module, which significantly reduce the model's parameters and computational complexity without sacrificing detection capabilities. Small object detection head is designed to further enhance spatial resolution and feature fusion to tackle the tiny object (4 pixels) detection. Comparison experiments and ablation studies on the VisDrone2019 benchmark demonstrate state-of-the-art performance of HierLight-YOLO.

cs.CV

A Globalized Semismooth Newton Method for Prox-regular Optimization Problems

We are concerned with a class of nonconvex and nonsmooth composite optimization problems, comprising a twice differentiable function and a prox-regular function. We establish a sufficient condition for the proximal mapping of a prox-regular function to be single-valued and locally Lipschitz continuous. By virtue of this property, we propose a hybrid of proximal gradient and semismooth Newton methods for solving these composite optimization problems, which is a globalized semismooth Newton method. The whole sequence is shown to converge to an $L$-stationary point under a Kurdyka-Łojasiewicz exponent assumption. Under an additional error bound condition and some other mild conditions, we prove that the sequence converges to a nonisolated $L$-stationary point at a superlinear convergence rate. Numerical comparison with several existing second order methods reveal that our approach performs comparably well in solving both the $\ell_q(0<q<1)$ quasi-norm regularized problems and the fused zero-norm regularization problems.

math.OC

M3D: Manifold-based Domain Adaptation with Dynamic Distribution for Non-Deep Transfer Learning in Cross-subject and Cross-session EEG-based Emotion Recognition

Emotion decoding using Electroencephalography (EEG)-based affective brain-computer interfaces (aBCIs) plays a crucial role in affective computing but is limited by challenges such as EEG's non-stationarity, individual variability, and the high cost of large labeled datasets. While deep learning methods are effective, they require extensive computational resources and large data volumes, limiting their practical application. To overcome these issues, we propose Manifold-based Domain Adaptation with Dynamic Distribution (M3D), a lightweight, non-deep transfer learning framework. M3D consists of four key modules: manifold feature transformation, dynamic distribution alignment, classifier learning, and ensemble learning. The data is mapped to an optimal Grassmann manifold space, enabling dynamic alignment of source and target domains. This alignment is designed to prioritize both marginal and conditional distributions, improving adaptation efficiency across diverse datasets. In classifier learning, the principle of structural risk minimization is applied to build robust classification models. Additionally, dynamic distribution alignment iteratively refines the classifier. The ensemble learning module aggregates classifiers from different optimization stages to leverage diversity and enhance prediction accuracy. M3D is evaluated on two EEG emotion recognition datasets using two validation protocols (cross-subject single-session and cross-subject cross-session) and a clinical EEG dataset for Major Depressive Disorder (MDD). Experimental results show that M3D outperforms traditional non-deep learning methods with a 4.47% average improvement and achieves deep learning-level performance with reduced data and computational requirements, demonstrating its potential for real-world aBCI applications.

cs.HC

An Inexact Variable Metric Proximal Gradient-subgradient Algorithm for a Class of Fractional Optimization Problems

In this paper, we study a class of fractional optimization problems, in which the numerator of the objective is the sum of a convex function and a differentiable function with a Lipschitz continuous gradient, while the denominator is a nonsmooth convex function. This model captures ratio-type formulations arising in scale-invariant sparse learning and related applications. To address this class of problems, we propose an inexact variable metric proximal gradient-subgradient algorithm (iVPGSA), which, to the best of our knowledge, is the first inexact proximal algorithm specifically designed for such type of fractional problems. By incorporating a variable metric proximal term and allowing for approximate subproblem solutions under a flexible error criterion, the proposed algorithm is highly adaptable to a broader range of problems while achieving favorable computational efficiency. Under suitable assumptions, we establish that any accumulation point of the generated sequence is a critical point of the target problem. Moreover, we develop a new Kurdyka-{\L}ojasiewicz (KL)-based analysis framework, relying only on the classical KL property and its associated exponent, to prove the global convergence of the entire sequence and characterize its convergence rate, \textit{without} requiring a strict sufficient descent property. Our results clarify how the classical KL exponent and inexactness jointly influence the convergence rate. Finally, numerical experiments on the $\ell_1/\ell_2$ Lasso problem and the constrained $\ell_1/\ell_2$ sparse optimization problem demonstrate the computational advantages of the iVPGSA over existing representative algorithms.

math.OC

Avoiding strict saddle points of nonconvex regularized problems

In this paper, we consider a class of non-convex and non-smooth sparse optimization problems, which encompass most existing nonconvex sparsity-inducing terms. We show the second-order optimality conditions only depend on the nonzeros of the stationary points. We propose two damped iterative reweighted algorithms including the iteratively reweighted $\ell_1$ algorithm (DIRL$_1$) and the iteratively reweighted $\ell_2$ (DIRL$_2$) algorithm, to solve these problems. For DIRL$_1$, we show the reweighted $\ell_1$ subproblem has support identification property so that DIRL$_1$ locally reverts to a gradient descent algorithm around a stationary point. For DIRL$_2$, we show the solution map of the reweighted $\ell_2$ subproblem is differentiable and Lipschitz continuous everywhere. Therefore, the map of DIRL$_1$ and DIRL$_2$ and their inverse are Lipschitz continuous, and the strict saddle points are their unstable fixed points. By applying the stable manifold theorem, these algorithms are shown to converge only to local minimizers with randomly initialization when the strictly saddle point property is assumed.

math.OC

Astronomical Knowledge Entity Extraction in Astrophysics Journal Articles via Large Language Models

Astronomical knowledge entities, such as celestial object identifiers, are crucial for literature retrieval and knowledge graph construction, and other research and applications in the field of astronomy. Traditional methods of extracting knowledge entities from texts face challenges like high manual effort, poor generalization, and costly maintenance. Consequently, there is a pressing need for improved methods to efficiently extract them. This study explores the potential of pre-trained Large Language Models (LLMs) to perform astronomical knowledge entity extraction (KEE) task from astrophysical journal articles using prompts. We propose a prompting strategy called Prompt-KEE, which includes five prompt elements, and design eight combination prompts based on them. Celestial object identifier and telescope name, two most typical astronomical knowledge entities, are selected to be experimental object. And we introduce four currently representative LLMs, namely Llama-2-70B, GPT-3.5, GPT-4, and Claude 2. To accommodate their token limitations, we construct two datasets: the full texts and paragraph collections of 30 articles. Leveraging the eight prompts, we test on full texts with GPT-4 and Claude 2, on paragraph collections with all LLMs. The experimental results demonstrated that pre-trained LLMs have the significant potential to perform KEE tasks in astrophysics journal articles, but there are differences in their performance. Furthermore, we analyze some important factors that influence the performance of LLMs in entity extraction and provide insights for future KEE tasks in astrophysical articles using LLMs.

astro-ph.IM

The Behavior of Error Bounds via Moreau Envelopes

In this paper, we first establish the equivalence of three types of error bounds: uniformized Kurdyka-Łojasiewicz (u-KL) property, uniformized level-set subdifferential error bound (u-LSEB) and uniformized Hölder error bound (u-HEB) for prox-regular functions. Then we study the behavior of the level-set subdifferential error bound (LSEB) and the local Hölder error bound (LHEB) which is expressed respectively by Moreau envelopes, under suitable assumptions. Finally, in order to illustrate our main results and to compare them with those of recent references, some examples are also given.

math.OC

Design and Characterization of a Novel Motion Conversion Element: Curved Groove Ball Bearing without Retainer

An unconventional and innovative mechanical transmission component, the Curved Groove Ball Bearing without Retainer (CGBBR), is introduced in this paper to facilitate the conversion between rotary and reciprocating motion. The CGBBR boasts several advantages over conventional motion conversion mechanisms, including a streamlined and compact structure, as well as the mitigation of high-order vibrations. This paper delves into the design methodology, structural attributes, kinematic principles, and dynamic response properties of the CGBBR. A novel design method named closed curve envelopment theory is proposed to generate the distinctive spatial surface structure of the CGBBR. Moreover, a design scheme is presented that employs the concept of diameter-stroke ratio for CGBBR implementation. A calculation methodology is also introduced for determining the number of rolling elements within the CGBBR, which serves as the foundation for subsequent optimization design. This paper deeply analyzes the kinematics principle of CGBBR and provides a new insight of motion conversion in mechanism.

physics.app-ph

Nonconvex and Nonsmooth Sparse Optimization via Adaptively Iterative Reweighted Methods

We propose a general formulation of nonconvex and nonsmooth sparse optimization problems with convex set constraint, which can take into account most existing types of nonconvex sparsity-inducing terms, bringing strong applicability to a wide range of applications. We design a general algorithmic framework of iteratively reweighted algorithms for solving the proposed nonconvex and nonsmooth sparse optimization problems, which solves a sequence of weighted convex regularization problems with adaptively updated weights. First-order optimality condition is derived and global convergence results are provided under loose assumptions, making our theoretical results a practical tool for analyzing a family of various reweighted algorithms. The effectiveness and efficiency of our proposed formulation and the algorithms are demonstrated in numerical experiments on various sparse optimization problems.

cs.IT

Sparse estimation via $\ell_q$ optimization method in high-dimensional linear regression

In this paper, we discuss the statistical properties of the $\ell_q$ optimization methods $(0<q\leq 1)$, including the $\ell_q$ minimization method and the $\ell_q$ regularization method, for estimating a sparse parameter from noisy observations in high-dimensional linear regression with either a deterministic or random design. For this purpose, we introduce a general $q$-restricted eigenvalue condition (REC) and provide its sufficient conditions in terms of several widely-used regularity conditions such as sparse eigenvalue condition, restricted isometry property, and mutual incoherence property. By virtue of the $q$-REC, we exhibit the stable recovery property of the $\ell_q$ optimization methods for either deterministic or random designs by showing that the $\ell_2$ recovery bound $O(ε^2)$ for the $\ell_q$ minimization method and the oracle inequality and $\ell_2$ recovery bound $O(λ^{\frac{2}{2-q}}s)$ for the $\ell_q$ regularization method hold respectively with high probability. The results in this paper are nonasymptotic and only assume the weak $q$-REC. The preliminary numerical results verify the established statistical property and demonstrate the advantages of the $\ell_q$ regularization method over some existing sparse optimization methods.

stat.ML

Convergence Rates of Subgradient Methods for Quasi-convex Optimization Problems

Quasi-convex optimization acts a pivotal part in many fields including economics and finance; the subgradient method is an effective iterative algorithm for solving large-scale quasi-convex optimization problems. In this paper, we investigate the iteration complexity and convergence rates of various subgradient methods for solving quasi-convex optimization in a unified framework. In particular, we consider a sequence satisfying a general (inexact) basic inequality, and investigate the global convergence theorem and the iteration complexity when using the constant, diminishing or dynamic stepsize rules. More importantly, we establish the linear (or sublinear) convergence rates of the sequence under an additional assumption of weak sharp minima of Hölderian order and upper bounded noise. These convergence theorems are applied to establish the iteration complexity and convergence rates of several subgradient methods, including the standard/inexact/conditional subgradient methods, for solving quasi-convex optimization problems under the assumptions of the Hölder condition and/or the weak sharp minima of Hölderian order.

math.OC

Incremental Subgradient Methods for Minimizing The Sum of Quasi-convex Functions

The sum of ratios problem has a variety of important applications in economics and management science, but it is difficult to globally solve this problem. In this paper, we consider the minimization problem of a sum of a number of nondifferentiable quasi-convex component functions over a closed and convex set, which includes the sum of ratios problem as a special case. The sum of quasi-convex component functions is not necessarily to be quasi-convex, and so, this study goes beyond quasi-convex optimization. Exploiting the structure of the sum-minimization problem, we propose a new incremental subgradient method for this problem and investigate its convergence properties to a global optimal solution when using the constant, diminishing or dynamic stepsize rules and under a homogeneous assumption and the Hölder condition of order $p$. To economize on the computation cost of subgradients of a large number of component functions, we further propose a randomized incremental subgradient method, in which only one component function is randomly selected to construct the subgradient direction at each iteration. The convergence properties are obtained in terms of function values and distances of iterates from the optimal solution set with probability 1. The proposed incremental subgradient methods are applied to solve the sum of ratios problem, as well as the multiple Cobb-Douglas productions efficiency problem, and the numerical results show that the proposed methods are efficient for solving the large sum of ratios problem.

math.OC

Linear convergence of inexact descent method and inexact proximal gradient algorithms for lower-order regularization problems

The $\ell_p$ regularization problem with $0< p< 1$ has been widely studied for finding sparse solutions of linear inverse problems and gained successful applications in various mathematics and applied science fields. The proximal gradient algorithm is one of the most popular algorithms for solving the $\ell_p$ regularisation problem. In the present paper, we investigate the linear convergence issue of one inexact descent method and two inexact proximal gradient algorithms (PGA). For this purpose, an optimality condition theorem is explored to provide the equivalences among a local minimum, second-order optimality condition and second-order growth property of the $\ell_p$ regularization problem. By virtue of the second-order optimality condition and second-order growth property, we establish the linear convergence properties of the inexact descent method and inexact PGAs under some simple assumptions. Both linear convergence to a local minimal value and linear convergence to a local minimum are provided. Finally, the linear convergence results of the inexact numerical methods are extended to the infinite-dimensional Hilbert spaces.

math.OC

Group sparse optimization via $\ell_{p,q}$ regularization

In this paper, we investigate a group sparse optimization problem via $\ell_{p,q}$ regularization in three aspects: theory, algorithm and application. In the theoretical aspect, by introducing a notion of group restricted eigenvalue condition, we establish some oracle property and a global recovery bound of order $O(λ^\frac{2}{2-q})$ for any point in a level set of the $\ell_{p,q}$ regularization problem, and by virtue of modern variational analysis techniques, we also provide a local analysis of recovery bound of order $O(λ^2)$ for a path of local minima. In the algorithmic aspect, we apply the well-known proximal gradient method to solve the $\ell_{p,q}$ regularization problems, either by analytically solving some specific $\ell_{p,q}$ regularization subproblems, or by using the Newton method to solve general $\ell_{p,q}$ regularization subproblems. In particular, we establish the linear convergence rate of the proximal gradient method for solving the $\ell_{1,q}$ regularization problem under some mild conditions. As a consequence, the linear convergence rate of proximal gradient method for solving the usual $\ell_{q}$ regularization problem ($0<q<1$) is obtained. Finally in the aspect of application, we present some numerical results on both the simulated data and the real data in gene transcriptional regulation.

math.OC