SearcharxivSearch

arXiv subjects

Lihua Yang

Publications and source records attributed to Lihua Yang.

13 recordsLinked to original sources

Convergence analysis of the Riemannian proximal gradient method with inexact oracle

We extend the notion of an inexact first-order oracle from the Euclidean setting to Riemannian optimization, and conduct a convergence analysis for the Riemannian proximal gradient method equipped with this oracle, which we refer to as RPG-IO. Under mild conditions on the oracle errors, we establish the global convergence of RPG-IO. Specifically, we prove that (i) the norm of the search direction converges to zero; (ii) the sequence of function values converges to the function value at any accumulation point of the iterates; and (iii) every accumulation point of the iterates is a stationary point. Under the additional assumption of the Riemannian Kurdyka--{\L}ojasiewicz (KL) property, we prove that the full sequence of iterates generated by RPG-IO converges to a single stationary point. Moreover, we derive explicit convergence rates when the KL exponent is specified. Finally, when a strong inexact oracle is employed, we establish the convergence rate of the sequence of function values to the optimal value.

math.OC

Local Convergence Analysis of ADMM for Nonconvex Composite Optimization

In this paper, we study the local convergence of the standard ADMM scheme for a class of nonconvex composite optimization problems motivated by applications in signal processing and machine learning. The problems are constrained by a closed convex set, while their objective is the sum of a continuously differentiable, possibly nonconvex, smooth term and a polyhedral convex nonsmooth term composed with a linear mapping. Motivated by recent works of Rockafellar, we first provide an elementary proof of a local strong convexity property of the Moreau envelope of polyhedral convex functions on the orthogonal complement of an appropriate subspace. Building on this property, we establish the strong variational sufficiency of the reduced augmented Lagrangian under an appropriate second-order condition. We then derive a descent inequality for the ADMM iterates that is analogous to the classical descent inequality for convex ADMM. For a sufficiently large penalty parameter, and under suitable initialization and local trajectory conditions, we prove that the ADMM sequence converges to a stationary primal-dual point. When the constraint set is polyhedral convex, we further show that the weighted distance of the primal-dual sequence to the local solution set converges Q-linearly, while the primal sequence converges R-linearly. Finally, we present three illustrative examples together with an application-oriented verification for a class of possibly nonconvex quadratic programs, illustrating the role of the second-order condition, the local nature of the convergence theory, and its applicability.

math.OC

On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions

Transformer networks have achieved remarkable empirical success across a wide range of applications, yet their theoretical expressive power remains insufficiently understood. In this paper, we study the expressive capabilities of Transformer architectures. We first establish an explicit approximation of maxout networks by Transformer networks while preserving comparable model complexity. As a consequence, Transformers inherit the universal approximation capability of ReLU networks under similar complexity constraints. Building on this connection, we develop a framework to analyze the approximation of continuous piecewise linear functions by Transformers and quantitatively characterize their expressivity via the number of linear regions, which grows exponentially with depth. Our analysis establishes a theoretical bridge between approximation theory for standard feedforward neural networks and Transformer architectures. It also yields structural insights into Transformers: self-attention layers implement max-type operations, while feedforward layers realize token-wise affine transformations.

cs.LG

A proximal algorithm incorporating difference of convex functions optimization for solving a class of single-ratio fractional programming

In this paper, we consider a class of single-ratio fractional minimization problems, where both the numerator and denominator of the objective are convex functions satisfying positive homogeneity. Many nonsmooth optimization problems on the sphere that are commonly encountered in application scenarios across different scientific fields can be converted into this equivalent fractional programming. We derive local and global optimality conditions of the problem and subsequently propose a proximal-subgradient-difference of convex functions algorithm (PS-DCA) to compute its critical points. When the DCA step is removed, PS-DCA reduces to the proximal-subgradient algorithm (PSA). Under mild assumptions regarding the algorithm parameters, it is shown that any accumulation point of the sequence produced by PS-DCA or PSA is a critical point of the problem. Moreover, for a typical class of generalized graph Fourier mode problems, we establish global convergence of the entire sequence generated by PS-DCA or PSA. Numerical experiments conducted on computing the generalized graph Fourier modes demonstrate that, compared to proximal gradient-type algorithms, PS-DCA integrates difference of convex functions (d.c.) optimization, rendering it less sensitive to initial points and preventing the sequence it generates from being trapped in low-quality local minimizers.

math.OC

An Equivalent Graph Reconstruction Model and its Application in Recommendation Prediction

Recommendation algorithm plays an important role in recommendation system (RS), which predicts users' interests and preferences for some given items based on their known information. Recently, a recommendation algorithm based on the graph Laplacian regularization was proposed, which treats the prediction problem of the recommendation system as a reconstruction issue of small samples of the graph signal under the same graph model. Such a technique takes into account both known and unknown labeled samples information, thereby obtaining good prediction accuracy. However, when the data size is large, solving the reconstruction model is computationally expensive even with an approximate strategy. In this paper, we propose an equivalent reconstruction model that can be solved exactly with extremely low computational cost. Finally, a final prediction algorithm is proposed. We find in the experiments that the proposed method significantly reduces the computational cost while maintaining a good prediction accuracy.

cs.IR

Perfect Reconstruction Two-Channel Filter Banks on Arbitrary Graphs

This paper extends the existing theory of perfect reconstruction two-channel filter banks from bipartite graphs to non-bipartite graphs. By generalizing the concept of downsampling/upsampling we establish the frame of two-channel filter bank on arbitrary connected, undirected and weighted graphs. Then the equations for perfect reconstruction of the filter banks are presented and solved under proper conditions. Algorithms for designing orthogonal and biorthogonal banks are given and two typical orthogonal two-channel filter banks are calculated. The locality and approximation properties of such filter banks are discussed theoretically and experimentally.

cs.IT

Smoothing algorithms for nonsmooth and nonconvex minimization over the stiefel manifold

We consider a class of nonsmooth and nonconvex optimization problems over the Stiefel manifold where the objective function is the summation of a nonconvex smooth function and a nonsmooth Lipschitz continuous convex function composed with an linear mapping. We propose three numerical algorithms for solving this problem, by combining smoothing methods and some existing algorithms for smooth optimization over the Stiefel manifold. In particular, we approximate the aforementioned nonsmooth convex function by its Moreau envelope in our smoothing methods, and prove that the Moreau envelope has many favorable properties. Thanks to this and the scheme for updating the smoothing parameter, we show that any accumulation point of the solution sequence generated by the proposed algorithms is a stationary point of the original optimization problem. Numerical experiments on building graph Fourier basis are conducted to demonstrate the efficiency of the proposed algorithms.

math.OC

Spline-Like Wavelet Filterbanks with Perfect Reconstruction on Arbitrary Graphs

In this work, we propose a class of spline-like wavelet filterbanks for graph signals. These filterbanks possess the properties of critical sampling and perfect reconstruction. Besides, the analysis filters are localized in the graph domain because they are polynomials of the normalized adjacency matrix of the graph. We generalize the spline-like filters in the literature so that they have the ability to annihilate signals of some specified frequencies. Optimization problems are posed for the analysis filters to approximate desired responses. We conduct some experiments to demonstrate the good locality of the proposed filters and the good performance of the filterbank in the denoising task.

eess.SP

Perfect Reconstruction Two-Channel Filter Banks on Arbitrary Graphs Based on an Optimization Model

In this paper, we propose the construction of critically sampled perfect reconstruction two-channel filterbanks on arbitrary undirected graphs.Inspired by the design of graphQMF proposed in the literature, we propose a general ``spectral folding property'' similar to that of bipartite graphs and provide sufficient conditions for constructing perfect reconstruction filterbanks based on a general graph Fourier basis, which is not the eigenvectors of the Laplacian matrix. To obtain the desired graph Fourier basis, we need to solve a series of quadratic equality constrained quadratic optimization problems (QECQPs) which are known to be non-convex and difficult to solve. We develop an algorithm to obtain the global optimal solution within a pre-specified tolerance. Multi-resolution analysis on real-world data and synthetic data are performed to validate the effectiveness of the proposed filterbanks.

eess.SP

Atomic Filter: a Weak Form of Shift Operator for Graph Signals

The shift operation plays a crucial role in the classical signal processing. It is the generator of all the filters and the basic operation for time-frequency analysis, such as windowed Fourier transform and wavelet transform. With the rapid development of internet technology and big data science, a large amount of data are expressed as signals defined on graphs. In order to establish the theory of filtering, windowed Fourier transform and wavelet transform in the setting of graph signals, we need to extend the shift operation of classical signals to graph signals. It is a fundamental problem since the vertex set of a graph is usually not a vector space and the addition operation cannot be defined on the vertex set of the graph. In this paper, based on our understanding on the core role of shift operation in classical signal processing we propose the concept of atomic filters, which can be viewed as a weak form of the shift operator for graph signals. Then, we study the conditions such that an atomic filter is norm-preserving, periodic, or real-preserving. The property of real-preserving holds naturally in the classical signal processing, but no the research has been reported on this topic in the graph signal setting. With these conditions we propose the concept of normal atomic filters for graph signals, which degenerates into the classical shift operator under mild conditions if the graph is circulant. Typical examples of graphs that have or have not normal atomic filters are given. Finally, as an application, atomic filters are utilized to construct time-frequency atoms which constitute a frame of the graph signal space.

eess.SP

Graph Fourier Transform Based on $\ell_1$ Norm Variation Minimization

The definition of the graph Fourier transform is a fundamental issue in graph signal processing. Conventional graph Fourier transform is defined through the eigenvectors of the graph Laplacian matrix, which minimize the $\ell_2$ norm signal variation. However, the computation of Laplacian eigenvectors is expensive when the graph is large. In this paper, we propose an alternative definition of graph Fourier transform based on the $\ell_1$ norm variation minimization. We obtain a necessary condition satisfied by the $\ell_1$ Fourier basis, and provide a fast greedy algorithm to approximate the $\ell_1$ Fourier basis. Numerical experiments show the effectiveness of the greedy algorithm. Moreover, the Fourier transform under the greedy basis demonstrates a similar rate of decay to that of Laplacian basis for simulated or real signals.

cs.IT

Quantum measurements and maps preserving strict convex combinations and pure states

In this paper, a characterization of maps between quantum states that preserve pure states and strict convex combinations is obtained. Based on this characterization, a structural theorem for maps between multipartite quantum states that preserve separable pure states and strict convex combinations is established. Then these results are applied to characterize injective (local) quantum measurements and answer some conjectures proposed in [J.Phys.A:Math.Theor. 45 (2012) 205305].

quant-ph

Hybrid Chaplygin gas and phantom divide crossing

Hybrid Chaplygin gas model is put forward, in which the gases play the role of dark energy. For this model the coincidence problem is greatly alleviated. The effective equation of state of the dark energy may cross the phantom divide $w=-1$. Furthermore, the crossing behaviour is decoupled from any gravity theories. In the present model, $w<-1$ is only a transient behaviour. There is a de Sitter attractor in the future infinity. Hence, the big rip singularity, which often afflicts the models with matter whose effective equation of state less than -1, is naturally disappear. There exist stable scaling solutions, both at the early universe and the late universe. We discuss the perturbation growth of this model. We find that the index is consistent with observations.

astro-ph