SearcharxivSearch

arXiv subjects

Xingju Cai

Publications and source records attributed to Xingju Cai.

8 recordsLinked to original sources

A new three-operator splitting method for the monotone inclusion problem

This paper studies a class of monotone inclusion problems in a real Hilbert space involving the sum of three operators, where two are maximal monotone and the third is cocoercive. The Davis--Yin three-operator splitting method extends the two-operator splitting methods -- namely the forward-backward method and the Douglas--Rachford method -- to the three-operator setting. In addition, two other common splitting methods for two-operator problems are the reflected forward-backward and forward-reflected-backward methods. While several three-operator extensions exist for each of these methods individually, a unified framework that generalizes both remains absent. This raises the question: can they be extended to the three-operator case within a single algorithm? To address this, we propose a new splitting algorithm that unifies the Douglas--Rachford, reflected forward-backward, and forward-reflected-backward methods as special cases. We prove its weak convergence and establish its sublinear convergence rate for convex optimization problems under appropriate stepsize conditions. Finally, we present numerical experiments to validate the theoretical properties and demonstrate the effectiveness of the proposed method.

math.OC

S-D-RSM: Stochastic Distributed Regularized Splitting Method for Large-Scale Convex Optimization Problems

This paper investigates the problems large-scale distributed composite convex optimization, with motivations from a broad range of applications, including multi-agent systems, federated learning, smart grids, wireless sensor networks, compressed sensing, and so on. Stochastic gradient descent (SGD) and its variants are commonly employed to solve such problems. However, existing algorithms often rely on vanishing step sizes, strong convexity assumptions, or entail substantial computational overhead to ensure convergence or obtain favorable complexity. To bridge the gap between theory and practice, we integrate consensus optimization and operator splitting techniques (see Problem Reformulation) to develop a novel stochastic splitting algorithm, termed the \emph{stochastic distributed regularized splitting method} (S-D-RSM). In practice, S-D-RSM performs parallel updates of proximal mappings and gradient information for only a randomly selected subset of agents at each iteration. By introducing regularization terms, it effectively mitigates consensus discrepancies among distributed nodes. In contrast to conventional stochastic methods, our theoretical analysis establishes that S-D-RSM achieves global convergence without requiring diminishing step sizes or strong convexity assumptions. Furthermore, it achieves an iteration complexity of $\mathcal{O}(1/ε)$ with respect to both the objective function value and the consensus error. Numerical experiments show that S-D-RSM achieves up to 2--3$\times$ speedup compared to state-of-the-art baselines, while maintaining comparable or better accuracy. These results not only validate the algorithm's theoretical guarantees but also demonstrate its effectiveness in practical tasks such as compressed sensing and empirical risk minimization.

math.OC

Research on the descent direction of prediction correction algorithms for pseudo-convex/convex optimization problems

Prediction-correction algorithms are a highly effective class of methods for solving pseudo-convex optimization problems. The descent direction of these algorithms can be viewed as an adjustment to the gradient direction based on the prediction step. This paper investigates the adjustment coefficients of these descent directions and offers explanations from the perspective of differential equations. Unlike existing algorithms where the adjustment coefficient is always set to 1, we establish that the range of the adjustment coefficient lies within (1/2,1] for pseudo-convex optimization problems, and [0,1] for convex optimization problems. We also provide rigorous convergence proofs for these proposed algorithms. Numerical experiment results show that the algorithms perform best when the value of the adjustment coefficient makes the algorithm approach or equal to those in differential equations with higher-order global discrete error.

math.OC

The Güler-type acceleration for proximal gradient, linearized augmented Lagrangian and linearized alternating direction method of multipliers

In this paper, we introduce the Güler-type acceleration technique and utilize it to propose three acceleration algorithms: the Güler-type accelerated proximal gradient method (GPGM), the Güler-type accelerated linearized augmented Lagrangian method (GLALM) and the Güler-type accelerated linearized alternating direction method of multipliers (GLADMM). The key idea behind these algorithms is to fully leverage the information of negative term \bm{$-\|x^k-\hat{x}^{k-1}\|^2$} in order to design the extrapolation step. This concept of using negative terms to improve acceleration can be extended to other algorithms as well. Moreover, the proposed GLALM and GLADMM enable simultaneous acceleration of both primal and dual variables. Additionally, GPGM and GLALM achieve the same convergence rate of $O(\frac{1}{k^2})$ with some existing results. Although GLADMM achieves the same total convergence rate of $O(\frac{1}{N})$ as in existing results, the partial convergence rate is improved from $O(\frac{1}{N^{3/2}})$ to $O(\frac{1}{N^2})$. To validate the effectiveness of our algorithms, we conduct numerical experiments on various problem instances, including the $\ell_1$ regularized logistic regression, quadratic programming, and compressive sensing. The experimental results indicate that our algorithms outperform existing methods in terms of efficiency. This also demonstrates the potential of the stochastic algorithmic versions of these algorithms in application areas such as statistics, machine learning, and data mining. Finally, it is worth noting that this paper aims to introduce how Güler's acceleration technique can be applied to gradient-based algorithms and to provide a unified and concise framework for their construction.

math.OC

An Inexact Proximal Framework for Nonsmooth Riemannian Difference-of-Convex Optimization

Nonsmooth Riemannian optimization has attracted increasing attention, especially in problems with sparse structures. While existing formulations typically involve convex nonsmooth terms, incorporating nonsmooth difference-of-convex (DC) penalties can enhance recovery accuracy. In this paper, we study a class of nonsmooth Riemannian optimization problems whose objective is the sum of a smooth function and a nonsmooth DC term. We establish, for the first time in the manifold setting, the equivalence between such DC formulations (with suitably chosen nonsmooth DC terms) and their $\ell_0$-regularized or $\ell_0$-constrained counterparts. To solve these problems, we propose an inexact Riemannian proximal DC (iRPDC) algorithmic framework, which returns an $ε$-Riemannian critical point within $\mathcal{O}(ε^{-2})$ outer iterations. Within this framework, we develop several practical algorithms based on different subproblem solvers. Among them, one achieves an overall iteration complexity of $\mathcal{O}(ε^{-3})$, which matches the best-known bound in the literature. In contrast, existing algorithms either lack provable overall complexity or require $\mathcal{O}(ε^{-3})$ iterations in both outer and overall complexity. A notable feature of the iRPDC algorithmic framework is a novel inexactness criterion that not only enables efficient subproblem solutions via first-order methods but also facilitates a linesearch procedure that adaptively captures the local curvature. Numerical results on sparse principal component analysis demonstrate the modeling flexibility of the DC formulaton and the competitive performance of the proposed algorithmic framework.

math.OC

A distributed Douglas-Rachford splitting method for solving linear constrained multi-block weakly convex problems

In recent years, a distributed Douglas-Rachford splitting method (DDRSM) has been proposed to tackle multi-block separable convex optimization problems. This algorithm offers relatively easier subproblems and greater efficiency for large-scale problems compared to various augmented-Lagrangian-based parallel algorithms. Building upon this, we explore the extension of DDRSM to weakly convex cases. By assuming weak convexity of the objective function and introducing an error bound assumption, we demonstrate the linear convergence rate of DDRSM. Some promising numerical experiments involving compressed sensing and robust alignment of structures across images (RASL) show that DDRSM has advantages over augmented-Lagrangian-based algorithms, even in weakly convex scenarios.

math.OC

Novel clustered federated learning based on local loss

This paper proposes LCFL, a novel clustering metric for evaluating clients' data distributions in federated learning. LCFL aligns with federated learning requirements, accurately assessing client-to-client variations in data distribution. It offers advantages over existing clustered federated learning methods, addressing privacy concerns, improving applicability to non-convex models, and providing more accurate classification results. LCFL does not require prior knowledge of clients' data distributions. We provide a rigorous mathematical analysis, demonstrating the correctness and feasibility of our framework. Numerical experiments with neural network instances highlight the superior performance of LCFL over baselines on several clustered federated learning benchmarks.

cs.LG

Understanding the convergence of the preconditioned PDHG method: a view of indefinite proximal ADMM

The primal-dual hybrid gradient (PDHG) algorithm is popular in solving min-max problems which are being widely used in a variety of areas. To improve the applicability and efficiency of PDHG for different application scenarios, we focus on the preconditioned PDHG (PrePDHG) algorithm, which is a framework covering PDHG, alternating direction method of multipliers (ADMM), and other methods. We give the optimal convergence condition of PrePDHG in the sense that the key parameters in the condition can not be further improved, which fills the theoretical gap in the-state-of-art convergence results of PrePDHG, and obtain the ergodic and non-ergodic sublinear convergence rates of PrePDHG. The theoretical analysis is achieved by establishing the equivalence between PrePDHG and indefinite proximal ADMM. Besides, we discuss various choices of the proximal matrices in PrePDHG and derive some interesting results. For example, the convergence condition of diagonal PrePDHG is improved to be tight, the dual stepsize of the balanced augmented Lagrangian method can be enlarged to $4/3$ from $1$, and a balanced augmented Lagrangian method with symmetric Gauss-Seidel iterations is also explored. Numerical results on the matrix game, projection onto the Birkhoff polytope, earth mover's distance, and CT reconstruction verify the effectiveness and superiority of PrePDHG.

math.OC