SearcharxivSearch

arXiv subjects

Shixiang Chen

Publications and source records attributed to Shixiang Chen.

At least 19 recordsLinked to original sources

Jointly Sparse Blind Deconvolution via Riemannian Optimization

Blind deconvolution has been widely applied in system identification and signal processing. While joint sparsity commonly arises in practical scenarios, effectively exploiting this structure to enhance recovery performance remains a challenging and largely open problem. In this paper, we propose a joint-sparsity-promoting optimization problem and develop a Riemannian optimization algorithm for its accurate and efficient solution. We further establish theoretical guarantees that characterize the non-asymptotic relationship between the estimation error and the sample complexity, showing that exploiting joint sparsity can significantly reduce the sample complexity required for successful recovery. Numerical experiments are provided that validate the theoretical results and demonstrate the effectiveness of the proposed approach.

math.OC

From Manifold Identification to Newton Acceleration on Intersections: Sparse Stiefel Optimization

We study a Newton acceleration for sparse composite optimization on the Stiefel manifold. The main difficulty is geometric: the active manifold identified by the nonsmooth regularizer may fail to intersect the Stiefel manifold transversely, which obstructs a Riemannian Newton step on the identified manifold. In the transverse case, we prove local identification of the ManPG tangent proximal mapping. For nontransverse cases, we introduce an off-diagonally perturbed Stiefel family that generically restores the identification geometry while yielding an \(O(\|\Delta\|_F)\)-KKT guarantee for the original problem. We also derive verifiable support-level conditions for clean intersection, which cover nontransverse sparse patterns and yield the smooth moving local models used by the Newton correction. Based on these results, we propose MIX, a safeguarded ManPG/Newton-CG method on moving identified intersections. In the general clean-intersection setting, we prove global descent and KKT-residual guarantees for MIX. In the transverse or generically perturbed cases, if the sequence has an accumulation point satisfying certain regularity assumptions and the second-order sufficient condition (SOSC), then the full sequence converges to that point, with finite active-manifold identification and a local Q-superlinear rate. Numerical experiments on compressed modes and sparse PCA show that MIX substantially improves efficiency while preserving solution quality. Beyond the Stiefel manifold, we also outline how the safeguarded global-convergence mechanism of MIX extends to general smooth equality-constrained manifolds.

math.OC

A Geometry-Adaptive Regularized Newton-Type Method for Manifold-Affine Intersection Problems

We propose Regularized Newton-SLRA (RN-SLRA), a regularized Newton-type method for local manifold--affine intersection problems motivated by structured low-rank approximation. Classical Newton-SLRA achieves fast local convergence under transversality, but its tangent-space intersection step may become ill-defined, singular, or severely ill-conditioned when transversality fails. RN-SLRA overcomes this difficulty by replacing the exact tangent-space intersection step with a regularized quadratic subproblem over the affine space. Under intrinsic transversality, we prove local linear convergence to the intersection. Under transversality, we show that a residual-dependent choice of the regularization parameter yields higher-order local convergence; in particular, the method converges quadratically for the linear residual rule. We also analyze an inexact variant based on quasioptimal manifold projections. When the quasioptimality constant is sufficiently accurate, the inexact method retains local residual convergence. Numerical experiments on constructed degenerate SLRA instances and Hankel-structured examples illustrate the robustness of RN-SLRA in settings where Newton-SLRA may fail, and show that the inexact variant can reduce the projection cost in large-scale problems.

math.OC

Retractions by Alternating Projections

Alternating projections and their variants are classical tools for computing points in intersections of sets. Existing analyses for smooth manifolds mainly focus on local convergence rates under transversality or related regularity conditions. In this work, we develop a unified framework for a broad class of (possibly inexact) alternating-projection-type methods on intersections of smooth manifolds. Specifically, under the assumption that two $C^{2,1}$ embedded submanifolds $\mathcal{M}_1, \mathcal{M}_2 \subset \mathbb{R}^n$ intersect cleanly, we show that the associated alternating mapping admits a well-defined local limiting map $\psi$ on the intersection manifold $\mathcal{M}=\mathcal{M}_1\cap \mathcal{M}_2$, and that $\psi$ is a retraction on $\mathcal{M}$. If, in addition, $\mathcal{M}_1$ and $\mathcal{M}_2$ are $C^{3,1}$, then $\psi$ is a second-order retraction. Furthermore, the standard NewtonSLRA scheme, which exhibits quadratic local behavior under transversality, can be understood as inducing a second-order retraction on \(\M\). This framework thus provides new retraction-based optimization tools for problems constrained to the intersection manifold.

math.OC

Descent-Net: Learning Descent Directions for Constrained Optimization

Deep learning approaches, known for their ability to model complex relationships and fast execution, are increasingly being applied to solve large optimization problems. However, existing methods often face challenges in simultaneously ensuring feasibility and achieving an optimal objective value. To address this issue, we propose Descent-Net, a neural network designed to learn an effective descent direction from a feasible solution. By updating the solution along this learned direction, Descent-Net improves the objective value while preserving feasibility. Our method demonstrates strong performance on both synthetic optimization tasks and the real-world AC optimal power flow problem, while also exhibiting effective scalability to large problems, as shown by portfolio optimization experiments with thousands of assets.

math.OC

Distributed Stochastic Proximal Algorithm on Riemannian Submanifolds for Weakly-convex Functions

This paper aims to investigate the distributed stochastic optimization problems on compact embedded submanifolds (in the Euclidean space) where the local cost functions are weakly-convex. To address the manifold structure, we propose a distributed Riemannian stochastic proximal algorithm framework by utilizing the retraction and Riemannian consensus protocol, and analyze three specific algorithms: the distributed Riemannian stochastic subgradient, proximal point, and prox-linear algorithms. When the initial points satisfy certain conditions, we show that the iterates generated by this framework converge to a nearly stationary point in expectation while achieving consensus. We further establish the convergence rate of the algorithm framework as $\mathcal{O}(\frac{1+\kappa_g}{\sqrt{k}})$ where $k$ denotes the number of iterations and $\kappa_g$ shows the impact of manifold geometry on the algorithm performance. Finally, numerical experiments are implemented to demonstrate the theoretical results and show the empirical performance.

math.OC

ADARL: Adaptive Low-Rank Structures for Robust Policy Learning under Uncertainty

Robust reinforcement learning (Robust RL) seeks to handle epistemic uncertainty in environment dynamics, but existing approaches often rely on nested min--max optimization, which is computationally expensive and yields overly conservative policies. We propose \textbf{Adaptive Rank Representation (AdaRL)}, a bi-level optimization framework that improves robustness by aligning policy complexity with the intrinsic dimension of the task. At the lower level, AdaRL performs policy optimization under fixed-rank constraints with dynamics sampled from a Wasserstein ball around a centroid model. At the upper level, it adaptively adjusts the rank to balance the bias--variance trade-off, projecting policy parameters onto a low-rank manifold. This design avoids solving adversarial worst-case dynamics while ensuring robustness without over-parameterization. Empirical results on MuJoCo continuous control benchmarks demonstrate that AdaRL not only consistently outperforms fixed-rank baselines (e.g., SAC) and state-of-the-art robust RL methods (e.g., RNAC, Parseval), but also converges toward the intrinsic rank of the underlying tasks. These results highlight that adaptive low-rank policy representations provide an efficient and principled alternative for robust RL under model uncertainty.

cs.LG

Spin Faraday pattern formation in a circular spin-orbit coupled Bose-Einstein condensate with stripe phase

We investigate the spin Faraday pattern formation in a periodically driven, pancake-shaped spin-orbit-coupled (SOC) Bose-Einstein condensate (BEC) prepared with stripe phase. By modulating atomic interactions using in-phase and out-of-phase protocols, we observe collective excitation modes with distinct rotational symmetries (L-fold). Crucially, at the critical modulation frequency, out-of-phase modulation destabilizes the L = 6 pattern, whereas in-phase modulation not only preserves high symmetry but also excites higher-order modes. Unlike conventional binary BECs, Faraday patterns emerge here without initial noise due to SOC-induced symmetry breaking, with all patterns exhibiting supersolid characteristics. Furthermore, we demonstrate control over pattern symmetry, radial nodes, and pattern radius by tuning the modulation frequency, providing a new approach for manipulating quantum fluid dynamics. This work establishes a platform for exploring supersolidity and nonlinear excitations in SOC systems with stripe phase.

cond-mat.quant-gas

Local Linear Convergence of Infeasible Optimization with Orthogonal Constraints

Many classical and modern machine learning algorithms require solving optimization tasks under orthogonality constraints. Solving these tasks with feasible methods requires a gradient descent update followed by a retraction operation on the Stiefel manifold, which can be computationally expensive. Recently, an infeasible retraction-free approach, termed the landing algorithm, was proposed as an efficient alternative. Motivated by the common occurrence of orthogonality constraints in tasks such as principle component analysis and training of deep neural networks, this paper studies the landing algorithm and establishes a novel linear convergence rate for smooth non-convex functions using only a local Riemannian P{\L} condition. Numerical experiments demonstrate that the landing algorithm performs on par with the state-of-the-art retraction-based methods with substantially reduced computational overhead.

math.OC

Retraction-Free Decentralized Non-convex Optimization with Orthogonal Constraints

In this paper, we investigate decentralized non-convex optimization with orthogonal constraints. Conventional algorithms for this setting require either manifold retractions or other types of projection to ensure feasibility, both of which involve costly linear algebra operations (e.g., SVD or matrix inversion). On the other hand, infeasible methods are able to provide similar performance with higher computational efficiency. Inspired by this, we propose the first decentralized version of the retraction-free landing algorithm, called \textbf{D}ecentralized \textbf{R}etraction-\textbf{F}ree \textbf{G}radient \textbf{T}racking (DRFGT). We theoretically prove that DRFGT enjoys the ergodic convergence rate of $\mathcal{O}(1/K)$, matching the convergence rate of centralized, retraction-based methods. We further establish that under a local Riemannian P{\L} condition, DRFGT achieves a much faster linear convergence rate. Numerical experiments demonstrate that DRFGT performs on par with the state-of-the-art retraction-based methods with substantially reduced computational overhead.

cs.LG

FedLALR: Client-Specific Adaptive Learning Rates Achieve Linear Speedup for Non-IID Data

Federated learning is an emerging distributed machine learning method, enables a large number of clients to train a model without exchanging their local data. The time cost of communication is an essential bottleneck in federated learning, especially for training large-scale deep neural networks. Some communication-efficient federated learning methods, such as FedAvg and FedAdam, share the same learning rate across different clients. But they are not efficient when data is heterogeneous. To maximize the performance of optimization methods, the main challenge is how to adjust the learning rate without hurting the convergence. In this paper, we propose a heterogeneous local variant of AMSGrad, named FedLALR, in which each client adjusts its learning rate based on local historical gradient squares and synchronized learning rates. Theoretical analysis shows that our client-specified auto-tuned learning rate scheduling can converge and achieve linear speedup with respect to the number of clients, which enables promising scalability in federated optimization. We also empirically compare our method with several communication-efficient federated optimization methods. Extensive experimental results on Computer Vision (CV) tasks and Natural Language Processing (NLP) task show the efficacy of our proposed FedLALR method and also coincides with our theoretical findings.

cs.LG

OmniForce: On Human-Centered, Large Model Empowered and Cloud-Edge Collaborative AutoML System

Automated machine learning (AutoML) seeks to build ML models with minimal human effort. While considerable research has been conducted in the area of AutoML in general, aiming to take humans out of the loop when building artificial intelligence (AI) applications, scant literature has focused on how AutoML works well in open-environment scenarios such as the process of training and updating large models, industrial supply chains or the industrial metaverse, where people often face open-loop problems during the search process: they must continuously collect data, update data and models, satisfy the requirements of the development and deployment environment, support massive devices, modify evaluation metrics, etc. Addressing the open-environment issue with pure data-driven approaches requires considerable data, computing resources, and effort from dedicated data engineers, making current AutoML systems and platforms inefficient and computationally intractable. Human-computer interaction is a practical and feasible way to tackle the problem of open-environment AI. In this paper, we introduce OmniForce, a human-centered AutoML (HAML) system that yields both human-assisted ML and ML-assisted human techniques, to put an AutoML system into practice and build adaptive AI in open-environment scenarios. Specifically, we present OmniForce in terms of ML version management; pipeline-driven development and deployment collaborations; a flexible search strategy framework; and widely provisioned and crowdsourced application algorithms, including large models. Furthermore, the (large) models constructed by OmniForce can be automatically turned into remote services in a few minutes; this process is dubbed model as a service (MaaS). Experimental results obtained in multiple search spaces and real-world use cases demonstrate the efficacy and efficiency of OmniForce.

cs.LG

Dynamic Regularized Sharpness Aware Minimization in Federated Learning: Approaching Global Consistency and Smooth Landscape

In federated learning (FL), a cluster of local clients are chaired under the coordination of the global server and cooperatively train one model with privacy protection. Due to the multiple local updates and the isolated non-iid dataset, clients are prone to overfit into their own optima, which extremely deviates from the global objective and significantly undermines the performance. Most previous works only focus on enhancing the consistency between the local and global objectives to alleviate this prejudicial client drifts from the perspective of the optimization view, whose performance would be prominently deteriorated on the high heterogeneity. In this work, we propose a novel and general algorithm {\ttfamily FedSMOO} by jointly considering the optimization and generalization targets to efficiently improve the performance in FL. Concretely, {\ttfamily FedSMOO} adopts a dynamic regularizer to guarantee the local optima towards the global objective, which is meanwhile revised by the global Sharpness Aware Minimization (SAM) optimizer to search for the consistent flat minima. Our theoretical analysis indicates that {\ttfamily FedSMOO} achieves fast $\mathcal{O}(1/T)$ convergence rate with low generalization bound. Extensive numerical studies are conducted on the real-world dataset to verify its peerless efficiency and excellent generality.

cs.LG

Decentralized Weakly Convex Optimization Over the Stiefel Manifold

We focus on a class of non-smooth optimization problems over the Stiefel manifold in the decentralized setting, where a connected network of $n$ agents cooperatively minimize a finite-sum objective function with each component being weakly convex in the ambient Euclidean space. Such optimization problems, albeit frequently encountered in applications, are quite challenging due to their non-smoothness and non-convexity. To tackle them, we propose an iterative method called the decentralized Riemannian subgradient method (DRSM). The global convergence and an iteration complexity of $\mathcal{O}(\varepsilon^{-2} \log^2(\varepsilon^{-1}))$ for forcing a natural stationarity measure below $\varepsilon$ are established via the powerful tool of proximal smoothness from variational analysis, which could be of independent interest. Besides, we show the local linear convergence of the DRSM using geometrically diminishing stepsizes when the problem at hand further possesses a sharpness property. Numerical experiments are conducted to corroborate our theoretical findings.

math.OC

AdaSAM: Boosting Sharpness-Aware Minimization with Adaptive Learning Rate and Momentum for Training Deep Neural Networks

Sharpness aware minimization (SAM) optimizer has been extensively explored as it can generalize better for training deep neural networks via introducing extra perturbation steps to flatten the landscape of deep learning models. Integrating SAM with adaptive learning rate and momentum acceleration, dubbed AdaSAM, has already been explored empirically to train large-scale deep neural networks without theoretical guarantee due to the triple difficulties in analyzing the coupled perturbation step, adaptive learning rate and momentum step. In this paper, we try to analyze the convergence rate of AdaSAM in the stochastic non-convex setting. We theoretically show that AdaSAM admits a $\mathcal{O}(1/\sqrt{bT})$ convergence rate, which achieves linear speedup property with respect to mini-batch size $b$. Specifically, to decouple the stochastic gradient steps with the adaptive learning rate and perturbed gradient, we introduce the delayed second-order momentum term to decompose them to make them independent while taking an expectation during the analysis. Then we bound them by showing the adaptive learning rate has a limited range, which makes our analysis feasible. To the best of our knowledge, we are the first to provide the non-trivial convergence rate of SAM with an adaptive learning rate and momentum acceleration. At last, we conduct several experiments on several NLP tasks, which show that AdaSAM could achieve superior performance compared with SGD, AMSGrad, and SAM optimizers.

cs.LG

Inducing Neural Collapse in Imbalanced Learning: Do We Really Need a Learnable Classifier at the End of Deep Neural Network?

Modern deep neural networks for classification usually jointly learn a backbone for representation and a linear classifier to output the logit of each class. A recent study has shown a phenomenon called neural collapse that the within-class means of features and the classifier vectors converge to the vertices of a simplex equiangular tight frame (ETF) at the terminal phase of training on a balanced dataset. Since the ETF geometric structure maximally separates the pair-wise angles of all classes in the classifier, it is natural to raise the question, why do we spend an effort to learn a classifier when we know its optimal geometric structure? In this paper, we study the potential of learning a neural network for classification with the classifier randomly initialized as an ETF and fixed during training. Our analytical work based on the layer-peeled model indicates that the feature learning with a fixed ETF classifier naturally leads to the neural collapse state even when the dataset is imbalanced among classes. We further show that in this case the cross entropy (CE) loss is not necessary and can be replaced by a simple squared loss that shares the same global optimality but enjoys a better convergence property. Our experimental results show that our method is able to bring significant improvements with faster convergence on multiple imbalanced datasets.

cs.LG

Penalized Proximal Policy Optimization for Safe Reinforcement Learning

Safe reinforcement learning aims to learn the optimal policy while satisfying safety constraints, which is essential in real-world applications. However, current algorithms still struggle for efficient policy updates with hard constraint satisfaction. In this paper, we propose Penalized Proximal Policy Optimization (P3O), which solves the cumbersome constrained policy iteration via a single minimization of an equivalent unconstrained problem. Specifically, P3O utilizes a simple-yet-effective penalty function to eliminate cost constraints and removes the trust-region constraint by the clipped surrogate objective. We theoretically prove the exactness of the proposed method with a finite penalty factor and provide a worst-case analysis for approximate error when evaluated on sample trajectories. Moreover, we extend P3O to more challenging multi-constraint and multi-agent scenarios which are less studied in previous work. Extensive experiments show that P3O outperforms state-of-the-art algorithms with respect to both reward improvement and constraint satisfaction on a set of constrained locomotive tasks.

cs.LG

Manifold Proximal Point Algorithms for Dual Principal Component Pursuit and Orthogonal Dictionary Learning

We consider the problem of maximizing the $\ell_1$ norm of a linear map over the sphere, which arises in various machine learning applications such as orthogonal dictionary learning (ODL) and robust subspace recovery (RSR). The problem is numerically challenging due to its nonsmooth objective and nonconvex constraint, and its algorithmic aspects have not been well explored. In this paper, we show how the manifold structure of the sphere can be exploited to design fast algorithms for tackling this problem. Specifically, our contribution is threefold. First, we present a manifold proximal point algorithm (ManPPA) for the problem and show that it converges at a sublinear rate. Furthermore, we show that ManPPA can achieve a quadratic convergence rate when applied to the ODL and RSR problems. Second, we propose a stochastic variant of ManPPA called StManPPA, which is well suited for large-scale computation, and establish its sublinear convergence rate. Both ManPPA and StManPPA have provably faster convergence rates than existing subgradient-type methods. Third, using ManPPA as a building block, we propose a new approach to solving a matrix analog of the problem, in which the sphere is replaced by the Stiefel manifold. The results from our extensive numerical experiments on the ODL and RSR problems demonstrate the efficiency and efficacy of our proposed methods.

math.OC