SearcharxivSearch

arXiv subjects

Masoud Ahookhosh

Publications and source records attributed to Masoud Ahookhosh.

At least 19 recordsLinked to original sources

Beyond Conventional Federated Learning via High-Order Regularization

Federated clients that perform several local optimization steps can return parameter displacements with widely different magnitudes. The quadratic regularization of FedProx grows linearly with displacement and therefore offers limited control over the contrast between ordinary and unusually large client movements. We here introduce HiFedProx, which replaces the quadratic penalty with a scale-matched power-type regularizer indexed by $p\geq2$. All powers have the same regularization-gradient magnitude at a reference displacement $R$, while every $p>2$ gives a weaker response below $R$ and a stronger response above it. An exact affine reference calculation shows that increasing $p$ compresses relative displacement disparities, although very large powers approach fixed-radius behavior and increase local curvature. HiFedProx combines this geometry with finite-budget stochastic client optimization and same-minibatch Armijo backtracking. In paired five-seed experiments on a frozen 60-writer FEMNIST subset, a common-parameter study over $p\in\{2,3,4,5,6,7,8\}$ shows similar clean-training performance but substantial gains under composite stress. The lowest moderate- and severe-stress losses occur at $p=7$ and $p=6$, improving over $p=2$ by $11.44\%$ and $23.16\%$, respectively. Although displacement-tail ratios continue to decrease through $p=8$, predictive performance peaks in an intermediate range and Armijo trial cost increases with $p$. These results indicate that the exponent should be calibrated rather than maximized. In our experiments, $p=5$--$7$ provides the most useful range.

cs.LG

Difference-of-Convex Optimization via Inexact Smoothing Descent Methods: Difference of High-Order Moreau Envelopes

This paper studies difference-of-convex (DC) optimization problems through smoothing descent techniques. In particular, we introduce the difference of high-order Moreau envelopes (HOME-DC) and establish its fundamental and differential properties. Approximating the underlying proximal points, we generate an inexact first-order oracle for HOME-DC and characterize its accuracy guarantees. Building upon this oracle, we propose a class of inexact descent methods for minimizing DC functions and provide a convergence analysis. The proposed framework extends the applicability of envelope-based optimization techniques to a broad class of structured nonconvex problems while accommodating inexact solutions to subproblems. Preliminary numerical experiments on a sparse clustering problem demonstrate the approach's practical potential and support the theoretical findings.

math.OC

Relative Weak Convexity and Projected Subgradient Methods: Analysis and Convergence

We introduce the class of relatively weakly convex functions, which extends the classical notion of weak convexity by measuring nonconvexity relative to a distance-generating function. We investigate the fundamental properties of this function class, establishing characterization results, calculus rules, and illustrative examples. We further analyze the associated optimization landscape and identify a neighborhood of the set of global minimizers that is free of saddle points. Motivated by this geometric structure, we propose the Projected SubGradient Algorithm (PSGA) with several step-size strategies. Under a sharpness error bound, we prove that, when initialized within this saddle-point-free neighborhood, the iterates generated by PSGA converge to a global minimizer for each of the proposed step-size strategies. Furthermore, linear convergence is established for the geometrically decaying step-size strategy.

math.OC

Speeding Up Nonsmooth Bayesian MCMC Sampling via Inexact Proximal Unadjusted Langevin Algorithm

We study sampling from posterior distributions with nonsmooth composite potentials, a setting in which proximal-based Langevin methods are theoretically appealing but in practice limited to simple functions with closed-form proximal operators. We introduce iPULA for composite potentials, an inexact proximal unadjusted Langevin algorithm that replaces exact proximal steps with controlled approximations. Our approach leverages the Moreau envelope to smooth the potential, while allowing inexact evaluation of its gradient through inexact proximal computations. We establish non-asymptotic convergence guarantees for iPULA, explicitly characterizing the impact of inexactness on the sampling error and showing that the inexactness preserves convergence rates up to a quantifiable bias. We demonstrate the practical relevance of iPULA on a medical image reconstruction task, where proximal operators cannot be computed exactly. Experiments demonstrate the effectiveness of iPULA and support our theoretical results.

math.OC

Robust Learning Meets Quasar-Convex Optimization: Inexact High-Order Proximal-Point Methods

Robust learning aims to maintain model performance under noise, corruption, and distributional shifts, which are prevalent in modern machine learning applications. This work shows that examples of robust learning problems can be formulated as (strongly) quasar-convex optimization problems, which admit a benign landscape with no saddle points. We then propose HiPPA, an inexact high-order proximal-point method that employs a model-value gap to control the inexactness of subproblem solutions. Notably, we prove global convergence of HiPPA to global minima and establish that it attains a (local) linear or superlinear convergence rate, depending on the regularization order and inexactness control. Our numerical experiments on robust feature-alignment distillation indicate strong empirical performance of HiPPA and results consistent with our theoretical findings.

math.OC

Quasar-Convex Optimization: Fundamental Properties and High-Order Proximal-Point Methods

We study the optimization of (strongly) quasar-convex functions, a class that arises naturally in many machine learning and data science applications due to its favorable properties. The fundamental properties of this class are first developed, including its stability under standard calculus operations, growth conditions, and the absence of spurious critical points, which together imply a benign global geometry with no saddle points. Motivated by these properties, a class of proximal-point algorithms (HiPPA) with high-order regularization of order $p>1$ is introduced. Conditions are identified under which the iterates converge globally to minimizers, and a unified convergence analysis is provided with explicit rates and iteration complexity bounds under appropriate regularity assumptions. The results reveal a sharp transition in behavior with respect to the order $p$: for $p\in(1,2)$, the method achieves local linear convergence with complexity $\mathcal{O}(\log(\varepsilon^{-1}))$ when initialized sufficiently close to a minimizer; for $p=2$, it attains global linear convergence with the same complexity; and for $p>2$, it exhibits superlinear convergence with complexity $\mathcal{O}(\log\log(\varepsilon^{-1}))$, where $\varepsilon>0$ denotes the target accuracy. The theory is complemented with preliminary numerical experiments on selected machine learning problems, which illustrate the effectiveness of the proposed methods and are consistent with the theoretical findings.

math.OC

Projected subgradient methods for paraconvex optimization: Application to robust low-rank matrix recovery

This paper is devoted to the class of paraconvex functions and presents some of its fundamental properties, characterization, and examples that can be used for their recognition and optimization. Next, the convergence analysis of the projected subgradient methods with several step-sizes (i.e., constant, nonsummable, square-summable but not summable, geometrically decaying, and Scaled Polyak's step-sizes) to global minima for this class of functions is studied. In particular, the convergence rate of the proposed methods is investigated under paraconvexity and the Hölderian error bound condition, where the latter is an extension of the classical error bound condition. The preliminary numerical experiments on several robust low-rank matrix recovery problems (i.e., robust matrix completion, image inpainting, robust nonnegative matrix factorization, robust matrix compression, and robust image deblurring) indicate promising behavior for these projected subgradient methods, validating our theoretical foundations.

math.OC

Moreau envelope and proximal-point methods under the lens of high-order regularization

This paper is devoted to investigating the fundamental properties of the high-order proximal operator (HOPE) and the high-order Moreau envelope (HOME) in the nonconvex setting, where the quadratic regularization ($p=2$) is replaced by a $p$-order regularizer with $p > 1$. After establishing several basic properties of HOPE and HOME, we study the differentiability and weak smoothness of HOME under $q$-prox-regularity with $q \geq 2$ and $p$-calmness for $p \in (1,2]$ and $2 \leq p \leq q$. Furthermore, we propose a high-order proximal-point algorithm (HiPPA) and analyze the convergence of the generated sequence to proximal fixed points. Our results pave the way for the development of a high-order smoothing theory with $p>1$ that can lead to new algorithmic advances in the nonconvex setting. To illustrate this potential for nonsmooth and nonconvex optimization, we apply HiPPA to the Nesterov-Chebyshev-Rosenbrock functions.

math.OC

Minimizing Smooth Kurdyka-{\L}ojasiewicz Functions via Generalized Descent Methods: Convergence Rate and Complexity

This paper introduces a generalized descent algorithm (DEAL) for minimizing smooth nonconvex functions. If the objective function is nonsmooth, a smoothing technique (e.g., forward-backward and high-order Moreau envelopes) is applied to generate a smooth counterpart. The proposed framework unifies several methods, such as gradient-based methods with constant step-sizes and Armijo line search, and several proximal splitting methods. The method is built around a generalized descent inequality that adapts the amount of decrease to the geometry of the objective function. Under the Kurdyka-{\L}ojasiewicz (KL) property, we establish global convergence of the generated sequence to critical points and provide a unified convergence rate analysis. In particular, we show that the convergence behavior depends jointly on the KL exponent and the descent order, and we identify a precise condition under which generalized descent methods achieve linear convergence. By choosing the order of high-order proximal regularization according to the KL exponent, our boosted high-order proximal-point method achieves linear convergence for arbitrary KL exponents. If the objective function satisfies a global KL inequality, we further strengthen the results by proving convergence to global minimizers and deriving explicit iteration-complexity bounds. Numerical experiments validate our theoretical foundation.

math.OC

On fundamental properties of high-order forward-backward envelope

This paper studies the fundamental properties of the high-order forward-backward splitting mapping (HiFBS) and its associated high-order forward-backward envelope (HiFBE) through the lens of high-order regularization for nonconvex composite functions. Specifically, we (i) establish the boundedness and uniform boundedness of HiFBS, along with the H\"older and Lipschitz continuity of HiFBE; (ii) derive an explicit form for the subdifferentials of HiFBE; and (iii) investigate necessary and sufficient conditions for the differentiability and weak smoothness of HiFBE under suitable assumptions. By leveraging the prox-regularity of $g$ and the concept of $p$-calmness, we further demonstrate the local single-valuedness and continuity of HiFBS, which in turn guarantee the differentiability of HiFBE in neighborhoods of calm points. This paves the way for the development of gradient-based algorithms tailored to nonconvex composite optimization problems.

math.OC

(Adaptive) Scaled gradient methods beyond locally Holder smoothness: Lyapunov analysis, convergence rate and complexity

This paper addresses the unconstrained minimization of smooth convex functions whose gradients are locally Holder continuous. Building on these results, we analyze the Scaled Gradient Algorithm (SGA) under local smoothness assumptions, proving its global convergence and iteration complexity. Furthermore, under local strong convexity and the Kurdyka-Lojasiewicz (KL) inequality, we establish linear convergence rates and provide explicit complexity bounds. In particular, we show that when the gradient is locally Lipschitz continuous, SGA attains linear convergence for any KL exponent. We then introduce and analyze an adaptive variant of SGA (AdaSGA), which automatically adjusts the scaling and step-size parameters. For this method, we show global convergence, and derive local linear rates under strong convexity.

math.OC

First-order majorization-minimization meets high-order majorant: Boosted inexact high-order forward-backward method

This paper introduces a first-order majorization-minimization framework based on a high-order majorant for continuous functions, incorporating a non-quadratic regularization term of degree $p>1$. Notably, it is shown to be valid if and only if the function is $p$-paraconcave, thus extending beyond Lipschitz and Hölder gradient continuity for $p \in (1,2]$, and implying concavity for $p>2$. In the smooth setting, this majorant recovers a variant of the classical descent lemma with quadratic regularization. Building on this foundation, we develop a high-order inexact forward-backward algorithm (HiFBA) and its line-search-accelerated variant, named Boosted HiFBA. For convergence analysis, we introduce a high-order forward-backward envelope (HiFBE), which serves as a Lyapunov function. We establish subsequential convergence under suitable inexactness conditions, and we prove global convergence with linear rates for functions satisfying the Kurdyka-Łojasiewicz inequality. Our preliminary experiments on linear inverse problems and regularized nonnegative matrix factorization highlight the efficiency of HiFBA and its boosted variant, demonstrating their potential for solving challenging nonconvex optimization problems.

math.OC

Inexact Levenberg-Marquardt methods under Hölder metric subregularity

This paper investigates two inexact Levenberg-Marquardt (LM) methods for solving systems of nonlinear equations. Both approaches compute approximate search directions by solving the LM linear system inexactly, subject to specific residual-based conditions. The first method uses an adaptive scheme to update the LM parameter, and we establish its local superlinear convergence under Hölder metric subregularity and local Hölder continuity of the gradient. The second method combines an inexact LM step with a nonmonotone quadratic regularization strategy. For this variant, we prove global convergence under the assumption of Lipschitz continuous gradients and derive a worst-case global complexity bound, showing that an approximate stationary point can be found in $\mathcal{O}(ε^{-2})$ function and gradient evaluations. Finally, we justify the use of the LSQR algorithm for efficiently solving the linear systems involved, which is used in our numerical experiment on several nonlinear systems, including those appearing in real-world biochemical reaction networks, monotone and nonlinear equations, and image deblurring problems.

math.OC

Second-order methods for provably escaping strict saddle points in composite nonconvex and nonsmooth optimization

This study introduces two second-order methods designed to provably avoid saddle points in composite nonconvex optimization problems: (i) a nonsmooth trust-region method and (ii) a curvilinear linesearch method. These developments are grounded in the forward-backward envelope (FBE), for which we analyze the local second-order differentiability around critical points and establish a novel equivalence between its second-order stationary points and those of the original objective. We show that the proposed algorithms converge to second-order stationary points of the FBE under a mild local smoothness condition on the proximal mapping of the nonsmooth term. Notably, for \( \C^2 \)-partly smooth functions, this condition holds under a standard strict complementarity assumption. To the best of our knowledge, these are the first second-order algorithms that provably escape nonsmooth strict saddle points of composite nonconvex optimization, regardless of the initialization. Our preliminary numerical experiments show promising performance of the developed methods, validating our theoretical foundations.

math.OC

Asymptotic Convergence Analysis of High-Order Proximal-Point Methods Beyond Sublinear Rates

This paper investigates the asymptotic convergence behavior of the high-order proximal-point algorithm (HiPPA) to global minimizers, extending existing analyses beyond sublinear convergence rates and complexity analysis. Specifically, we study the proximal operator of a proper lower semicontinuous function augmented with a $p$th-order regularization for $p>1$, and establish the convergence of HiPPA to a global minimizer with a particular focus on its convergence rate. To this end, we focus on minimizing functions in the class of uniformly quasiconvex functions, which includes strongly convex, uniformly convex, and strongly quasiconvex functions as special cases. Our analysis reveals the following convergence behaviors of HiPPA when the uniform quasiconvexity modulus $\phi$ admits a power function of degree $q$ as a lower bound, i.e., $\phi(t) \geq c t^q$ for some $c>0$, on an interval $\mathcal{I}$: (i) for $q\in (1,2)$ and $\mathcal{I}=[0,1)$, HiPPA exhibits a local linear rate for $p\in [q,2)$; (ii) HiPPA converges linearly when $p=2$, $q=2$, and also when $p=q>2$, provided that $\mathcal{I}=[0,\infty)$; (iii) for $q\geq 2$ and $\mathcal{I}=[0,\infty)$, HiPPA achieves a superlinear rate for $p>q$. Notably, to our knowledge, some of these results are novel, even in the context of strongly or uniformly convex functions, offering new insights into optimizing generalized convex problems.

math.OC

ItsDEAL: Inexact two-level smoothing descent algorithms for weakly convex optimization

This paper deals with nonconvex optimization problems via a two-level smoothing framework in which the high-order Moreau envelope (HOME) is applied to generate a smooth approximation of weakly convex cost functions. As such, the differentiability and weak smoothness of HOME are further studied, as is necessary for developing inexact first-order methods for finding its critical points. Building on the concept of the inexact two-level smoothing optimization (ItsOPT), the proposed scheme offers a versatile setting, called Inexact two-level smoothing DEscent ALgorithm (ItsDEAL), for developing inexact first-order methods: (i) solving the proximal subproblem approximately to provide an inexact first-order oracle of HOME at the lower-level; (ii) developing an upper inexact first-order method at the upper-level. In particular, parameter-free inexact descent methods (i.e., dynamic step-sizes and an inexact nonmonotone Armijo line search) are studied that effectively leverage the weak smooth property of HOME. Although the subsequential convergence of these methods is investigated under some mild inexactness assumptions, the global convergence and the linear rates are studied under the extra Kurdyka-Łojasiewicz (KL) property. In order to validate the theoretical foundation, preliminary numerical experiments for robust sparse recovery problems are provided which reveal a promising behavior of the proposed methods.

math.OC

ItsOPT: An inexact two-level smoothing framework for nonconvex optimization via high-order Moreau envelope

This paper introduces ItsOPT, an {\it inexact two-level smoothing optimization framework} designed to find first-order critical points of nonsmooth and nonconvex functions. The framework consists of two levels of methodologies: at the upper level, a zeroth-, first-, or second-order method can be tailored to minimize a smooth approximation; at the lower level, the high-order proximal auxiliary problems are solved inexactly, generating an inexact oracle for the smooth function. As a smoothing technique, we introduce the high-order Moreau envelope (HOME) and study its fundamental properties under standard assumptions. Next, by combining a boosted high-order proximal-point algorithm (Boosted HiPPA) at the upper level with the inexact oracle from the lower level, we obtain a zeroth-order instance of ItsOPT. Global convergence rates are established under the Kurdyka-{\L}ojasiewicz (KL) property of the cost and envelope functions, together with reasonable conditions on the accuracy of the proximal terms. Surprisingly, for any KL exponent $\theta\in (0,1)$ of the original cost, setting the regularization order $p=\frac{1}{1-\theta}$ ensures that Boosted HiPPA converges linearly to a proximal fixed point. This is the first algorithm with this property for KL functions. Preliminary numerical experiments on a robust low-rank matrix recovery problem demonstrate the promising performance of the proposed algorithm, supporting our theoretical foundations.

math.OC

Matrix Completion via Nonsmooth Regularization of Fully Connected Neural Networks

Conventional matrix completion methods approximate the missing values by assuming the matrix to be low-rank, which leads to a linear approximation of missing values. It has been shown that enhanced performance could be attained by using nonlinear estimators such as deep neural networks. Deep fully connected neural networks (FCNNs), one of the most suitable architectures for matrix completion, suffer from over-fitting due to their high capacity, which leads to low generalizability. In this paper, we control over-fitting by regularizing the FCNN model in terms of the $\ell_{1}$ norm of intermediate representations and nuclear norm of weight matrices. As such, the resulting regularized objective function becomes nonsmooth and nonconvex, i.e., existing gradient-based methods cannot be applied to our model. We propose a variant of the proximal gradient method and investigate its convergence to a critical point. In the initial epochs of FCNN training, the regularization terms are ignored, and through epochs, the effect of that increases. The gradual addition of nonsmooth regularization terms is the main reason for the better performance of the deep neural network with nonsmooth regularization terms (DNN-NSR) algorithm. Our simulations indicate the superiority of the proposed algorithm in comparison with existing linear and nonlinear algorithms.

cs.IT