SearcharxivSearch

arXiv subjects

Morteza Kimiaei

Publications and source records attributed to Morteza Kimiaei.

11 recordsLinked to original sources

Subspace methods for min-max problems

This paper introduces four groups of subspace methods for nonlinear monotone equations, with applications to large-scale machine learning problems. The methods use Jacobian-free subspace ({\tt JFS}) directions of conjugate-gradient type, combined with either fixed step sizes or variable step sizes generated by the projected method of Solodov and Svaiter. To ensure convergence independently of the specific algebraic form of the subspace directions, we impose an angle condition together with an explicit scaling rule controlling the effective search directions. Under monotonicity and Lipschitz continuity of the operator, we establish global convergence for both the line-search and fixed-step frameworks, as well as a best-iterate residual rate $O(\ell^{-1/2})$. Under a local error bound, the distance to the solution set satisfies the sharper decay $o(\ell^{-1/2})$. If the operator is continuously differentiable and its Jacobian is nonsingular at a solution, the required local error bound and local isolation follow, yielding $R$-linear local convergence. The residual sequence then converges geometrically and hence satisfies the last-iterate rate $o(\ell^{-1})$, without strong monotonicity. We also derive iteration and residual-evaluation complexity bounds: $O(\varepsilon^{-2})$ for the baseline best-iterate guarantee and $O(\log(\varepsilon^{-1}))$ in the local linear regime, together with a uniform bound on backtracking residual evaluations. Under additional asymptotic assumptions, the proposed {\tt JFS} directions and several classical update directions admit related optimistic gradient descent--ascent (\texttt{OGDA})-type residual--memory representations. Numerical experiments illustrate the robustness and efficiency of the methods on representative min--max problems.

math.OC

From a Scalar to a Matrix Setting for the Dai--Liao Parameter

As is well known, both the numerical performance and the theoretical properties of the Dai--Liao conjugate gradient algorithm are highly dependent on the adjustment of its key parameter. Here, we first employ the well-known Harmonic--Geometric--Arithmetic--Quadratic mean inequality to reinforce the optimality of two previously proposed scalar adaptive settings of the Dai--Liao parameter. In other words, we show how these parameter choices are capable of enhancing the well-conditioning of the Dai--Liao search direction matrix by shrinking the intervals containing its singular values. A similar analysis is also carried out for scaled memoryless quasi--Newton updating formulas in order to further justify the optimality of two classical scaling parameters associated with these updates. Then, as the main contribution of this work, we move from the classical scalar setting of the Dai--Liao parameter to a matrix setting aimed at enhancing flexibility and diversity within optimization methods. In particular, we show that, under such a matrix formulation of the parameter, a well-known three-term conjugate gradient algorithm emerges as a member of the proposed extended Dai--Liao class of algorithms while enjoying several computationally attractive properties. Among these, the well-conditioning of the associated search direction matrix is especially noteworthy, as well as the ability to make more explicit use of the second-order information of the model. Finally, to provide practical evidence supporting the proposed matrix setting of the Dai--Liao parameter, we conduct a series of numerical experiments on standard benchmark test problems and report the results in detail. Generally speaking, the proposed framework is shown to retain both theoretical soundness and computational reliability.

math.OC

A Direction Adaptation Evaluation Strategy for Noisy Derivative-Free Optimization

In this paper, we develop a direction adaptation evolution strategy (DAES) -- a new MAES-type method -- for noisy derivative-free optimization, designed to reconcile the population-based search mechanisms of evolution strategies with rigorous complexity analysis. Unlike standard MAES schemes, DAES fixes the adaptation matrix to the identity and replaces matrix adaptation with a structured direction-generation mechanism based on symmetric sampling, joint sorting-selection of noisy function values, three-group recombination, and a new triangular search direction. Specifically, candidates are sampled along paired positive and negative directions, their inexact function values are jointly ranked, and the reordered directions are partitioned into three groups to construct three recombination points whose geometry defines the triangular direction. A signed sufficient-decrease search and extrapolation mechanism is then applied along this direction. This structure yields a population-based MAES-type algorithm that retains competitive practical behavior while being amenable to nonasymptotic analysis under noisy evaluations. We establish high-probability complexity bounds for nonconvex, convex, and strongly convex objective functions and derive corresponding guarantees at the noise-limited accuracy level. To the best of our knowledge, these results provide the first high-probability complexity guarantees for a noisy MAES-type derivative-free method with this direction-adaptation structure. Finally, numerical experiments on the 655 prince test problems from the BARON collection compare DAES with the advanced MAES-type solver MADFO and show a favorable trade-off between evaluation efficiency and ultimate robustness.

math.OC

An Approximate Conjugate Subgradient Algorithm with Matrix Parameter for Derivative-Free Nonsmooth Optimization Problems

We propose a derivative-free matrix conjugate-subgradient method for unconstrained nonsmooth optimization of locally Lipschitz functions. The method constructs discrete gradients using only function values and forms a finite sampled model of the Goldstein subdifferential. A minimal-norm element of the convex hull of the sampled discrete gradients is then computed and used both as a stationarity measure and as the reference vector for generating descent-oriented directions. To improve robustness beyond the basic steepest-descent direction, we introduce a matrix memory correction together with coefficient damping, diagonal scaling, bounded-angle correction, and matrix-stability safeguards. A two-point line-search procedure with enrichment is used to obtain either a serious step or an improved local model. Under suitable consistency assumptions on the discrete-gradient approximation and line-search sampling, the method generates directions satisfying a safeguarded descent property and computes approximate Goldstein stationary points. Numerical experiments on nonsmooth test problems with dimensions up to 1000 show that both proposed variants are robust for lower and medium accuracy requirements, while the matrix conjugate-subgradient variant remains the most reliable under the strictest tolerance.

math.OC

Reservoir Zero-Coordinatewise Projected Subspace Search for Minimization Over Sparse Symmetric Sets in Machine Learning

We study a class of nonconvex cardinality-constrained optimization problems arising in sparse learning. These problems are NP-hard due to the combinatorial nature of sparsity constraints. We introduce a Reservoir Zero-Coordinatewise Projected Subspace Search (RZCW-PSS) algorithm, a simplex-style method on sparse manifolds that integrates coordinatewise search, symmetry-aware swap-based support updates, randomized low-dimensional subspace exploration, and zero-coordinatewise reservoir injection. The proposed method augments classical coordinate and swap moves with sparse-compatible subspace searches constructed from a dynamically maintained reservoir of previously accepted feasible points. A key feature of the approach is a refined reservoir initialization strategy that embeds sparse projection directly into a uniform sampling procedure, preserving geometric diversity within the feasible set. The algorithm also includes an optional support-identification safeguard that enforces full-support stabilization under a fixed support-change decrease threshold. We establish that, under the stated regularity, sampling, and subproblem-accuracy assumptions, every full-support accumulation point of the RZCW-PSS iterates is Beck--Hallak zero-coordinatewise stationary almost surely; with the safeguard and full-support initialization, this conclusion applies to all accumulation points. We further prove a conditional local linear convergence rate after support stabilization and derive the corresponding logarithmic local iteration complexity. Numerical experiments on synthetic sparse learning problems demonstrate that RZCW-PSS improves robustness and solution quality while remaining computationally competitive with Partial Simplex Search, Basic Feasible Search, and Zero-Coordinatewise Search methods.

math.OC

Diagonal Hessian Approximation Based on Conjugacy Condition for Noisy Derivative-Free Optimization Problems in High Dimensions

We consider large-scale noisy derivative-free optimization (DFO) problems in which only function values are available and gradient or subgradient information cannot be reliably estimated. Matrix-adaptation evolution strategies (MAES) and their limited-memory variants are among the most robust DFO methods under noise; however, their performance may deteriorate when the noise level is large. In such regimes, sorting and selection may misidentify informative sampled points, making the recombination step less reliable and weakening the scaling information used by affine or matrix-adaptation mechanisms. This can substantially reduce the efficiency of MAES-type methods, especially in high-dimensional settings. To address this limitation, we propose a DFO method that replaces the full affine-scaling matrix with a diagonal approximation constructed from conjugacy-type conditions. The proposed mechanism does not attempt to estimate gradients, subgradients, or interpolation models, nor does it learn dense covariance information from noisy rankings. Instead, it uses consecutive normalized recombination displacements in a conservative diagonal update, thereby limiting the influence of unreliable selection information while preserving the derivative-free structure of the underlying evolutionary framework. As a result, the method is computationally cheaper than full matrix-adaptation schemes and limited-memory affine-scaling variants, while providing a stable scaling mechanism in noisy environments. Numerical experiments on noisy benchmark problems show that the proposed method is competitive with, and often more efficient than, MAES-type baselines, particularly when the noise level is large and ranking-based selection becomes unreliable.

math.OC

Towards Real Time Control of Water Engineering with Nonlinear Hyperbolic Partial Differential Equations

This paper examines aspirational requirements for software addressing mixed-integer optimization problems constrained by the nonlinear Shallow Water partial differential equations (PDEs), motivated by applications such as river-flow management in hydropower cascades. Realistic deployment of such software would require the simultaneous treatment of nonlinear and potentially non-smooth PDE dynamics, limited theoretical guarantees on the existence and regularity of control-to-state mappings under varying boundary conditions, and computational performance compatible with operational decision-making. In addition, practical settings motivate consideration of uncertainty arising from forecasts of demand, inflows, and environmental conditions. At present, the theoretical foundations, numerical optimization methods, and large-scale scientific computing tools required to address these challenges in a unified and tractable manner remain the subject of ongoing research across the associated research communities. Rather than proposing a complete solution, this work uses the problem as a case study to identify and organize the mathematical, algorithmic, and computational components that would be necessary for its realization. The resulting framework highlights open challenges and intermediate research directions, and may inform both more circumscribed related problems and the design of future large-scale collaborative efforts aimed at addressing such objectives.

math.OC

An efficient penalty decomposition algorithm for minimization over sparse symmetric sets

This paper proposes an improved quasi-Newton penalty decomposition algorithm for the minimization of continuously differentiable functions, possibly nonconvex, over sparse symmetric sets. The method solves a sequence of penalty subproblems approximately via a two-block decomposition scheme: the first subproblem admits a closed-form solution without sparsity constraints, while the second subproblem is handled through an efficient sparse projection over the symmetric feasible set. Under a new assumption on the gradient of the objective function, weaker than global Lipschitz continuity from the origin, we establish that accumulation points of the outer iterates are basic feasible and cardinality-constrained Mordukhovich stationarity points. To ensure robustness and efficiency in finite-precision arithmetic, the algorithm incorporates several practical enhancements, including an enhanced line search strategy based on either backtracking or extrapolation, and four inexpensive diagonal Hessian approximations derived from differences of previous iterates and gradients or from eigenvalue-distribution information. Numerical experiments on a diverse benchmark of $30$ synthetic and data-driven test problems, including machine-learning datasets from the UCI repository and sparse symmetric instances with dimensions ranging from $10$ to $500$, demonstrate that the proposed algorithm is competitive with several state-of-the-art methods in terms of efficiency, robustness, and strong stationarity.

math.OC

Machine Learning Algorithms for Improving Exact Classical Solvers in Mixed Integer Continuous Optimization

Integer and mixed-integer nonlinear programming (INLP, MINLP) are central to logistics, energy, and scheduling, but remain computationally challenging. This survey examines how machine learning and reinforcement learning can enhance exact optimization methods-particularly branch-and-bound (BB)-without compromising global optimality. We cover discrete, continuous, and mixed-integer formulations, and highlight applications such as vehicle routing, hydropower planning, and crew scheduling. We introduce a unified BB framework that embeds learning-based strategies into branching, cut selection, node ordering, and parameter control. Classical algorithms are augmented using supervised, imitation, and reinforcement learning models to accelerate convergence while maintaining correctness. We conclude with a taxonomy of learning methods by solver class and learning paradigm, and outline open challenges in generalization, hybridization, and scaling intelligent solvers.

math.OC

Machine Learning Algorithms for Improving Black Box Optimization Solvers

Black-box optimization (BBO) addresses problems where objectives are accessible only through costly queries without gradients or explicit structure. Classical derivative-free methods -- line search, direct search, and model-based solvers such as Bayesian optimization -- form the backbone of BBO, yet often struggle in high-dimensional, noisy, or mixed-integer settings. Recent advances use machine learning (ML) and reinforcement learning (RL) to enhance BBO: ML provides expressive surrogates, adaptive updates, meta-learning portfolios, and generative models, while RL enables dynamic operator configuration, robustness, and meta-optimization across tasks. This paper surveys these developments, covering representative algorithms such as NNs with the modular model-based optimization framework (mlrMBO), zeroth-order adaptive momentum methods (ZO-AdaMM), automated BBO (ABBO), distributed block-wise optimization (DiBB), partition-based Bayesian optimization (SPBOpt), the transformer-based optimizer (B2Opt), diffusion-model-based BBO, surrogate-assisted RL for differential evolution (Surr-RLDE), robust BBO (RBO), coordinate-ascent model-based optimization with relative entropy (CAS-MORE), log-barrier stochastic gradient descent (LB-SGD), policy improvement with black-box (PIBB), and offline Q-learning with Mamba backbones (Q-Mamba). We also review benchmark efforts such as the NeurIPS 2020 BBO Challenge and the MetaBox framework. Overall, we highlight how ML and RL transform classical inexact solvers into more scalable, robust, and adaptive frameworks for real-world optimization.

cs.LG

Worst case complexity bounds for linesearch-type derivative-free algorithms

This paper is devoted to the analysis of worst case complexity bounds for linesearch-type derivative-free algorithms for the minimization of general non-convex smooth functions. We prove that two linesearch-type algorithms enjoy the same complexity properties which have been proved for pattern and direct search algorithms. In particular, we consider two derivative-free algorithms based on two different linesearch techniques and manage to prove that the number of iterations and of function evaluations required to drive the norm of the gradient of the objective function below a given threshold $ε$ is ${\cal O}(ε^{-2})$ in the worst case.

math.OC