SearcharxivSearch

arXiv subjects

Lara Zebiane

Publications and source records attributed to Lara Zebiane.

4 recordsLinked to original sources

A Gradient Sampling Algorithm for Noisy Nonsmooth Nonconvex Optimization

An algorithm is proposed, analyzed, and tested for minimizing locally Lipschitz objective functions that may be nonconvex and/or nonsmooth. The algorithm, which is built upon the gradient-sampling methodology, is designed specifically for cases when objective and generalized gradient values might be subject to bounded, uncontrollable errors. A novel assumption pertaining to such errors in nonsmooth settings is introduced that decouples the noise that may be present due to typical errors in derivative estimates with noise that may be due to misidentification of segments of differentiability of the objective. Similarly to state-of-the-art guarantees for noisy smooth optimization, it is proved for the algorithm that, with probability one, either the sequence of objective values will decrease without bound or the algorithm will generate an iterate at which a measure of stationarity is below a threshold that depends proportionally on the error bounds for the objective function and generalized gradient values. The results of numerical experiments are presented, which show that the algorithm can indeed perform approximate optimization robustly despite errors in objective and generalized gradient values.

math.OC

Low-Order Explicit Hessian Imitation Method for Large-Scale Supervised Machine Learning

An algorithm is proposed for solving optimization problems arising in neural network training for supervised learning. The unique feature of the algorithm is the use of an auxiliary loss, in addition to the original loss employed for model training. The purpose of the auxiliary loss is to provide a mechanism for creating a low-order Hessian-type approximation for the original loss. The proposed algorithm employs the resulting low-order second-derivative approximation terms in place of the second-order momentum terms (i.e., squared elements of the gradient of the loss function) in an overall scheme that has computational cost on par with an Adam-type approach. Whereas the squared elements of a gradient vector do not necessarily approximate second-order derivatives well, by careful construction of the auxiliary loss, second-order derivative-type approximations for the original loss can be computed and employed by the algorithm in an efficient manner. A convergence guarantee is provided for the proposed algorithm that is on par with guarantees available for similar stochastic diagonal-scaling methods. The results of numerical experiments show situations when the proposed algorithm outperforms Adam and other popular modern optimizers.

math.OC

Active-Set Identification in Noisy and Stochastic Optimization

Identifying active constraints from a point near an optimal solution is important both theoretically and practically in constrained continuous optimization, as it can help identify optimal Lagrange multipliers and essentially reduces an inequality-constrained problem to an equality-constrained one. Traditional active-set identification guarantees have been proved under assumptions of smoothness and constraint qualifications, and assume exact function and derivative values. This work extends these results to settings when both objective and constraint function and derivative values have deterministic or stochastic noise. Two strategies are proposed that, under mild conditions, are proved to identify the active set of a local minimizer correctly when a point is close enough to the local minimizer and the noise is sufficiently small. Guarantees are also stated for the use of active-set identification strategies within a stochastic algorithm. We demonstrate our findings with two simple illustrative examples and a more realistic constrained neural-network training task.

math.OC

NonOpt: Nonconvex, Nonsmooth Optimizer

NonOpt, a C++ software package for minimizing locally Lipschitz objective functions, is presented. The software is intended primarily for minimizing objective functions that are nonconvex and/or nonsmooth. The package has implementations of two main algorithmic strategies: a gradient-sampling and a proximal-bundle method. Each algorithmic strategy can employ quasi-Newton techniques for accelerating convergence in practice. The main computational cost in each iteration is solving a subproblem with a quadratic objective function, a linear equality constraint, and bound constraints. The software contains dual active-set and interior-point subproblem solvers that are designed specifically for solving these subproblems efficiently. The results of numerical experiments with various test problems are provided to demonstrate the speed and reliability of the software.

math.OC