SearcharxivSearch

arXiv subjects

Mirai Tanaka

Publications and source records attributed to Mirai Tanaka.

11 recordsLinked to original sources

Majorization-minimization Bregman proximal gradient algorithms for NMF with the Kullback--Leibler divergence

Nonnegative matrix factorization (NMF) is a popular method in machine learning and signal processing to decompose a given nonnegative matrix into two nonnegative matrices. In this paper, we propose new algorithms, called majorization-minimization Bregman proximal gradient algorithm (MMBPG) and MMBPG with extrapolation (MMBPGe) to solve NMF. These iterative algorithms minimize the objective function and its potential function monotonically. Assuming the Kurdyka--\L{}ojasiewicz property, we establish that a sequence generated by MMBPG(e) globally converges to a stationary point. We apply MMBPG and MMBPGe to the Kullback--Leibler (KL) divergence-based NMF. While most existing KL-based NMF methods update two blocks or each variable alternately, our algorithms update all variables simultaneously. MMBPG and MMBPGe for KL-based NMF are equipped with a separable Bregman distance that satisfies the smooth adaptable property and that makes its subproblem solvable in closed form. Using this fact, we guarantee that a sequence generated by MMBPG(e) globally converges to a Karush--Kuhn--Tucker (KKT) point of KL-based NMF. In numerical experiments, we compare proposed algorithms with existing algorithms on synthetic data and real-world data.

math.OC

A unified Euler--Lagrange system for analyzing continuous-time accelerated gradient methods

This paper presents an Euler--Lagrange system for a continuous-time model of the accelerated gradient methods in smooth convex optimization and proposes an associated Lyapunov-function-based convergence analysis framework. Recently, ordinary differential equations (ODEs) with dumping terms have been developed to intuitively interpret the accelerated gradient methods, and the design of unified model describing the various individual ODE models have been examined. In existing reports, the Lagrangian, which results in the Euler-Lagrange equation, and the Lyapunov function for the convergence analysis have been separately proposed for each ODE. This paper proposes a unified Euler--Lagrange system and its Lyapunov function to cover the existing various models. In the convergence analysis using the Lyapunov function, a condition that parameters in the Lagrangian and Lyapunov function must satisfy is derived, and a parameter design for improving the convergence rate naturally results in the mysterious dumping coefficients. Especially, a symmetric Bregman divergence can lead to a relaxed condition of the parameters and a resulting improved convergence rate. As an application of this study, a slight modification in the Lyapunov function establishes the similar convergence proof for ODEs with smooth approximation in nondifferentiable objective function minimization.

math.OC

Convergence Rate Analysis of Continuous- and Discrete-Time Smoothing Gradient Algorithms

This paper addresses the gradient flow -- the continuous-time representation of the gradient method -- with the smooth approximation of a non-differentiable objective function and presents convergence analysis framework. Similar to the gradient method, the gradient flow is inapplicable to the non-differentiable function minimization; therefore, this paper addresses the smoothing gradient method, which exploits a decreasing smoothing parameter sequence in the smooth approximation. The convergence analysis is presented using conventional Lyapunov-function-based techniques, and a Lyapunov function applicable to both strongly convex and non-strongly convex objective functions is provided by taking into consideration the effect of the smooth approximation. Based on the equivalence of the stepsize in the smoothing gradient method and the discretization step in the forward Euler scheme for the numerical integration of the smoothing gradient flow, the sample values of the exact solution of the smoothing gradient flow are compared with the state variable of the smoothing gradient method, and the equivalence of the convergence rates is shown.

math.OC

On a minimization problem of the maximum generalized eigenvalue: properties and algorithms

We study properties and algorithms of a minimization problem of the maximum generalized eigenvalue of symmetric-matrix-valued affine functions, which is nonsmooth and quasiconvex, and has application to eigenfrequency optimization of truss structures. We derive an explicit formula of the Clarke subdifferential of the maximum generalized eigenvalue and prove the maximum generalized eigenvalue is a pseudoconvex function, which is a subclass of a quasiconvex function, under suitable assumptions. Then, we consider smoothing methods to solve the problem. We introduce a smooth approximation of the maximum generalized eigenvalue and prove the convergence rate of the smoothing projected gradient method to a global optimal solution in the considered problem. Also, some heuristic techniques to reduce the computational costs, acceleration and inexact smoothing, are proposed and evaluated by numerical experiments.

math.OC

Inverse-Optimization-Based Uncertainty Set for Robust Linear Optimization

We consider solving linear optimization (LO) problems with uncertain objective coefficients. For such problems, we often employ robust optimization (RO) approaches by introducing an uncertainty set for the unknown coefficients. Typical RO approaches require observations or prior knowledge of the unknown coefficient to define an appropriate uncertainty set. However, such information may not always be available in practice. In this study, we propose a novel uncertainty set for robust linear optimization (RLO) problems without prior knowledge of the unknown coefficients. Instead, we assume to have data of known constraint parameters and corresponding optimal solutions. Specifically, we derive an explicit form of the uncertainty set as a polytope by applying techniques of inverse optimization (IO). We prove that the RLO problem with the proposed uncertainty set can be equivalently reformulated as an LO problem. Numerical experiments show that the RO approach with the proposed uncertainty set outperforms classical IO in terms of performance stability.

math.OC

Blind Deconvolution with Non-smooth Regularization via Bregman Proximal DCAs

Blind deconvolution is a technique to recover an original signal without knowing a convolving filter. It is naturally formulated as a minimization of a quartic objective function under some assumption. Because its differentiable part does not have a Lipschitz continuous gradient, existing first-order methods are not theoretically supported. In this paper, we employ the Bregman-based proximal methods, whose convergence is theoretically guaranteed under the $L$-smooth adaptable ($L$-smad) property. We first reformulate the objective function as a difference of convex (DC) functions and apply the Bregman proximal DC algorithm (BPDCA). This DC decomposition satisfies the $L$-smad property. The method is extended to the BPDCA with extrapolation (BPDCAe) for faster convergence. When our regularizer has a sufficiently simple structure, each iteration is solved in a closed-form expression, and thus our algorithms solve large-scale problems efficiently. We also provide the stability analysis of the equilibrium and demonstrate the proposed methods through numerical experiments on image deblurring. The results show that BPDCAe successfully recovered the original image and outperformed other existing algorithms.

math.OC

A Gradient Method for Multilevel Optimization

Although application examples of multilevel optimization have already been discussed since the 1990s, the development of solution methods was almost limited to bilevel cases due to the difficulty of the problem. In recent years, in machine learning, Franceschi et al. have proposed a method for solving bilevel optimization problems by replacing their lower-level problems with the $T$ steepest descent update equations with some prechosen iteration number $T$. In this paper, we have developed a gradient-based algorithm for multilevel optimization with $n$ levels based on their idea and proved that our reformulation asymptotically converges to the original multilevel problem. As far as we know, this is one of the first algorithms with some theoretical guarantee for multilevel optimization. Numerical experiments show that a trilevel hyperparameter learning model considering data poisoning produces more stable prediction results than an existing bilevel hyperparameter learning model in noisy data settings.

math.OC

New Bregman proximal type algorithms for solving DC optimization problems

Difference of Convex (DC) optimization problems have objective functions that are differences between two convex functions. Representative ways of solving these problems are the proximal DC algorithms, which require that the convex part of the objective function have $L$-smoothness. In this article, we propose the Bregman Proximal DC Algorithm (BPDCA) for solving large-scale DC optimization problems that do not possess $L$-smoothness. Instead, it requires that the convex part of the objective function has the $L$-smooth adaptable property that is exploited in Bregman proximal gradient algorithms. In addition, we propose an accelerated version, the Bregman Proximal DC Algorithm with extrapolation (BPDCAe), with a new restart scheme. We show the global convergence of the iterates generated by BPDCA(e) to a limiting critical point under the assumption of the Kurdyka-{\L}ojasiewicz property or subanalyticity of the objective function and other weaker conditions than those of the existing methods. We applied our algorithms to phase retrieval, which can be described both as a nonconvex optimization problem and as a DC optimization problem. Numerical experiments showed that BPDCAe outperformed existing Bregman proximal-type algorithms because the DC formulation allows for larger admissible step sizes.

math.OC

Semi-flat minima and saddle points by embedding neural networks to overparameterization

We theoretically study the landscape of the training error for neural networks in overparameterized cases. We consider three basic methods for embedding a network into a wider one with more hidden units, and discuss whether a minimum point of the narrower network gives a minimum or saddle point of the wider one. Our results show that the networks with smooth and ReLU activation have different partially flat landscapes around the embedded point. We also relate these results to a difference of their generalization abilities in overparameterized realization.

cs.LG

Extension of the LP-Newton method to SOCPs via semi-infinite representation

The LP-Newton method solves the linear programming problem (LP) by repeatedly projecting a current point onto a certain relevant polytope. In this paper, we extend the algorithmic framework of the LP-Newton method to the second-order cone programming problem (SOCP) via a linear semi-infinite programming (LSIP) reformulation of the given SOCP. In the extension, we produce a sequence by projection onto polyhedral cones constructed from LPs obtained by finitely relaxing the LSIP. We show the global convergence property of the proposed algorithm under mild assumptions, and investigate its efficiency through numerical experiments comparing the proposed approach with the primal-dual interior-point method for the SOCP.

math.OC

Efficient Preconditioning for Noisy Separable NMFs by Successive Projection Based Low-Rank Approximations

The successive projection algorithm (SPA) can quickly solve a nonnegative matrix factorization problem under a separability assumption. Even if noise is added to the problem, SPA is robust as long as the perturbations caused by the noise are small. In particular, robustness against noise should be high when handling the problems arising from real applications. The preconditioner proposed by Gillis and Vavasis (2015) makes it possible to enhance the noise robustness of SPA. Meanwhile, an additional computational cost is required. The construction of the preconditioner contains a step to compute the top-$k$ truncated singular value decomposition of an input matrix. It is known that the decomposition provides the best rank-$k$ approximation to the input matrix; in other words, a matrix with the smallest approximation error among all matrices of rank less than $k$. This step is an obstacle to an efficient implementation of the preconditioned SPA. To address the cost issue, we propose a modification of the algorithm for constructing the preconditioner. Although the original algorithm uses the best rank-$k$ approximation, instead of it, our modification uses an alternative. Ideally, this alternative should have high approximation accuracy and low computational cost. To ensure this, our modification employs a rank-$k$ approximation produced by an SPA based algorithm. We analyze the accuracy of the approximation and evaluate the computational cost of the algorithm. We then present an empirical study revealing the actual performance of the SPA based rank-$k$ approximation algorithm and the modified preconditioned SPA.

math.NA