Searcharxiv⌕ Search

arXiv subjects

Nobuo Yamashita

Publications and source records attributed to Nobuo Yamashita.

At least 19 recordsLinked to original sources

Extension of the safeguarding stepsize interval in Adaptive Gradient Descent

In this paper, inspired by the idea of Adaptive Gradient Descent (AdGD) by Malitsky and Mishchenko, we propose a method that adaptively provides a stepsize interval, which represents the condition for stepsizes ensuring global convergence. By projecting the Barzilai Borwein stepsize onto this interval, we can ensure global convergence and expect further acceleration of convergence. Furthermore, we propose enlarging the stepsize interval obtained by AdGD and show that global convergence can be guaranteed within this enlarged stepsize interval. This extension of the interval enables us to adopt the full Barzilai-Borwein (BB) stepsize more frequently. We report the results of numerical experiments on the stepsize that combines the proposed interval with the BB stepsize.

math.OC↗

Augmented Lagrangian methods for convex optimization with priority constraints via an infeasibility control framework

We consider convex optimization problems with prioritized equality constraints, which may be infeasible. In many applications, such as network optimization and image reconstruction, it is often desirable to compute solutions that satisfy higher-priority constraints as much as possible even when no feasible solution exists. To address this issue, we introduce a new solution framework based on the notion of a hierarchically optimal shift, which captures the hierarchy among constraints by sequentially minimizing constraint violations according to their priorities. Based on this concept, we define a hierarchically optimal solution as an optimal solution of a suitably shifted problem, thereby providing a well-defined notion of optimality even in the absence of feasibility. Furthermore, we propose a novel augmented Lagrangian method equipped with a framework for infeasibility control. The core component is an infeasibility control problem, which generates a sequence of approximate shifts converging to the hierarchically optimal shift. This approach enables explicit and systematic handling of prioritized constraint violations, in contrast to existing methods that treat all constraints uniformly. Under suitable assumptions, we show that the generated sequence of shifts converges to the hierarchically optimal shift, and that any accumulation point of the primal iterates is a hierarchically optimal solution. Numerical experiments show that the proposed method achieves solutions consistent with the prescribed constraint hierarchy for both feasible and infeasible cases.

math.OC↗

An uncertainty model for positive-valued parameters with application to robust optimization

Many practical optimization problems involve uncertain parameters that are strictly positive. However, the most common uncertainty sets used in robust optimization are the box and the ellipsoidal sets, which may include non-positive values when the level of uncertainty is large. This can lead to overly conservative solutions or make the corresponding robust counterpart infeasible. To overcome this, in this paper, we propose a new uncertainty-set model that not only preserves positivity but is also computationally tractable. The proposed set uses a particular convex function that measures the variation of uncertain parameters from their nominal values. We can also write the dual reformulation of the associated robust problem. For the theoretical results, we show several properties of the proposed model, including analytical bounds that guide the choice of the uncertainty level, as well as a probabilistic guarantee result. To check the validity of our proposal, we consider photovoltaic-battery operation planning problems and support vector machines in the numerical experiments. For these problems, standard uncertainty models may lead to infeasibility of the robust counterpart, while the proposed uncertainty set gives a tractable dual reformulation.

math.OC↗

Block Coordinate Descent Network Simplex Methods for Optimal Transport

We propose the Block Coordinate Descent Network Simplex (BCDNS) method for solving large-scale discrete Optimal Transport (OT) problems. BCDNS integrates the Network Simplex (NS) algorithm with a block coordinate descent (BCD) strategy, decomposing the full problem into smaller subproblems per iteration and reusing basis variables to ensure feasibility. We prove that BCDNS terminates in a finite number of iterations with an exact optimal solution, and we characterize its per-iteration complexity as O(s N), where s is a user-defined parameter in (0,1) and N is the total number of variables. Numerical experiments demonstrate that BCDNS matches the classical NS method in solution accuracy, reduces memory footprint compared to the Sinkhorn algorithm, achieves speed-ups of up to tens of times over the classical NS method, and exhibits runtime comparable to a high-precision Sinkhorn implementation.

math.OC↗

A strong second-order sequential optimality condition for nonlinear programming problems

Most numerical methods developed for solving nonlinear programming problems are designed to find points that satisfy certain optimality conditions. While the Karush-Kuhn-Tucker conditions are well-known, they become invalid when constraint qualifications (CQ) are not met. Recent advances in sequential optimality conditions address this limitation in both first- and second-order cases, providing genuine optimality guarantees at local optima, even when CQs do not hold. However, some second-order sequential optimality conditions still require some restrictive conditions on constraints in the recent literature. In this paper, we propose a new strong second-order sequential optimality condition without CQs. We also show that a penalty-type method and an augmented Lagrangian method generate points satisfying these new optimality conditions.

math.OC↗

Convergence analysis of a regularized Newton method with generalized regularization terms for convex optimization problems

This paper presents a regularized Newton method (RNM) with generalized regularization terms for unconstrained convex optimization problems. The generalized regularization includes quadratic, cubic, and elastic net regularizations as special cases. Therefore, the proposed method serves as a general framework that includes not only the classical and cubic RNMs but also a novel RNM with elastic net regularization. We show that the proposed RNM has the global $\mathcal{O}(k^{-2})$ and local superlinear convergence, which are the same as those of the cubic RNM.

math.OC↗

Monotonicity for Multiobjective Accelerated Proximal Gradient Methods

Accelerated proximal gradient methods, which are also called fast iterative shrinkage-thresholding algorithms (FISTA) are known to be efficient for many applications. Recently, Tanabe et al. proposed an extension of FISTA for multiobjective optimization problems. However, similarly to the single-objective minimization case, the objective functions values may increase in some iterations, and inexact computations of subproblems can also lead to divergence. Motivated by this, here we propose a variant of the FISTA for multiobjective optimization, that imposes some monotonicity of the objective functions values. In the single-objective case, we retrieve the so-called MFISTA, proposed by Beck and Teboulle. We also prove that our method has global convergence with rate $O(1/k^2)$, where $k$ is the number of iterations, and show some numerical advantages in requiring monotonicity.

math.OC↗

An accelerated proximal gradient method for multiobjective optimization

This paper presents an accelerated proximal gradient method for multiobjective optimization, in which each objective function is the sum of a continuously differentiable, convex function and a closed, proper, convex function. Extending first-order methods for multiobjective problems without scalarization has been widely studied, but providing accelerated methods with accurate proofs of convergence rates remains an open problem. Our proposed method is a multiobjective generalization of the accelerated proximal gradient method, also known as the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA), for scalar optimization. The key to this successful extension is solving a subproblem with terms exclusive to the multiobjective case. This approach allows us to demonstrate the global convergence rate of the proposed method ($O(1 / k^2)$), using a merit function to measure the complexity. Furthermore, we present an efficient way to solve the subproblem via its dual representation, and we confirm the validity of the proposed method through some numerical experiments.

math.OC↗

New merit functions for multiobjective optimization and their properties

A merit (gap) function is a map that returns zero at the solutions of problems and strictly positive values otherwise. Its minimization is equivalent to the original problem by definition, and it can estimate the distance between a given point and the solution set. Ideally, this function should have some properties, including the ease of computation, continuity, differentiability, boundedness of the level set, and error boundedness. In this work, we propose new merit functions for multiobjective optimization with lower semicontinuous objectives, convex objectives, and composite objectives, and we show that they have such desirable properties under reasonable assumptions.

math.OC↗

Convergence analysis and acceleration of the smoothing methods for solving extensive-form games

The extensive-form game has been studied considerably in recent years. It can represent games with multiple decision points and incomplete information, and hence it is helpful in formulating games with uncertain inputs, such as poker. We consider an extended-form game with two players and zero-sum, i.e., the sum of their payoffs is always zero. In such games, the problem of finding the optimal strategy can be formulated as a bilinear saddle-point problem. This formulation grows huge depending on the size of the game, since it has variables representing the strategies at all decision points for each player. To solve such large-scale bilinear saddle-point problems, the excessive gap technique (EGT), a smoothing method, has been studied. This method generates a sequence of approximate solutions whose error is guaranteed to converge at $\mathcal{O}(1/k)$, where $k$ is the number of iterations. However, it has the disadvantage of having poor theoretical bounds on the error related to the game size. This makes it inapplicable to large games. Our goal is to improve the smoothing method for solving extensive-form games so that it can be applied to large-scale games. To this end, we make two contributions in this work. First, we slightly modify the strongly convex function used in the smoothing method in order to improve the theoretical bounds related to the game size. Second, we propose a heuristic called centering trick, which allows the smoothing method to be combined with other methods and consequently accelerates the convergence in practice. As a result, we combine EGT with CFR+, a state-of-the-art method for extensive-form games, to achieve good performance in games where conventional smoothing methods do not perform well. The proposed smoothing method is shown to have the potential to solve large games in practice.

cs.GT↗

Distributionally Robust Expected Residual Minimization for Stochastic Variational Inequality Problems

The stochastic variational inequality problem (SVIP) is an equilibrium model that includes random variables and has been widely applied in various fields such as economics and engineering. Expected residual minimization (ERM) is an established model for obtaining a reasonable solution for the SVIP, and its objective function is an expected value of a suitable merit function for the SVIP. However, the ERM is restricted to the case where the distribution is known in advance. We extend the ERM to ensure the attainment of robust solutions for the SVIP under the uncertainty distribution (the extended ERM is referred to as distributionally robust expected residual minimization (DRERM), where the worst-case distribution is derived from the set of probability measures in which the expected value and variance take the same sample mean and variance, respectively). Under suitable assumptions, we demonstrate that the DRERM can be reformulated as a deterministic convex nonlinear semidefinite programming to avoid numerical integration.

math.OC↗

A Stochastic Variance Reduced Gradient using Barzilai-Borwein Techniques as Second Order Information

In this paper, we consider to improve the stochastic variance reduce gradient (SVRG) method via incorporating the curvature information of the objective function. We propose to reduce the variance of stochastic gradients using the computationally efficient Barzilai-Borwein (BB) method by incorporating it into the SVRG. We also incorporate a BB-step size as its variant. We prove its linear convergence theorem that works not only for the proposed method but also for the other existing variants of SVRG with second-order information. We conduct the numerical experiments on the benchmark datasets and show that the proposed method with constant step size performs better than the existing variance reduced methods for some test problems.

math.OC↗

A globally convergent fast iterative shrinkage-thresholding algorithm with a new momentum factor for single and multi-objective convex optimization

Convex-composite optimization, which minimizes an objective function represented by the sum of a differentiable function and a convex one, is widely used in machine learning and signal/image processing. Fast Iterative Shrinkage Thresholding Algorithm (FISTA) is a typical method for solving this problem and has a global convergence rate of $O(1 / k^2)$. Recently, this has been extended to multi-objective optimization, together with the proof of the $O(1 / k^2)$ global convergence rate. However, its momentum factor is classical, and the convergence of its iterates has not been proven. In this work, introducing some additional hyperparameters $(a, b)$, we propose another accelerated proximal gradient method with a general momentum factor, which is new even for the single-objective cases. We show that our proposed method also has a global convergence rate of $O(1/k^2)$ for any $(a,b)$, and further that the generated sequence of iterates converges to a weak Pareto solution when $a$ is positive, an essential property for the finite-time manifold identification. Moreover, we report numerical results with various $(a,b)$, showing that some of these choices give better results than the classical momentum factors.

math.OC↗

Convergence rates analysis of a multiobjective proximal gradient method

Many descent algorithms for multiobjective optimization have been developed in the last two decades. Tanabe et al. (Comput Optim Appl 72(2):339--361, 2019) proposed a proximal gradient method for multiobjective optimization, which can solve multiobjective problems, whose objective function is the sum of a continuously differentiable function and a closed, proper, and convex one. Under reasonable assumptions, it is known that the accumulation points of the sequences generated by this method are Pareto stationary. However, the convergence rates were not established in that paper. Here, we show global convergence rates for the multiobjective proximal gradient method, matching what is known in scalar optimization. More specifically, by using merit functions to measure the complexity, we present the convergence rates for non-convex ($O(\sqrt{1 / k})$), convex ($O(1 / k)$), and strongly convex ($O(r^k)$ for some $r \in (0, 1)$) problems. We also extend the so-called Polyak-Łojasiewicz (PL) inequality for multiobjective optimization and establish the linear convergence rate for multiobjective problems that satisfy such inequalities ($O(r^k)$ for some $r \in (0, 1)$).

math.OC↗

An equivalent nonlinear optimization model with triangular low-rank factorization for semidefinite programs

In this paper, we propose a new nonlinear optimization model to solve semidefinite optimization problems (SDPs), providing some properties related to local optimal solutions. The proposed model is based on another nonlinear optimization model given by Burer and Monteiro (2003), but it has several nice properties not seen in the existing one. Firstly, the decision variable of the proposed model is a triangular low-rank matrix, and hence the dimension of its decision variable space is smaller. Secondly, the existence of a strict local optimum of the proposed model is guaranteed under some conditions, whereas the existing model has no strict local optimum. In other words, it is difficult to construct solution methods equipped with fast convergence using the existing model. Some numerical results are also presented to examine the efficiency of the proposed model.

math.OC↗

A Regularized Limited Memory BFGS method for Large-Scale Unconstrained Optimization and its Efficient Implementations

The limited memory BFGS (L-BFGS) method is one of the popular methods for solving large-scale unconstrained optimization. Since the standard L-BFGS method uses a line search to guarantee its global convergence, it sometimes requires a large number of function evaluations. To overcome the difficulty, we propose a new L-BFGS with a certain regularization technique. We show its global convergence under the usual assumptions. In order to make the method more robust and efficient, we also extend it with several techniques such as nonmonotone technique and simultaneous use of the Wolfe line search. Finally, we present some numerical results for test problems in CUTEst, which show that the proposed method is robust in terms of solving number of problems.

math.OC↗

Duality of optimization problems with gauge functions

Recently, Yamanaka and Yamashita proposed the so-called positively homogeneous optimization problem, which includes many important problems, such as the absolute-value and the gauge optimizations. They presented a closed form of the dual formulation for the problem, and showed weak duality and the equivalence to the Lagrangian dual under some conditions. In this work, we focus on a special positively homogeneous optimization problem, whose objective function and constraints consist of some gauge and linear functions. We prove not only weak duality but also strong duality. We also study necessary and sufficient optimality conditions associated to the problem. Moreover, we give sufficient conditions under which we can recover a primal solution from a Karush-Kuhn-Tucker point of the dual formulation. Finally, we discuss how to extend the above results to general convex optimization problems by considering the so-called perspective functions.

math.OC↗

Alternating Direction Method of Multipliers with Variable Metric Indefinite Proximal Terms for Convex Optimization

This paper studies a proximal alternating direction method of multipliers (ADMM) with variable metric indefinite proximal terms for linearly constrained convex optimization problems. The proximal ADMM plays an important role in many application areas, since the subproblems of the method are easy to solve. Recently, it is reported that the proximal ADMM with a certain fixed indefinite proximal term is faster than that with a positive semidefinite term, and still has the global convergence property. On the other hand, Gu and Yamashita studied a variable metric semidefinite proximal ADMM whose proximal term is generated by the BFGS update. They reported that a slightly indefinite matrix also makes the algorithm work well in their numerical experiments. Motivated by this fact, we consider a variable metric indefinite proximal ADMM, and give sufficient conditions on the proximal terms for the global convergence. Moreover, we propose a new indefinite proximal term based on the BFGS update which can satisfy the conditions for the global convergence.

math.OC↗