Searcharxiv⌕ Search

arXiv subjects

Kansei Ushiyama

Publications and source records attributed to Kansei Ushiyama.

6 recordsLinked to original sources

A $\sqrt{2}$-accelerated FISTA for composite strongly convex problems

In this paper, we propose a novel accelerated forward-backward splitting algorithm for minimizing composite convex functions expressed as the sum of a smooth function and a possibly nonsmooth function. When the composite objective is strongly convex, the proposed method achieves a linear convergence rate in objective value that improves the leading constant in the exponent by a factor of $\sqrt{2}$ relative to FISTA and also improves upon the best previously known convergence rate for this problem class. Our convergence analysis remains valid even when one of the two component functions is weakly convex, provided that their sum remains convex. The proposed algorithm is derived by discretizing a continuous-time model of the Information-Theoretic Exact Method (ITEM), an optimal first-order method for unconstrained smooth strongly convex minimization.

math.OC↗

A Restart-Free Accelerated Algorithm for Non-Convex Minimization: Continuous and Discrete Analysis

We propose two novel first-order methods for minimizing nonconvex functions with Lipschitz-continuous gradients and Hessians. These algorithms attain an $\varepsilon$-approximate first-order stationary point in $\mathrm{O}(\varepsilon^{-7/4})$ function and gradient evaluations, without using $\varepsilon$ as an input parameter. While existing methods rely on restart mechanisms to achieve this complexity, our methods do not. Consequently, the first algorithm enjoys a simple implementation, making its last iterate differentiable with respect to the initial point. By estimating the Lipschitz constants adaptively, we develop the second algorithm that does not require prior knowledge of the constants. This algorithm exhibits better numerical performance than existing parameter-free methods for certain problems, which can be attributed to its restart-free design. Both algorithms are derived by discretizing a newly introduced continuous-time model represented by an ordinary differential equation, and their continuous- and discrete-time convergence analyses proceed in a parallel manner under the Performance Estimation Problem framework.

math.OC↗

Essential Convergence Rates of Continuous-Time Models for Optimization Methods

Designing and analyzing optimization methods via continuous-time models expressed as ordinary differential equations (ODEs) is a promising approach for its intuitiveness and simplicity. A key concern, however, is that the convergence rates of such models can be arbitrarily modified by time rescaling, rendering the task of seeking ODEs with ``fast'' convergence meaningless. To eliminate this ambiguity of the rates, we introduce the notion of the essential convergence rate. We justify this notion by proving that, under appropriate assumptions on discretization, no method obtained by discretizing an ODE can achieve a faster rate than its essential convergence rate.

math.OC↗

Differentiable Pareto-Smoothed Weighting for High-Dimensional Heterogeneous Treatment Effect Estimation

There is a growing interest in estimating heterogeneous treatment effects across individuals using their high-dimensional feature attributes. Achieving high performance in such high-dimensional heterogeneous treatment effect estimation is challenging because in this setup, it is usual that some features induce sample selection bias while others do not but are predictive of potential outcomes. To avoid losing such predictive feature information, existing methods learn separate feature representations using inverse probability weighting (IPW). However, due to their numerically unstable IPW weights, these methods suffer from estimation bias under a finite sample setup. To develop a numerically robust estimator by weighted representation learning, we propose a differentiable Pareto-smoothed weighting framework that replaces extreme weight values in an end-to-end fashion. Our experimental results show that by effectively correcting the weight values, our proposed method outperforms the existing ones, including traditional weighting schemes. Our code is available at https://github.com/ychika/DPSW.

stat.ML↗

A new unified framework for designing convex optimization methods with prescribed theoretical convergence estimates: A numerical analysis approach

We propose a new unified framework for describing and designing gradient-based convex optimization methods from a numerical analysis perspective. There the key is the new concept of weak discrete gradients (weak DGs), which is a generalization of DGs standard in numerical analysis. Via weak DG, we consider abstract optimization methods, and prove unified convergence rate estimates that hold independent of the choice of weak DGs except for some constants in the final estimate. With some choices of weak DGs, we can reproduce many popular existing methods, such as the steepest descent and Nesterov's accelerated gradient method, and also some recent variants from numerical analysis community. By considering new weak DGs, we can easily explore new theoretically-guaranteed optimization methods; we show some examples. We believe this work is the first attempt to fully integrate research branches in optimization and numerical analysis areas, so far independently developed.

math.OC↗

Essential convergence rate of ordinary differential equations appearing in optimization

Some continuous optimization methods can be connected to ordinary differential equations (ODEs) by taking continuous limits, and their convergence rates can be explained by the ODEs. However, since such ODEs can achieve any convergence rate by time scaling, the correspondence is not as straightforward as usually expected, and deriving new methods through ODEs is not quite direct. In this letter, we pay attention to stability restriction in discretizing ODEs and show that acceleration by time scaling basically implies deceleration in discretization; they balance out so that we can define an attainable unique convergence rate which we call an "essential convergence rate".

math.NA↗