Searcharxiv⌕ Search

arXiv subjects

Ken'ichiro Tanaka

Publications and source records attributed to Ken'ichiro Tanaka.

At least 19 recordsLinked to original sources

Convergence analysis of dynamical systems for optimization by an improved Lyapunov framework

We study the convergence analysis of continuous-time dynamical systems associated with optimization methods for strongly convex functions. Recent works have proposed systematic constructions of Lyapunov functions for such analysis, while also revealing limitations of the Lyapunov analysis. Aujol--Dossal--Rondepierre (2023) have proposed a technique to address this issue by reorganizing Lyapunov functions so as to evaluate a quantity $f(x(t)) - f_* - g(t)\|x(t)-x_*\|^2$ rather than $f(x(t)) - f_*$. By combining this technique with our computer-assisted framework to discover Lyapunov functions, we develop an improved method that reproduces an existing convergence rate or yields better rates than previous studies.

math.OC↗

On solving nonlinear simultaneous equations arising from the double-exponential Sinc-collocation method for initial value problems

The double-exponential Sinc-collocation method is known as a super-accurate method for solving initial value problems of ordinary differential equations, for which the error decreases almost exponentially as a function of the number of sample points in the temporal direction, $N$. However, this method requires solving nonlinear simultaneous equations in $nN$ variables when the problem dimension is $n$. Recently, Ogata pointed out that Gauss-Seidel type fixed-point iteration works surprisingly well for solving these equations, typically reducing the error by one or two orders of magnitude at each iteration. In this paper, we analyze the convergence of this iteration and give a sufficient condition for its global convergence. We also provide an upper bound on its convergence factor, which explains the efficiency of this iteration. Some numerical examples that illustrate the validity of our analysis are also provided.

math.NA↗

Computer-Assisted Search for Differential Equations Corresponding to Optimization Methods and Their Convergence Rates

Let $f:\mathbb{R}^n \to \mathbb{R}$ be a continuously differentiable convex function with its minimizer denoted by $x_*$ and optimal value $f_* = f(x_*)$. Optimization algorithms such as the gradient descent method can often be interpreted in the continuous-time limit as differential equations known as continuous dynamical systems. Analyzing the convergence rate of $f(x) - f_*$ in such systems often relies on constructing appropriate Lyapunov functions. However, these Lyapunov functions have been designed through heuristic reasoning rather than a systematic framework. Several studies have addressed this issue. In particular, Suh, Roh, and Ryu (2022) proposed a constructive approach that involves introducing dilated coordinates and applying integration by parts. Although this method significantly improves the process of designing Lyapunov functions, it still involves arbitrary choices among many possible options, and thus retains a heuristic nature in identifying Lyapunov functions that yield the best convergence rates. In this study, we propose a systematic framework for exploring these choices computationally. More precisely, we propose a brute-force approach using symbolic computation by computer algebra systems to explore every possibility. By formulating the design of Lyapunov functions for continuous dynamical systems as an optimization problem, we aim to optimize the Lyapunov function itself. As a result, our framework successfully reproduces many previously reported results and, in several cases, discovers new convergence rates that have not been shown in the existing studies.

math.OC↗

Computation of the exponential function of matrices by a formula without oscillatory integrals on infinite intervals

We propose a quadrature-based formula for computing the exponential function of matrices with a non-oscillatory integral on an infinite interval and an oscillatory integral on a finite interval. In the literature, existing quadrature-based formulas are based on the inverse Laplace transform or the Fourier transform. We show these expressions are essentially equivalent in terms of complex integrals and choose the former as a starting point to reduce computational cost. By choosing a simple integral path, we derive an integral expression mentioned above. Then, we can easily apply the double-exponential formula and the Gauss-Legendre formula, which have rigorous error bounds. As numerical experiments show, the proposed formula outperforms the existing formulas when the imaginary parts of the eigenvalues of matrices have large absolute values.

math.NA↗

How sharp are error bounds? --lower bounds on quadrature worst-case errors for analytic functions

Numerical integration over the real line for analytic functions is studied. Our main focus is on the sharpness of the error bounds. We first derive two general lower estimates for the worst-case integration error, and then apply these to establish lower bounds for various quadrature rules. These bounds turn out to be either novel or improve upon existing results, leading to lower bounds that closely match upper bounds for various formulas. Specifically, for the suitably truncated trapezoidal rule, we improve upon general lower bounds on the worst-case error obtained by Sugihara [\textit{Numer. Math.}, 75 (1997), pp.~379--395] and provide exceptionally sharp lower bounds apart from a polynomial factor, and in particular show that the worst-case error for the trapezoidal rule by Sugihara is not improvable by more than a polynomial factor. Additionally, our research reveals a discrepancy between the error decay of the trapezoidal rule and Sugihara's lower bound for general numerical integration rules, introducing a new open problem. Moreover, Gauss--Hermite quadrature is proven sub-optimal under the decay conditions on integrands we consider, a result not deducible from upper-bound arguments alone. Furthermore, to establish the near-optimality of the suitably scaled Gauss--Legendre and Clenshaw--Curtis quadratures, we generalize a recent result of Trefethen [\textit{SIAM Rev.}, 64 (2022), pp.~132--150] for the upper error bounds in terms of the decay conditions.

math.NA↗

Accelerated gradient descent method for functionals of probability measures by new convexity and smoothness based on transport maps

We consider problems of minimizing functionals $\mathcal{F}$ of probability measures on the Euclidean space. To propose an accelerated gradient descent algorithm for such problems, we consider gradient flow of transport maps that give push-forward measures of an initial measure. Then we propose a deterministic accelerated algorithm by extending Nesterov's acceleration technique with momentum. This algorithm do not based on the Wasserstein geometry. Furthermore, to estimate the convergence rate of the accelerated algorithm, we introduce new convexity and smoothness for $\mathcal{F}$ based on transport maps. As a result, we can show that the accelerated algorithm converges faster than a normal gradient descent algorithm. Numerical experiments support this theoretical result.

math.OC↗

Skeleton structure inherent in discrete-time quantum walks

In this paper, we claim that a common underlying structure--a skeleton structure--is present behind discrete-time quantum walks (QWs) on a one-dimensional lattice with a homogeneous coin matrix. This skeleton structure is independent of the initial state, and partially, even of the coin matrix. This structure is best interpreted in the context of quantum-walk-replicating random walks (QWRWs), i.e., random walks that replicate the probability distribution of quantum walks, where this newly found structure acts as a simplified formula for the transition probability. Additionally, we construct a random walk whose transition probabilities are defined by the skeleton structure and demonstrate that the resultant properties of the walkers are similar to both the original QWs and QWRWs.

math-ph↗

Convergence analysis of approximation formulas for analytic functions via duality for potential energy minimization

We investigate the approximation formulas that were proposed by Tanaka & Sugihara (2019), in weighted Hardy spaces, which are analytic function spaces with certain asymptotic decay. Under the criterion of minimum worst error of $n$-point approximation formulas, we demonstrate that the formulas are nearly optimal. We also obtain the upper bounds of the approximation errors that coincide with the existing heuristic bounds in asymptotic order by duality theorem for the minimization problem of potential energy.

math.NA↗

Convergence rates for energies of interacting particles whose distribution spreads out as their number increases

We consider a class of particle systems which appear in various applications such as approximation theory, plasticity, potential theory and space-filling designs. The positions of the particles on the real line are described as a global minimum of an interaction energy, which consists of a nonlocal, repulsive interaction part and a confining part. Motivated by the applications, we cover non-standard scenarios in which the confining potential weakens as the number of particles increases. This results in a large area over which the particles spread out. Our aim is to approximate the particle interaction energy by a corresponding continuum interacting energy. Our main results are bounds on the corresponding energy difference and on the difference between the related potential values. We demonstrate that these bounds are useful to problems in approximation theory and plasticity. The proof of these bounds relies on convexity assumptions on the interaction and confining potentials. It combines recent advances in the literature with a new upper bound on the minimizer of the continuum interaction energy.

math.AP↗

Wavelet characterization of exponentially weighted Besov space with dominating mixed smoothness and its application to function approximation

Although numerous studies have focused on normal Besov spaces, limited studies have been conducted on exponentially weighted Besov spaces. Therefore, we define exponentially weighted Besov space $VB_{p,q}^{δ,w}(\mathbb{R}^d)$ whose smoothness includes normal Besov spaces, Besov spaces with dominating mixed smoothness, and their interpolation. Furthermore, we obtain wavelet characterization of $VB_{p,q}^{δ,w}(\mathbb{R}^d)$. Next, approximation formulas such as sparse grids are derived using the determined formula. The results of this study are expected to provide considerable insight into the application of exponentially weighted Besov spaces with mixed smoothness.

math.FA↗

Sparse solutions of the kernel herding algorithm by improved gradient approximation

The kernel herding algorithm is used to construct quadrature rules in a reproducing kernel Hilbert space (RKHS). While the computational efficiency of the algorithm and stability of the output quadrature formulas are advantages of this method, the convergence speed of the integration error for a given number of nodes is slow compared to that of other quadrature methods. In this paper, we propose a modified kernel herding algorithm whose framework was introduced in a previous study and aim to obtain sparser solutions while preserving the advantages of standard kernel herding. In the proposed algorithm, the negative gradient is approximated by several vertex directions, and the current solution is updated by moving in the approximate descent direction in each iteration. We show that the convergence speed of the integration error is directly determined by the cosine of the angle between the negative gradient and approximate gradient. Based on this, we propose new gradient approximation algorithms and analyze them theoretically, including through convergence analysis. In numerical experiments, we confirm the effectiveness of the proposed algorithms in terms of sparsity of nodes and computational efficiency. Moreover, we provide a new theoretical analysis of the kernel quadrature rules with fully-corrective weights, which realizes faster convergence speeds than those of previous studies.

math.NA↗

Yet another DE-Sinc indefinite integration formula

Based on the Sinc approximation combined with the tanh transformation, Haber derived an approximation formula for numerical indefinite integration over the finite interval (-1, 1). The formula uses a special function for the basis functions. In contrast, Stenger derived another formula, which does not use any special function but does include a double sum. Subsequently, Muhammad and Mori proposed a formula, which replaces the tanh transformation with the double-exponential transformation in Haber's formula. Almost simultaneously, Tanaka et al. proposed another formula, which was based on the same replacement in Stenger's formula. As they reported, the replacement drastically improves the convergence rate of Haber's and Stenger's formula. In addition to the formulas above, Stenger derived yet another indefinite integration formula based on the Sinc approximation combined with the tanh transformation, which has an elegant matrix-vector form. In this paper, we propose the replacement of the tanh transformation with the double-exponential transformation in Stenger's second formula. We provide a theoretical analysis as well as a numerical comparison.

math.NA↗

Sparser Kernel Herding with Pairwise Conditional Gradients without Swap Steps

The Pairwise Conditional Gradients (PCG) algorithm is a powerful extension of the Frank-Wolfe algorithm leading to particularly sparse solutions, which makes PCG very appealing for problems such as sparse signal recovery, sparse regression, and kernel herding. Unfortunately, PCG exhibits so-called swap steps that might not provide sufficient primal progress. The number of these bad steps is bounded by a function in the dimension and as such known guarantees do not generalize to the infinite-dimensional case, which would be needed for kernel herding. We propose a new variant of PCG, the so-called Blended Pairwise Conditional Gradients (BPCG). This new algorithm does not exhibit any swap steps, is very easy to implement, and does not require any internal gradient alignment procedures. The convergence rate of BPCG is basically that of PCG if no drop steps would occur and as such is no worse than PCG but improves and provides new rates in many cases. Moreover, we observe in the numerical experiments that BPCG's solutions are much sparser than those of PCG. We apply BPCG to the kernel herding setting, where we derive nice quadrature rules and provide numerical results demonstrating the performance of our method.

math.OC↗

Monte Carlo construction of cubature on Wiener space

In this paper, we investigate application of mathematical optimization to construction of a cubature formula on Wiener space, which is a weak approximation method of stochastic differential equations introduced by Lyons and Victoir (Cubature on Wiener Space, Proc. R. Soc. Lond. A 460, 169--198). After giving a brief review of the cubature theory on Wiener space, we show that a cubature formula of general dimension and degree can be obtained through a Monte Carlo sampling and linear programming. This paper also includes an extension of stochastic Tchakaloff's theorem, which technically yields the proof of our primary result.

math.PR↗

Kernel quadrature by applying a point-wise gradient descent method to discrete energies

We propose a method for generating nodes for kernel quadrature by a point-wise gradient descent method. For kernel quadrature, most methods for generating nodes are based on the worst case error of a quadrature formula in a reproducing kernel Hilbert space corresponding to the kernel. In typical ones among those methods, a new node is chosen among a candidate set of points in each step by an optimization problem with respect to a new node. Although such sequential methods are appropriate for adaptive quadrature, it is difficult to apply standard routines for mathematical optimization to the problem. In this paper, we propose a method that updates a set of points one by one with a simple gradient descent method. To this end, we provide an upper bound of the worst case error by using the fundamental solution of the Laplacian on $\mathbf{R}^{d}$. We observe the good performance of the proposed method by numerical experiments.

math.NA↗

Kernel-based interpolation at approximate Fekete points

We construct approximate Fekete point sets for kernel-based interpolation by maximising the determinant of a kernel Gram matrix obtained via truncation of an orthonormal expansion of the kernel. Uniform error estimates are proved for kernel interpolants at the resulting points. If the kernel is Gaussian we show that the approximate Fekete points in one dimension are the solution to a convex optimisation problem and that the interpolants converge with a super-exponential rate. Numerical examples are provided for the Gaussian kernel.

math.NA↗

Generation of point sets by convex optimization for interpolation in reproducing kernel Hilbert spaces

We propose algorithms to take point sets for kernel-based interpolation of functions in reproducing kernel Hilbert spaces (RKHSs) by convex optimization. We consider the case of kernels with the Mercer expansion and propose an algorithm by deriving a second-order cone programming (SOCP) problem that yields $n$ points at one sitting for a given integer $n$. In addition, by modifying the SOCP problem slightly, we propose another sequential algorithm that adds an arbitrary number of new points in each step. Numerical experiments show that in several cases the proposed algorithms compete with the $P$-greedy algorithm, which is known to provide nearly optimal points.

math.NA↗

Construction of conformal maps based on the locations of singularities for improving the double exponential formula

The double exponential formula, or the DE formula, is a high-precision integration formula using a change of variables called a DE transformation; whereas there is a disadvantage that it is sensitive to singularities of an integrand near the real axis. To overcome this disadvantage, Slevinsky and Olver (SIAM J. Sci. Comput., 2015) attempted to improve it by constructing conformal maps based on the locations of singularities. Based on their ideas, we construct a new transformation formula. Our method employs special types of the Schwarz-Christoffel transformations for which we can derive their explicit form. Then, the new transformation formula can be regarded as a generalization of the DE transformations. We confirm its effectiveness by numerical experiments.

math.NA↗