SearcharxivSearch

arXiv subjects

Sho Shimoyama

Publications and source records attributed to Sho Shimoyama.

4 recordsLinked to original sources

Mirror flows and $p$-Laplacian eigenvalue problems on metric measure spaces

We study mirror flows as Banach-space counterparts of Hilbertian gradient flows and show that they play an essential role in $p$-Laplacian eigenvalue problems on metric measure spaces. Under a Rellich--Kondrachov type compactness assumption and for $p\ge2$, we establish a Ljusternik--Schnirelman type existence theorem without assuming $C^1$-regularity of the associated energy. More precisely, we prove that every element $\lambda$ of the Krasnoselskii spectrum is an eigenvalue of the $p$-Laplacian $\Delta_p$; that is, there exists a nontrivial solution $f$ to $\Delta_p f=-\lambda |f|^{p-2}f$.We also investigate the large time behavior of the corresponding mirror flow and prove its convergence to an eigenfunction when the eigenvalue is simple and isolated.

math.AP

On gradient descent-ascent flows in metric spaces

Gradient descent-ascent (GDA) flows play a central role in finding saddle points of bivariate functionals, with applications in optimization, game theory, and robust control. While they are well-understood in Hilbert and Banach spaces via maximal monotone operator theory, their extension to general metric spaces, particularly Wasserstein spaces, has remained largely unexplored. In this paper, we develop a mathematical theory of GDA flows on the product of two complete metric spaces, formulating them as solutions to a system of evolution variational inequalities (EVIs) driven by a proper, closed functional $ϕ$. Under mild convex-concave and regularity assumptions on $ϕ$, we prove the existence, uniqueness, and stability of the flows via a novel minimizing-maximizing movement scheme and a minimax theorem on metric spaces. We establish a $λ$-contraction property, derive a quantitative error estimate for the discrete scheme, and demonstrate regularization effects analogous to classical gradient flows. Moreover, we obtain an exponential decay bound for the Nikaidô--Isoda duality gap along the flow. Focusing on Wasserstein spaces over Hilbert spaces, we show the global existence in time and the exponential convergence of the Wasserstein GDA flow to the unique saddle point for strongly convex-concave functionals. Our framework unifies and extends existing analyses, offering a metric-geometric perspective on GDA dynamics in nonlinear and non-smooth settings.

math.FA

Transformation of $p$-gradient flows to $p'$-gradient flows in metric spaces

We explicitly construct parameter transformations between gradient flows in metric spaces, called curves of maximal slope, having different exponents when the associated function satisfies a suitable convexity condition. These transformations induce the uniqueness of gradient flows for all exponents under a natural assumption which is satisfied in many examples. We also prove the regularizing effects of gradient flows. To establish these results, we directly deal with gradient flows instead of using variational discrete approximations which are often used in the study of gradient flows.

math.AP

Why Guided Dialog Policy Learning performs well? Understanding the role of adversarial learning and its alternative

Dialog policies, which determine a system's action based on the current state at each dialog turn, are crucial to the success of the dialog. In recent years, reinforcement learning (RL) has emerged as a promising option for dialog policy learning (DPL). In RL-based DPL, dialog policies are updated according to rewards. The manual construction of fine-grained rewards, such as state-action-based ones, to effectively guide the dialog policy is challenging in multi-domain task-oriented dialog scenarios with numerous state-action pair combinations. One way to estimate rewards from collected data is to train the reward estimator and dialog policy simultaneously using adversarial learning (AL). Although this method has demonstrated superior performance experimentally, it is fraught with the inherent problems of AL, such as mode collapse. This paper first identifies the role of AL in DPL through detailed analyses of the objective functions of dialog policy and reward estimator. Next, based on these analyses, we propose a method that eliminates AL from reward estimation and DPL while retaining its advantages. We evaluate our method using MultiWOZ, a multi-domain task-oriented dialog corpus.

cs.CL