SearcharxivSearch

arXiv subjects

Konatsu Miyamoto

Publications and source records attributed to Konatsu Miyamoto.

2 recordsLinked to original sources

Path-dependent Poisson random measures and stochastic integrals constructed from general point processes

In this paper, we consider an extension of the Poisson random measure for the formulation of continuous-time reinforcement learning, such that both the frequency and the width of the jumps depend on the path. Starting from a general point process, we define a new Poisson random measure as limit of the linear sum of these counting processes, and name it the Mesgaki random measure. We also construct its Stochastic integral and Itô's formula.

math.PR

Convergence of Q-value in case of Gaussian rewards

In this paper, as a study of reinforcement learning, we converge the Q function to unbounded rewards such as Gaussian distribution. From the central limit theorem, in some real-world applications it is natural to assume that rewards follow a Gaussian distribution , but existing proofs cannot guarantee convergence of the Q-function. Furthermore, in the distribution-type reinforcement learning and Bayesian reinforcement learning that have become popular in recent years, it is better to allow the reward to have a Gaussian distribution. Therefore, in this paper, we prove the convergence of the Q-function under the condition of $E[r(s,a)^2]<\infty$, which is much more relaxed than the existing research. Finally, as a bonus, a proof of the policy gradient theorem for distributed reinforcement learning is also posted.

math.OC