SearcharxivSearch

arXiv subjects

Jingmin He

Publications and source records attributed to Jingmin He.

2 recordsLinked to original sources

Reinforcement Learning for Dividend Optimization in Partially Observed Regime-Switching Diffusion Model

This paper studies the optimal dividend problem with a bounded payout rate in a partially observed regime-switching diffusion model, where, in practice, the market regime is unobserved and key model parameters are unknown. To address this partial-information setting, we propose a continuous-time reinforcement learning (RL) approach within an exploratory (entropy-regularized) stochastic control framework for discounted dividends under regime switching. The associated exploratory Hamilton-Jacobi-Bellman (HJB) system admits semi-analytical characterizations of the value function and the optimal exploratory dividend policy, determined by two unknown functions solving two ordinary differential equations (ODEs) together with positive real roots of the induced quadratic equations. Exploiting this structure, we introduce parametric families for both the value function and the policy, using low-degree polynomial approximations to the ODE solutions. We then develop an actor-critic RL algorithm to learn the optimal exploratory policy through interactions with the market environment: it performs belief-state filtering from observed data and iterates policy evaluation and policy improvement online to refine the policy. Numerical experiments demonstrate strong out-of-sample performance of the learned dividend policies.

math.OC

Optimal Dividend Control with Transaction Costs under Exponential Parisian Ruin for a Refracted Levy Risk Model

This paper concerns an optimal impulse control problem associated with a refracted L\'{e}vy process, involving the reduction of reserves to a predetermined level whenever they exceed a specified threshold. The ruin time is determined by Parisian exponential delays and limited by a lower ultimate bankrupt barrier. We initially obtained the necessary and sufficient conditions for the value function and the optimal impulse control policy. Given a candidate for the optimal strategy, the corresponding expected discounted dividend function is subsequently formulated in terms of the Parisian refracted scale function, which is employed to measure the expected discounted utility of the impulse control. Then, the optimality of the proposed impulse control is verified using the HJB inequalities, and a monotonicity-based criterion is established to identify the admissible region of optimal thresholds, which serves as the basis for the numerical computation of their optimal levels. Finally, we present applications and numerical examples related to Brownian risk process and Cram\'{e}r-Lundberg process with exponential claims, demonstrating the uniqueness of the optimal impulse strategy and exploring its sensitivity to parameters.

math.OC