SearcharxivSearch

arXiv subjects

Ruolan He

Publications and source records attributed to Ruolan He.

2 recordsLinked to original sources

A Principal-Agent Mean-Field Game Model of Insurance with Risk Interdependence

We study an insurance contract-design problem under moral hazard, endogenous participation, and strategic risk interdependence. Because the resulting $N$-agent game suffers from the curse of dimensionality, we approximate the strategic interactions via a heterogeneous mean-field game. We rigorously establish the existence of a lower-level mean-field Nash equilibrium using measurable selection arguments and the Kakutani fixed-point theorem. By proving the $L^1$-Lipschitz continuity of the aggregate participation threshold, we further establish equilibrium uniqueness via a contraction mapping. We then embed this mean-field response into the insurer's upper-level Stackelberg optimization problem. We formulate the objective through general performance envelopes to accommodate potential equilibrium multiplicity, proving the existence of upper-level $\varepsilon$-optimal contracts, and demonstrating the existence of an exact Stackelberg equilibrium under the uniqueness regime. We conclude by extending the model to finite contract menus, providing numerical evidence that multi-contract screening improves the principal's expected payoff in interdependent risk environments.

math.OC

Thompson Sampling Algorithm for Stochastic Games

We study a stochastic differential game with $N$ competitive players in a linear-quadratic framework with ergodic cost, where $d$-dimensional diffusion processes govern the state dynamics with an unknown common drift (matrix). Assuming a Gaussian prior on the drift, we use filtering techniques to update its posterior estimates. Based on these estimates, we propose a Thompson-sampling-based algorithm with dynamic episode lengths to approximate strategies. We show that the Bayesian regret for each player has an error bound of order $O(\sqrt{T\log(T)})$, where $T$ is the time-horizon, independent of the number of players. This implies that average regret per unit time goes to zero. Finally, we prove that the algorithm results in a Nash equilibrium.

math.OC