SearcharxivSearch

arXiv subjects

Shanyu Han

Publications and source records attributed to Shanyu Han.

4 recordsLinked to original sources

Robust Bayesian Dynamic Programming for On-policy Risk-sensitive Reinforcement Learning

We propose a novel framework for risk-sensitive reinforcement learning (RSRL) that incorporates robustness against transition uncertainty. We define two distinct yet coupled risk measures: an inner risk measure addressing state and cost randomness and an outer risk measure capturing transition dynamics uncertainty. Our framework unifies and generalizes most existing RL frameworks by permitting general coherent risk measures for both inner and outer risk measures. Within this framework, we construct a risk-sensitive robust Markov decision process (RSRMDP), derive its Bellman equation, and provide error analysis under a given posterior distribution. We further develop a Bayesian Dynamic Programming (Bayesian DP) algorithm that alternates between posterior updates and value iteration. The approach employs an estimator for the risk-based Bellman operator that combines Monte Carlo sampling with convex optimization, for which we prove strong consistency guarantees. Furthermore, we demonstrate that the algorithm converges to a near-optimal policy in the training environment and analyze both the sample complexity and the computational complexity under the Dirichlet posterior and CVaR. Finally, we validate our approach through two numerical experiments. The results exhibit excellent convergence properties while providing intuitive demonstrations of its advantages in both risk-sensitivity and robustness. Empirically, we further demonstrate the advantages of the proposed algorithm through an application on option hedging.

q-fin.RM

Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions

We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.

q-fin.MF

Equilibrium in Style: A Modeling Framework on the Cash Flow and the Life Cycle of a Consumer Store

The consumer store is ubiquitous and plays an important role in our everyday lives. It is an open question why stores usually have such short life cycles (typically around 3 years in China). This paper proposes a theoretical framework based on an equilibrium in style supply of stores and style demand of consumers to characterize store cash flow (revenue), leading to a strong explanation of this puzzle. In our model, we derive that the preference shifting of consumers is the main reason for the cash flow decreasing to its break-even line over time, while the visibility broadening leads to initial growth, resulting in rainbow-shaped cash flow and its life cycle. Moreover, the intensified spatial competition will lead to an unexpected decrease in the store's cash flow, or even closure. We calibrate our model with proprietary data of three Chinese stores from three representative industries and study the relationship between customers' preference shifting and cash flow. To our knowledge, there have been no prior attempts to quantitatively model the life cycle of the store.

econ.TH

An experimental and theoretical investigation of the N + C2 reaction at low temperature

Rate constants for the N + C2 reaction have been measured in a continuous supersonic flow reactor over the range 57 K to 296 K by the relative rate technique employing the N + OH - H + NO reaction as a reference. Excess concentrations of atomic nitrogen were produced by the microwave discharge method and C2 and OH radicals were created by the in-situ pulsed laser photolysis of precursor molecules C2Br4 and H2O2 respectively. In parallel, quantum dynamics calculations were performed based on an accurate global potential energy surfaces for the three lowest lying quartet states of the C2N molecule. The 14A" potential energy surface is barrierless, having two deep potential wells corresponding to the NCC and CNC intermediates. Both the experimental and theoretical work show that the rate constant decreases to low temperature, although the experimentally measured values fall more rapidly than the theoretical ones except at the lowest temperatures. Astrochemical simulations indicate that this reaction could be the dominant source of CN in dense interstellar clouds.

astro-ph.GA