SearcharxivSearch

arXiv subjects

Huiwen Yan

Publications and source records attributed to Huiwen Yan.

5 recordsLinked to original sources

Markov Potential Game and Multi-Agent Reinforcement Learning for Autonomous Driving

Autonomous driving (AD) requires safe and reliable decision-making among interacting agents, e.g., vehicles, bicycles, and pedestrians. Multi-agent reinforcement learning (MARL) modeled by Markov games (MGs) provides a suitable framework to characterize such agents' interactions during decision-making. Nash equilibria (NEs) are often the desired solution in an MG. However, it is typically challenging to compute an NE in general-sum games, unless the game is a Markov potential game (MPG), which ensures the NE attainability under a few learning algorithms such as gradient play. However, it has been an open question how to construct an MPG and whether these construction rules are suitable for AD applications. In this paper, we provide sufficient conditions under which an MG is an MPG and show that these conditions can accommodate general driving objectives for autonomous vehicles (AVs) using highway forced merge scenarios as illustrative examples. A parameter-sharing neural network (NN) structure is designed to enable decentralized policy execution. The trained driving policy from MPGs is evaluated in both simulated and naturalistic traffic datasets. Comparative studies with single-agent RL and with human drivers whose behaviors are recorded in the traffic datasets are reported, respectively.

eess.SY

Directed-MAML: Meta Reinforcement Learning Algorithm with Task-directed Approximation

Model-Agnostic Meta-Learning (MAML) is a versatile meta-learning framework applicable to both supervised learning and reinforcement learning (RL). However, applying MAML to meta-reinforcement learning (meta-RL) presents notable challenges. First, MAML relies on second-order gradient computations, leading to significant computational and memory overhead. Second, the nested structure of optimization increases the problem's complexity, making convergence to a global optimum more challenging. To overcome these limitations, we propose Directed-MAML, a novel task-directed meta-RL algorithm. Before the second-order gradient step, Directed-MAML applies an additional first-order task-directed approximation to estimate the effect of second-order gradients, thereby accelerating convergence to the optimum and reducing computational cost. Experimental results demonstrate that Directed-MAML surpasses MAML-based baselines in computational efficiency and convergence speed in the scenarios of CartPole-v1, LunarLander-v2 and two-vehicle intersection crossing. Furthermore, we show that task-directed approximation can be effectively integrated into other meta-learning algorithms, such as First-Order Model-Agnostic Meta-Learning (FOMAML) and Meta Stochastic Gradient Descent(Meta-SGD), yielding improved computational efficiency and convergence speed.

cs.LG

Markov Potential Game Construction and Multi-Agent Reinforcement Learning with Applications to Autonomous Driving

Markov games (MGs) provide a mathematical foundation for multi-agent reinforcement learning (MARL), enabling self-interested agents to learn their optimal policies while interacting with others in a shared environment. However, due to the complexities of an MG problem, seeking (Markov perfect) Nash equilibrium (NE) is often very challenging for a general-sum MG. Markov potential games (MPGs), which are a special class of MGs, have appealing properties such as guaranteed existence of pure NEs and guaranteed convergence of gradient play algorithms, thereby leading to desirable properties for many MARL algorithms in their NE-seeking processes. However, the question of how to construct MPGs has remained open. This paper provides sufficient conditions on the reward design and on the Markov decision process (MDP), under which an MG is an MPG. Numerical results on autonomous driving applications are reported.

eess.SY

Dynkin Game of Convertible Bonds and Their Optimal Strategy

This paper studies the valuation and optimal strategy of convertible bonds as a Dynkin game by using the reflected backward stochastic differential equation method and the variational inequality method. We first reduce such a Dynkin game to an optimal stopping time problem with state constraint, and then in a Markovian setting, we investigate the optimal strategy by analyzing the properties of the corresponding free boundary, including its position, asymptotics, monotonicity and regularity. We identify situations when call precedes conversion, and vice versa. Moreover, we show that the irregular payoff results in the possibly non-monotonic conversion boundary. Surprisingly, the price of the convertible bond is not necessarily monotonic in time: it may even increase when time approaches maturity.

q-fin.MF

Indifference Pricing and Hedging in a Multiple-Priors Model with Trading Constraints

This paper considers utility indifference valuation of derivatives under model uncertainty and trading constraints, where the utility is formulated as an additive stochastic differential utility of both intertemporal consumption and terminal wealth, and the uncertain prospects are ranked according to a multiple-priors model of Chen and Epstein (2002). The price is determined by two optimal stochastic control problems (mixed with optimal stopping time in the case of American option) of forward-backward stochastic differential equations. By means of backward stochastic differential equation and partial differential equation methods, we show that both bid and ask prices are closely related to the Black-Scholes risk-neutral price with modified dividend rates. The two prices will actually coincide with each other if there is no trading constraint or the model uncertainty disappears. Finally, two applications to European option and American option are discussed.

q-fin.MF