SearcharxivSearch

arXiv subjects

Feihong Yang

Publications and source records attributed to Feihong Yang.

3 recordsLinked to original sources

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population whose induced game well approximates the full game. Existing PSRO variants typically expand the population using best responses to meta-strategies computed from restricted-game payoffs, which can lead to inefficient expansions that provide limited global improvement. We propose to guide population expansion by directly evaluating the post-expansion population quality. Specifically, we adopt Population Exploitability (PE) to measure how well a restricted strategy set represents the full game, and introduce a two-phase exploration--selection framework that explicitly minimizes PE during expansion. We instantiate this framework as Global PSRO, a practical DRL-based algorithm that efficiently generates candidate responses and estimates PE via parameter-sharing conditional neural networks. Experiments across multiple two-player zero-sum games show that Global PSRO achieves lower exploitability and approximates Nash equilibria with significantly fewer policy iterations than prior PSRO methods.

cs.AI

Distributed Scheduling at Non-Signalized Intersections with Mixed Cooperative and Non-Cooperative Vehicles

Intersection management with mixed cooperative and non-cooperative vehicles is crucial in next-generation transportation systems. For fully non-cooperative systems, a minimax scheduling framework was established, while it is inefficient in mixed systems as the benefit of cooperation is not exploited. This letter focuses on the efficient scheduling in mixed systems and proposes a two-stage policy that makes full use of the cooperation relation. Specifically, a long-horizon self-organization policy is first developed to optimize the passing order of cooperative vehicles in a distributed manner, which is proved convergent when inbound roads are sufficiently long. Then a short-horizon trajectory planning policy is proposed to improve the efficiency when an ego-vehicle faces both cooperative and non-cooperative vehicles, and its safety and efficiency are theoretically validated. Furthermore, numerical simulations verify that the proposed policies can effectively reduce the scheduling cost and improve the throughput for cooperative vehicles.

eess.SY

A Minimax Framework for Two-Agent Scheduling with Inertial Constraints

Autonomous agents are promising in applications such as intelligent transportation and smart manufacturing, and scheduling of agents has to take their inertial constraints into consideration. Most current researches require the obedience of all agents, which is hard to achieve in non-dedicated systems such as traffic intersections. In this article, we establish a minimax framework for the scheduling of two inertially constrained agents with no cooperation assumptions. Specifically, we first provide a unified and sufficient representation for various types of situation information, and define a state value function characterizing the agent's preference of states under a given situation. Then, the minimax control policy along with the calculation methods is proposed which optimizes the worst-case state value function at each step, and the safety guarantee of the policy is also presented. Furthermore, several generalizations are introduced on the applicable scenarios of the proposed framework. Numerical simulations show that the minimax control policy can reduce the largest scheduling cost by $13.4\%$ compared with queueing and following policies. Finally, the effects of decision period, observation period and inertial constraints are also numerically discussed.

eess.SY