SearcharxivSearch

arXiv subjects

Zongying Shi

Publications and source records attributed to Zongying Shi.

9 recordsLinked to original sources

Reach-avoid games for players with damped double integrator dynamics

This paper investigates a reach-avoid game between two players with damped double integrator dynamics. An optimal state-feedback strategy is derived using a differential game framework combined with geometric analysis. To facilitate the analysis, we introduce the concept of multiple reachable region by characterizing the motion of players with damped double integrator dynamics. Based on this, a new type of the attackers dominance region is introduced. We show that distinct strategies are required depending on the location of the terminal position within different areas of the attacker's dominance region. Furthermore, we prove that the proposed strategies satisfy the necessary condition for optimality. Numerical simulations are provided to illustrate the conclusions.

eess.SY

Multiplayer Homicidal Chauffeur Reach-Avoid Games: A Pursuit Enclosure Function Approach

This paper presents a multiplayer Homicidal Chauffeur reach-avoid differential game, which involves Dubins-car pursuers and simple-motion evaders. The goal of the pursuers is to cooperatively protect a planar convex region from the evaders, who strive to reach the region. We propose a cooperative strategy for the pursuers based on subgames for multiple pursuers against one evader and optimal task allocation. We introduce pursuit enclosure functions (PEFs) and propose a new enclosure region pursuit (ERP) winning approach that supports forward analysis for the strategy synthesis in the subgames. We show that if a pursuit coalition is able to defend the region against an evader under the ERP winning, then no more than two pursuers in the coalition are necessarily needed. We also propose a steer-to-ERP approach to certify the ERP winning and synthesize the ERP winning strategy. To implement the strategy, we introduce a positional PEF and provide the necessary parameters, states, and strategies that ensure the ERP winning for both one pursuer and two pursuers against one evader. Additionally, we formulate a binary integer program using the subgame outcomes to maximize the captured evaders in the ERP winning for the pursuit task allocation. Finally, we propose a multiplayer receding-horizon strategy where the ERP winnings are checked in each horizon, the task is allocated, and the strategies of the pursuers are determined. Numerical examples are provided to illustrate the results.

eess.SY

Evaluation and Learning in Two-Player Symmetric Games via Best and Better Responses

Artificial intelligence and robotic competitions are accompanied by a class of game paradigms in which each player privately commits a strategy to a game system which simulates the game using the collected joint strategy and then returns payoffs to players. This paper considers the strategy commitment for two-player symmetric games in which the players' strategy spaces are identical and their payoffs are symmetric. First, we introduce two digraph-based metrics at a meta-level for strategy evaluation in two-agent reinforcement learning, grounded on sink equilibrium. The metrics rank the strategies of a single player and determine the set of strategies which are preferred for the private commitment. Then, in order to find the preferred strategies under the metrics, we propose two variants of the classical learning algorithm self-play, called strictly best-response and weakly better-response self-plays. By modeling learning processes as walks over joint-strategy response digraphs, we prove that the learnt strategies by two variants are preferred under two metrics, respectively. The preferred strategies under both two metrics are identified and adjacency matrices induced by one metric and one variant are connected. Finally, simulations are provided to illustrate the results.

cs.GT

Multiplayer Homicidal Chauffeur Reach-Avoid Games via Guaranteed Winning Strategies

This paper studies a planar multiplayer Homicidal Chauffeur reach-avoid differential game, where each pursuer is a Dubins car and each evader has simple motion. The pursuers aim to protect a goal region cooperatively from the evaders. Due to the high-dimensional strategy space among pursuers, we decompose the whole game into multiple one-pursuer-one-evader subgames, each of which is solved in an analytical approach instead of solving Hamilton-Jacobi-Isaacs equations. For each subgame, an evasion region (ER) is introduced, based on which a pursuit strategy guaranteeing the winning of a simple-motion pursuer under specific conditions is proposed. Motivated by the simple-motion pursuer, a strategy for a Dubins-car pursuer is proposed when the pursuer-evader configuration satisfies separation condition (SC) and interception orientation (IO). The necessary and sufficient condition on capture radius, minimum turning radius and speed ratio to guarantee the pursuit winning is derived. When the IO is not satisfied (Non-IO), a heading adjustment pursuit strategy is proposed, and the condition to achieve IO within a finite time, is given. Then, a two-step pursuit strategy is proposed for the SC and Non-IO case. A non-convex optimization problem is introduced to give a condition guaranteeing the winning of the pursuer. A polynomial equation gives a lower bound of the non-convex problem, providing a sufficient and efficient pursuit winning condition. Finally, these pairwise outcomes are collected for the pursuer-evader matching. Simulations are provided to illustrate the theoretical results.

cs.GT

Policy Evaluation and Seeking for Multi-Agent Reinforcement Learning via Best Response

This paper introduces two metrics (cycle-based and memory-based metrics), grounded on a dynamical game-theoretic solution concept called sink equilibrium, for the evaluation, ranking, and computation of policies in multi-agent learning. We adopt strict best response dynamics (SBRD) to model selfish behaviors at a meta-level for multi-agent reinforcement learning. Our approach can deal with dynamical cyclical behaviors (unlike approaches based on Nash equilibria and Elo ratings), and is more compatible with single-agent reinforcement learning than alpha-rank which relies on weakly better responses. We first consider settings where the difference between largest and second largest underlying metric has a known lower bound. With this knowledge we propose a class of perturbed SBRD with the following property: only policies with maximum metric are observed with nonzero probability for a broad class of stochastic games with finite memory. We then consider settings where the lower bound for the difference is unknown. For this setting, we propose a class of perturbed SBRD such that the metrics of the policies observed with nonzero probability differ from the optimal by any given tolerance. The proposed perturbed SBRD addresses the opponent-induced non-stationarity by fixing the strategies of others for the learning agent, and uses empirical game-theoretic analysis to estimate payoffs for each strategy profile obtained due to the perturbation.

cs.GT

Matching-Based Capture Strategies for 3D Heterogeneous Multiplayer Reach-Avoid Differential Games

This paper studies a 3D multiplayer reach-avoid differential game with a goal region and a play region. Multiple pursuers defend the goal region by consecutively capturing multiple evaders in the play region. The players have heterogeneous moving speeds and the pursuers have heterogeneous capture radii. Since this game is hard to analyze directly, we decompose the whole game as many subgames which involve multiple pursuers and only one evader. Then, these subgames are used as a building block for the pursuer-evader matching. First, for multiple pursuers and one evader, we introduce an evasion space (ES) method characterized by a potential function to construct a guaranteed pursuer winning strategy. Then, based on this strategy, we develop conditions to determine whether a pursuit team can guard the goal region against one evader. It is shown that in 3D, if a pursuit team is able to defend the goal region against an evader, then at most three pursuers in the team are necessarily needed. We also compute the value function of the Hamilton-Jacobi-Isaacs (HJI) equation for a special subgame of degree. To capture the maximum number of evaders in the open-loop sense, we formulate a maximum bipartite matching problem with conflict graph (MBMC). We show that the MBMC is NP-hard and design a polynomial-time constant-factor approximation algorithm to solve it. Finally, we propose a receding horizon strategy for the pursuit team where in each horizon an MBMC is solved and the strategies of the pursuers are given. We also extend our results to the case of a bounded convex play region where the evaders escape through an exit. Two numerical examples are provided to demonstrate the obtained results.

math.OC

Guarding a Subspace in High-Dimensional Space with Two Defenders and One Attacker

This paper considers a subspace guarding game in high-dimensional space which consists of a play subspace and a target subspace. Two faster defenders cooperate to protect the target subspace by capturing an attacker which strives to enter the target subspace from the play subspace without being captured. A closed-form solution is provided from the perspectives of kind and degree. Contributions of the work include the use of the attack subspace (AS) method to construct the barrier, by which the game winner can be perfectly predicted before the game starts. In addition to this inclusion, with the priori information about the game result, a critical payoff function is designed when the defenders can win the game. Then, the optimal strategy for each player is explicitly reformulated as a saddle-point equilibrium. Finally, we apply these theoretical results to a half-space guarding game in three-dimensional space. Since the whole achieved developments are analytical, they require a little memory without computational burden and allow for real-time updates, beyond the capacity of traditional Hamilton-Jacobi-Isaacs method. It is worth noting that this is the first time in the current work to consider the target guarding games for arbitrary high-dimensional space, and in a fully analytical form.

eess.SY

Construction of the Barrier for Reach-Avoid Differential Games in Three-Dimensional Space with Four Equal-Speed Players

This paper considers a reach-avoid differential game in three-dimensional space with four equal-speed players. A plane divides the game space into a play subspace and a goal subspace. The evader aims at entering the goal subspace while three pursuers cooperate to prevent that by capturing the evader. A complete, closed-form barrier for this differential game is provided, by which the game winner can be perfectly predicted before the game starts. All possible cooperations among three pursuers are considered and thus the guaranteed winning for each team is a prior. Furthermore, an algorithm is designed to compute the barrier for multiple pursuers of any numbers and any initial configurations. More realistically, since the whole achieved developments are analytical, they require a little memory without computational burden and allow for real-time updates, beyond the capacity of traditional Hamilton-Jacobi-Isaacs method.

math.OC

Task Assignment for Multiplayer Reach-Avoid Games in Convex Domains via Analytical Barriers

This work considers a multiplayer reach-avoid game between two adversarial teams in a general convex domain which consists of a target region and a play region. The evasion team, initially lying in the play region, aims to send as many its team members into the target region as possible, while the pursuit team with its team members initially distributed in both play region and target region, strives to prevent that by capturing the evaders. We aim at investigating a task assignment about the pursuer-evader matching, which can maximize the number of the evaders who can be captured before reaching the target region safely when both teams play optimally. To address this, two winning regions for a group of pursuers to intercept an evader are determined by constructing an analytical barrier which divides these two parts. Then, a task assignment to guarantee the most evaders intercepted is provided by solving a simplified 0-1 integer programming instead of a non-deterministic polynomial problem, easing the computation burden dramatically. It is worth noting that except the task assignment, the whole analysis is analytical. Finally, simulation results are also presented.

math.OC