SearcharxivSearch

arXiv subjects

Yunhan Huang

Publications and source records attributed to Yunhan Huang.

At least 19 recordsLinked to original sources

M2fNet: Multi-modal Forest Monitoring Network on Large-scale Virtual Dataset

Forest monitoring and education are key to forest protection, education and management, which is an effective way to measure the progress of a country's forest and climate commitments. Due to the lack of a large-scale wild forest monitoring benchmark, the common practice is to train the model on a common outdoor benchmark (e.g., KITTI) and evaluate it on real forest datasets (e.g., CanaTree100). However, there is a large domain gap in this setting, which makes the evaluation and deployment difficult. In this paper, we propose a new photorealistic virtual forest dataset and a multimodal transformer-based algorithm for tree detection and instance segmentation. To the best of our knowledge, it is the first time that a multimodal detection and segmentation algorithm is applied to large-scale forest scenes. We believe that the proposed dataset and method will inspire the simulation, computer vision, education, and forestry communities towards a more comprehensive multi-modal understanding.

cs.GR

LiDAR-Forest Dataset: LiDAR Point Cloud Simulation Dataset for Forestry Application

The popularity of LiDAR devices and sensor technology has gradually empowered users from autonomous driving to forest monitoring, and research on 3D LiDAR has made remarkable progress over the years. Unlike 2D images, whose focused area is visible and rich in texture information, understanding the point distribution can help companies and researchers find better ways to develop point-based 3D applications. In this work, we contribute an unreal-based LiDAR simulation tool and a 3D simulation dataset named LiDAR-Forest, which can be used by various studies to evaluate forest reconstruction, tree DBH estimation, and point cloud compression for easy visualization. The simulation is customizable in tree species, LiDAR types and scene generation, with low cost and high efficiency.

cs.RO

The Inverse Problem of Linear-Quadratic Differential Games: When is a Control Strategies Profile Nash?

This paper aims to formulate and study the inverse problem of non-cooperative linear quadratic games: Given a profile of control strategies, find cost parameters for which this profile of control strategies is Nash. We formulate the problem as a leader-followers problem, where a leader aims to implant a desired profile of control strategies among selfish players. In this paper, we leverage frequency-domain techniques to develop a necessary and sufficient condition on the existence of cost parameters for a given profile of stabilizing control strategies to be Nash under a given linear system. The necessary and sufficient condition includes the circle criterion for each player and a rank condition related to the transfer function of each player. The condition provides an analytical method to check the existence of such cost parameters, while previous studies need to solve a convex feasibility problem numerically to answer the same question. We develop an identity in frequency-domain representation to characterize the cost parameters, which we refer to as the Kalman equation. The Kalman equation reduces redundancy in the time-domain analysis that involves solving a convex feasibility problem. Using the Kalman equation, we also show the leader can enforce the same Nash profile by applying penalties on the shared state instead of penalizing the player for other players' actions to avoid the impression of unfairness.

math.OC

A Pursuit-Evasion Differential Game with Strategic Information Acquisition

This paper studies a two-person linear-quadratic-Gaussian pursuit-evasion differential game with costly but controlled information. One player can decide when to observe the other player's state. However, one observation of another player's state comes with two costs: the direct cost of observing and the implicit cost of exposing his state. We call games of this type a Pursuit-Evasion-Exposure-Concealment (PEEC) game. The PEEC game constitutes two types of strategies: The control strategies and the observation strategies. We fully characterize the Nash control strategies of the PEEC game using techniques such as completing squares and the calculus of variations. We show that the derivation of the Nash observation strategies and the Nash control strategies can be decoupled. We develop a set of necessary conditions that facilitate the numerical computation of the Nash observation strategies. We show, in theory, that players with less maneuverability prefer concealment to exposure. We also show that when the game's horizon goes to infinity, the Nash observation strategy is to observe periodically, and the expected distance between the pursuer and the evader goes to zero with a bounded second moment. We conducted a series of numerical experiments to study the proposed PEEC game. We illustrate the numerical results using both figures and animation. Numerical results show that the pursuer can maintain high-grade performance even when the number of observations is limited. We also show that an evader with low maneuverability can still escape if the evader increases his stealthiness.

eess.SY

Reinforcement Learning for Linear Quadratic Control is Vulnerable Under Cost Manipulation

In this work, we study the deception of a Linear-Quadratic-Gaussian (LQG) agent by manipulating the cost signals. We show that a small falsification of the cost parameters will only lead to a bounded change in the optimal policy. The bound is linear on the amount of falsification the attacker can apply to the cost parameters. We propose an attack model where the attacker aims to mislead the agent into learning a `nefarious' policy by intentionally falsifying the cost parameters. We formulate the attack's problem as a convex optimization problem and develop necessary and sufficient conditions to check the achievability of the attacker's goal. We showcase the adversarial manipulation on two types of LQG learners: the batch RL learner and the other is the adaptive dynamic programming (ADP) learner. Our results demonstrate that with only 2.296% of falsification on the cost data, the attacker misleads the batch RL into learning the 'nefarious' policy that leads the vehicle to a dangerous position. The attacker can also gradually trick the ADP learner into learning the same `nefarious' policy by consistently feeding the learner a falsified cost signal that stays close to the actual cost signal. The paper aims to raise people's awareness of the security threats faced by RL-enabled control systems.

eess.SY

Game-Theoretic Frameworks for Epidemic Spreading and Human Decision Making: A Review

This review presents and reviews various solved and open problems in developing, analyzing, and mitigating epidemic spreading processes under human decision-making. We provide a review of a range of epidemic models and explain the pros and cons of different epidemic models. We exhibit the art of coupling epidemic models and decision models in the existing literature. {More specifically, we provide answers to fundamental questions in human decision-making amid epidemics, including what interventions to take to combat the disease, who are decision-makers, when to take interventions, and how to make interventions.} Among many decision models, game-theoretic models have become increasingly crucial in modeling human responses/behavior amid epidemics in the last decade. In this review, we motivate the game-theoretic approach to human decision-making amid epidemics. This review provides an overview of the existing literature by developing a multi-dimensional taxonomy, which categorizes existing literature based on multiple dimensions, including 1) types of games, such as differential games, stochastic games, evolutionary games, and static games; 2) types of interventions, such as social distancing, vaccination, quarantine, taking antidotes, etc.; 3) the types of decision-makers, such as individuals, adversaries, and central authorities at different hierarchical levels. A fine-grained dynamic game framework is proposed to capture the essence of game-theoretic decision-making amid epidemics. We showcase three representative {frameworks} with unique ways of integrating game-theoretic decision-making into the epidemic models from a vast body of literature. {Each of the three framework has a unique way of modeling, conducting analytical analysis, and deriving results.} In the end, we identify several main open problems and research gaps left to be addressed and filled.

eess.SY

Reinforcement Learning for Feedback-Enabled Cyber Resilience

Digitization and remote connectivity have enlarged the attack surface and made cyber systems more vulnerable. As attackers become increasingly sophisticated and resourceful, mere reliance on traditional cyber protection, such as intrusion detection, firewalls, and encryption, is insufficient to secure the cyber systems. Cyber resilience provides a new security paradigm that complements inadequate protection with resilience mechanisms. A Cyber-Resilient Mechanism (CRM) adapts to the known or zero-day threats and uncertainties in real-time and strategically responds to them to maintain critical functions of the cyber systems in the event of successful attacks. Feedback architectures play a pivotal role in enabling the online sensing, reasoning, and actuation process of the CRM. Reinforcement Learning (RL) is an essential tool that epitomizes the feedback architectures for cyber resilience. It allows the CRM to provide sequential responses to attacks with limited or without prior knowledge of the environment and the attacker. In this work, we review the literature on RL for cyber resilience and discuss cyber resilience against three major types of vulnerabilities, i.e., posture-related, information-related, and human-related vulnerabilities. We introduce three application domains of CRMs: moving target defense, defensive cyber deception, and assistive human security technologies. The RL algorithms also have vulnerabilities themselves. We explain the three vulnerabilities of RL and present attack models where the attacker targets the information exchanged between the environment and the agent: the rewards, the state observations, and the action commands. We show that the attacker can trick the RL agent into learning a nefarious policy with minimum attacking effort. Lastly, we discuss the future challenges of RL for cyber security and resilience and emerging applications of RL-based CRMs.

cs.CR

Optimal Curing Strategy for Competing Epidemics Spreading over Complex Networks

Optimal curing strategy of suppressing competing epidemics spreading over complex networks is a critical issue. In this paper, we first establish a framework to capture the coupling between two epidemics, and then analyze the system's equilibrium states by categorizing them into three classes, and deriving their stability conditions. The designed curing strategy globally optimizes the trade-off between the curing cost and the severity of epidemics in the network. In addition, we provide structural results on the predictability of epidemic spreading by showing the existence and uniqueness of the solution. We also demonstrate the robustness of curing strategy by showing the continuity of epidemic severity with respect to the applied curing effort. A gradient descent algorithm based on a fixed-point iterative scheme is proposed to find the optimal curing strategy. Depending on the system parameters, the curing strategy can lead to switching between equilibria of the epidemic network as the control cost varies. Finally, we use case studies to corroborate and illustrate the obtained theoretical results.

eess.SY

Cross-Layer Coordinated Attacks on Cyber-Physical Systems: A LQG Game Framework with Controlled Observations

This work establishes a game-theoretic framework to study cross-layer coordinated attacks on cyber-physical systems (CPSs). The attacker can interfere with the physical process and launch jamming attacks on the communication channels simultaneously. At the same time, the defender can dodge the jamming by dispensing with observations. The generic framework captures a wide variety of classic attack models on CPSs. Leveraging dynamic programming techniques, we fully characterize the Subgame Perfect Equilibrium (SPE) control strategies. We also derive the SPE observation and jamming strategies and provide efficient computational methods to compute them. The results demonstrate that the physical and cyber attacks are coordinated and depend on each other. On the one hand, the control strategies are linear in the state estimate, and the estimate error caused by jamming attacks will induce performance degradation. On the other hand, the interactions between the attacker and the defender in the physical layer significantly impact the observation and jamming strategies. Numerical examples illustrate the interactions between the defender and the attacker through their observation and jamming strategies.

eess.SY

Transactive Resilience in Renewable Microgrids: A Contract-Theoretic Approach

Renewable energy-based microgrids play a critical role in future smart grids. Due to the uncertainties of renewable generations, the microgrids face potential risk of load shedding during operation. To address this problem, we propose a contract-based approach to enhance the resilience of microgrids. Specifically, in the framework, the microgrids who may not be self-efficient to meet their local demands can purchase the needed power from their connected microgrids by signing a contract that specifies the power price in advance. We leverage a principal-agent model to capture the energy trading relationships between the microgrids through a resilience as a service (RaaS) paradigm. By focusing on the incentive compatible and individual rational constraints of the service requester, the service provider designs the optimal contracts for the transactive resilience that yields the largest payoff despite the incomplete information. We characterize analytical solutions of the optimal contracts for several scenarios where the service requester has various options on its hidden actions. Numerical simulations are used to illustrate and corroborate the obtained results.

eess.SY

Assets Defending Differential Games with Partial Information and Selected Observations

In this paper, we consider a linear-quadratic-Gaussian defending assets differential game (DADG) where the attacker and the defender do not know each other's state information while they know the trajectory of a moving asset. Both players can choose to observe the other player's state information by paying a cost. The defender and the attacker have to craft both control strategies and observation strategies. We obtain a closed-form feedback solution that characterizes the Nash control strategies. We show that the trajectory of the asset does not affect both players' observation choices. Moreover, we show that the observation choices of the defender and the attacker can be decoupled and the Nash observation strategies can be found by solving two independent optimization problems. A set of necessary conditions is developed to characterize the optimal observation instances. Based on the necessary conditions, an effective algorithm is proposed to numerically compute the optimal observation instances. A case study is presented to demonstrate the effectiveness of the optimal observation instances.

eess.SY

Self-Triggered Markov Decision Processes

In this paper, we study Markov Decision Processes (MDPs) with self-triggered strategies, where the idea of self-triggered control is extended to more generic MDP models. This extension broadens the application of self-triggering policies to a broader range of systems. We study the co-design problems of the control policy and the triggering policy to optimize two pre-specified cost criteria. The first cost criterion is introduced by incorporating a pre-specified update penalty into the traditional MDP cost criteria to reduce the use of communication resources. Under this criteria, a novel dynamic programming (DP) equation called DP equation with optimized lookahead to proposed to solve for the self-triggering policy under this criteria. The second self-triggering policy is to maximize the triggering time while still guaranteeing a pre-specified level of sub-optimality. Theoretical underpinnings are established for the computation and implementation of both policies. Through a gridworld numerical example, we illustrate the two policies' effectiveness in reducing sources consumption and demonstrate the trade-offs between resource consumption and system performance.

eess.SY

Infinite-Horizon Linear-Quadratic-Gaussian Control with Costly Measurements

In this paper, we consider an infinite horizon Linear-Quadratic-Gaussian control problem with controlled and costly measurements. A control strategy and a measurement strategy are co-designed to optimize the trade-off among control performance, actuating costs, and measurement costs. We address the co-design and co-optimization problem by establishing a dynamic programming equation with controlled lookahead. By leveraging the dynamic programming equation, we fully characterize the optimal control strategy and the measurement strategy analytically. The optimal control is linear in the state estimate that depends on the measurement strategy. We prove that the optimal measurement strategy is independent of the measured state and is periodic. And the optimal period length is determined by the cost of measurements and system parameters. We demonstrate the potential application of the co-design and co-optimization problem in an optimal self-triggered control paradigm. Two examples are provided to show the effectiveness of the optimal measurement strategy in reducing the overhead of measurements while keeping the system performance.

eess.SY

Manipulating Reinforcement Learning: Poisoning Attacks on Cost Signals

This chapter studies emerging cyber-attacks on reinforcement learning (RL) and introduces a quantitative approach to analyze the vulnerabilities of RL. Focusing on adversarial manipulation on the cost signals, we analyze the performance degradation of TD($λ$) and $Q$-learning algorithms under the manipulation. For TD($λ$), the approximation learned from the manipulated costs has an approximation error bound proportional to the magnitude of the attack. The effect of the adversarial attacks on the bound does not depend on the choice of $λ$. In $Q$-learning, we show that $Q$-learning algorithms converge under stealthy attacks and bounded falsifications on cost signals. We characterize the relation between the falsified cost and the $Q$-factors as well as the policy learned by the learning agent which provides fundamental limits for feasible offensive and defensive moves. We propose a robust region in terms of the cost within which the adversary can never achieve the targeted policy. We provide conditions on the falsified cost which can mislead the agent to learn an adversary's favored policy. A case study of TD($λ$) learning is provided to corroborate the results.

cs.LG

Optimal Demand Response and Supply Schedule Under Market-Driven Price: A Stackelberg Dynamic Game Approach

In this work, we use a Stackelberg infinite discrete-time dynamic game model to study the optimal supply schedule and the optimal demand response under a market-driven dynamic price. A two-layer optimization framework is established. At the lower layer, for each user, different appliances are scheduled for energy consumption. For enegy provider, different generators are utilized for energy generation. At the upper level, with the supplier acting as a leader and the users acting as followers, a Stackelberg dynamic game is used to capture the interaction among the energy provider and the users where the energy provider and the users care only about their own cost. We analyze the one-leader-N-followers Stackelberg dynamic game and characterize the Stackelberg equilibrium. We provide a closed-form Nash solution of the optimal dynamic demand response problem when the supply is announced. A set of linear constraints is developed to characterize the Stackeberg equilibrium. Simulation results show that the price is driven to a reasonable value. Also, the total demand and the supply is balanced.

eess.SY

Dynamic Games for Secure and Resilient Control System Design

Modern control systems are featured by their hierarchical structure composing of cyber, physical, and human layers. The intricate dependencies among multiple layers and units of modern control systems require an integrated framework to address cross-layer design issues related to security and resilience challenges. To this end, game theory provides a bottom-up modeling paradigm to capture the strategic interactions among multiple components of the complex system and enables a holistic view to understand and design cyber-physical-human control systems. In this review, we first provide a multi-layer perspective toward increasingly complex and integrated control systems and then introduce several variants of dynamic games for modeling different layers of control systems. We present game-theoretic methods for understanding the fundamental tradeoffs of robustness, security, and resilience and developing a clean-slate cross-layer approach to enhance the system performance in various adversarial environments. This review also includes three quintessential research problems that represent three research directions where dynamic game approaches can bridge between multiple research areas and make significant contributions to the design of modern control systems. The paper is concluded with a discussion on emerging areas of research that crosscut dynamic games and control systems.

eess.SY

Deceptive Reinforcement Learning Under Adversarial Manipulations on Cost Signals

This paper studies reinforcement learning (RL) under malicious falsification on cost signals and introduces a quantitative framework of attack models to understand the vulnerabilities of RL. Focusing on $Q$-learning, we show that $Q$-learning algorithms converge under stealthy attacks and bounded falsifications on cost signals. We characterize the relation between the falsified cost and the $Q$-factors as well as the policy learned by the learning agent which provides fundamental limits for feasible offensive and defensive moves. We propose a robust region in terms of the cost within which the adversary can never achieve the targeted policy. We provide conditions on the falsified cost which can mislead the agent to learn an adversary's favored policy. A numerical case study of water reservoir control is provided to show the potential hazards of RL in learning-based control systems and corroborate the results.

cs.LG

Continuous-Time Markov Decision Processes with Controlled Observations

In this paper, we study a continuous-time discounted jump Markov decision process with both controlled actions and observations. The observation is only available for a discrete set of time instances. At each time of observation, one has to select an optimal timing for the next observation and a control trajectory for the time interval between two observation points. We provide a theoretical framework that the decision maker can utilize to find the optimal observation epochs and the optimal actions jointly. Two cases are investigated. One is gated queueing systems in which we explicitly characterize the optimal action and the optimal observation where the optimal observation is shown to be independent of the state. Another is the inventory control problem with Poisson arrival process in which we obtain numerically the optimal action and observation. The results show that it is optimal to observe more frequently at a region of states where the optimal action adapts constantly.

math.OC