SearcharxivSearch

arXiv subjects

Zhikun She

Publications and source records attributed to Zhikun She.

7 recordsLinked to original sources

MSACL: Multi-Step Actor-Critic Learning with Lyapunov Certificates for Exponentially Stabilizing Control

For stabilizing control tasks, model-free reinforcement learning (RL) approaches face numerous challenges, particularly regarding the issues of effectiveness and efficiency in complex high-dimensional environments with limited training data. To address these challenges, we propose Multi-Step Actor-Critic Learning with Lyapunov Certificates (MSACL), a novel approach that integrates exponential stability into off-policy maximum entropy reinforcement learning (MERL). In contrast to existing RL-based approaches that depend on elaborate reward engineering and single-step constraints, MSACL adopts intuitive reward design and exploits multi-step samples to enable exploratory actor-critic learning. Specifically, we first introduce Exponential Stability Labels (ESLs) to categorize training samples and propose a $\lambda$-weighted aggregation mechanism to learn Lyapunov certificates. Based on these certificates, we further design a stability-aware advantage function to guide policy optimization, thereby promoting rapid Lyapunov descent and robust state convergence. We evaluate MSACL across six benchmarks, comprising four stabilizing and two high-dimensional tracking tasks. Experimental results demonstrate its consistent performance improvements over both standard RL baselines and state-of-the-art Lyapunov-based RL algorithms. Beyond rapid convergence, MSACL exhibits robustness against environmental uncertainties and generalization to unseen reference signals. The source code and benchmarking environments are available at \href{https://github.com/YuanZhe-Xing/MSACL}{https://github.com/YuanZhe-Xing/MSACL}.

cs.LG

Explaining human cooperation through a dual mechanism of individual and social learning

Cooperation on social networks is crucial for understanding human survival and development. Although network structure has been found to significantly influence cooperation, human experiments have observed different cooperation phenomena under similar conditions. While evidence suggests that these differences arise from human exploration, our understanding of its impact mechanisms and characteristics remains limited. Here, we seek to formalize human exploration as an individual learning process involving trial and reflection, and integrate social learning to examine how their interdependence shapes cooperation. We find that individual learning can alter neighbor imitation tendencies, and the resulting shifts in the local cooperative environment feed back into the experiential cognition that guides individual learning. This coupled dynamic makes the ability of social networks to promote cooperation largely dependent on whether individuals focus on long-term payoffs, and exhibits a series of characteristics that can explain previously unexplained and seemingly contradictory cooperation phenomena. Surprisingly, individual learning can promote cooperation more than social learning when its probability is negatively correlated with payoffs, a mechanism rooted in the psychological tendency to avoid trial-and-error when individuals are satisfied with their current payoffs. These results explain the contradictory cooperation phenomenon by accounting for decision preferences and cognitive processes underlying exploration, bridging the gap between theoretical research and reality.

physics.soc-ph

Navigating Robot Swarm Through a Virtual Tube with Flow-Adaptive Distribution Control

With the rapid development of robot swarm technology and its diverse applications, navigating robot swarms through complex environments has emerged as a critical research direction. To ensure safe navigation and avoid potential collisions with obstacles, the concept of virtual tubes has been introduced to define safe and navigable regions. However, current control methods in virtual tubes face the congestion issues, particularly in narrow ones with low throughput. To address these challenges, we first propose a novel control method that combines a modified artificial potential field (APF) for swarm navigation and density feedback control for distribution regulation. Then we generate a global velocity field that not only ensures collision-free navigation but also achieves locally input-to-state stability (LISS) for density tracking. Finally, numerical simulations and realistic applications validate the effectiveness and advantages of the proposed method in navigating robot swarms through narrow virtual tubes.

cs.RO

Urban traffic resilience control -- An ecological resilience perspective

Urban traffic resilience has gained increased attention, with most studies adopting an engineering perspective that assumes a single optimal equilibrium and prioritizes local recovery. On the other hand, systems may possess multiple metastable states, and ecological resilience is the ability to switch between these states according to perturbations. Control strategies from these two resilience perspectives yield distinct outcomes. In fact, ecological resilience oriented control has rarely been viewed in urban traffic, despite the fact that traffic system is a complex system in highly uncertain environment with possible multiple metastable states. This absence highlights the necessity for urban traffic ecological resilience definition. To bridge this gap, we defines urban traffic ecological resilience as the ability to absorb uncertain perturbations by shifting to alternative states. The goal is to generate a system with greater adaptability, without necessarily returning to the original equilibrium. Our control framework comprises three aspects: portraying the recoverable scopes; designing alternative steady states; and controlling system to shift to alternative steady states for adapting large disturbances. Among them, the recoverable scopes are portrayed by attraction region; the alternative steady states are set close to the optimal state and outside the attraction region of the original equilibrium; the controller needs to ensure the local stability of the alternative steady states, without changing the trajectories inside the attraction region of the original equilibrium. Comparisons with classical engineering resilience oriented urban traffic resilience control schemes show that, proposed ecological resilience oriented control schemes can generate greater resilience. These results will contribute to the fundamental theory of future resilient intelligent transportation system.

nlin.AO

High-Speed Interception Multicopter Control by Image-based Visual Servoing

In recent years, reports of illegal drones threatening public safety have increased. For the invasion of fully autonomous drones, traditional methods such as radio frequency interference and GPS shielding may fail. This paper proposes a scheme that uses an autonomous multicopter with a strapdown camera to intercept a maneuvering intruder UAV. The interceptor multicopter can autonomously detect and intercept intruders moving at high speed in the air. The strapdown camera avoids the complex mechanical structure of the electro-optical pod, making the interceptor multicopter compact. However, the coupling of the camera and multicopter motion makes interception tasks difficult. To solve this problem, an Image-Based Visual Servoing (IBVS) controller is proposed to make the interception fast and accurate. Then, in response to the time delay of sensor imaging and image processing relative to attitude changes in high-speed scenarios, a Delayed Kalman Filter (DKF) observer is generalized to predict the current image position and increase the update frequency. Finally, Hardware-in-the-Loop (HITL) simulations and outdoor flight experiments verify that this method has a high interception accuracy and success rate. In the flight experiments, a high-speed interception is achieved with a terminal speed of 20 m/s.

cs.RO

Possible origin for the similar phase transitions in k-core and interdependent networks

The models of $k$-core percolation and interdependent networks (IN) have been extensively studied in their respective fields. A recent study has revealed that they share several common critical exponents. However, several newly discovered exponents in IN have not been explored in $k$-core percolation, and the origin of the similarity still remains unclear. Here, we investigate k-core percolation in random networks. We find that for k-core percolation,the fractality of the giant component fluctuations is manifested by a fractal fluctuation dimension, $\widetilde d_f = 3/4$, within a correlation \emph{size} $N'$ that scales as $N' \propto (p-p_c)^{-\widetildeν}$, with $\widetildeν= 2$, same as found in IN. Indeed, here, $\widetildeν\equiv d\cdot ν'$ and $\widetilde{d}_f \equiv d'_f/d$, where $ν'$ and $d'_f$ are respectively the same as the correlation \emph{length} exponent and the fractal fluctuation dimension observed in $d$-dimensional IN spatial networks. These two new exponents found here for $k$-core percolation demonstrate the same scaling behaviors as found for IN with the same critical exponents, reinforcing the similarity between the two models. Furthermore, we suggest that these two models are similar since both have two types of interactions: short-range (SR) connectivity and long-range (LR) influences. In IN the LR are the influences of dependency links while in k-core we find here that for $k=1$ and $k=2$ the influences are short range while for $k\geq3$ the influence is long range. In addition, analytical arguments for a universal hyper-scaling relation for the fractal fluctuation dimension of the $k$-core giant component and for IN as well as for any mixed-order transition are established.Our analysis enhances the comprehension of k-core percolation and supports the generalization of the concept of fractal fluctuations in mixed-order phase transitions.

physics.soc-ph

Synthesizing Robust Domains of Attraction for State-Constrained Perturbed Polynomial Systems

In this paper we propose a novel semi-definite programming based method to compute robust domains of attraction for state-constrained perturbed polynomial systems. A robust domain of attraction is a set of states such that every trajectory starting from it will approach an equilibrium while never violating a specified state constraint, regardless of the actual perturbation. The semi-definite program is constructed by relaxing a generalized Zubov's equation. The existence of solutions to the constructed semi-definite program is guaranteed and there exists a sequence of solutions such that their strict one sub-level sets inner-approximate the interior of the maximal robust domain of attraction in measure under appropriate assumptions. Some illustrative examples demonstrate the performance of our method.

eess.SY